Metadata-Version: 2.5
Name: llm-markdownify
Version: 0.6.0
Summary: Convert PDFs, images to high-quality Markdown using Vision LLMs.
Project-URL: Homepage, https://github.com/sethupavan12/Markdownify
Project-URL: Repository, https://github.com/sethupavan12/Markdownify
Project-URL: Issues, https://github.com/sethupavan12/Markdownify/issues
Author: Sethu Pavan Venkata Reddy Pastula
License: Apache-2.0
License-File: LICENSE
License-File: NOTICE
Keywords: docx,image,jpeg,jpg,litellm,llm,markdown,ocr,pdf,png,vision
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Text Processing :: Markup
Classifier: Topic :: Utilities
Requires-Python: >=3.10
Requires-Dist: anthropic<2,>=1.0
Requires-Dist: litellm<2,>=1.100
Requires-Dist: openai<3,>=2.0
Requires-Dist: pillow>=10.3.0
Requires-Dist: pydantic>=2.7.0
Requires-Dist: pypdfium2<6,>=4.30.0
Requires-Dist: tenacity>=8.2.0
Requires-Dist: tqdm>=4.66.0
Requires-Dist: typer>=0.12.3
Provides-Extra: dev
Requires-Dist: pre-commit>=3.6.0; extra == 'dev'
Requires-Dist: pytest-cov>=5.0.0; extra == 'dev'
Requires-Dist: pytest>=8.2.0; extra == 'dev'
Requires-Dist: ruff>=0.5.6; extra == 'dev'
Provides-Extra: docx
Requires-Dist: docx2pdf>=0.1.8; extra == 'docx'
Description-Content-Type: text/markdown

# llm-markdownify

[![PyPI](https://img.shields.io/pypi/v/llm-markdownify)](https://pypi.org/project/llm-markdownify/)
[![License](https://img.shields.io/badge/license-Apache--2.0-blue)](LICENSE)
[![CI](https://github.com/sethupavan12/Markdownify/actions/workflows/ci.yml/badge.svg)](https://github.com/sethupavan12/Markdownify/actions/workflows/ci.yml)

Turn PDFs, scans and images into clean Markdown with the vision model you already pay for.
Built for RAG pipelines and AI agents.

```bash
pip install llm-markdownify
markdownify report.pdf -o report.md --model gpt-5.4-mini
```

Each page goes to a vision LLM with a prompt we measure on a public benchmark, so the Markdown comes
back the way a retriever or an agent wants it:

- the text exactly as printed, in reading order, including multi-column layouts
- no running headers, footers or page numbers polluting your chunks
- math as LaTeX (`$...$`, `$$...$$`)
- tables as Markdown, or as HTML when they have merged cells, and stitched back together when they
  break across pages
- charts as a description plus the data values, diagrams as Mermaid
- old scans and handwriting transcribed, not summarized

It works with any provider: OpenAI, Anthropic, Gemini, DeepSeek, Azure, OpenRouter, or a model running
on your own machine through Ollama, LM Studio or vLLM.

![Handwritten notes converted to Markdown](examples/image.png)

## How good is it?

We measure every change on [olmOCR-Bench](https://huggingface.co/datasets/allenai/olmOCR-bench), the
benchmark most PDF-to-Markdown tools publish. It checks 1,403 real pages with about 7,000 pass/fail
tests: is this sentence present, is the page footer gone, is paragraph A before paragraph B, is this
table cell next to that one, does this equation render the same.

**llm-markdownify 0.5 with `gpt-6-luna` scores 81.6 ± 1.0 on the full benchmark.**

| System | olmOCR-Bench | Source |
|---|---|---|
| Chandra 2 | 85.9 | [Datalab](https://huggingface.co/datalab-to/chandra-ocr-2) (self-reported) |
| Mistral OCR 4 | 85.2 | [Mistral](https://mistral.ai/news/ocr-4/) (self-reported) |
| olmOCR 2 (7B model trained for this benchmark) | 82.4 | [Ai2](https://github.com/allenai/olmocr) |
| **llm-markdownify 0.5 + gpt-6-luna** | **81.6** | measured with [`evals/olmocr_bench`](evals/olmocr_bench) |
| Marker 2 (balanced) | 76.0 | [Datalab](https://github.com/datalab-to/marker) (self-reported) |
| Docling | 50.3 | [Marker's benchmark](https://github.com/datalab-to/marker) |

By category: tables 89.1, headers/footers 86.6, old scans with math 82.5, multi-column 81.4, arXiv
math 79.4, tiny text 88.7, old scans 45.1. Old, faded scans are the weak spot.

What the library adds on top of the model, measured on a fixed 105-page stratified subset:

| Same pages | Score | Median time per page |
|---|---|---|
| `gpt-6-luna` with a bare "convert this page to Markdown" prompt | 72.6 | 11.9 s |
| llm-markdownify 0.4 + `gpt-6-luna` | 72.5 (18 pages failed) | 14.8 s |
| llm-markdownify 0.5 + `gpt-5.4-mini` | 80.4 | 4.8 s |
| llm-markdownify 0.5 + `gpt-6-luna` | 84.9 | 14.6 s |

The biggest single difference is headers and footers. With the bare prompt, 24% of the checks that
running headers, footers and page numbers are gone pass (88% with llm-markdownify). Left in, that
furniture ends up in every RAG chunk. Other projects' numbers are what they
published; ours come from running the benchmark's own scorer on our output.

Reproduce it yourself with [`evals/olmocr_bench`](evals/olmocr_bench). The harness scores the Markdown
this library writes, through the same `convert()` call you would use.

## Quickstart

```bash
export OPENAI_API_KEY="sk-..."

markdownify input.pdf -o output.md --model gpt-5.4-mini   # PDF
markdownify scan.png -o scan.md --model gpt-5.4-mini      # PNG, JPG, WEBP, TIFF (multi-page), BMP, GIF
```

From Python:

```py
from llm_markdownify import convert

convert("input.pdf", "output.md", model="gpt-5.4-mini")
```

## Use any provider

The model name decides where the request goes (via [LiteLLM](https://docs.litellm.ai/docs/providers),
100+ providers). Set that provider's key and pick a vision-capable model.

| Provider | Key | Example `--model` |
|---|---|---|
| OpenAI | `OPENAI_API_KEY` | `gpt-5.4-mini` |
| Anthropic | `ANTHROPIC_API_KEY` | `anthropic/claude-sonnet-5` |
| Google Gemini | `GEMINI_API_KEY` | `gemini/gemini-2.5-flash` |
| DeepSeek | `DEEPSEEK_API_KEY` | `deepseek/deepseek-flash` |
| OpenRouter | `OPENROUTER_API_KEY` | `openrouter/anthropic/claude-sonnet-5` |
| Azure OpenAI | `AZURE_API_KEY`, `AZURE_API_BASE`, `AZURE_API_VERSION` | `azure/<deployment>` |

Anything that speaks the OpenAI API works too, including local servers with no key at all:

```bash
# Ollama, LM Studio, vLLM, llama.cpp, or a hosted OpenAI-compatible provider
markdownify input.pdf -o output.md --model openai/<model-name> --api-base http://localhost:11434/v1
```

## Large jobs, low latency

Thousands of documents and no one waiting? `markdownify-batch` sends them through the OpenAI Batch
API or Anthropic Message Batches at about half the price, finished within 24 hours:

```bash
markdownify-batch submit ./archive -o ./archive-md --model gpt-5.4-mini   # or anthropic/claude-opus-5
markdownify-batch status ./archive-md
markdownify-batch collect ./archive-md --retry-failed
```

Someone waiting on the answer? Use a fast model and skip the cross-page check:

```bash
markdownify doc.pdf -o doc.md --model gpt-5.4-mini --no-grouping --concurrency 16   # ~5 s per page
```

No API at all? Serve a vision model locally with LM Studio, Ollama, vLLM or llama.cpp and pass
`--api-base`. Measured speed and quality trade-offs, and the local setup guide, are in
[docs/speed-and-cost.md](docs/speed-and-cost.md).

## Built for production

- **Safe to call from threads.** Use `convert()` from a web server or a worker pool. (Retry, cache and
  rate-limit settings are currently process-wide, so give concurrent calls the same settings.)
- **Rate limits are waited out, real errors fail fast.** Rate limits, timeouts and 5xx responses are
  retried with backoff. A bad key or a bad request fails on the first attempt with one clear line,
  and never prints your API key.
- **Oversized pages are handled.** Page images are capped at 2048 px, below provider limits, so big
  scans don't get rejected.
- **Re-runs are free** with `--cache`: responses are cached on disk and keyed on the exact request.
- **Throughput and cost controls:** `--concurrency`, `--rate-limit`, `--max-image-px`,
  `--reasoning-effort`.

## Options

| Flag | Default | What it does |
|---|---|---|
| `--model` | `gpt-4.1-mini` or `$LLM_MARKDOWNIFY_MODEL` | any LiteLLM model name |
| `--profile` | `generic` | prompt profile: `generic`, `contracts`, or a JSON file with your own prompts |
| `--dpi` | 200 | PDF render resolution (ignored for images) |
| `--max-image-px` | 2048 | longest side of each page image sent to the model |
| `--max-group-pages` | 3 | max pages merged when a table or chart continues onto the next page |
| `--no-grouping` | | skip cross-page detection (one fewer model call per page) |
| `--temperature`, `--max-tokens`, `--reasoning-effort` | provider defaults, 16000 | generation settings |
| `--api-base` | | any OpenAI-compatible endpoint |
| `--concurrency`, `--grouping-concurrency`, `--rate-limit` | 4, same, none | throughput |
| `--cache`, `--cache-dir` | off, `~/.cache/llm-markdownify` | response cache |
| `-q`, `-v`, `--version` | | quiet, verbose, version |

Custom prompts: copy a built-in profile from
[`prompt_profiles.py`](src/llm_markdownify/prompt_profiles.py) into a JSON file with the fields `name`,
`continuation_system`, `continuation_user`, `markdown_system` and `markdown_user`, then pass
`--profile my_profile.json`.

DOCX input goes through Microsoft Word (macOS/Windows): `pip install "llm-markdownify[docx]"` and pass
`--allow-docx`. Exporting to PDF yourself is more reliable.

## Examples

The [gallery](examples/gallery.md) has about 80 inputs with their Markdown output: receipts, charts,
handwriting, formulas, forms, screenshots, scene text.

## Roadmap

Next up: an in-memory API that returns per-page results, stdout output and page ranges for agents, an
MCP server, a hybrid mode that uses the PDF's own text layer to cut cost, and presets for small local
models. Details and current numbers are in [docs/ROADMAP.md](docs/ROADMAP.md).

## Markdownify Cloud

To run Markdownify in production on your own infrastructure, with your own LLMs, or tuned for your
documents, see [markdownify.xyz](https://www.markdownify.xyz/). The cloud version adds features beyond
the open-source library and comes with hands-on integration support.

## Contributing

```bash
uv sync --all-extras --dev
uv run pytest
uv run ruff check src tests
```

See [CONTRIBUTING.md](CONTRIBUTING.md). Maintainers and coding agents: start with [AGENTS.md](AGENTS.md).

## License

Apache 2.0. If you distribute this project, keep the `LICENSE` and `NOTICE` files intact, crediting
the original author, Sethu Pavan Venkata Reddy Pastula.
