Metadata-Version: 2.4
Name: mistocr
Version: 0.7.0
Summary: Batch OCR for PDFs with heading restoration and visual content integration
Author-email: Solveit <nobody@fast.ai>
License-Expression: Apache-2.0
Project-URL: Repository, https://github.com/franckalbinet/mistocr
Project-URL: Documentation, https://franckalbinet.github.io/mistocr
Keywords: nbdev,jupyter,notebook,python
Classifier: Natural Language :: English
Classifier: Intended Audience :: Developers
Classifier: Development Status :: 4 - Beta
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: fastcore
Requires-Dist: mistralai>=2
Requires-Dist: pillow
Requires-Dist: dotenv
Requires-Dist: python-fastllm
Requires-Dist: aidialog
Requires-Dist: PyPDF2
Requires-Dist: pycachy
Dynamic: license-file

# Mistocr


<!-- WARNING: THIS FILE WAS AUTOGENERATED! DO NOT EDIT! -->

mistocr converts PDFs into clean markdown for RAG and LLM pipelines. It runs Mistral’s OCR, then uses an LLM to fix the heading hierarchy and to describe the images in the document.

OCR is often treated as a trivial step in AI pipelines. In practice, a poorly converted PDF gives every downstream step poor input. mistocr came out of months of processing large collections of niche-domain PDFs. It handles three problems that raw OCR output leaves:

- **Heading hierarchy.** OCR often gets heading levels wrong in long documents. mistocr sends all the headings to an LLM, which returns the corrected levels.
- **Visual content.** mistocr classifies each image as informative or decorative, and adds a description of it to the markdown. Charts, figures and diagrams become searchable text.
- **Cost.** mistocr sends a single PDF straight to Mistral’s OCR API. It sends several PDFs through Mistral’s batch API, which costs half as much: \$1 instead of \$2 per 1,000 pages.

mistocr uses Mistral’s `mistral-ocr-latest` model for OCR. It uses [fastllm](https://github.com/AnswerDotAI/fastllm) for the LLM steps, which supports Anthropic, OpenAI and Gemini models.

> [!NOTE]
>
> This [tutorial](https://share.solve.it.com/d/97f75412ca949af76a5945b4dfc443c7) shows mistocr on a real PDF. It also shows how the fixed headings let you navigate a long document by its structure.

## Get Started

Install mistocr from [PyPI](https://pypi.org/project/mistocr). It needs Python 3.10 or later.

``` sh
$ pip install mistocr
```

Set your API keys as environment variables. The OCR step needs `MISTRAL_API_KEY`. The heading and image steps need the key for your LLM provider: `ANTHROPIC_API_KEY` for the default Claude model, `OPENAI_API_KEY` for OpenAI models, or `GEMINI_API_KEY` for Gemini models.

``` python
import os
os.environ['MISTRAL_API_KEY'] = 'your-key-here'
os.environ['ANTHROPIC_API_KEY'] = 'your-key-here'
```

### Complete Pipeline

#### Single file processing

Run OCR, heading fixes and image descriptions on one PDF:

``` python
from mistocr.pipeline import pdf_to_md
await pdf_to_md('files/test/resnet.pdf', 'files/test/md_test')
```

    mistocr.pipeline - INFO - Step 1/3: Running OCR on files/test/resnet.pdf...
    mistocr.core - INFO - Waiting for batch job 4ec899ca-ada8-4fa7-8894-0191ff6ac4e5 (initial status: QUEUED)
    mistocr.core - DEBUG - Job 4ec899ca-ada8-4fa7-8894-0191ff6ac4e5 status: QUEUED (elapsed: 0s)
    mistocr.core - DEBUG - Job 4ec899ca-ada8-4fa7-8894-0191ff6ac4e5 status: RUNNING (elapsed: 2s)
    mistocr.core - DEBUG - Job 4ec899ca-ada8-4fa7-8894-0191ff6ac4e5 status: RUNNING (elapsed: 2s)
    mistocr.core - DEBUG - Job 4ec899ca-ada8-4fa7-8894-0191ff6ac4e5 status: RUNNING (elapsed: 2s)
    mistocr.core - INFO - Job 4ec899ca-ada8-4fa7-8894-0191ff6ac4e5 completed with status: SUCCESS
    mistocr.pipeline - INFO - Step 2/3: Fixing heading hierarchy...
    mistocr.pipeline - INFO - Step 3/3: Adding image descriptions...

    Describing 12 images...

    mistocr.pipeline - INFO - Done!

    Saved descriptions to /tmp/tmp62c7_ac1/resnet/img_descriptions.json
    Adding descriptions to 12 pages...
    Done! Enriched pages saved to files/test/md_test

[`pdf_to_md`](https://franckalbinet.github.io/mistocr/pipeline.html#pdf_to_md) does the following:

1.  Runs OCR on the PDF with Mistral’s OCR API.
2.  Fixes the heading hierarchy.
3.  Describes the images, such as charts and diagrams, and adds the descriptions to the markdown.
4.  Saves the result to `files/test/md_test`.

The output folder looks like this:

    files/test/md_test/
    ├── img/
    │   ├── img-0.jpeg
    │   ├── img-1.jpeg
    │   └── ...
    ├── page_1.md
    ├── page_2.md
    └── ...

Each image link in a page is followed by its description:

``` markdown
![Figure 1](img/img-0.jpeg)
AI-generated image description:
___
A residual learning block...
___
```

[`read_pgs`](https://franckalbinet.github.io/mistocr/core.html#read_pgs) reads the processed document back as one markdown string:

``` python
from mistocr.core import read_pgs
md = read_pgs('files/test/md_test')
print(md[:500])
```

    # Deep Residual Learning for Image Recognition  ... page 1

    Kaiming He Xiangyu Zhang Shaoqing Ren Jian Sun<br>Microsoft Research<br>\{kahe, v-xiangz, v-shren, jiansun\}@microsoft.com


    ## Abstract ... page 1

    Deeper neural networks are more difficult to train. We present a residual learning framework to ease the training of networks that are substantially deeper than those used previously. We explicitly reformulate the layers as learning residual functions with reference to the layer inputs, ins

[`read_pgs`](https://franckalbinet.github.io/mistocr/core.html#read_pgs) joins all pages by default. Pass `join=False` to get a list of pages. Each heading ends with `... page N`, the page it comes from.

### Advanced Usage

**Batch OCR for entire folders:**

``` python
from mistocr.core import ocr_pdf

# OCR all PDFs in a folder using Mistral's batch API
output_dirs = ocr_pdf('path/to/pdf_folder', dst='output_folder')
```

**Custom models and prompts for heading fixes:**

``` python
from mistocr.refine import fix_hdgs

# Use a different model or custom prompt
await fix_hdgs('ocr_output/doc1', 
               model='gpt-4o',
               prompt=your_custom_prompt)
```

**Custom image description with rate limiting:**

``` python
from mistocr.refine import add_img_descs

# Use a different model, and change how fast images are sent
await add_img_descs('ocr_output/doc1',
                    model='claude-opus-4-6',
                    semaphore=5,  # concurrent requests, default 2
                    delay=0.5)    # seconds to wait after each request, default 1
```

For complete control over each pipeline step, see the [core](https://fr.anckalbi.net/mistocr/core.html), [refine](https://fr.anckalbi.net/mistocr/refine.html), and [pipeline](https://fr.anckalbi.net/mistocr/pipeline.html) module documentation.

## Developer Guide

mistocr is developed with [nbdev](https://nbdev.fast.ai). If you’re new to nbdev, start with its [getting started guide](https://nbdev.fast.ai/getting_started.html).

### Install mistocr in Development mode

``` sh
# install mistocr in development mode
$ pip install -e .

# edit the notebooks in nbs/

# export the notebooks to mistocr, run the tests and clean the notebooks
$ nbdev-prepare
```

### Documentation

The documentation is on the [repository](https://github.com/franckalbinet/mistocr)’s [GitHub Pages](https://franckalbinet.github.io/mistocr/). The package is on [PyPI](https://pypi.org/project/mistocr/).
