Metadata-Version: 2.5
Name: file2records
Version: 0.2.2
Summary: Turn scientific papers (PDF, JATS and Elsevier XML, HTML, Word) into a structured dataset: one model extracts, a second audits, you review against the source text.
Project-URL: Homepage, https://github.com/Enkhnyam/chemistry-data-extractor-toolkit
Project-URL: Repository, https://github.com/Enkhnyam/chemistry-data-extractor-toolkit
Project-URL: Issues, https://github.com/Enkhnyam/chemistry-data-extractor-toolkit/issues
Author: Enkhnyam Battulga
License-Expression: MIT
License-File: LICENSE
Keywords: chemistry,data extraction,elsevier,jats,llm,pdf,scientific papers
Classifier: Development Status :: 4 - Beta
Classifier: Framework :: FastAPI
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Chemistry
Requires-Python: >=3.10
Requires-Dist: fastapi>=0.115.0
Requires-Dist: httpx>=0.27
Requires-Dist: litellm>=1.91.0
Requires-Dist: lxml>=5.0
Requires-Dist: markdown-it-py>=3.0.0
Requires-Dist: pydantic>=2.0
Requires-Dist: python-docx>=1.1
Requires-Dist: python-dotenv>=1.0.1
Requires-Dist: python-multipart>=0.0.12
Requires-Dist: tenacity>=8.0
Requires-Dist: uvicorn[standard]>=0.32.0
Provides-Extra: pdf
Requires-Dist: docling>=2.110.0; extra == 'pdf'
Description-Content-Type: text/markdown

# file2records

file2records builds a dataset from a folder of scientific papers. You say which fields a
record has. A language model reads each paper and fills them in, a second pass checks every
record against the paper, and you review the result with the source passage next to each
value.

![How file2records works: parse, extract, judge, review](https://raw.githubusercontent.com/Enkhnyam/chemistry-data-extractor-toolkit/main/docs/img/pipeline.png)

It reads PDF, JATS XML, Elsevier XML, HTML, and Word files, and works with any
OpenAI-compatible AI service, such as your university's. Everything stays on your computer
except the paper text sent to that service.

**Documentation: <https://enkhnyam.github.io/chemistry-data-extractor-toolkit/>**

## Install

```bash
pip install file2records            # XML, HTML, Word, Markdown (about 350 MB)
pip install "file2records[pdf]"     # also PDF (adds PyTorch, a few GB)
```

## Try it

```bash
file2records serve                  # a finished example in your browser, no key needed
file2records serve my-project       # your own project
```

## Use it from Python

```python
from pathlib import Path
import file2records as fr

project = fr.Project("my-project")
project.add("papers")                    # PDF, XML, HTML, Word, Markdown

project.schema = {                       # the columns of your dataset
    "catalyst": ("string", "Catalyst as the paper names it"),
    "temperature_c": ("number", "Reaction temperature in °C"),
}
project.prompt = Path("prompt.txt")      # what counts as one record

project.extract(only=r"hydrogenation")   # only papers whose text matches
project.judge()                          # a second pass checks every record
project.export("dataset.csv")            # one row per record, with its DOI
```

The model comes from a `.env` file with three values from your AI service:

```bash
FILE2RECORDS_ENDPOINT=https://chat.kiconnect.nrw/api/v1    # RWTH KI:connect, as an example
FILE2RECORDS_API_KEY=your-key
FILE2RECORDS_MODEL=gpt-oss-120b
```

[Get started](https://enkhnyam.github.io/chemistry-data-extractor-toolkit/get-started/) walks
through the whole pipeline once, in the browser and in Python, with three real papers.

## Development

```bash
uv sync                                       # includes docling for PDF tests
uv run python -m unittest discover -s tests   # no network, no key
uvx zensical serve                            # the docs, at http://localhost:8000
```

Releasing: [RELEASING.md](RELEASING.md). Licence: the code is MIT; the demo paper is CC BY
([NOTICE](src/file2records/demo/NOTICE.md)). Please cite: [CITATION.cff](CITATION.cff).
