Metadata-Version: 2.4
Name: deepresearch-flow
Version: 0.14.1
Summary: Workflow tools for paper extraction, review, and research automation.
Author-email: DengQi <dengqi935@gmail.com>
License: MIT License
        
        Copyright (c) 2025 DengQi
        
        Permission is hereby granted, free of charge, to any person obtaining a copy
        of this software and associated documentation files (the "Software"), to deal
        in the Software without restriction, including without limitation the rights
        to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
        copies of the Software, and to permit persons to whom the Software is
        furnished to do so, subject to the following conditions:
        
        The above copyright notice and this permission notice shall be included in all
        copies or substantial portions of the Software.
        
        THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
        IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
        FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
        AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
        LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
        OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
        SOFTWARE.
        
Project-URL: Homepage, https://github.com/nerdneilsfield/ai-deepresearch-flow
Project-URL: Repository, https://github.com/nerdneilsfield/ai-deepresearch-flow
Project-URL: Issues, https://github.com/nerdneilsfield/ai-deepresearch-flow/issues
Keywords: research,papers,pdf,ocr,llm,workflow
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Requires-Python: >=3.12
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: anthropic>=0.77.0
Requires-Dist: click>=8.1.7
Requires-Dist: coloredlogs>=15.0.1
Requires-Dist: dashscope>=1.25.10
Requires-Dist: lancedb>=0.20.0
Requires-Dist: google-auth>=2.48.0
Requires-Dist: google-genai>=1.60.0
Requires-Dist: httpx>=0.27.0
Requires-Dist: jinja2>=3.1.3
Requires-Dist: json-repair>=0.55.1
Requires-Dist: jsonschema>=4.26.0
Requires-Dist: markdown-it-py>=3.0.0
Requires-Dist: fastmcp>=3.4.2
Requires-Dist: mdit-py-plugins>=0.4.0
Requires-Dist: pyarrow>=18.0.0
Requires-Dist: pypdf>=6.6.2
Requires-Dist: pylatexenc>=2.10
Requires-Dist: pybtex>=0.24.0
Requires-Dist: python-multipart>=0.0.20
Requires-Dist: rich>=14.3.1
Requires-Dist: rumdl>=0.1.6
Requires-Dist: starlette>=0.52.1
Requires-Dist: tiktoken>=0.9.0
Requires-Dist: tqdm>=4.67.2
Requires-Dist: uvicorn>=0.27.1
Dynamic: license-file

<p align="center">
  <img src=".github/assets/logo.png" width="140" alt="ai-deepresearch-flow logo" />
</p>

<h3 align="center">ai-deepresearch-flow</h3>

<p align="center">
  <em>From documents to deep research insight — automatically.</em>
</p>

<p align="center">
  <a href="README.md">English</a> | <a href="README_ZH.md">中文</a>
</p>

<p align="center">
  <a href="https://github.com/nerdneilsfield/ai-deepresearch-flow/actions">
    <img src="https://img.shields.io/github/actions/workflow/status/nerdneilsfield/ai-deepresearch-flow/push-to-pypi.yml?style=flat-square" />
  </a>
  <a href="https://pypi.org/project/deepresearch-flow/">
    <img src="https://img.shields.io/pypi/v/deepresearch-flow?style=flat-square" />
  </a>
  <a href="https://pypi.org/project/deepresearch-flow/">
    <img src="https://img.shields.io/pypi/pyversions/deepresearch-flow?style=flat-square" />
  </a>
  <a href="https://hub.docker.com/r/nerdneils/deepresearch-flow">
    <img src="https://img.shields.io/docker/v/nerdneils/deepresearch-flow?style=flat-square" />
  </a>
  <a href="https://github.com/nerdneilsfield/ai-deepresearch-flow/pkgs/container/deepresearch-flow">
    <img src="https://img.shields.io/badge/ghcr.io-nerdneilsfield%2Fdeepresearch-flow-0f172a?style=flat-square" />
  </a>
  <a href="https://github.com/nerdneilsfield/ai-deepresearch-flow/blob/main/LICENSE">
    <img src="https://img.shields.io/github/license/nerdneilsfield/ai-deepresearch-flow?style=flat-square" />
  </a>
  <a href="https://github.com/nerdneilsfield/ai-deepresearch-flow/stargazers">
    <img src="https://img.shields.io/github/stars/nerdneilsfield/ai-deepresearch-flow?style=flat-square" />
  </a>
  <a href="https://pypi.org/project/deepresearch-flow">
    <img alt="PyPI - Version" src="https://img.shields.io/pypi/v/deepresearch-flow">
  </a>
  <a href="https://github.com/nerdneilsfield/ai-deepresearch-flow/issues">
    <img src="https://img.shields.io/github/issues/nerdneilsfield/ai-deepresearch-flow?style=flat-square" />
  </a>
</p>

---

## Core Pain Points

- **OCR Chaos**: Raw markdown from OCR tools is often broken — tables drift, formulas break, references are non-clickable.
- **Translation Nightmares**: Translating technical papers often destroys code blocks, LaTeX formulas, and table structures.
- **Information Overload**: Extracting structured insights (authors, venues, summaries) from hundreds of PDFs manually is impossible.
- **Context Switching**: Managing PDFs, summaries, and translations in different windows kills focus.

## Solution

DeepResearch Flow provides a unified pipeline to **Repair**, **Translate**, **Extract**, and **Serve** your research library.

## Key Features

- **Smart Extraction** — Turn unstructured Markdown into schema-enforced JSON (summaries, metadata, Q&A) using LLMs.
- **Precision Translation** — Translate OCR Markdown to Chinese/Japanese while freezing formulas, code, tables, and references.
- **Local Knowledge DB** — Web UI with Split View (Source/Translation/Summary), full-text search, and multi-dimensional filtering.
- **Snapshot + API Serve** — Production-ready SQLite snapshot with static assets and read-only JSON API.
- **OCR Post-Processing** — Fix broken references, merge split paragraphs, repair LaTeX and Mermaid diagrams.
- **Semantic Search** — LanceDB-backed vector search with hybrid recall and cloud reranking.
- **MCP Integration** — FastMCP server for AI agent access with bounded read tools, static-bearer Streamable HTTP/SSE, and GitHub OAuth at `/oauth/mcp`.

---

## Quick Start

### 1) Installation

```bash
uv pip install deepresearch-flow
# or: pip install deepresearch-flow
```

### 2) Configuration

```bash
cp config.example.toml config.toml
```

Minimal config with weighted multi-provider routing:

```toml
main_model = [
  { model = "openai/gpt-4o-mini", weight = 4 },
  { model = "claude/claude-sonnet-4-5-20250929", weight = 1 }
]

[[providers]]
name = "openai"
type = "openai_compatible"
base = [
  { url = "https://api.openai.com/v1", weight = 1, key = [
    { value = "env:OPENAI_API_KEY", weight = 4 }
  ] }
]
models = [
  { model_name = "gpt-4o-mini", is_support_json_schema = true }
]

[[providers]]
name = "claude"
type = "claude"
base = [
  { url = "https://api.anthropic.com", weight = 1, key = [
    { value = "env:ANTHROPIC_API_KEY", weight = 1 }
  ] }
]
models = [
  { model_name = "claude-sonnet-4-5-20250929" }
]
```

Keys use `env:VAR_NAME` syntax to keep secrets out of config files. Multiple providers (Ollama, Gemini, DashScope, Azure OpenAI) are supported. For full configuration options (embedding, rerank, translator defaults, search), see `config.example.toml`.

### 3) The "Zero to Hero" Workflow

Start with `./pdfs/` and, optionally, `./papers.bib`. You do not need an
existing JSON library, SQLite database, or processed Markdown directory.

The workflow produces these roots:

```text
pdfs/ + papers.bib
  → ocr_output/
  → md_simple/            # local image files
  → md_base64/            # images embedded as data URLs
  ├─ summary_json/<template>.json
  └─ md_base64_translated/
```

#### Run the data-root script

Place PDFs in `DATA_ROOT/pdf/`. From the repository root, prepare `config.toml`
and `ocr.toml`, then install the Node dependencies with `npm install` for Mermaid validation.
Run the full workflow:

```bash
uv run python scripts/process_data_root.py DATA_ROOT \
  --config config.toml --ocr-config ocr.toml --model openai/gpt-4o-mini
```

Replace the model with a configured `provider/model`. Outputs are `ocr/`,
`md_simple/`, `md_base64/`, `simple.json`, `deep_read.json`, and
`md_base64_translated/`, all under `DATA_ROOT`. The script repairs OCR before
organizing Markdown. After extraction and translation, it repairs formatting and
formulas in both JSON files and translated Markdown, then repairs Mermaid only
in `deep_read.json`. Translation defaults to Chinese; use `--target-lang` to change it.

OCR and organize can run independently without an LLM model:

```bash
uv run python scripts/process_data_root.py DATA_ROOT --steps ocr --ocr-config ocr.toml
uv run python scripts/process_data_root.py DATA_ROOT --steps organize
```

Select steps with `--steps`, or run from one step through the end with `--from-step`:

```bash
# Requires existing md_simple/ and md_base64/
uv run python scripts/process_data_root.py DATA_ROOT \
  --model openai/gpt-4o-mini --steps simple,deep_read,translate

# Requires md_base64/ and both JSON files to exist
uv run python scripts/process_data_root.py DATA_ROOT \
  --model openai/gpt-4o-mini --from-step translate
```

Available steps, in execution order:

```text
ocr, fix, fix-math, organize, simple, deep_read, translate,
fix-simple, fix-math-simple, fix-deep_read, fix-math-deep_read,
fix-translated, fix-math-translated, fix-mermaid-deep_read
```

The two selection options are mutually exclusive. `--steps` accepts comma-separated
names, runs them in workflow order, and runs duplicate names only once. It does not
add upstream steps. Inputs must exist or be produced by an earlier selected step.
Only selected steps require their configuration files and tools; translation alone
needs neither PDFs nor OCR configuration nor `mmdc`. `--model` is required only for extraction, translation, formula repair, and Mermaid repair.
OCR, formatting fixes, and organize can run without it.
Add `--dry-run` to preview commands without creating files or calling providers.

Error reports and intermediate artifacts stay under `DATA_ROOT/logs/`.
Commands keep the caller working directory, except extraction, which runs under
`logs/` to contain its relative intermediate-output directory.
Each command displays its output, errors, and native progress bars directly in the
current terminal; output is not redirected to files. `progress.json` records the
latest run, and `pipeline.log` records step start, completion, and failure events.
Configuration paths resolve
against the caller's directory; use absolute paths for custom file paths inside configs.
Nonzero command exits stop the workflow. Some commands return zero despite per-file
failures, so check error reports and terminal output even when progress says `completed`.
Reruns execute the selected steps using each command's existing-output skip rules,
not the progress file. Repairs modify generated outputs in place, not source PDFs.

#### Step 1: OCR PDFs or Images

Copy and configure the OCR settings:

```bash
cp ocr.example.toml ocr.toml
# Set: export PADDLE_OCR_TOKEN=xxx
# The example uses PaddleOCR-VL-1.6's asynchronous Job API.
# Adjust poll_interval_seconds and job_timeout_seconds in ocr.toml if needed.

uv run deepresearch-flow recognize ocr ./pdfs \
  --config ocr.toml \
  --output-dir ./ocr_output
# Processes up to 4 files concurrently by default; override with --workers 2.
```

The backend writes MinerU-compatible layouts: one `full.md` and `images/`
directory per document. The configured timeout stops local polling only; it does
not cancel the remote PaddleOCR job.

#### Step 2: Repair Nested OCR Outputs

Each OCR document is nested below `ocr_output/`, so both repair commands must
use `-r`:

```bash
# Repair Markdown structure in every OCR document
uv run deepresearch-flow recognize fix \
  --input ./ocr_output -r --in-place

# Repair LaTeX formulas in every OCR document
uv run deepresearch-flow recognize fix-math \
  --input ./ocr_output -r \
  --model openai/gpt-4o-mini \
  --in-place
```

<p align="center">
  <img src=".github/assets/fix.png" width="70%" alt="fix" />
</p>

<p align="center">
  <img src=".github/assets/fix-math.png" width="70%" alt="fix math" />
</p>

#### Step 3: Organize Source Markdown

Create both source representations in one pass. `organize` also needs `-r`
to discover nested OCR layouts. Do not pass `--fix`: Step 2 has already
repaired the OCR source.

```bash
uv run deepresearch-flow recognize organize \
  --input ./ocr_output -r \
  --output-simple ./md_simple \
  --output-base64 ./md_base64
```

`md_simple/` keeps image files under `md_simple/images/`; `md_base64/`
embeds images, so it is the translation input.

#### Step 4: Generate Structured Summaries

Generate one JSON bundle per selected prompt template. This example uses
`deep_read`; repeat it for every template you need, naming each output
`./summary_json/<template>.json`.

```bash
uv run deepresearch-flow paper extract \
  --input ./md_simple \
  --model openai/gpt-4o-mini \
  --prompt-template deep_read \
  --output ./summary_json/deep_read.json
```

<p align="center">
  <img src=".github/assets/extract.png" width="70%" alt="extract" />
</p>

#### Step 4.1: Verify and Retry Summary Fields

Keep verification reports outside `summary_json/` so JSON repair scans only
summary bundles. `paper db verify` validates the JSON bundle; it does not
require a database. Repeat this unit for every selected template.

```bash
uv run deepresearch-flow paper db verify \
  --input-json ./summary_json/deep_read.json \
  --prompt-template deep_read \
  --output-json ./summary_verify/deep_read.json

uv run deepresearch-flow paper extract \
  --input ./md_simple \
  --model openai/gpt-4o-mini \
  --prompt-template deep_read \
  --output ./summary_json/deep_read.json \
  --retry-list-json ./summary_verify/deep_read.json
```

<p align="center">
  <img src=".github/assets/verify.png" width="70%" alt="verify" />
</p>

#### Step 5: Translate Base64 Markdown

```bash
uv run deepresearch-flow translator translate \
  --input ./md_base64 \
  --target-lang zh \
  --model openai/gpt-4o-mini \
  --fix-level moderate \
  --output-dir ./md_base64_translated
```

#### Step 6: Repair Generated Artifacts

Repair every summary JSON after extraction and retry. JSON inputs require
`--json`; keep `-r` because the directory can contain multiple template
bundles.

```bash
uv run deepresearch-flow recognize fix \
  --input ./summary_json --json -r --in-place

uv run deepresearch-flow recognize fix-math \
  --input ./summary_json --json -r \
  --model openai/gpt-4o-mini \
  --in-place

uv run deepresearch-flow recognize fix-mermaid \
  --input ./summary_json --json -r \
  --model openai/gpt-4o-mini \
  --in-place
```

<p align="center">
  <img src=".github/assets/fix-mermaid.png" width="70%" alt="fix mermaid" />
</p>

Repair the translated Markdown separately. Mermaid repair is only part of the
summary JSON branch.

```bash
uv run deepresearch-flow recognize fix \
  --input ./md_base64_translated -r --in-place

uv run deepresearch-flow recognize fix-math \
  --input ./md_base64_translated -r \
  --model openai/gpt-4o-mini \
  --in-place
```

#### Step 7: Build a Snapshot Database or Serve Locally

Both commands consume the repaired summary JSON. Add one `--input` option for
each additional file in `summary_json/`; neither command consumes the other
command's output.

Build a persistent SQLite snapshot and static assets:

```bash
uv run deepresearch-flow paper db snapshot build \
  --input ./summary_json/deep_read.json \
  --bibtex ./papers.bib \
  --md-root ./md_simple \
  --md-translated-root ./md_base64_translated \
  --pdf-root ./pdfs \
  --output-db ./dist/paper_snapshot.db \
  --static-export-dir ./dist/paper-static
```

Or start the local web UI directly from the same inputs:

```bash
uv run deepresearch-flow paper db serve \
  --input ./summary_json/deep_read.json \
  --bibtex ./papers.bib \
  --md-root ./md_simple \
  --md-translated-root ./md_base64_translated \
  --pdf-root ./pdfs \
  --host 127.0.0.1
```

If you have no BibTeX file, omit `--bibtex ./papers.bib`.

#### Step 8: Add Semantic Search (Optional)

Build a LanceDB vector index from the same repaired summaries and Markdown
roots:

```bash
uv run deepresearch-flow paper embed \
  --config ./config.toml \
  --input ./summary_json/deep_read.json \
  --md-root ./md_simple \
  --md-translated-root ./md_base64_translated \
  --max-concurrency 4 \
  --document-window 8 \
  --output-embed-db ./paper_vectors
```

Serve with semantic search enabled:

```bash
uv run deepresearch-flow paper db serve \
  --input ./summary_json/deep_read.json \
  --bibtex ./papers.bib \
  --md-root ./md_simple \
  --md-translated-root ./md_base64_translated \
  --pdf-root ./pdfs \
  --embed-db ./paper_vectors \
  --search-access-token "your-token"
```

#### Step 9: MCP Integration (Optional)

The project exposes bounded MCP tools for AI agent access via FastMCP. See the [MCP documentation](docs/en/api-and-mcp.md#mcp) for endpoint, auth, and tool reference.

---

## Further Reading

- **[Advanced Workflows](docs/en/workflow.md)** — Incremental builds, merging JSON/BibTeX, supplementing templates
- **[Deployment](docs/en/deployment.md)** — CDN serving, Nginx/Caddy config, Docker, Compose
- **[API & MCP](docs/en/api-and-mcp.md)** — Admin API, push/push-semantic, MCP endpoints, auth, and tools
- **[Reference](docs/en/reference.md)** — Translator, Extract, DB & Recognize in detail
- **[Snapshot Management](docs/en/snapshot-management.md)** — Snapshot migration, supplement, update

---

<p align="center">
  Built with love for the Open Science community.
</p>
