Metadata-Version: 2.5
Name: flint-slating
Version: 0.1.9
Summary: MCP server that reads PDFs (metadata, TOC, Markdown, text, images, tables) for downstream LLM consumers.
Author-email: Gary Frattarola <garyf@parkviewlab.ai>
License-Expression: MIT OR Apache-2.0
License-File: LICENSE-APACHE
License-File: LICENSE-MIT
Requires-Python: >=3.13
Requires-Dist: anyio
Requires-Dist: docling>=2.0
Requires-Dist: fastapi
Requires-Dist: httpx
Requires-Dist: mcp[cli]<2,>=1.27
Requires-Dist: pdfplumber>=0.11
Requires-Dist: pydantic
Requires-Dist: pypdf>=5.0
Requires-Dist: starlette
Requires-Dist: torch
Requires-Dist: torchvision
Requires-Dist: uvicorn[standard]
Description-Content-Type: text/markdown

<!--
SPDX-FileCopyrightText: 2026 Gary Frattarola <garyf@parkviewlab.ai>

SPDX-License-Identifier: MIT OR Apache-2.0
-->

# flint-slating

MCP server that reads PDFs and exposes them to LLM consumers as
structured Markdown, plus the usual ancillaries: metadata, outline,
images, tables.

Designed to pair with a separate "wiki" MCP server that handles the
*writing* side — an agent calls `flint-slating` to read PDFs and another
MCP to persist notes about them into a frontmattered-markdown knowledge
base.

## What it does

Built on a permissive-license PDF stack:

| Library | License | Role |
|---|---|---|
| [Docling](https://github.com/docling-project/docling) | MIT | PDF → Markdown with heading hierarchy, multi-column reading order, and Markdown tables |
| [pypdf](https://github.com/py-pdf/pypdf) | BSD-3 | metadata, TOC, page count, encryption checks, image enumeration |
| [pdfplumber](https://github.com/jsvine/pdfplumber) | MIT | per-page table extraction |

**There is no PyMuPDF, no MuPDF, no AGPL or GPL anywhere in the
dependency tree.** The `licenses` job of `license-check.yml` rejects PRs that pull in
copyleft transitive deps.

## Transports

Two transports off the same MCP server, selected via `--transport`:

| Transport | Run via | Use case |
|---|---|---|
| **Streamable-HTTP** (default) | `uvx flint-slating` or `--transport http` | Long-lived local daemon, container, or shared service. |
| **stdio** | `uvx flint-slating --transport stdio` | The standard MCP integration shape — drop into `claude_desktop_config.json` or any `mcp.json`. |

## Run

### As an HTTP daemon (default)

```bash
uvx flint-slating                    # listens on PORT (default 35833)
curl http://127.0.0.1:35833/health
```

Or pin it:

```bash
uv tool install flint-slating
flint-slating
```

### As a stdio MCP server

```bash
uvx flint-slating --transport stdio
```

Wire into Claude Code's MCP config:

```json
{
  "mcpServers": {
    "flint-slating": {
      "command": "uvx",
      "args": ["flint-slating", "--transport", "stdio"]
    }
  }
}
```

### Docker

```bash
docker run --rm \
  -p 35833:35833 \
  -v $(pwd)/pdfs:/pdfs:ro \
  -v flint-slating-data:/data \
  ghcr.io/parkviewlab/flint-slating:latest
```

Or use [`docker-compose.yml`](docker-compose.yml) for a persistent stack.

## MCP tools

All PDF tools take a `source` argument with one of:

- `{"path": "/abs/path/to/file.pdf"}` — local file
- `{"url": "https://..."}` — streamed to a content-addressed cache
- `{"bytes_b64": "...", "filename": "x.pdf"}` — base64 upload (size-capped)

| Tool | What it does |
|---|---|
| `pdf_info` | `{page_count, metadata, is_encrypted, sha256}` |
| `pdf_toc` | flat outline `[{level, title, page}]` |
| `pdf_read_text` | plain text by page range (fast — pypdf, no ML) |
| `pdf_read_markdown` | high-quality Markdown via Docling (hybrid sync/async — see below) |
| `pdf_read_chunks` | per-page Markdown chunks with tables/images/toc_items (hybrid sync/async) |
| `pdf_list_images` | enumerate images: `[{page, index, name, width, height, ext}]` |
| `pdf_extract_image` | base64 bytes of one image |
| `pdf_find_tables` | per-page Markdown tables via pdfplumber |
| `get_job_status` | poll a background job |
| `get_job_result` | fetch a finished job's artifact |
| `cancel_job` | cancel a running job |

### Hybrid sync/async

`pdf_read_markdown` and `pdf_read_chunks` run inline when
the number of pages requested is at most `SYNC_PAGE_THRESHOLD` (default 20;
all pages when `pages` is omitted, so the count is the PDF's `page_count`
then). For larger requests they
queue a background job and return a `job_id` — poll `get_job_status`
until `state=="done"`, then call `get_job_result` (or, in HTTP mode,
fetch `output_url` directly).

**stdio mode** transparently waits for the job inline — there's no HTTP
server to download from, so the originating tool call blocks until the
result is ready and returns it directly.

## HTTP endpoints (HTTP mode only)

- `GET /health` — `{ok, version, uptime_seconds}`
- `GET /admin/version` — package and dependency versions, and Docling status (`docling_model_loaded` reports whether the converter has been built, not whether a model is loaded)
- `GET /admin/jobs` — recent job list
- `GET /outputs/{job_id}/result.md` — finished Markdown
- `GET /outputs/{job_id}/result.json` — finished chunked output
- `GET /outputs/{job_id}/log.jsonl` — append-only job log
- `POST /sse` — MCP Streamable-HTTP transport

## Configuration

| Env var | Default (daemon) | Default (container) | Purpose |
|---|---|---|---|
| `PORT` | `35833` | `35833` | HTTP bind port |
| `HOST` | `0.0.0.0` | `0.0.0.0` | HTTP bind address |
| `OUTPUT_ROOT` | `./output` | `/data/output` | Per-job output dirs |
| `CACHE_ROOT` | `./cache` | `/data/cache` | Materialized URL / base64 PDFs |
| `OUTPUT_EXPIRY_DAYS` | `7` | `7` | Hourly sweep (HTTP mode only) of job directories older than N days by directory mtime; `0` disables |
| `MAX_INLINE_PDF_BYTES` | `25 MB` | `25 MB` | Cap on base64 upload size |
| `MAX_URL_PDF_BYTES` | `200 MB` | `200 MB` | Cap on URL download size |
| `SYNC_PAGE_THRESHOLD` | `20` | `20` | Inline-vs-job cutoff for Markdown conversion |
| `JOB_HISTORY_MAX` | `100` | `100` | Caps the job history kept and the `limit` of `/admin/jobs` |
| `MAX_IMAGE_EXTRACT_BYTES` | `8 MiB` | `8 MiB` | Cap on the size of one image returned by `pdf_extract_image` |
| `DOCLING_ARTIFACTS_PATH` | unset | unset | Directory of pre-downloaded Docling models; see Resource notes |
| `PUBLIC_BASE_URL` | `http://localhost:35833` | `http://localhost:35833` | Used to build `output_url` |

## Resource notes

- Docling downloads its layout models (~200–500 MB) on the first
  conversion unless `DOCLING_ARTIFACTS_PATH` is set. The container image
  does **not** pre-fetch them (pre-fetching dominated multi-arch build
  time under QEMU), so the first Markdown conversion pays the download.
  If `DOCLING_ARTIFACTS_PATH` is set, it must point at a directory that
  already holds every model (`docling-tools models download -o <dir>`,
  for instance into a mounted volume); Docling never downloads into it.
  Default: unset.
- OCR follows Docling's default options, which run OCR; there is no
  setting for it.
- pypdf, pdfplumber, and the URL / base64 paths are fast and have no ML
  overhead — use `pdf_info`, `pdf_toc`, `pdf_read_text`, and
  `pdf_find_tables` whenever Markdown isn't strictly needed.

## Releasing

Tag-driven CI publishes to both PyPI (`flint-slating`) and GHCR (`ghcr.io/parkviewlab/flint-slating`):

The release runs from the CLI in the `flint-slating-main` worktree, as the
ParkviewLab handbook's `releases.md` gives it under "Cutting a release" and
"The release's last step":

```bash
git pull --ff-only                        # sync main
git -C ../flint-slating-develop pull --ff-only   # sync develop
git merge --no-ff develop                 # promote develop to main
git bump <patch|minor|major|release|X.Y.Z>
git release                               # annotated tag, derived from pyproject.toml
git push --follow-tags                    # the tag push fires the release workflow
git back-merge                            # the release's last step
```

The release workflow refuses a tag that does not match `pyproject.toml`'s `version`, that still carries a dev marker (`.devN`), that is not on `origin/main`, or that is not greater than the previous release tag.

### Commit message convention

After the publish jobs, a **changelog** job generates the new `CHANGELOG.md` section — an LLM-written "Highlights" paragraph plus a categorized list written by dev-tools' `generate-changelog` — commits it back to `main`, and creates the GitHub Release with the same content as its body. Categorization uses [Conventional Commits](https://www.conventionalcommits.org/) prefixes (the full list is in the ParkviewLab handbook's `commits-and-changelogs.md`):

| Title | Group in the notes | Notes |
|---|---|---|
| any type with `!` after it (`feat!:`), or a breaking-change footer | Breaking changes | listed there once, whatever its type |
| `feat:` | Features | user-visible |
| `fix:` | Bug fixes | user-visible |
| `perf:` | Performance | user-visible |
| `refactor:` | Refactor | |
| `docs:` | Docs | |
| `test:` | Tests | |
| `revert:` | Reverts | GitHub's Revert button titles a PR `Revert "…"`, which has no type |
| `build:` / `chore:` / `ci:` / `style:` | Maintenance | |
| any other title | Other changes | the whole title |
| a commit with no pull request | Direct commits | its subject and short hash |

A title without a recognised type is not dropped: it is listed whole under Other changes. So prefix your PR titles, and correct a title before the merge, since retitling afterwards does not change the commit. The groups appear in the order above, and an empty group is left out.

## License

Licensed under either of

- Apache License, Version 2.0 ([LICENSE-APACHE](LICENSE-APACHE) or
  <http://www.apache.org/licenses/LICENSE-2.0>), or
- MIT license ([LICENSE-MIT](LICENSE-MIT) or
  <http://opensource.org/licenses/MIT>)

at your option. In SPDX terms: `MIT OR Apache-2.0`.

Unless you explicitly state otherwise, any contribution intentionally submitted
for inclusion in this work by you shall be dual-licensed as above, without any
additional terms or conditions. See [LICENSING.md](LICENSING.md).

flint-slating only depends on permissive-licensed libraries; the `licenses` job
of `license-check.yml` enforces this on every PR. torch and torchvision are
pinned to the [CPU-only PyTorch wheel index](https://download.pytorch.org/whl/cpu)
for the GHCR image and for `uv sync` from a checkout (`uv.lock` resolves torch
to the CPU wheels), so those do not bundle NVIDIA's proprietary CUDA libraries.
A PyPI install (`uvx flint-slating`, `uv tool install flint-slating`) does not
read that pin and gets PyPI's torch, which on Linux pulls NVIDIA's CUDA
libraries unless torch is installed from the CPU index. Inference runs on CPU
on Linux/Windows and on MPS (Metal) on Apple Silicon. See
[THIRD_PARTY_LICENSES.md](THIRD_PARTY_LICENSES.md) for the per-dependency
license breakdown.

---
<sub>© 2026 Gary Frattarola · Licensed under [MIT](LICENSE-MIT) OR [Apache-2.0](LICENSE-APACHE) · part of [ParkviewLab](https://github.com/ParkviewLab)</sub>
