Metadata-Version: 2.5
Name: gtrex
Version: 0.1.0a1
Summary: Grounded text extraction you can check: verifiable quotation grounding for LLM extraction pipelines, with declarative normalization ladders, repair-and-reject cascades, and calibrated acceptance thresholds.
License-Expression: Apache-2.0
License-File: LICENSE
Keywords: attribution,conformal-prediction,grounded-extraction,grounding,hallucination,information-extraction,llm,provenance,reproducibility
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Text Processing :: Linguistic
Classifier: Typing :: Typed
Requires-Python: >=3.10
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == 'dev'
Requires-Dist: ruff>=0.4; extra == 'dev'
Description-Content-Type: text/markdown

# gtrex

*Grounded text extraction you can check.*

> **Status: alpha (0.1.0a1).** The API and output formats may still change.
> Known limitations are listed near the end of this file, so that they travel
> with the package.

**For fixed extractions, the reported grounding rate depends on the
verification policy.** "The quotation appears in the source" always involves
choices: which normalizations to apply, what to do with an ellipsis, whether
stitched fragments count, which tier to call grounded. A rate reported without
those choices cannot be compared with anyone else's. `gtrex` makes the policy
explicit, versioned and fingerprinted, so a grounding claim is something
another lab can check.

```bash
pip install gtrex        # no runtime dependencies, Python >= 3.10
```

```python
from gtrex import Policy, Verifier

verifier = Verifier(filing_text, Policy())
cert = verifier.verify(model_quote)

cert.grounded   # True
cert.tier       # 'typography'  -- how much normalization it took
cert.start      # 48213         -- offsets into the original text, at every tier
cert.span_text  # the document's own characters
```

---

## What it does

**1. A declarative normalization ladder.** Five cumulative tiers, each
declaring whether it is *sound*, i.e. whether it can merge two source texts
that say different things. Verification reports the tightest tier that
contains the whole quotation, not a boolean. On the eight claims of
`examples/quickstart.py`:

```
tier          at tier   cumulative     share
--------------------------------------------
exact               1            1   12.50%
whitespace          0            1   12.50%
typography          1            2   25.00%
case                1            3   37.50%
skeleton            1            4   50.00%   <- UNSOUND, not counted
stitched            1            -   12.50%
unlocated           3            -   37.50%
```

The `skeleton` tier is marked unsound because the collision is real, not
hypothetical:

```
"Operating loss was ($10.0 million)."  ->  operatinglosswas100million
"Operating loss was $100 million."     ->  operatinglosswas100million
```

Accepting it requires passing `allow_unsound=True` explicitly.

**2. A repair-and-reject cascade.** When a quotation is a near-miss, the
cascade does not accept the model's wording: it locates the source span and
**substitutes the document's own characters**, keeping the model's original
wording and the certificates before and after. Anything that cannot be located
above the threshold you set is rejected. There is no clipping stage: an
invented tail is not repaired by cutting it off.

```
verify -> fragment salvage -> locate-and-substitute -> reject
```

**3. Two guards for edits no threshold can catch.** A changed figure and a
dropped negation both score high, because they share almost every token with
the source, and both would yield a record that is *verbatim source text and
still does not say what the model extracted*. From the quickstart:

```
claim: "Operating loss was $100 million for the year ended ..."    score 0.87
span:  "Operating loss was ($10.0 million) for the year ended ..." -> numeric_guard

claim: "We believe that we will be able to compete effectively ..."     score 0.97
span:  "We believe that we will not be able to compete effectively ..." -> polarity_guard
```

These are checks, not weights, and they are partial (see the limitations
below): they block two common catastrophic edits; they do not prove that a
substitution preserves meaning.

**4. A threshold you calibrate, not guess.** The substitution threshold has no
default, and repairing without one is an error. `gtrex calibrate` estimates
one with Learn-then-Test from labelled repairs, and states the assumption the
guarantee rests on (see "Guarantees, precisely" below).

**5. A manifest.** Policy, fingerprint, package version, every tier's
soundness rationale, calibration, and the ladder profile, written next to the
data so the assumptions travel with it.

---

## Why a ladder instead of a number

Three failure modes a single grounding rate cannot distinguish:

| | What it looks like | What it is |
|---|---|---|
| **Stitched** | `"We depend on suppliers ... could materially affect results"` | Both fragments are genuine; the quotation is not. |
| **Skeleton collision** | `"Operating loss was $100 million"` | Matches the source only after punctuation is deleted, against a figure that is ten times smaller and negative. |
| **Parser drift** | quotation fails everywhere | The source text contains `[[HEADING_CANDIDATE]]` markers the model never saw. Declare them with `marker_pattern`. |

A verifier that reports one number will call the first two "grounded" and the
third "hallucinated". All three verdicts would be wrong.

---

## CLI

```bash
gtrex explain                       # print the policy and every assumption
gtrex verify   claims.jsonl -o certified.jsonl --manifest run.json
gtrex repair   claims.jsonl -o repaired.jsonl --threshold 0.97 --drop-rejected
gtrex sensitivity claims.jsonl      # the same claims under three policies
gtrex calibrate labelled.jsonl --alpha 0.05 --compare 0.65
```

(`0.97` stands for a threshold you have calibrated; there is no default.)

Input is JSON Lines. Each record needs a `claim` and either an inline `source`
or a `source_path` (read as UTF-8, byte for byte). Every other field is passed
through untouched and in input order; the result is added under one key,
`gtrex`. A record that already has a `gtrex` field is refused unless you pass
`--replace-existing`.

```jsonl
{"id": "rf-04", "claim": "We depend on a limited number of suppliers", "source_path": "filings/acme-2023/item1a.txt"}
```

`gtrex repair` stores the source's own text in `claim`, and keeps the model's
wording in `gtrex.claim_original`, with `gtrex.before` (the certificate of the
claim as given) and `gtrex.after` (the certificate of the stored text).

`sensitivity` is the experiment worth running first on your own data. On the
quickstart's eight claims:

```
policy       fingerprint      accept_through   grounded
--------------------------------------------------------
default      aceccc83810752ce case              37.50%
strict       e5d0e7e668dd009a exact             12.50%
permissive   5c7d533662b87aca skeleton          50.00%

Same 8 claims, same sources. The reported grounding rate varies by 37.5
percentage points across these three verification policies.
```

The three policies differ in more than normalization (ellipsis handling and
contiguity too); `gtrex explain --policy strict` prints each one in full.

---

## Guarantees, precisely

`calibrate_risk` implements Learn-then-Test with fixed-sequence testing over a
descending threshold grid, using the exact binomial tail as the p-value. The
controlled quantity is the joint risk `P(accepted and unfaithful)`. The
guarantee, `P(risk <= alpha) >= 1 - delta` over the calibration draw, holds
**only if the calibration examples are independent draws from the claims you
will repair**. Several examples from one document, or several perturbations of
one quotation, are not independent: give each example a `group` (document id)
and `calibrate_risk` refuses repeated groups. The rate of unfaithful repairs
among accepted ones is reported as a point estimate.

`calibrate_coverage` is split conformal on `1 - similarity`, giving
`P(a genuine repair is accepted) >= 1 - alpha` for exchangeable scores.

A threshold calibrated on synthetic perturbations is calibrated for those
perturbations. Use the certified value at full precision; rounding it down
admits scores the certificate excluded. Both functions raise
`CalibrationError` rather than return an uncertifiable threshold.

---

## What it does not check

- **Entailment.** A verbatim quotation can be attached to the wrong entity,
  relation, or direction. Grounding is necessary, not sufficient.
- **Recall.** Nothing here measures what the model failed to extract.
- **Attribute fields.** Verify the quotation, then validate your schema's
  other fields against it separately.

**A rejected claim is not a detected hallucination.** It is a claim the policy
will not vouch for, which also covers splicing, reordering, paraphrase, parser
drift, and a source that was not the document the model saw.

---

## Known limitations

This is an alpha. Before relying on it, know that:

- **The guards are partial.** The numeric guard compares digit strings, so it
  ignores units and scale words (`$1.5 million` vs `$1.5 thousand`), sign
  conventions, figure order and spelled-out numbers. The polarity guard counts
  watched words, so it misses contractions (`don't`), a negation moved within
  a sentence, and direction words (`increased` vs `decreased`); it also counts
  the month "May" as a modal.
- **The fingerprint covers the configuration**, plus a verifier revision that
  is bumped when verdicts change. It is not a hash of the code: report the
  package version with it.
- **Unicode.** Normalization is per character, so a decomposed accent does not
  match its precomposed form. Tokenization assumes space-separated scripts.
- **Performance.** Each source is normalized once per tier; very large corpora
  are slower than they need to be.
- **Repair.** Substitution trusts the aligner's best span above your
  threshold. Keep `gtrex.claim_original` and review substituted records where
  it matters.

---

## Documentation

`docs/DECISIONS.md` (in the source distribution and repository) documents every
parameter: what it assumes, what it costs, and when to change it. Read it
before reporting a number. `examples/quickstart.py` is a self-contained tour.

## Development

From a source checkout:

```bash
pip install -e ".[dev]"
pytest
```

## Citation

If you use this in published work, report your `Policy.fingerprint` and the
package version alongside any grounding rate.

## License

Apache-2.0.
