Metadata-Version: 2.5
Name: gtrex
Version: 0.1.0b3
Summary: Grounded text extraction you can check: verifiable quotation grounding for LLM extraction pipelines, with declarative normalization ladders, repair-and-reject cascades, and calibrated acceptance thresholds.
License-Expression: Apache-2.0
License-File: LICENSE
Keywords: attribution,conformal-prediction,grounded-extraction,grounding,hallucination,information-extraction,llm,provenance,reproducibility
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Text Processing :: Linguistic
Classifier: Typing :: Typed
Requires-Python: >=3.10
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == 'dev'
Requires-Dist: ruff>=0.4; extra == 'dev'
Description-Content-Type: text/markdown

# gtrex

*Grounded text extraction you can check.*

> **Status: beta (0.1.0b3).** The API and output formats may still change.
> Known limitations are listed near the end of this file, so that they travel
> with the package. 0.1.0b1 added lexical-boundary checks and legal-text typography support;
> 0.1.0b2 builds normalization tiers on demand for large-corpus verification;
> 0.1.0b3 adds exact-offset checking for stored provenance. See `CHANGELOG.md` for
> what changed, and for how to migrate policies and calibrations written by
> 0.1.0a1 through 0.1.0a5. Repair faithfulness has not been established on
> held-out human-adjudicated alignment-stage claims.

**For fixed extractions, the reported grounding rate depends on the
verification policy.** "The quotation appears in the source" always involves
choices: which normalizations to apply, what to do with an ellipsis, whether
stitched fragments count, which tier to call grounded. A rate reported without
those choices cannot be compared with anyone else's. `gtrex` makes the policy
explicit, versioned and fingerprinted, so a grounding claim is something
another lab can check.

```bash
pip install gtrex        # no runtime dependencies, Python >= 3.10
```

```python
from gtrex import Policy, Verifier

verifier = Verifier(filing_text, Policy())
cert = verifier.verify(model_quote)

cert.grounded   # True
cert.tier       # 'typography'  -- how much normalization it took
cert.start      # 48213         -- offsets into the original text, at every tier
cert.span_text  # the document's own characters
```

If the extraction already stores source offsets, bind the quotation to that
specific occurrence: `verifier.verify_at(model_quote, char_start, char_end)`.
Offsets refer to the exact source string supplied to the verifier. A match
elsewhere in the document cannot validate an incorrect stored location.

---

## What it does

**1. A declarative normalization ladder.** Five cumulative tiers, each
declaring whether it is *sound*: a recorded assumption that the tier forgives
only presentation, not content, in your domain. It is an assumption, not a
proof (case folding merges `mW` and `MW`), which is why it travels with the
data. Verification reports the tightest tier that contains the whole
quotation, not a boolean, and a match must cover whole source characters: `1`
taken out of `¼` does not count. On the eight claims of
`examples/quickstart.py`:

```
tier          at tier   cumulative     share
--------------------------------------------
exact               1            1   12.50%
whitespace          0            1   12.50%
typography          1            2   25.00%
case                1            3   37.50%
skeleton            1            4   50.00%   <- UNSOUND, not counted
stitched            1            -   12.50%
unlocated           3            -   37.50%
```

The `skeleton` tier is marked unsound because the collision is real, not
hypothetical:

```
"Operating loss was ($10.0 million)."  ->  operatinglosswas100million
"Operating loss was $100 million."     ->  operatinglosswas100million
```

Accepting it requires passing `allow_unsound=True` explicitly.

**2. A repair-and-reject cascade.** When a quotation is a near-miss, the
cascade (in its default `substitute` mode) does not accept the model's
wording: it locates the source span and **substitutes the document's own
characters**, keeping the model's original wording and the certificates before
and after. Anything that cannot be located above the threshold you set is
rejected. `annotate` mode keeps the model's wording and attaches the located
span as evidence; `off` only verifies.

```
verify -> [fragment salvage, opt-in] -> locate-and-substitute -> reject
```

Verification always checks the whole claim; nothing is clipped before it is
compared, so a verbatim opening with an invented tail is not grounded. A
substitution may then store a source span *shorter* than the claim: the
stored text is certified by its own `after` certificate, and the claim as
given keeps its `before` certificate, which says it did not verify.

**3. Two guards for edits no threshold can catch.** A changed figure and a
dropped negation both score high, because they share almost every token with
the source, and both would yield a record that is *verbatim source text and
still does not say what the model extracted*. From the quickstart:

```
claim: "Operating loss was $100 million for the year ended ..."    score 0.87
span:  "Operating loss was ($10.0 million) for the year ended ..." -> numeric_guard

claim: "We believe that we will be able to compete effectively ..."     score 0.97
span:  "We believe that we will not be able to compete effectively ..." -> polarity_guard
```

These are checks, not weights, and they are partial (see the limitations
below): they block two common catastrophic edits; they do not prove that a
substitution preserves meaning.

**4. A threshold you calibrate, not guess.** The substitution threshold has no
default, and repairing without one is an error. `gtrex calibrate` estimates
one with Learn-then-Test from labelled repairs, states the assumption the
guarantee rests on (see "Guarantees, precisely" below), and binds the result
to the scorer and pipeline it was made for, so that `gtrex repair
--calibration` refuses to use it anywhere else.

**5. A manifest.** Policy, fingerprints, package version, Unicode database
version, every tier's soundness rationale, calibration, and the ladder
profile, written next to the data so the assumptions travel with it. Every
part must come from the manifest's own policy; mixing is refused.

---

## Why a ladder instead of a number

Three failure modes a single grounding rate cannot distinguish:

| | What it looks like | What it is |
|---|---|---|
| **Stitched** | `"We depend on suppliers ... could materially affect results"` | Both fragments are genuine; the quotation is not. |
| **Skeleton collision** | `"Operating loss was $100 million"` | Matches the source only after punctuation is deleted, against a figure that is ten times smaller and negative. |
| **Parser drift** | quotation fails everywhere | The source text contains `[[HEADING_CANDIDATE]]` markers the model never saw. Declare them with `marker_pattern`. |

A verifier that reports one number will call the first two "grounded" and the
third "hallucinated". All three verdicts would be wrong.

---

## CLI

```bash
gtrex explain                       # print the policy and every assumption
gtrex verify   claims.jsonl -o certified.jsonl --manifest run.json
gtrex calibrate labelled.jsonl -o calibration.json --alpha 0.05 --compare 0.65
gtrex repair   claims.jsonl -o repaired.jsonl --calibration calibration.json --drop-rejected
gtrex sensitivity claims.jsonl      # the same claims under three policies
```

`repair` needs a threshold: `--calibration` takes the one `calibrate` wrote,
after checking it was calibrated for this policy's scorer and pipeline, and
`--threshold` sets one by hand. `--repair-mode annotate` keeps the model's
wording; `--salvage longest_grounded` opts into fragment salvage.

Input is JSON Lines. Each record needs a string `claim` and exactly one of an inline
`source` or a `source_path` (read as UTF-8, byte for byte). Every other field
is passed through untouched and in input order; the result is added under one
key, `gtrex`. A record that already has a `gtrex` field is refused unless you
pass `--replace-existing`.

```jsonl
{"id": "rf-04", "claim": "We depend on a limited number of suppliers", "source_path": "filings/acme-2023/item1a.txt"}
```

`gtrex repair` stores the record's text in `claim` (in `substitute` mode, the
source's own text), and keeps the model's wording in `gtrex.claim_original`,
with `gtrex.before` (the certificate of the claim as given) and `gtrex.after`
(the certificate of the stored text, when it differs). `gtrex.accepted` says
the record was retained; `gtrex.grounded` says its text verifies. An
annotation is accepted and not grounded.

Exit status is 0 on success and 1 on invalid input (the message names the
file and line) or a failed calibration; JSON goes to stdout only when stdout
is the requested output, and is always UTF-8, whatever the console encoding.

`sensitivity` is the experiment worth running first on your own data. On the
quickstart's eight claims:

```
policy       fingerprint      accept_through   grounded  contiguous
--------------------------------------------------------------------
default      63c0146e38cdb3cd case              37.50%     37.50%
strict       ddd428e220a5db79 exact             12.50%     12.50%
permissive   2d19400f2feff9cc skeleton          50.00%     50.00%

Same 8 claims, same sources. The reported grounding rate varies by 37.5
percentage points across these three verification policies.
```

`grounded` is the share of claims the policy accepts, the mean of the
individual verdicts; `contiguous` is the share contained as one run through
the accepted tier. They differ only when a policy accepts stitched
quotations. The three policies differ in more than normalization (ellipsis
handling and contiguity too); `gtrex explain --policy strict` prints each one
in full.

---

## Guarantees, precisely

A repair is accepted when `score >= threshold`, and calibration uses the same
rule. `calibrate_risk` implements Learn-then-Test with fixed-sequence testing
over the observed scores (plus 1.0, or a grid you supply), using the exact
binomial tail as the p-value, and returns the least conservative certifiable
threshold among those candidates. The controlled quantity is the joint risk

    P(score >= threshold AND the repair is unfaithful)

for claims like the calibration claims that reach the alignment-score filter.
It is **not** `P(unfaithful | accepted)`, which is reported only as a point
estimate, and it does not cover claims accepted at verification or by
fragment salvage. The guarantee, `P(risk <= alpha) >= 1 - delta` over the
calibration draw (`delta` is the failure probability; the confidence is
`1 - delta`), holds **only if the calibration examples are independent draws
from the claims you will repair**. Several examples from one document, or
several perturbations of one quotation, are not independent: give each
example a `group` (document id) and `calibrate_risk` refuses repeated groups.
Overriding that is recorded in the calibration as `uncertified`.

`calibrate_coverage` is split conformal on `1 - similarity`, giving
`P(a genuine repair is accepted) >= 1 - alpha` for exchangeable scores at the
same stage; the guards may still block some of those repairs afterwards.

A calibration records its target, stage, assumptions, a digest of its data,
and the identity of the scorer and pipeline its scores came from; a manifest
or `repair --calibration` refuses one made for another configuration. The
loader also rejects incompatible method revisions, acceptance rules, targets,
stages and unknown fields. This validates the artifact's declared format; it
does not authenticate a file against deliberate editing. A
threshold calibrated on synthetic perturbations is calibrated for those
perturbations. Use the certified value at full precision (scores and
thresholds are exported unrounded); rounding it down admits scores the
certificate excluded. Both functions raise `CalibrationError` rather than
return an uncertifiable threshold.

---

## What it does not check

- **Entailment.** A verbatim quotation can be attached to the wrong entity,
  relation, or direction. Grounding is necessary, not sufficient.
- **Recall.** Nothing here measures what the model failed to extract.
- **Attribute fields.** Verify the quotation, then validate your schema's
  other fields against it separately.

**A rejected claim is not a detected hallucination.** It is a claim the policy
will not vouch for, which also covers splicing, reordering, paraphrase, parser
drift, and a source that was not the document the model saw.

---

## Known limitations

This is a beta. Before relying on it, know that:

- **Lexical boundaries are heuristic.** The default rejects quotes clipped
  from inside a word, hyphenated term, signed figure, decimal, grouped number
  or time. It decodes HTML entities before checking the edges and allows
  apostrophe-suffixed names and sentence punctuation. Scripts that do not
  separate words with spaces need a language-specific tokenizer; this check
  does not infer their word boundaries. Set `match_boundaries=False` to
  reproduce historical substring behavior, but that switch alone does not
  undo the new hyphen normalization step.
- **"Sound" is an assumption.** A sound tier is one the policy declares to
  forgive only presentation. Case folding merges `mW` and `MW`; typography
  folding merges a minus sign with a hyphen. Choose the ladder for your
  domain.
- **The guards are partial.** The numeric guard compares numbers read by the
  English convention (`,` groups thousands, `.` marks decimals; a malformed
  number, or one written with a fraction character such as `1½`, must match
  verbatim; a superscript is a separate figure), so it ignores units and
  most units, uncommon scale words, many sign conventions,
  and spelled-out numbers. It checks figure order and adjacent unary signs
  on decimal figures, including compatible Unicode forms and a sign clipped
  from the start of an aligned span.
  The polarity guard compares watched words in the claim with the located
  span and nearby negative words before it. It folds format and compatibility
  characters and line-end hyphenation, and checks common contractions and
  financial direction words. It still misses some shifted negation scopes;
  it also counts the month "May" as a modal. These guards run on alignment
  repair, not on claims that already verify as literal source substrings.
  Alignment substitutions additionally require lexical boundaries and a
  grounded certificate for the substituted source span.
- **The fingerprint covers the configuration**, including every
  normalization step's id, version and parameters, plus a verifier revision
  that is bumped when verdicts change. It is not a hash of the code: report
  the package version with it.
- **Unicode.** Normalization is per character, so a decomposed accent does not
  match its precomposed form. Verdicts depend on the Python's Unicode
  database, which the manifest records; the same fingerprint under another
  database is not a claim of identical verdicts. Tokenization assumes
  space-separated scripts.
- **Calibration is synthetic until you label your own repairs.** The
  guarantees hold for the population the calibration examples were drawn
  from, at the alignment stage only.
- **Performance.** Each source is normalized once per tier; very large corpora
  are slower than they need to be.
- **Repair.** Substitution trusts the aligner's best span above your
  threshold, and may store a span shorter than the claim. Keep
  `gtrex.claim_original` and review substituted records where it matters.
  Opted-in salvage keeps one verified fragment and records what it dropped;
  what was dropped is not certified.

---

## Documentation

`examples/quickstart.py` is a self-contained tour in the source archive.
The detailed policy guide, beta evaluations, testing record and research-effort
ledger remain in the project repository for the planned public GitHub release.
They are research records, not files required to install or run GTREX.

## Development

From a source checkout:

```bash
pip install -e ".[dev]"
pytest
```

## Citation

If you use this in published work, report your `Policy.fingerprint`, the
package version and the Unicode database version (all in the manifest)
alongside any grounding rate.

## License

Apache-2.0.
