Metadata-Version: 2.4
Name: quire-grammar-model
Version: 0.3.1
Summary: Bundled phase-2 edit-tagger for quire-grammar (data only, no code).
Author: Roger Cooper
License-Expression: Apache-2.0
Project-URL: Homepage, https://github.com/jutreuter/quire-grammar
Requires-Python: >=3.11
Description-Content-Type: text/markdown

# quire-grammar-model

The bundled phase-2 edit-tagger for
[quire-grammar](https://github.com/jutreuter/quire-grammar) — **data only, no
code**.

```
pip install quire-grammar[model]     # pulls this + onnxruntime + tokenizers
```

`quire_grammar.check()` runs rules-only without it and activates the model
when it's present — "off if missing", like a spell dictionary.

## What's inside

| file | what |
|---|---|
| `tagger.onnx` | GECToR-style token tagger, int8, ~66 MB |
| `tokenizer.json` | the matching `tokenizers` (WordPiece) definition |
| `tag_vocab.json` | the edit-tag label list |

`quire_grammar_model.path()` returns the directory; `available()` says whether
every file is present.

## Provenance

Fine-tuned from `distilbert-base-cased` (Apache-2.0; pretrained on Wikipedia +
BooksCorpus) on synthetic native-writer errors injected over public-domain
prose (Project Gutenberg). No non-commercial or research-only corpus is in the
training set. Licensed **Apache-2.0**.

The model file is regenerated from the training pipeline in the main repo
(`train/` → `train.export --out-dir packages/quire-grammar-model/src/quire_grammar_model/`),
not edited by hand.
