Metadata-Version: 2.5
Name: bayan-tokenizer
Version: 0.1.0
Summary: Bayan — morphological, root-cognizant tokenizer: Arabic words as (root, pattern) pairs, zero dependencies
Project-URL: Homepage, https://github.com/CodeWithJuber/mizan
Project-URL: Source, https://github.com/CodeWithJuber/mizan/tree/main/packages/bayan
Project-URL: Documentation, https://github.com/CodeWithJuber/mizan/tree/main/packages/bayan#readme
Project-URL: Issues, https://github.com/CodeWithJuber/mizan/issues
Author: MIZAN Contributors
License: MIT
License-File: LICENSE
Keywords: arabic,linguistics,morphology,nlp,quranic,root,tokenizer
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Text Processing :: Linguistic
Requires-Python: >=3.9
Description-Content-Type: text/markdown

# Bayan — morphological tokenizer (مُبَيِّن)

Bayan is Ruh's morphological tokenizer: Arabic words are analysed into
**(root, pattern)** pairs instead of BPE subword fragments. English words are
bridged through a concept map to shared roots.

```python
from bayan import BayanTokenizer

tok = BayanTokenizer()
tok.tokenize("الكتاب")
# [{'surface': 'الكتاب', 'root': 'كتب', 'root_id': ..., 'pattern': 'NOUN', 'pattern_id': ...}]

tok.roots("العلم نور")     # ['علم', 'نور']
tok.patterns("العلم نور")  # ['NOUN', 'NOUN']
tok.fertility("العلم نور") # 1.0 — one morphological token per word, by construction
```

> "He created man. He taught him al-Bayan (clear expression)." — Quran 55:3-4

## Why

Subword tokenizers (BPE/WordPiece) shred Arabic morphology: one root scatters
across dozens of unrelated fragments. Bayan keeps the root — the unit Arabic
grammar has used for 1,400+ years — so models see *كتب* in *كتاب*, *كاتب*,
*مكتبة* as the same root, not three coincidences.

## Install

```bash
pip install bayan-tokenizer
```

Zero dependencies. Pure standard library — no torch, no numpy, no downloads.
Works on Python 3.9+.

## API

| Method | Returns |
|---|---|
| `tokenize(text)` | list of `{surface, root, root_id, pattern, pattern_id}` — one per word |
| `roots(text)` | resolved linguistic roots (stopwords and `<ENG:…>` placeholders excluded) |
| `patterns(text)` | morphological pattern name per word |
| `fertility(text)` | tokens-per-word (1.0 by construction; reported for honest BPE comparison) |
| `vocab_size` | `n_roots × n_patterns` |
| `supported()` | `True` when the tables/lexicon path loaded |

Stopwords yield `root == ""` and pattern `"STOPWORD"`.

## Honest limits

- The bundled analyzer is **rule-based** (tables + lexicon). It has **not**
  been benchmarked against CAMeL Tools or Farasa — do not claim SOTA.
- Unknown English words surface as `<ENG:bucket>` placeholders, never as
  invented roots.
- `BayanTokenizer.with_model(path)` is an explicit stub and raises
  `NotAvailableError`: model-backed Bayan is not implemented in this release.

## Development

The package vendors the tokenizer from the Mizan monorepo
(`ruh_model/tokenizer/` + `backend/qca/roots.py`). After changing the
upstream sources, re-vendor and rebuild:

```bash
python packages/bayan/scripts/sync_vendor.py
cd packages/bayan && python -m build && twine check dist/*
```

See `../../docs/BAYAN_RELEASE.md` for the full release checklist.

## License

MIT — see [LICENSE](LICENSE).
