Metadata-Version: 2.4
Name: morphbpe-pl
Version: 0.5.0
Summary: MorphBPE: Morphologically-aware BPE tokenizer for Polish
Author: Piotr Styła
License-Expression: MIT
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Natural Language :: Polish
Requires-Python: >=3.10
Description-Content-Type: text/markdown
Requires-Dist: numpy>=1.24
Requires-Dist: tqdm>=4.60
Provides-Extra: morfeusz
Requires-Dist: morfeusz2; extra == "morfeusz"
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"

# MorphBPE

**Morphologically-aware BPE tokenizer for Polish.**

MorphBPE extends Byte-Pair Encoding with morphological boundary penalties.
Merges that cross morpheme boundaries receive a frequency penalty controlled by `alpha`,
preserving suffix (or prefix) structure during tokenization.

Based on: Styła (2026), *MorphBPE: morfologicznie informowany BPE dla polszczyzny*.

## Installation

```bash
pip install morphbpe-pl

# With Morfeusz2 support (recommended):
pip install morphbpe-pl[morfeusz]
```

## Quick Start

### Python API

```python
from morphbpe import MorphBPETokenizer

# Train
tok = MorphBPETokenizer(vocab_size=32000, alpha=2.0, constraint="suffix")
tok.train(["list of", "training texts", ...])

# Encode / decode
ids = tok.encode("prezydent podpisał ustawę")
text = tok.decode(ids)

# With special tokens (for BERT-style models)
ids = tok.encode("prezydent podpisał ustawę", add_special_tokens=True)
# → [2, ...token_ids..., 3]  (<cls> ... <sep>)

# Subword tokens
tokens = tok.encode_word("ustawodawstwo")
# ['ustaw', 'odaw', 'stwo'] (example)

# Batch encoding with padding
batch = tok.batch_encode(["text one", "longer text two"], padding=True)
# → {"input_ids": [[...], [...]], "attention_mask": [[...], [...]]}

# Save / load
tok.save("my-tokenizer")
tok = MorphBPETokenizer.load("my-tokenizer")

# Metrics
print(tok.fertility(test_texts))
print(tok.morph_edit_distance(test_texts))
mean, lo, hi = tok.morph_edit_distance_ci(test_texts)
```

### HuggingFace Integration

```python
from morphbpe.hf_tokenizer import MorphBPEHFTokenizer

tok = MorphBPEHFTokenizer.from_pretrained("my-tokenizer")

# Single text
result = tok("prezydent podpisał ustawę")
# → {"input_ids": [...], "attention_mask": [...]}

# Batch
result = tok(["text one", "text two"], padding=True, return_tensors="pt")

# Special tokens
tok.pad_token_id   # 0
tok.unk_token_id   # 1
tok.cls_token_id   # 2
tok.sep_token_id   # 3
tok.mask_token_id  # 4
```

### CLI

```bash
# Train
morphbpe train corpus.txt --vocab-size 32000 --alpha 2.0 --constraint suffix -o my-model

# Encode
echo "prezydent podpisał ustawę" | morphbpe encode my-model

# Info
morphbpe info my-model
```

## Key Parameters

| Parameter | Default | Description |
|-----------|---------|-------------|
| `vocab_size` | 1000 | Target vocabulary size |
| `alpha` | 1.0 | Penalty strength: 0=BPE, 1=block, >1=over-penalization |
| `constraint` | `"suffix"` | Boundary type: `suffix`, `prefix`, or `both` |

**Recommended for Polish:** `alpha=2.0, constraint="suffix"` (see paper, §Penalty Sweep).

## Features

### Special Tokens

IDs 0–4 are reserved for special tokens:

| Token | ID | Purpose |
|-------|----|---------|
| `<pad>` | 0 | Padding |
| `<unk>` | 1 | Unknown (only via byte fallback) |
| `<cls>` | 2 | Classification start |
| `<sep>` | 3 | Separator |
| `<mask>` | 4 | Masked LM |

### Byte-Level Fallback

Any UTF-8 input can be tokenized without `<unk>`. Unknown characters are encoded as their UTF-8 bytes (IDs 5–260). This guarantees:

- ✅ Emoji: `🎉🌍` → byte tokens
- ✅ CJK: `你好, こんにちは` → byte tokens
- ✅ Cyrillic: `привет` → byte tokens
- ✅ Polish diacritics: `zażółć gęślą jaźń` → byte tokens

### Pre-Tokenizer

Input text is pre-tokenized using a regex that splits punctuation, numbers, and words before BPE:

```python
from morphbpe import pre_tokenize

pre_tokenize("Hello, world! Rok 2024.")
# → ['Hello', ',', ' ', 'world', '!', ' ', 'Rok', ' ', '2024', '.']
```

### Trie Encoder

Encoding uses a trie data structure for O(word_length) lookup, replacing the previous O(n×merges) string-replace approach. **~36× faster** on real corpora.

### Batch Encoding

```python
batch = tok.batch_encode(
    ["text one", "longer text two"],
    padding=True,           # pad to longest in batch
    max_length=128,         # truncate if longer
    return_attention_mask=True,
)
```

### Subword Regularization

Random segmentation sampling for robustness (Kudo 2018). Each call may produce a different tokenization — useful for data augmentation during training:

```python
# Deterministic (greedy)
ids = tok.encode("ustawodawstwo")

# Randomized (for augmentation)
ids = tok.sample_encode("ustawodawstwo", temperature=0.1)

# Reproducible
ids = tok.sample_encode("ustawodawstwo", temperature=0.1, seed=42)
```

| Temperature | Behavior |
|---|---|
| 0.0 | Greedy (identical to `encode`) |
| 0.1 | Slight variation (recommended) |
| 0.5 | Moderate randomization |
| 1.0 | Uniform over valid token lengths |

### Corpus Filtering with Laya

Use [laya](https://github.com/PiotrStyla/laya) (local classifier, ~300ms on GPU) to filter training data by topic before training. Produces cleaner, domain-specific tokens:

```bash
# Filter Sejm speeches by topic
python scripts/filter_corpus_with_laya.py

# Compare filtered vs unfiltered tokenizer
python scripts/compare_filtered_tokenizer.py
```

**Cost comparison** (8k texts):

| Method | Cost | Speed | Quality |
|---|---|---|---|
| **laya (CPU)** | $0 | ~33 hours | Good |
| **laya (GPU)** | $0 | ~40 min | Good |
| **GPT-4o** | ~$5 | ~20 min | Best |
| **Rule-based** | $0 | Instant | Poor |

## How It Works

During BPE training, each candidate merge is scored by its corpus frequency.
MorphBPE multiplies the frequency of boundary-crossing merges by `(1 - alpha)`:

- `alpha=0.0` → no penalty → standard BPE
- `alpha=0.5` → 50% frequency reduction for crossing merges
- `alpha=1.0` → full block (frequency = 0)
- `alpha=2.0` → negative score → crossing merges actively deprioritized

Boundaries are detected using [Morfeusz2](http://morfeusz.sgjp.pl/) (Polish morphological analyzer)
or a heuristic fallback for environments without Morfeusz2.

## Results

### Intrinsic Evaluation (from paper, v=1000)

| Metric | BPE | MorphBPE α=1.0 | Δ |
|--------|-----|----------------|---|
| Fertility (v=1000) | 1.935 | 1.935 | +0.1% |
| MorphED (v=1000) | 1.838 | 1.712 | **−6.8%** |
| CE LSTM (v=1000) | 3.744 | 3.646 | **−2.6%** |
| CE Transformer (v=1000) | 3.525 | 3.431 | **−2.7%** |

Over-penalization (`alpha=2.0`) reduces CE by an additional 1.6–3.1%.

### Extrinsic Evaluation (vocab 8k–32k, fine-grained alpha sweep)

Corpus: Sejm speeches (8000 train / 500 test), seq_len=32, 3 epochs. See `plots/` for visualizations.

| Vocab | α | Model | CE | PPL |
|-------|---|-------|----|-----|
| 8k | 0.0 (BPE) | LSTM | 4.1073 | 60.78 |
| 8k | 0.0 (BPE) | Transformer | 3.9140 | 50.10 |
| 8k | 0.5 | LSTM | 4.0433 | 57.01 |
| 8k | 0.5 | Transformer | 3.8552 | 47.24 |
| 8k | 1.0 | LSTM | 3.9243 | 50.62 |
| 8k | 1.0 | Transformer | 3.7743 | 43.57 |
| 8k | 1.5 | LSTM | 3.8849 | 48.66 |
| 8k | 1.5 | Transformer | 3.7112 | 40.90 |
| 8k | 2.0 | LSTM | 3.7758 | 43.63 |
| 8k | 2.0 | Transformer | 3.6319 | 37.79 |
| 8k | 2.5 | LSTM | 3.7871 | 44.13 |
| 8k | 2.5 | Transformer | **3.6210** | **37.37** |
| 16k | 0.0 (BPE) | LSTM | 4.1494 | 63.40 |
| 16k | 0.0 (BPE) | Transformer | 3.9473 | 51.80 |
| 16k | 0.5 | LSTM | 4.1132 | 61.14 |
| 16k | 0.5 | Transformer | 3.8966 | 49.24 |
| 16k | 1.0 | LSTM | 3.9870 | 53.89 |
| 16k | 1.0 | Transformer | 3.7789 | 43.77 |
| 16k | 1.5 | LSTM | 3.9287 | 50.84 |
| 16k | 1.5 | Transformer | 3.7188 | 41.21 |
| 16k | 2.0 | LSTM | 3.7984 | 44.63 |
| 16k | 2.0 | Transformer | 3.6272 | 37.61 |
| 16k | 2.5 | LSTM | 3.8305 | 46.09 |
| 16k | 2.5 | Transformer | 3.6276 | 37.62 |
| 32k | 0.0 (BPE) | LSTM | 4.2140 | 67.62 |
| 32k | 0.0 (BPE) | Transformer | 3.9605 | 52.49 |
| 32k | 0.5 | LSTM | 4.1380 | 62.68 |
| 32k | 0.5 | Transformer | 3.9135 | 50.07 |
| 32k | 1.0 | LSTM | 4.0106 | 55.18 |
| 32k | 1.0 | Transformer | 3.7803 | 43.83 |
| 32k | 1.5 | LSTM | 3.9687 | 52.92 |
| 32k | 1.5 | Transformer | 3.7210 | 41.31 |
| 32k | 2.0 | LSTM | 3.8427 | 46.65 |
| 32k | 2.0 | Transformer | 3.6347 | 37.89 |
| 32k | 2.5 | LSTM | 3.8777 | 48.31 |
| 32k | 2.5 | Transformer | 3.6403 | 38.10 |

**Key findings:**

- **Monotonic improvement up to α=2.0**: CE decreases monotonically with α up to 2.0 for all three vocab sizes (8k, 16k, 32k) across both models
- **α=2.0 is the sweet spot**: α=2.5 shows diminishing returns — improvements plateau or reverse (especially at 16k and 32k), confirming α=2.0 as the optimal penalty strength
- **Best overall**: 8k α=2.5 Transformer (CE=3.6210, PPL=37.37) edges out 8k α=2.0 (CE=3.6319), but 16k and 32k α=2.5 are worse than α=2.0
- **α=1.5 sweet spot for efficiency**: 8k α=1.5 Transformer (CE=3.7112, PPL=40.90) beats 32k BPE Transformer (CE=3.9605, PPL=52.49) — 4x smaller vocab, 22% lower PPL. *Caveat: CE/PPL across different vocab sizes are not strictly comparable without bits-per-character normalization (see Methodological caveats below).*
- **CE reduction vs BPE**: α=0.5 → 0.9–1.8%; α=1.0 → 3.6–4.8%; α=1.5 → 5.2–6.0%; α=2.0 → 7.2–8.8%; α=2.5 → 7.2–8.1% (plateau)
- **Transformer > LSTM**: 3.8–6.2% CE reduction across all configs
- **Vocab size**: 8k and 16k perform similarly at α=2.0; 32k slightly worse due to sparse coverage on the small pilot corpus

### Cost-Benefit Analysis

Comparing MorphBPE 8k α=2.0 against a 4× larger BPE 32k tokenizer reveals a **trade-off**, not uniform dominance:

| Metric | BPE 32k | MorphBPE 8k α=2.0 | Δ | Winner |
|--------|---------|---------------------|---|--------|
| Vocab size | 32,000 | 8,000 | **4× smaller** | MorphBPE |
| Fertility (↓ better) | 1.020 | 1.149 | **+12.6%** | BPE 32k |
| MorphED (↓ better) | 0.389 | 0.456 | **+17.2%** | BPE 32k |
| CE Transformer (↓ better) | 3.9605 | 3.632 | **−8.3%** | MorphBPE |
| PPL Transformer (↓ better) | 52.49 | 37.79 | **−28.0%** | MorphBPE |

**Interpretation**: MorphBPE 8k α=2.0 wins decisively on language-model quality (CE/PPL) and vocabulary size, but at a cost — its fertility is 12.6% higher and its MorphED is 17.2% higher (worse) than BPE 32k. This is **not Pareto dominance**: the two configs trade off intrinsic segmentation quality (where the larger BPE 32k vocab wins on MorphED/fertility) against downstream LM performance and memory footprint (where MorphBPE 8k wins).

The higher MorphED and fertility for MorphBPE 8k are partly expected: it has 4× fewer tokens, so it must split words into more (and less morphologically-clean) pieces than a 32k vocab. The interesting result is that despite worse intrinsic metrics, the morphologically-informed 8k tokenizer still yields **lower LM cross-entropy** than 32k BPE.

> **Methodological caveats** (this is a small-corpus pilot — single runs, no seeds):
> 1. **CE/PPL are not directly comparable across vocab sizes.** Different tokenization granularity (fertility) affects per-token entropy independently of segmentation quality. A rigorous "8k beats 32k" claim requires bits-per-character, not raw CE/PPL. *(Planned: add bits-per-character.)*
> 2. **Small corpus (8k sentences) favors smaller vocabularies** regardless of morphology — 32k is under-trained with rare tokens on this little data, which may partly explain MorphBPE 8k's edge over BPE 32k.
> 3. **No repeats/seeds** — differences on the order of thousandths of CE (e.g. 8k vs 16k at α=2.0) have no estimated variance.
> 4. Conclusions about transfer to production-scale models are premature with an 8k-sentence corpus and seq_len=32.

See `plots/` for visualizations: `pareto_ce_vs_vocab.png`, `fertility_vs_morphed.png`, `efficiency_ratio.png`.

## License

MIT
