Metadata-Version: 2.4
Name: sasakscript
Version: 0.8.0
Summary: A Linguistically-Aware Bidirectional Transliteration Library and Benchmark for Sasak Language and Script
Author: SasakScript Working Group
Author-email: kodetr <contact@kodetr.com>
License-Expression: MIT
Project-URL: Homepage, https://kodetr.com
Project-URL: Repository, https://github.com/kodetr/sasakscript
Project-URL: Documentation, https://github.com/kodetr/sasakscript#readme
Project-URL: Bug Tracker, https://github.com/kodetr/sasakscript/issues
Keywords: sasak,nlp,transliteration,unicode,balinese-script,low-resource-nlp
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Text Processing :: Linguistic
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pyyaml>=6.0
Provides-Extra: api
Requires-Dist: fastapi>=0.100.0; extra == "api"
Requires-Dist: uvicorn>=0.20.0; extra == "api"
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"
Requires-Dist: pytest-cov>=4.0; extra == "dev"
Requires-Dist: httpx>=0.24.0; extra == "dev"
Dynamic: license-file

# SasakScript: A Linguistically-Aware Bidirectional Transliteration Library and Benchmark for Sasak

[![PyPI version](https://img.shields.io/badge/version-0.8.0-blue.svg)](https://pypi.org/project/sasakscript/)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
[![Python 3.9+](https://img.shields.io/badge/python-3.9+-blue.svg)](https://www.python.org/downloads/)
[![Test Coverage](https://img.shields.io/badge/coverage-96%25-brightgreen.svg)](#)

---

## 1. Introduction

**SasakScript** is a Python library and standardized evaluation benchmark for bidirectional transliteration between the Latin alphabet and the traditional Sasak script (*Aksara Sasak*, encoded in the Balinese Unicode block `U+1B00`–`U+1B7F`). 

To the best of our knowledge, existing computational studies on Sasak script have predominantly focused on rule-based Latin-to-Sasak transliteration and image-based character recognition, while a reproducible, linguistically-aware, bidirectional transliteration library accompanied by a dedicated benchmark remains insufficiently explored.

SasakScript integrates phonological mutation recovery, morphological boundary parsing, maximal onset syllabification, dialectal demonstrative variation, and an extensive multi-tier evaluation framework.

---

## 2. Motivation

Sasak (*basa Sasak*) is an Austronesian language spoken by approximately 3 million people on the island of Lombok, Indonesia. While traditionally inscribed onto palm-leaf manuscripts (*lontar*), the script faces severe endangerment:
- **Low-Resource Scarcity:** Computational linguistic tools and standardized digital corpora for Aksara Sasak have remained minimal.
- **Orthographic Nuance:** Aksara Sasak is an abugida where consonants possess an inherent vowel [a], vowels are written as diacritics (*pangangge*), consonant codas take specialized signs (*cecek*, *surang*, *bisah*), and clusters require conjunct consonants (*gantungan*) formed via the virama (*adeg-adeg* `U+1B44`).
- **Phonological Complexity:** Active verb derivation invokes homorganic nasal substitution, epenthesis before liquids/glides, and morpheme boundary sandhi that naive transliterators fail to capture.

---

## 3. Installation

### From PyPI
```bash
pip install sasakscript
```

### From Source (Editable Mode)
```bash
git clone https://github.com/kodetr/sasakscript.git
cd sasakscript
pip install -e .
```

### With REST API Support
```bash
pip install -e ".[api]"
```

---

## 4. Quick Start

### Python API
```python
import sasakscript as sk

# 1. Latin to Aksara Sasak
res_sasak = sk.latin_to_sasak("tiang lalo jok pasar")
print(res_sasak.output)
# ᬢᬶᬬᬂ ᬮᬮᭀ ᬚᭀᬓ᭄ ᬧᬲᬃ

# 2. Aksara Sasak to Latin
res_latin = sk.sasak_to_latin("ᬩᬮᬾ")
print(res_latin.output)
# bale

# 3. Bidirectional Round-Trip Validation
rt = sk.round_trip("kanak bajang nulis surat")
print(f"Exact match: {rt.exact_match} | CER: {rt.cer:.4f}")

# 4. Morphological Analysis
morph = sk.analyze_morphology("pemangan")
print(morph)
# {'word': 'pemangan', 'root': 'mangan', 'prefix': ['pe-'], ...}

# 5. Pipeline Explainability Trace
explained = sk.transliterate("bale", explain=True)
print(explained.explanation["stages"])

# 6. Bilingual Sasak -> Indonesian Dictionary & Glossing
print(sk.translate_word("bale"))    # "rumah, tempat tinggal"
print(sk.translate_word("mangan"))  # "makan"
```

### Command-Line Interface (CLI)
```bash
# Transliterate text
sasakscript translate "tiang lalo jok pasar"
# ᬢᬶᬬᬂ ᬮᬮᭀ ᬚᭀᬓ᭄ ᬧᬲᬃ

# Lookup Sasak word meaning in Indonesian
sasakscript lookup bale
# Kata:       bale
# Arti (ID):  rumah, tempat tinggal

# Lookup with JSON output
sasakscript lookup mangan --json

# Transliterate file
sasakscript translate input.txt --direction latin-to-sasak --output output.txt

# Morphological analysis
sasakscript analyze pemangan --json

# Run benchmark suite
sasakscript benchmark

# Inspect Unicode characters
sasakscript validate-unicode "ᬩᬮᬾ"
```

### FastAPI REST Service
Launch the local API server:
```bash
uvicorn sasakscript.api.app:app --host 0.0.0.0 --port 8000 --reload
```
Interactive Swagger API documentation will be available at `http://localhost:8000/docs`.

---

## 5. Linguistic Pipeline

SasakScript processes text through eight deterministic, explainable stages:

```
INPUT
  ↓
[Stage 1: NORMALIZATION]       (Unicode NFC, apostrophe → q, diacritic standardization)
  ↓
[Stage 2: TOKENIZATION]        (Word boundaries, punctuation, numbers, whitespace)
  ↓
[Stage 3: MORPHOLOGY]          (Root recovery, prefix/suffix/infix/circumfix decomposition)
  ↓
[Stage 4: PHONOLOGY]           (Active nasal mutations, liquid epenthesis, hiatus resolution)
  ↓
[Stage 5: SYLLABIFICATION]     (Maximal Onset Principle: V, CV, CVC, CCV, CCVC, VC)
  ↓
[Stage 6: ORTHOGRAPHIC RULES]  (Prioritized YAML rules: coda signs, initial vowels, gantungan)
  ↓
[Stage 7: UNICODE SYNTHESIS]   (Canonical Balinese codepoint assembly, virama U+1B44)
  ↓
OUTPUT
```

---

## 6. Unicode Architecture

Sasak script is encoded in the **Balinese Unicode Block** (`U+1B00`–`U+1B7F`). 

### Disputed Codepoints Policy (`U+1B45`–`U+1B4B`)
In accordance with Unicode Technical Note #51 (*Balinese and Sasak Orthography in Unicode*, Perdana 2024), seven dedicated codepoints proposed in Unicode 5.0 for Sasak are marked **`UNVERIFIED`** by default because historical Sasak lontar manuscripts predominantly employ the standard consonant bases combined with the *Rerekan* diacritic (`U+1B34`). SasakScript strictly prevents the injection of invented rules without primary literature verification.

---

## 7. Dialect Module

SasakScript models the 5 canonical Sasak dialect clusters identified by Austronesian dialectology (Teeuw 1958; Mahsun 2006, 2007):
1. **`meno-mene`:** Central Sasak (Praya, Janapria, Kopang) — **VERIFIED Standard Reference**.
2. **`ngeno-ngene`:** East Sasak (Selong, Masbagik, Keruak).
3. **`ngeto-ngete`:** Northeast Sasak (Sembalun, Suela, Sambelia).
4. **`mriaq-mriku`:** South Sasak (Pujut, Kedaro, Sekotong).
5. **`kuto-kute`:** North Sasak (Bayan, Gangga, Tanjung).

> [!NOTE]
> Per Section 9 of the project specification, dialect transformation rules lacking primary empirical attestation raise `NotImplementedError` unless `allow_unverified=True` is explicitly passed.

---

## 8. Dataset: SasakScript-Bench

`SasakScript-Bench` is an annotated benchmark suite partitioned into six core linguistic categories:
- **`lexical`:** Base roots, kinships, and high-frequency terms.
- **`morphological`:** Productive affixation, nasalization, and reduplication.
- **`sentence`:** Complete declarative and interrogative sentences.
- **`dialect`:** Dialectal demonstrative pronouns.
- **`cultural`:** Classical lontar manuscripts and *seloka* (pantun).
- **`hard_case`:** Complex onset clusters, Rerekan loans, hiatus, and glottal codas.

### Quality and Zero-Leakage Guarantee
- **No Synthetic Gold:** Synthetic data is strictly barred from the gold test set.
- **Lexical-Disjoint Split:** Guarantees zero root morpheme overlap between train and test sets to evaluate true generalization rather than vocabulary memorization.

---

## 9. Benchmark & Baselines

SasakScript provides implementations for 5 standard baselines (Section 16):
- **Baseline 1:** Naive 1-to-1 character mapping table.
- **Baseline 2:** Standard rule-based syllabic engine.
- **Baseline 3:** Morphology-aware rule engine.
- **Baseline 4:** Phonology-aware rule engine.
- **Proposed:** Full SasakScript unified pipeline.

### Reproducibility
Execute the complete benchmark suite with one command:
```bash
python3 experiments/run_all.py
```
This automatically produces:
- `results/baseline.csv`
- `results/ablation.csv`
- `results/dialect.csv`
- `results/roundtrip.csv`
- `results/error_analysis.csv`

---

## 10. Limitations

1. **Dialect Phonetic Generalization:** While demonstratives are empirically verified, generative dialectal vocalic harmony across arbitrary vocabulary is marked `UNVERIFIED` pending field survey recordings.
2. **Manuscript Paleographic Variance:** The current engine targets the standard orthography established by the Department of Education & Culture (Depdikbud NTB) and primary lontar traditions; highly regional scribal idiosyncratic ligatures may require custom rule additions.

---

## 11. Citation

If you use SasakScript or SasakScript-Bench in your research, please cite:

```bibtex
@software{sasakscript2026,
  author    = {kodetr and SasakScript Working Group},
  title     = {SasakScript: A Linguistically-Aware Bidirectional Transliteration Library and Benchmark for Sasak Language and Script},
  year      = {2026},
  url       = {https://github.com/kodetr/sasakscript},
  note      = {Website: \url{https://kodetr.com}},
  version   = {0.8.0}
}
```

---

## 12. Authors & Maintainers

- **Author & Maintainer:** **kodetr** ([https://kodetr.com](https://kodetr.com))
- **Research & Linguistic Contributors:** SasakScript Working Group

---

## 13. License

This project is licensed under the [MIT License](LICENSE).
The dataset annotations are released under [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/).
