Metadata-Version: 2.4
Name: nusanlp
Version: 0.2.3
Summary: An open-source NLP toolkit for regional languages of West Nusa Tenggara (Sasak, Samawa, Mbojo) and Bahasa Indonesia.
Author: kodetr
License: MIT
Project-URL: Homepage, https://kodetr.com
Project-URL: Repository, https://github.com/kodetr/nusanlp
Project-URL: Documentation, https://github.com/kodetr/nusanlp#readme
Project-URL: Facebook, https://facebook.com/kodetr
Project-URL: Instagram, https://instagram.com/kodetr
Keywords: nlp,nusanlp,sasak,samawa,sumbawa,mbojo,bima,indonesian,bahasa-indonesia,indonesian-languages,computational-linguistics,morphology,stemmer,stopwords,low-resource-language
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Science/Research
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Natural Language :: Indonesian
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Text Processing :: Linguistic
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pydantic>=2.0.0
Provides-Extra: dev
Requires-Dist: pytest>=7.0.0; extra == "dev"
Requires-Dist: pytest-cov>=4.0.0; extra == "dev"
Requires-Dist: flake8>=6.0.0; extra == "dev"
Requires-Dist: mypy>=1.0.0; extra == "dev"
Provides-Extra: api
Requires-Dist: fastapi>=0.100.0; extra == "api"
Requires-Dist: uvicorn>=0.20.0; extra == "api"
Provides-Extra: ml
Requires-Dist: scikit-learn>=1.0.0; extra == "ml"
Dynamic: license-file

# 🏝️ NusaNLP: Low-Resource NLP Toolkit for West Nusa Tenggara, Indonesia

[![PyPI Version](https://img.shields.io/pypi/v/nusanlp.svg?color=blue)](https://pypi.org/project/nusanlp/)
[![Website](https://img.shields.io/badge/Website-kodetr.com-059669?logo=googlechrome&logoColor=white)](https://kodetr.com)
[![Facebook](https://img.shields.io/badge/Facebook-kodetr-1877F2?logo=facebook&logoColor=white)](https://facebook.com/kodetr)
[![Instagram](https://img.shields.io/badge/Instagram-@kodetr-E4405F?logo=instagram&logoColor=white)](https://instagram.com/kodetr)
[![Python 3.9+](https://img.shields.io/badge/python-3.9+-blue.svg)](https://www.python.org/downloads/)
[![License: MIT](https://img.shields.io/badge/License-MIT-green.svg)](https://opensource.org/licenses/MIT)
[![Tests: 147 passed](https://img.shields.io/badge/tests-147%20passed-brightgreen.svg)]()
[![Coverage: 82%](https://img.shields.io/badge/coverage-82%25-brightgreen.svg)]()

**NusaNLP** is an open-source Natural Language Processing (NLP) toolkit engineered specifically for the regional languages of West Nusa Tenggara (*Nusa Tenggara Barat* / NTB) and Bahasa Indonesia:

1. **Bahasa Sasak** (`sasak`) — Spoken by ~3 million people across the island of Lombok.
2. **Bahasa Samawa / Sumbawa** (`samawa`) — Spoken in Sumbawa and West Sumbawa regencies.
3. **Bahasa Mbojo / Bima** (`mbojo`) — Spoken in Bima Regency, Bima City, and Dompu Regency.
4. **Bahasa Indonesia** (`indonesian`) — Fully integrated with tokenization, clitic handling, stopword filtering, morphological stemming, language identification (100% test accuracy), and bilingual translation alignment.

NusaNLP is built to be **production-ready, modular, extensible via a plugin language registry, and research-grade**.

---

## 📦 Installation

### Via PyPI (Recommended)

Install NusaNLP directly from PyPI using `pip`:

```bash
pip install nusanlp
```

#### Optional Feature Sets
```bash
# REST API (FastAPI + Uvicorn server)
pip install "nusanlp[api]"

# Machine Learning additions (Scikit-Learn vectorizers & metrics)
pip install "nusanlp[ml]"

# Development, testing, and linting tools
pip install "nusanlp[dev]"

# All optional dependencies
pip install "nusanlp[api,ml,dev]"
```

### From Source (For Development & Research)
```bash
git clone https://github.com/kodetr/nusanlp.git
cd nusanlp
pip install -e ".[dev]"
```

---

## 🚀 Quick Start

### 1. Unified Pipeline (with Auto-Detection)
```python
import nusanlp
from nusanlp import Pipeline

# Initialize pipeline with automatic language detection
nlp = Pipeline(language="auto")

result = nlp.process("Tiang lalo mangan leq bale")
print(result)
```

**Output:**
```json
{
  "language": "sasak",
  "original_text": "Tiang lalo mangan leq bale",
  "normalized_text": "tiang lalo mangan leq bale",
  "tokens": ["tiang", "lalo", "mangan", "leq", "bale"],
  "tokens_without_stopwords": ["lalo", "mangan", "bale"],
  "analysis": [
    {"word": "tiang", "root": "tiang", "prefix": null, "suffix": null, "features": []},
    {"word": "lalo", "root": "lalo", "prefix": null, "suffix": null, "features": []},
    {"word": "mangan", "root": "ang", "prefix": "m", "suffix": "an", "features": ["SUFF:an", "PREF:m"]},
    {"word": "leq", "root": "leq", "prefix": null, "suffix": null, "features": []},
    {"word": "bale", "root": "bal", "prefix": null, "suffix": "e", "features": ["SUFF:e"]}
  ],
  "pos_tags": [
    ["tiang", "NOUN"],
    ["lalo", "VERB"],
    ["mangan", "VERB"],
    ["leq", "ADP"],
    ["bale", "NOUN"]
  ]
}
```

### 2. Language Detection
Detect whether a sentence is Sasak, Samawa, Mbojo, or Indonesian:
```python
import nusanlp

res = nusanlp.detect_language("Tiang lalo leq peken")
print(res)
# {'language': 'sasak', 'confidence': 1.0}

res_smw = nusanlp.detect_language("Kaji lalo tama ko peken ling ina")
print(res_smw)
# {'language': 'samawa', 'confidence': 1.0}

res_mbj = nusanlp.detect_language("Nahu ngaha di uma labo ita")
print(res_mbj)
# {'language': 'mbojo', 'confidence': 1.0}
```

### 3. Text Normalization
```python
# Lowercasing, whitespace, repeated character & slang normalization
norm = nusanlp.normalize("AQ   LEKOOOOK  GATI", language="sasak")
print(norm)
# Output: "aku lekook gati"
```

### 4. Modular Tokenizer & Clitic Separation
```python
tokens = nusanlp.tokenize("Kaji lalo-mo ko peken", language="samawa")
print(tokens)
# Output: ['Kaji', 'lalo', '-mo', 'ko', 'peken']
```

### 5. Stopword Removal
```python
clean = nusanlp.remove_stopwords("nahu ngaha di uma labo ita", language="mbojo")
print(clean)
# Output: ['ngaha', 'uma']
```

### 6. Multilingual Dictionary Lookup
```python
entry = nusanlp.lookup("mangan", language="sasak")
print(entry)
# {'word': 'mangan', 'language': 'sasak', 'translations': {'indonesian': 'makan'}, 'pos': 'VERB', 'lemma': 'mangan'}

entry_smw = nusanlp.lookup("balong", language="samawa")
print(entry_smw)
# {'word': 'balong', 'language': 'samawa', 'translations': {'indonesian': 'baik / bagus', 'english': 'good'}, 'pos': 'ADJ', 'lemma': 'balong'}
```

### 7. Machine Learning & TF-IDF Integration (Scikit-Learn)
NusaNLP seamlessly integrates into `TfidfVectorizer` and `CountVectorizer`, providing morphological lemmatization and low-resource stopword removal to prevent vocabulary explosion:

```python
from sklearn.feature_extraction.text import TfidfVectorizer
import nusanlp

corpus = [
    "Tiang lalo mangan leq bale adat Sasak.",
    "Kaji mangan leq bale adat Samawa.",
    "Nahu ngaha di uma rime Mbojo."
]

# Method A: Use NusaNLP custom tokenizer callable directly
vectorizer = TfidfVectorizer(tokenizer=nusanlp.get_tfidf_tokenizer(), lowercase=False)
X = vectorizer.fit_transform(corpus)
print("TF-IDF Features:", vectorizer.get_feature_names_out())

# Method B: Preprocess texts beforehand
clean_docs = [nusanlp.preprocess_for_tfidf(doc) for doc in corpus]
# clean_docs: ['mangan bale adat sasak', 'mangan bale adat samawa', 'ngaha uma rime mbojo']
```

---

## 🏛️ Architecture & Project Structure

```text
nusanlp/
│
├── src/
│   ├── nusanlp/
│   │   ├── __init__.py                 # Unified API exports
│   │   ├── core/
│   │   │   ├── pipeline.py             # Unified Pipeline engine
│   │   │   ├── language.py             # BaseLanguage & LanguageRegistry
│   │   │   ├── config.py               # Pipeline & Normalizer configs
│   │   │   └── types.py                # Pydantic models & enums
│   │   │
│   │   ├── languages/
│   │   │   ├── base.py                 # AbstractLanguageHandler
│   │   │   ├── sasak/                  # Sasak modules
│   │   │   ├── samawa/                 # Samawa (Sumbawa) modules
│   │   │   ├── mbojo/                  # Mbojo (Bima) modules
│   │   │   └── indonesian/             # Indonesian language module
│   │   │
│   │   ├── preprocessing/
│   │   │   ├── normalize.py            # Multi-language normalization
│   │   │   └── clean.py                # Text cleaning utilities
│   │   │
│   │   ├── tokenization/
│   │   │   └── tokenizer.py            # RuleBasedTokenizer
│   │   │
│   │   ├── morphology/
│   │   │   └── analyzer.py             # MorphologicalAnalyzer
│   │   │
│   │   ├── lexicon/
│   │   │   ├── dictionary.py           # Multilingual dictionary
│   │   │   └── loader.py               # JSON/CSV loaders
│   │   │
│   │   ├── language_detection/
│   │   │   ├── detector.py             # detect_language API
│   │   │   └── model.py                # N-gram Statistical Classifier
│   │   │
│   │   ├── cli/
│   │   │   └── main.py                 # CLI entry point
│   │   │
│   │   ├── api/
│   │   │   └── app.py                  # FastAPI REST API
│   │   │
│   │   └── utils/
│   │       ├── loaders.py              # Resource discovery
│   │       └── validators.py           # Pydantic data validators
│   │
│   └── sasaknlp/                       # Backward-compatibility layer
│
├── resources/
│   ├── sasak/                          # Balai Bahasa NTB lexicon & stopwords
│   ├── samawa/                         # Sumbawa lexicon & stopwords
│   └── mbojo/                          # Bima/Dompu lexicon & stopwords
│
├── datasets/
├── training/                           # Language detector training scripts
├── tests/                              # 134 automated unit tests
├── examples/                           # Usage examples
└── pyproject.toml
```

---

## 💻 Command Line Interface (CLI)

```bash
# Language Detection
nusanlp detect "Tiang lalo leq peken"

# Tokenize
nusanlp tokenize --lang sasak "Tiang lalo leq peken"

# Normalize
nusanlp normalize --lang mbojo "NAHU MAI DI KAMPUNG"

# Dictionary Lookup
nusanlp dictionary lookup --lang samawa "bale"

# End-to-end Process
nusanlp process --lang auto "Tiang lalo mangan leq bale" --json
```

---

## 🌐 REST API

Run the optional FastAPI server:
```bash
uvicorn nusanlp.api.app:app --host 127.0.0.1 --port 8000 --reload
```

Endpoints available:
- `POST /detect-language`
- `POST /tokenize`
- `POST /normalize`
- `POST /process`
- `GET /dictionary/{language}/{word}`
- `GET /docs` (Interactive Swagger UI)

---

## 🧪 Testing & Evaluation
 
Run test suite with `pytest`:
```bash
pytest tests -v
```

Run unified scientific evaluation benchmark across all 3 languages:
```bash
python3 scripts/evaluate_nusanlp_all.py
```

### 📊 Scientific Evaluation Highlights
NusaNLP has undergone rigorous empirical validation across 30,000 morphological pairs and 3,000 language detection sentences:

| Evaluation Dimension | Dataset | Metric | Score |
| :--- | :--- | :--- | :---: |
| **Sasak Stemmer** | Benchmark 10k | Exact Match Accuracy | **93.23%** (8,539 words/s) |
| **Samawa Stemmer** | Benchmark 10k | Exact Match Accuracy | **94.63%** (41,727 words/s) |
| **Mbojo Stemmer** | Benchmark 10k | Exact Match Accuracy | **93.20%** (42,028 words/s) |
| **Language Identification** | Test Set 3,000 sentences | Global Accuracy | **99.90%** (Macro-F1: 99.90%) |
| **Pipeline Latency** | Authentic Sentences | Mean Latency | **0.046 - 0.358 ms / sentence** |
| **Downstream TF-IDF ML** | Multi-dialect Corpus | Vocabulary Compression | **14.1% feature reduction (99.4% Acc)** |

See the complete evaluation report in [docs/laporan_evaluasi_nusanlp.md](docs/laporan_evaluasi_nusanlp.md).

---

### 📈 Scientific Figures & Benchmark Visualizations

#### Figure 1: Morphology Accuracy & Affix Decomposition
![Figure 1: Morphology Accuracy](docs/figures/fig1_morphology_accuracy.png)
*Figure 1: Exact-match stemming accuracy and affix breakdown across Sasak (93.23%), Samawa (94.63%), and Mbojo (93.20%) on 30,000 empirical evaluation pairs.*

#### Figure 2: Dialect Performance & Cross-Dialect Robustness
![Figure 2: Dialect Performance](docs/figures/fig2_dialect_performance.png)
*Figure 2: Morphological stemmer resilience across 11 standardized regional dialects of West Nusa Tenggara.*

#### Figure 3: Language Identification Confusion Matrix & Classification Breakdown
![Figure 3: Confusion Matrix](docs/figures/fig3_error_taxonomy.png)
*Figure 3: Four-class normalized confusion matrix and language classification performance across 3,000 test sentences (Global Accuracy: 99.90%).*

#### Figure 4: Pipeline Throughput & Latency Distribution
![Figure 4: Pipeline Benchmark](docs/figures/fig4_pipeline_benchmark.png)
*Figure 4: Computational throughput (sentences/second) and sub-millisecond execution latency across all regional language pipelines.*

#### Figure 5: Dataset Acquisition Architecture & Corpus Pipeline
![Figure 5: Dataset Acquisition Pipeline](docs/figures/fig5_dataset_acquisition_pipeline.png)
*Figure 5: Multi-tier corpus engineering workflow, dictionary validation, and benchmark synthesis pipeline encompassing 959,279+ records.*

---

## 📚 Datasets & Resources
- **Total Dataset Volume**: **959,279 records** (Sentence corpora, morphology benchmarks, language detection, and official dictionaries).
- **Official Dictionaries (Balai Bahasa NTB)**: 3,352 Sasak, 659 Samawa, 2,777 Mbojo entries (100% Pydantic compliant).
- **Curated Stopwords**: **440 functional words** (159 Sasak, 147 Samawa, 134 Mbojo).
- See detailed documentation in [datasets/README.md](datasets/README.md) and [docs/PANDUAN_LENGKAP.md](docs/PANDUAN_LENGKAP.md).

---

## 📖 Citation

If you use NusaNLP in your academic work or research, please cite:

```bibtex
@software{nusanlp2026,
  author = {kodetr},
  title = {NusaNLP: An Open-Source Natural Language Processing Toolkit for Low-Resource Languages of West Nusa Tenggara, Indonesia},
  year = {2026},
  url = {https://github.com/kodetr/nusanlp}
}
```

---

## 🌐 Author & Connect

Developed and maintained by **kodetr**:
- 🌐 **Website**: [https://kodetr.com](https://kodetr.com)
- 📘 **Facebook**: [https://facebook.com/kodetr](https://facebook.com/kodetr)
- 📸 **Instagram**: [https://instagram.com/kodetr](https://instagram.com/kodetr)
- 💻 **GitHub**: [https://github.com/kodetr](https://github.com/kodetr)
- 📦 **PyPI**: [https://pypi.org/project/nusanlp/](https://pypi.org/project/nusanlp/)

---

## 📄 License

Distributed under the **MIT License**. See `LICENSE` for details.
