Metadata-Version: 2.4
Name: quran-toolkit
Version: 0.1.0
Summary: Detect, tag, tokenize, diacritize and retrieve Qur'anic quotations inside Arabic/Persian text.
Author-email: Abbas Safardoost <a.safardoust@gmail.com>
License-Expression: MIT
Project-URL: Homepage, https://github.com/asdoost/quran-toolkit
Project-URL: Repository, https://github.com/asdoost/quran-toolkit
Project-URL: Issues, https://github.com/asdoost/quran-toolkit/issues
Keywords: quran,ayah,arabic,persian,nlp,text-processing,islamic-studies
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Natural Language :: Arabic
Classifier: Natural Language :: Persian
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.8
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Religion
Classifier: Topic :: Text Processing :: Linguistic
Requires-Python: >=3.8
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: arabing>=0.1.0
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"
Dynamic: license-file

# quran-toolkit 📖
﷽

**quran-toolkit** is a Python library for detecting and tagging Qur'anic quotations embedded within Arabic or Persian text. It scans free text, matches sequences of words against a reference Qur'an corpus, and replaces/annotates/reconstructs them — making it easy to identify, cite, diacritize, or retrieve Qur'anic content inside articles, books, sermons, transcripts, or any Arabic/Persian text corpus.

---

## Features

- 🔍 **Sequence matching** — detects runs of consecutive words that correspond to real, contiguous Qur'anic text (not just isolated word hits).
- 🕌 **Multi-ayah spans** — correctly tags quotations that cross ayah boundaries (e.g. `الرحمن الرحیم مالک یوم الدین` → `[1:3-4]`).
- ✨ **Script normalization** — powered by [`arabing`](https://pypi.org/project/arabing/), robust to diacritics (tashkeel), tatweel, Qur'anic recitation/waqf/structural marks, and common orthographic variants (`أ/إ/آ/ا`, `ي/ى/ی`, `ة/ه`, `ك/ک`, `ؤ/و`, `ئ/ی`).
- 📝 **Multiple spellings per word** — the reference CSV can list several accepted written forms per word (e.g. `الکتاب`/`الکتب`, `الصراط`/`الصرط`), all treated as equivalent.
- ✍️ **Disjoint `و` handling** — correctly detects quotations even when Persian writing conventions split the conjunction `و` from the following word (e.g. `و لا` instead of the Qur'anic `ولا`), including multiple/back-to-back occurrences anywhere in the quotation, and quotations that begin at a word whose Qur'anic form has a leading `و` the writer naturally omits.
- 🔖 **Three tagging modes** — replace quotations with `[surah:ayah]`, `[SurahName:ayah]`, or a fixed 3-letter placeholder like `[QRN]`.
- 🔢 **Persian digit output** — surah/ayah numbers are rendered in Persian digits by default (e.g. `[۱:۲]`), with an option to use Latin digits.
- 📋 **Tokenize view** — render each detected ayah on its own labeled line while keeping all surrounding text exactly where it was.
- 🕋 **Diacritization** — reconstruct the fully diacritized Uthmani wording of any detected (bare) quotation.
- 📖 **Retrieval** — fetch the diacritized text of the specific ayah(s) cited, or the entire surah(s) they belong to.
- ⚖️ **Basmalah / repeated-ayah disambiguation** — heuristics to reduce mis-attribution of very common phrases that occur identically in multiple places in the Qur'an.
- 📦 **Structured results** — get back rich `Match` objects (surah, ayah range, per-word positions, character offsets, matched text), not just a rewritten string.
- 🖥️ **CLI included** — tag, tokenize, diacritize, or retrieve from a text file without writing any code.
- ♻️ **Reusable detector object** — load and index the Qur'an once, run detection on many texts efficiently.

---

## Installation

```bash
pip install quran-toolkit
```

Or, from source:

```bash
git clone https://github.com/asdoost/quran-toolkit.git
cd quran-toolkit
pip install -e .
```

quran-toolkit depends on [`arabing`](https://pypi.org/project/arabing/) for script-aware normalization and digit conversion; it's installed automatically as a dependency.

---

## Reference Files

quran-toolkit bundles two reference files under `quran_toolkit/data/`, used for different purposes:

### 1. `quran.csv` — word-matching index

Used to detect quotations. Each row represents **one word position** in the Mushaf:

```
surah_index, ayah_index, word_position, form_1, form_2, ...
```

- `surah_index` — the surah number (1–114)
- `ayah_index` — the ayah number within the surah
- `word_position` — the position of the word within the ayah (1-based)
- `form_1, form_2, ...` — one or more accepted spellings of that word

Example (excerpt of Surah Al-Fatiha):

```csv
1,1,1,بسم
1,1,2,الله
1,1,3,الرحمن
1,1,4,الرحیم
1,2,1,الحمد
1,2,2,لله
1,2,3,رب
1,2,4,العالمین,العلمین
1,4,1,مالک,ملک
```

### 2. `quran-uthmani.txt` — diacritized text

Used by `diacritize()` and `retrieve()` to reconstruct the fully-voweled Uthmani wording. One line per ayah, pipe-delimited:

```
surah|ayah|diacritized_text
```

Example:

```
1|1|بِسْمِ ٱللَّهِ ٱلرَّحْمَٰنِ ٱلرَّحِيمِ
1|2|ٱلْحَمْدُ لِلَّهِ رَبِّ ٱلْعَٰلَمِينَ
```

Full versions of both files, covering all 114 surahs, are required for real-world use. Small samples covering Surah Al-Fatiha and part of Al-Baqarah are included under `tests/data/` for testing and experimentation. You can point `AyahDetector` at your own files via `quran_csv_path` / `quran_uthmani_path`.

---

## Quick Start

```python
from quran_toolkit import AyahDetector

# Loads the bundled reference files by default
detector = AyahDetector()

text = (
    "قال الخطیب: الحمد لله رب العالمین، ثم ذکر آیة "
    "الرحمن الرحیم مالک یوم الدین لیوضح معناها."
)

print(detector.tag(text))
```

Output:

```
قال الخطیب: [۱:۲]، ثم ذکر آیة [۱:۳-۴] لیوضح معناها.
```

---

## Usage

### 1. `tag()` — replace quotations with a reference

Three formatting modes are available via `ref_mode`:

```python
# 1. "index" (default): [surah_number:ayah_number]
detector.tag(text)
# "... [۱:۲] ..."

# 2. "name": [SurahName:ayah_number]
detector.tag(text, ref_mode="name")
# "... [الفاتحة:۲] ..."

# 3. "tag": a fixed 3-letter placeholder, default "[QRN]"
detector.tag(text, ref_mode="tag")
# "... [QRN] ..."

# customize the placeholder (must be exactly 3 capital letters in brackets)
detector.tag(text, ref_mode="tag", tag_placeholder="[COR]")
# "... [COR] ..."
```

By default, surah/ayah numbers are rendered in **Persian digits**. Pass `persian_digits=False` for Latin digits:

```python
detector.tag(text, persian_digits=False)
# "... [1:2] ..."
```

### 2. `tokenize()` — list each ayah in place

Keeps all surrounding text untouched, and puts each detected ayah on its own line in the form `[matched words] surah:ayah`:

```python
print(detector.tokenize(text, persian_digits=False))
```

```
قال الخطیب: 
[الحمد لله رب العالمین] 1:2
، ثم ذکر آیة 
[الرحمن الرحیم] 1:3
[مالک یوم الدین] 1:4
 لیوضح معناها.
```

Quotations spanning multiple ayat are automatically split into one line per ayah. `ref_mode` and `tag_placeholder` work the same way as in `tag()`.

### 3. `diacritize()` — reconstruct full Uthmani wording

Replaces each detected (bare/undiacritized) quotation with its fully diacritized Qur'anic text, pulled from `quran-uthmani.txt`:

```python
print(detector.diacritize("الحمد لله رب العالمین"))
# ٱلْحَمْدُ لِلَّهِ رَبِّ ٱلْعَٰلَمِينَ
```

If a quotation starts at a word whose Qur'anic form has a leading `و` that the writer omitted, `diacritize()` restores the full authentic wording, including that `و`.

### 4. `retrieve()` — fetch cited or full-surah text

```python
# Just the specific ayah(s) actually cited
print(detector.retrieve(text, persian_digits=False))
```
```
=== الفاتحة (1:2) ===
2: ٱلْحَمْدُ لِلَّهِ رَبِّ ٱلْعَٰلَمِينَ

=== الفاتحة (1:3-4) ===
3: ٱلرَّحْمَٰنِ ٱلرَّحِيمِ
4: مَٰلِكِ يَوْمِ ٱلدِّينِ
```

```python
# The entire surah(s) containing the cited ayah(s)
print(detector.retrieve(text, full_surah=True, persian_digits=False))
```
```
=== الفاتحة (1) ===
1: بِسْمِ ٱللَّهِ ٱلرَّحْمَٰنِ ٱلرَّحِيمِ
2: ٱلْحَمْدُ لِلَّهِ رَبِّ ٱلْعَٰلَمِينَ
...
7: صِرَٰطَ ٱلَّذِينَ أَنْعَمْتَ عَلَيْهِمْ غَيْرِ ٱلْمَغْضُوبِ عَلَيْهِمْ وَلَا ٱلضَّآلِّينَ
```

### 5. Getting structured matches

For programmatic use (analytics, UI highlighting, exporting, etc.), use `find()` to get a list of `Match` objects instead of a rewritten string:

```python
matches = detector.find(text)

for m in matches:
    print(m.surah, m.ayah_start, m.ayah_end, m.char_start, m.char_end, m.text)
```

Each `Match` exposes:

| Attribute            | Description                                                    |
|----------------------|------------------------------------------------------------------|
| `surah`              | Surah number of the match                                       |
| `ayah_start`         | First ayah number covered by the match                          |
| `ayah_end`           | Last ayah number covered by the match                            |
| `word_start`         | Index of the first matched word in the Qur'an token list        |
| `word_end`           | Index of the last matched word (inclusive)                       |
| `char_start`         | Start character offset in the original text                     |
| `char_end`           | End character offset in the original text                        |
| `text`               | The actual matched substring, as found in the input              |
| `words`              | List of `WordSpan` objects (per-word ayah, position, offsets, text) |
| `reference(ref_mode)`| Formatted reference string for the given mode                    |
| `ayah_lines(ref_mode)`| Per-ayah `"[words] surah:ayah"` lines (used by `tokenize()`)     |
| `diacritized_text(reference)` | Fully diacritized Uthmani wording of the match           |

### 6. Custom reference formatting

Pass your own `template` function to `tag()` to fully control how tags are rendered:

```python
def custom_template(m):
    return f"(Qur'an {m.surah}:{m.ayah_start})"

detector.tag(text, template=custom_template)
```

### 7. Tuning the minimum match length

By default, only runs of **4 or more consecutive Qur'anic words** are treated as quotations (to avoid false positives on common short phrases). Adjust with `min_words`:

```python
detector = AyahDetector(min_words=3)
# or per-call:
detector.tag(text, min_words=5)
```

---

## Handling Disjoint `و` (Persian Writing Convention)

In Qur'anic orthography, the conjunction `و` ("and") is always joined to the following word (e.g. `والذین`, `ولا`, `وقالوا`). In everyday Persian writing, however, `و` is conventionally written as a separate word. quran-toolkit transparently handles this in three ways:

1. **Detached `و` mid-quotation**: `و لا` in the input matches the Qur'anic `ولا`, anywhere in the sequence, any number of times — including back-to-back occurrences.
2. **Dropped leading `و` at the start of a quotation**: if a writer starts quoting from a word whose Qur'anic form has a leading `و` (e.g. starting at `قالوا` when the actual Qur'anic word is `وقالوا`), the quotation is still correctly matched and attributed.
3. **No truncation**: these merges never desynchronize the sequence — even when a disjoint `و` occurs right before the final word of the ayah, the entire ayah (including the last word) is still captured.

```python
detector.tag("قالوا لا تخف و لا تحزن انا منجوک")
# correctly detected and tagged as one complete ayah, despite the disjoint 'و'
```

---

## Command-Line Interface

```bash
quran-toolkit input.txt [options]
```

| Option              | Description                                                        |
|---------------------|----------------------------------------------------------------------|
| `input`             | Path to input text file, or `-` for stdin                           |
| `--quran`           | Path to Qur'an word-matching CSV (default: bundled `quran.csv`)     |
| `--quran-uthmani`   | Path to diacritized Uthmani text (default: bundled `quran-uthmani.txt`); used by `--diacritize`/`--retrieve` |
| `--min-words`       | Minimum quotation length in words (default: 4)                      |
| `--ref-mode`        | `index` (default), `name`, or `tag`                                  |
| `--tag-placeholder` | Placeholder for `--ref-mode=tag`; 3 capital letters in brackets (default `[QRN]`) |
| `--latin-digits`    | Use Latin digits instead of Persian digits in output                |
| `--tokenize`        | Output one line per matched ayah, keeping surrounding text in place |
| `--diacritize`      | Replace quotations with fully diacritized Uthmani wording            |
| `--retrieve`        | Output the diacritized ayah(s)/surah(s) cited                        |
| `--full-surah`      | With `--retrieve`, output entire surah(s) instead of just cited ayah(s) |
| `-o, --output`      | Output file path (default: stdout)                                   |

Examples:

```bash
echo "الحمد لله رب العالمین" | quran-toolkit - --ref-mode name
quran-toolkit input.txt --ref-mode tag --tag-placeholder "[QRN]" -o output.txt
quran-toolkit input.txt --tokenize --latin-digits
quran-toolkit input.txt --diacritize
quran-toolkit input.txt --retrieve --full-surah
```

---

## How It Works

1. **Normalization** (`arabing`-powered) — both the Qur'an reference and the input text are normalized (diacritics, tatweel, Qur'anic recitation/waqf/structural marks removed, letter variants unified) so spelling differences don't prevent matches. Tokens that normalize to an empty string (e.g. stray punctuation or a lone diacritic) are dropped before matching so they can't spuriously interrupt a sequence.
2. **Indexing** — every normalized word form in the Qur'an is indexed to the list of positions where it occurs, allowing fast lookup of candidate starting points.
3. **Sequence extension** — for each word in the input, candidate Qur'anic positions are looked up and the match is greedily extended forward, transparently absorbing disjoint/dropped `و` as described above. Two independent counters (raw-token position and Qur'anic-word position) are tracked throughout to keep offsets correct even across multiple merges.
4. **Disambiguation** — when multiple Qur'anic locations tie for the same match length (e.g. the Basmalah, which appears identically in many surahs), continuity with the surah of the previous match is preferred, falling back to the earliest occurrence in Mushaf order.
5. **Output** — matches at or above `min_words` are replaced (`tag`), listed in place (`tokenize`), fully diacritized (`diacritize`), or used to fetch reference text (`retrieve`); everything else in the input is left untouched.

---

## Limitations & Notes

- Detection quality depends entirely on the completeness and accuracy of the Qur'an CSV you provide (all 114 surahs, with sufficient alternate spellings).
- Very short or highly repeated phrases (e.g. `الله اکبر`, `الم`) are intentionally filtered out via `min_words` to reduce false positives — lower this threshold with care.
- The current matching algorithm does not tolerate misspelled/omitted words *within* a quotation beyond the `و` handling described above (no general fuzzy matching yet) — see [Roadmap](#roadmap).
- Overlap resolution is greedy (left-to-right); in rare cases with overlapping candidate quotations, this may not pick the globally optimal set of matches.
- `diacritize()`/`retrieve()` require `quran-uthmani.txt` to have full coverage of whatever ayat you expect to detect; missing ayat simply fall back to the original matched text (`diacritize`) or are omitted from the output (`retrieve`).

---

## Roadmap

- [ ] Aho–Corasick / automaton-based indexing for large-scale corpus processing
- [ ] Optimal (non-greedy) overlap resolution via weighted interval scheduling
- [ ] Fuzzy matching mode (tolerate word substitutions/omissions, useful for OCR'd or paraphrased text)
- [ ] `require_full_ayah` strict mode (only tag matches aligned exactly to ayah boundaries)
- [ ] Export matches to JSON / pandas DataFrame
- [ ] Generalize disjoint-proclitic handling beyond `و` (e.g. `ف`)
- [ ] Pluggable normalizers for other scripts/languages

---

## Contributing

Contributions are welcome! Please open an issue or pull request on GitHub. When contributing:

1. Add/update tests under `tests/` — sample reference files are provided at `tests/data/sample_quran.csv` and `tests/data/sample_quran_uthmani.txt`.
2. Run `pytest tests/ -v` before submitting.
3. Keep normalization changes backward-compatible where possible, since they affect matching behavior globally.

---

## License

MIT License. See [LICENSE](LICENSE) for details.

---

## Acknowledgements

Built to help researchers, developers, and content platforms automatically recognize, cite, and reconstruct Qur'anic text within Arabic and Persian documents. Script normalization powered by [`arabing`](https://pypi.org/project/arabing/).
