Metadata-Version: 2.4
Name: text2rel
Version: 1.0
Summary: An HTML cleaner customized for kinship ties extraction from large text corpra.
Requires-Python: >=3.10,<3.13
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Provides-Extra: plots
Requires-Dist: beautifulsoup4 (>=4.12,<5.0)
Requires-Dist: jellyfish
Requires-Dist: matplotlib ; extra == "plots"
Requires-Dist: numpy (==1.26.4)
Requires-Dist: openai
Requires-Dist: pandas (==2.2.2)
Requires-Dist: pymupdf (>=1.24,<2)
Requires-Dist: rapidfuzz
Requires-Dist: regex
Requires-Dist: scipy ; extra == "plots"
Requires-Dist: tiktoken
Description-Content-Type: text/markdown

![Alt text](images/Kinship_Ties_Extraction.png)
![PyPI version](https://img.shields.io/pypi/v/html-cleaner-kinship)



This is the package implementation of the code for the kinship ties and relational information from large text corpora, such as genealogies, biographies, and historical dictionaries. 

### Important information: 
- For example package usage, see ```test_workflow.ipynb```
- For information about individual functions before the package implementation and release, see ```examples``` folder. 

### Expected input files

Each document should normally have these three files together in the same
folder under `documents_root`:

```text
0_GenelogiesAndBiographies/
├── Example genealogy/
│   ├── Example genealogy_mod.htm       # source HTML used for the inventory
│   ├── Example genealogy_mod.pdf       # PDF with a readable OCR text layer
│   └── Example genealogy_Original.pdf  # original PDF without OCR
└── Biographies/
    └── Example biography/
        ├── Example biography_mod.htm
        ├── Example biography_mod.pdf
        └── Example biography_Original.pdf
```

The filenames must share the same base name. The package uses the HTML for
inventory creation, prefers `*_mod.pdf` for page matching, and falls back to
`*_Original.pdf`, which can be OCRed with Tesseract when necessary.

### Current release: 
Current release includes inventory creation, setting text bounds, assigning font usage and selecting appropriate text chunks from the files (including restriction to "MainText" fonts and JumPJumP insertion). 

### Files structure: 

1. #### cleaner.py 
Contains the HTMLCleaner class and thus the main logic. 

2. #### utils.py
Contains helper functions (like JumPJumP insertion)

3. #### io.py
Contains the file processing logic, such as loading and saving the JSON inventory. 

4. #### page_matching.py
Adds one-based PDF `start_page` and `end_page` values to chunk metadata using
exact and fuzzy text matching. Missing PDFs are skipped with null page values
and explicit filename guidance.

### Add PDF page numbers

If you already created an `HTMLCleaner`, use `cleaner.add_pdf_pages()`. If you
only have an inventory JSON file, use `add_pages_to_inventory_file()` directly.

```python
summary = cleaner.add_pdf_pages(
    documents_root="../0_GenelogiesAndBiographies",
    output_path="exact_html_inventory_new_ids_cleaned_pages.json",
)
```

The default fuzzy threshold is `0.80`. The matcher prefers `*_mod.pdf`, falls
back to `_Original.pdf` when needed, and can use Tesseract for image-only
originals when Tesseract is installed.

### Human Labeler Allocation

When text chunks are allocated to labelers, each generated TXT filename
indicates its task type:

- `Task_0` contains chunks shared across all labelers.
- `Task_1` through `Task_9` contain labeler-specific, unshared chunks.

### New section 4 assessment workflow

The reusable assessment follows sections 4.2.2.1 and 4.2.2.2 of
`240426 Information extraction system 260801.ipynb`. Install the optional plot
dependencies, load the processed human and AI tables, and then create the
assessment tables before drawing anything:

```bash
pip install -e ".[plots]"
```

```python
from text2rel import (
    ReliabilityAssessment,
    build_sweep_report,
    plot_sweep_report,
    plot_threshold_curves,
)

assessment = ReliabilityAssessment("inventory.json")
assessment.load_relations_data(relations_csv="human_relations.csv")
assessment.load_chatgpt_relations("machine_relations.csv")

# one_to_one is the package's stricter default for headline metrics.
relation_results = assessment.assess_llm_against_humans(data="relations")
plot_threshold_curves(relation_results["threshold_metrics"], direction="one_to_one")

# directional_best reproduces the notebook's two reusable-candidate views.
diagnostics = assessment.assess_llm_against_humans(
    data="relations", matching_mode="directional_best"
)
fn_report = build_sweep_report(
    diagnostics["human_to_llm_best_pairs"], data="relations"
)
plot_sweep_report(fn_report)
```

`assess_llm_against_humans()` returns the selected pairs, threshold metrics,
false-negative and false-positive review tables, structural exclusions, and
configuration metadata. Plot functions only render these returned tables; they
do not change matching or scores. See the separate **New section 4 assessment
plots and tables** section at the end of `test_workflow.ipynb` for relations
and events, notebook-style titles, confidence/length diagnostics, and review
summaries.

### Labelbox event construction modes

`labelbox_events_to_dataframe(..., construction_mode="graph")` is the default.
It follows section 2.4.3.2 of the 260801 notebook: connected annotations are
rebuilt around one Verb, safe reversed arrows are corrected, and incomplete or
ambiguous structures remain auditable. Use `construction_mode="legacy_simple"`
only when explicitly reproducing the package's earlier permissive conversion.
Relation conversion uses its audited endpoint joins separately.

