Metadata-Version: 2.4
Name: mtminepy
Version: 0.2.0
Summary: MTMinePy - A Multilingual Text Mining Platform with Pluggable Tokenizers and LLM-Assisted Segmentation Verification
Author: EasyCam
License-Expression: GPL-3.0-only
Project-URL: Homepage, https://github.com/EasyCam/MTMinePy
Project-URL: Repository, https://github.com/EasyCam/MTMinePy
Project-URL: Issues, https://github.com/EasyCam/MTMinePy/issues
Keywords: text-mining,nlp,multilingual,flask,academic,text-analysis,tokenizer,word-segmentation,llm,ollama,stylometry,digital-humanities
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Science/Research
Classifier: Intended Audience :: Education
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Classifier: Topic :: Text Processing :: Linguistic
Classifier: Framework :: Flask
Classifier: Operating System :: OS Independent
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: flask>=2.0
Requires-Dist: flask-session>=0.5
Requires-Dist: pandas>=1.5
Requires-Dist: numpy>=1.23
Requires-Dist: scikit-learn>=1.2
Requires-Dist: scipy>=1.10
Requires-Dist: langdetect>=1.0
Requires-Dist: jieba>=0.42
Requires-Dist: nltk>=3.8
Requires-Dist: matplotlib>=3.6
Requires-Dist: networkx>=3.0
Requires-Dist: openpyxl>=3.0
Requires-Dist: statsmodels>=0.14
Requires-Dist: waitress>=2.1
Provides-Extra: tokenizers
Requires-Dist: jieba>=0.42; extra == "tokenizers"
Requires-Dist: snownlp>=0.12; extra == "tokenizers"
Requires-Dist: pkuseg>=0.0.25; extra == "tokenizers"
Requires-Dist: thulac>=0.2; extra == "tokenizers"
Requires-Dist: hanlp>=2.1; extra == "tokenizers"
Requires-Dist: ltp>=4.2; extra == "tokenizers"
Requires-Dist: sudachipy>=0.6; extra == "tokenizers"
Requires-Dist: sudachidict-core>=20230109; extra == "tokenizers"
Requires-Dist: janome>=0.5; extra == "tokenizers"
Requires-Dist: fugashi>=1.3; extra == "tokenizers"
Requires-Dist: unidic-lite>=1.0; extra == "tokenizers"
Requires-Dist: tinysegmenter>=0.3; extra == "tokenizers"
Requires-Dist: konlpy>=0.6; extra == "tokenizers"
Requires-Dist: soynlp>=0.0.493; extra == "tokenizers"
Requires-Dist: pymorphy3>=2.0; extra == "tokenizers"
Requires-Dist: pymorphy3-dicts-ru>=2.4; extra == "tokenizers"
Requires-Dist: razdel>=0.5; extra == "tokenizers"
Requires-Dist: natasha>=1.6; extra == "tokenizers"
Requires-Dist: pythainlp>=4.0; extra == "tokenizers"
Requires-Dist: underthesea>=6.0; extra == "tokenizers"
Requires-Dist: spacy>=3.5; extra == "tokenizers"
Requires-Dist: stanza>=1.6; extra == "tokenizers"
Requires-Dist: blingfire>=0.1.8; extra == "tokenizers"
Requires-Dist: sacremoses>=0.1; extra == "tokenizers"
Requires-Dist: nltk>=3.8; extra == "tokenizers"
Provides-Extra: llm
Requires-Dist: openai>=1.0; extra == "llm"
Requires-Dist: ollama>=0.2; extra == "llm"
Requires-Dist: transformers>=4.40; extra == "llm"
Provides-Extra: full
Requires-Dist: janome>=0.5; extra == "full"
Requires-Dist: spacy>=3.5; extra == "full"
Requires-Dist: hanlp>=2.1; extra == "full"
Requires-Dist: ltp>=4.2; extra == "full"
Requires-Dist: umap-learn>=0.5; extra == "full"
Requires-Dist: boruta>=0.3; extra == "full"
Requires-Dist: prince>=0.13; extra == "full"
Requires-Dist: lexical-diversity>=0.1; extra == "full"
Requires-Dist: seaborn>=0.12; extra == "full"
Requires-Dist: wordcloud>=1.9; extra == "full"
Requires-Dist: xlsxwriter>=3.0; extra == "full"
Requires-Dist: pymorphy3>=2.0; extra == "full"
Requires-Dist: pymorphy3-dicts-ru>=2.4; extra == "full"
Requires-Dist: razdel>=0.5; extra == "full"
Requires-Dist: natasha>=1.6; extra == "full"
Requires-Dist: sudachipy>=0.6; extra == "full"
Requires-Dist: sudachidict-core>=20230109; extra == "full"
Requires-Dist: fugashi>=1.3; extra == "full"
Requires-Dist: unidic-lite>=1.0; extra == "full"
Requires-Dist: tinysegmenter>=0.3; extra == "full"
Requires-Dist: soynlp>=0.0.493; extra == "full"
Requires-Dist: stanza>=1.6; extra == "full"
Requires-Dist: sacremoses>=0.1; extra == "full"
Requires-Dist: pythainlp>=4.0; extra == "full"
Requires-Dist: underthesea>=6.0; extra == "full"
Requires-Dist: openai>=1.0; extra == "full"
Provides-Extra: dev
Requires-Dist: build; extra == "dev"
Requires-Dist: twine; extra == "dev"
Requires-Dist: pytest>=7.0; extra == "dev"
Dynamic: license-file

# MTMinePy - Multilingual Text Miner with Python

[![PyPI version](https://badge.fury.io/py/mtminepy.svg)](https://pypi.org/project/mtminepy/)
[![License: GPL v3](https://img.shields.io/badge/License-GPLv3-blue.svg)](https://www.gnu.org/licenses/gpl-3.0)

**MTMinePy** is a Python-based academic text mining platform inspired by MTMineR. It is a comprehensive Flask web application designed for powerful text mining and analysis, supporting interactive visualization and advanced modeling.

## Key Features
- **Advanced NLP**: Integrated with `jieba`, `HanLP`, `LTP`, `Spacy`, and `NLTK`.
- **Multilingual**: Native support for 25+ languages including Chinese, English, Japanese, Korean, and Russian.
- **All-in-one text pipeline**: segmentation → POS tagging → full statistics → n-gram → vectorization.
- **Interactive Visualization**: Powered by **ECharts**, supporting responsive force-directed networks, dynamic word clouds, and interactive scatter plots.
- **Academic Metric Analysis**: Supports advanced distance functions (Hsim, Close, Esim) and high-end visualization.
- **Advanced Modeling**: Comprehensive suite of Unsupervised (Clustering, Topic Modeling) and Supervised learning algorithms.

## Text Segmentation & Full Statistics

Following MTMineR's idea of *"Java tools for tokenizing / counting / tagging / discriminating /
vectorizing, then R for statistics and visualization"*, MTMinePy folds the whole chain into one
tool. See the **Text Segmentation & Metrics** page in the sidebar, or the CLI below.

### Pluggable multilingual tokenizers

| Language | Supported tokenizers |
|---|---|
| Chinese | jieba (default), HanLP, LTP, pkuseg, THULAC, SnowNLP |
| Japanese | SudachiPy, fugashi (MeCab + UniDic), Janome, TinySegmenter |
| Korean | KoNLPy (Kkma / Okt / Komoran / Hannanum), soynlp |
| Russian | pymorphy3 + razdel, pymorphy2, Natasha, razdel, spaCy, Stanza |
| Thai | PyThaiNLP (Thai has no spaces — a dedicated tool is required) |
| Vietnamese | underthesea (compound-word aware) |
| Multi | spaCy, Stanza (70+ languages), NLTK, BlingFire, Sacremoses |
| Fallback | Unicode regex (works for any language) |

Missing backends are greyed out in the UI with an install hint (e.g. "missing model
ru_core_news_sm, run python -m spacy download ru_core_news_sm"); a failing backend
automatically falls back without breaking the pipeline. **Every backend is covered by
an automated self-test** (see below).

### LLM refinement of tokenization (iterative, multi-round)

After segmentation you can plug in **any model** to check and improve the tokens:

- **Local Ollama** (zero config: `qwen2.5:7b`, `granite4:350m`, …)
- **OpenAI and any OpenAI-compatible endpoint** (DeepSeek / Moonshot / vLLM / LM Studio / …)
- **Local Transformers models** (HuggingFace weights)

Two strategies, iterated round after round until it converges:

| Strategy | How it works | Best for |
|---|---|---|
| `rewrite` | The model rewrites the whole segmentation | Large models, best quality |
| `judge` | The program numbers adjacent pairs; the model only answers which to merge | **Works even with 350M models** |
| `auto` (default) | Try rewrite, fall back to judge when validation fails | General |

**Four safety gates** so the model can never damage your data:
1. **Character invariant** — each round must reproduce the text character-for-character;
2. **Anti-degeneration** — max token length + minimum token-count ratio block "merge everything into one word";
3. **Candidate filtering** — only content-word pairs are offered to the model (particles, prepositions, adverbs, pronouns, punctuation and function words are excluded);
4. **Fully traceable** — every edit (merge/split) and the model's raw reply are logged; a review-only mode is available.

### Self-test

```bash
mtminepy selftest --all        # run every tokenizer of every language
mtminepy selftest --lang ru
```

### POS (Part of Speech)

After segmentation you can filter by part of speech — e.g. keep only nouns for topic
analysis, or only adjectives for sentiment. Use either group names
(`noun` / `verb` / `adj` / `adv` / `num` / `pron` / `prep` / `conj` …) or raw tags from
any tagset (`n` / `nr` / `NN` / `VB` / `NOUN` / `名詞` …).

### Statistics

Characters, byte length (UTF-8 / GBK / UTF-16), lines / paragraphs / sentences, average and
max sentence length, token count, vocabulary size, average / max word length, word-length
distribution, POS distribution, and a family of lexical-richness measures
(TTR, Guiraud, Carroll CTTR, Herdan Log-TTR, MSTTR, hapax, Sichel, Yule's K).

### n-gram

Sliding-window counts of n adjacent units: `n=1` single word, `n=2` adjacent pairs,
`n=3` triples… Outputs `n / ngram / freq / rel_freq / doc_freq` and exports to CSV.

### Vectorization

Turns text into a document × feature matrix (BoW / TF-IDF, with n-gram features,
`max_features`, `min_df`), reports the top feature weights, and exports the matrix to CSV.

## Screenshots

### Chinese Analysis
| Co-occurrence Network | Word Cloud | Clustering |
|:---:|:---:|:---:|
| ![Chinese Network](images/中文共现网络.png) | ![Chinese WordCloud](images/中文词云.png) | ![Chinese Clustering](images/中文聚类.png) |

### English Analysis
| Co-occurrence Network | Word Cloud | Clustering |
|:---:|:---:|:---:|
| ![English Network](images/英文共现网络.png) | ![English WordCloud](images/英文词云.png) | ![English Clustering](images/英文聚类.png) |

## Installation

### Install from PyPI (Recommended)

```bash
pip install mtminepy
```

To install with all optional NLP backends (Janome, spaCy, HanLP, LTP, UMAP, Boruta, etc.):

```bash
pip install mtminepy[full]
```

### Install from source

```bash
git clone https://github.com/EasyCam/MTMinePy.git
cd MTMinePy
pip install -e .
```

## Usage

### Run from command line

After installation, run directly:

```bash
mtminepy
```

Access the dashboard at `http://localhost:5000`.

### Command-line options

```bash
mtminepy --help
mtminepy --port 8080          # Custom port
mtminepy --host 127.0.0.1    # Bind to localhost only
mtminepy --debug              # Flask debug mode
mtminepy --version            # Show version
```

### Command line: all-in-one text processing

```bash
# List available tokenizers per language
mtminepy backends --lang ru
mtminepy backends --all

# List selectable POS groups
mtminepy pos --lang cn

# Segment + POS-tag, output MTMineR-style "word/POS" tagged text
mtminepy segment Lu_AQ.txt -o Lu_AQ_Taged.txt --lang cn --keep-punct

# Full statistics (tokens / word length / byte length / sentence length / richness)
mtminepy stats Lu_AQ.txt --lang cn --pos noun

# n-gram frequencies (1..3, top 30 per level)
mtminepy ngram Lu_AQ.txt --lang cn -n 1 --max-n 3 -k 30 -o ngram.csv

# Vectorize (TF-IDF / BoW) and export the matrix
mtminepy vectorize Lu_AQ.txt --lang cn --pos noun \
    --method tfidf --ngram 1,1 --max-features 500 -o vectors.csv
```

Common options: `--lang`, `--engine` (tokenizer, default `auto`), `--pos`, `--stopwords`,
`--split-lines` (one document per line), `-o`. Input may be a single file or a directory
(each `.txt` / `.md` inside becomes one document).

### Command line: LLM refinement

```bash
# List pluggable model providers and their models
mtminepy llm
mtminepy llm --test --provider ollama --model qwen2.5:7b

# Refine tokenization with a local Ollama model (auto strategy, 2 rounds)
mtminepy refine --lang cn --provider ollama --model qwen2.5:7b \
    --rounds 2 --strategy auto --text "阿Q没有家，住在未庄的土谷祠里。"

# Any OpenAI-compatible endpoint (DeepSeek / vLLM / LM Studio …)
mtminepy refine Lu_AQ.txt -o refined.txt \
    --provider openai_compatible --base-url https://api.deepseek.com/v1 \
    --model deepseek-chat --api-key $DEEPSEEK_API_KEY --rounds 3

# Review only (report the model's suggestions without changing anything)
mtminepy refine --provider ollama --model qwen2.5:7b --review --text "…"
```

Environment variables are honoured as well: `MTMINEPY_LLM_PROVIDER`, `MTMINEPY_LLM_MODEL`,
`MTMINEPY_LLM_BASE_URL`, `MTMINEPY_LLM_API_KEY`, `OLLAMA_HOST`, `OPENAI_API_KEY`,
`OPENAI_BASE_URL`.

### Tests

```bash
pytest tests/ -v          # tokenizer matrix + refinement logic (offline, Mock model)
python tests/test_tokenizers.py
python tests/test_refine.py
```

### Run from Python

```python
from mtminepy.app import create_app

app = create_app()
app.run(host='0.0.0.0', port=5000)
```

## Advanced Capabilities

### Modeling Algorithms
MTMinePy supports a wide range of standard machine learning algorithms for text analysis:
*   **Feature Engineering**: TF-IDF, Bag of Words (CountVectorizer), N-gram support.
*   **Unsupervised Learning**:
    *   **Topic Modeling**: Latent Dirichlet Allocation (LDA), Non-negative Matrix Factorization (NMF), STM (Structural Topic Model).
    *   **Clustering**: K-Means, Agglomerative (Hierarchical), DBSCAN, Spectral Clustering.
    *   **Dimensionality Reduction**: PCA, t-SNE, UMAP, Factor Analysis.
*   **Supervised Learning** (Classification):
    *   Support Vector Machines (SVM)
    *   Random Forest
    *   Linear Discriminant Analysis (LDA)
    *   Quadratic Discriminant Analysis (QDA)
    *   Logistic Regression (Elastic Net)

### Mathematical Models (Distance & Similarity)
MTMinePy supports advanced metrics for academic research:

#### Advanced Custom Similarity Measures
1.  **Hsim (Yang Fengzhao, 2007)**
    $$ Hsim(x_i, x_j) = \frac{1}{n} \sum_{k=1}^n \frac{1}{1+|x_{ik}-x_{jk}|} $$
    
2.  **Close (Shao Changsheng, et al., 2011)**
    $$ Close(x_i, x_j) = \frac{1}{n} \sum_{k=1}^n e^{-|x_{ik}-x_{jk}|} $$

3.  **Esim (Wang Xiaoyang, et al., 2013)**
    $$ Esim(x_{ik}, x_{jk}) = \frac{1}{n} \sum_{k=1}^d \omega_k e^{-\frac{|x_{ik}-x_{jk}|}{|x_{ik}-x_{jk}|+|x_{ik}+x_{jk}|/2}} $$
