Metadata-Version: 2.4
Name: tidyforge
Version: 0.1.0.dev0
Summary: A modular, extensible, high-performance Python library for Pandas DataFrame cleaning with a clean, chainable API.
Author: Ashwin Ashok Chavan
License: MIT
Project-URL: Homepage, https://github.com/Ash7540/TidyForge
Project-URL: Repository, https://github.com/Ash7540/TidyForge
Project-URL: Issues, https://github.com/Ash7540/TidyForge/issues
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pandas>=2.0.0
Requires-Dist: numpy>=1.24.0
Requires-Dist: contractions>=0.1.73
Requires-Dist: demoji>=1.1.0
Provides-Extra: dev
Requires-Dist: pytest>=7.0.0; extra == "dev"
Requires-Dist: pytest-cov>=4.0.0; extra == "dev"
Requires-Dist: ruff>=0.1.0; extra == "dev"
Requires-Dist: mypy>=1.0.0; extra == "dev"
Dynamic: license-file

# TidyForge

[![CI](https://github.com/Ash7540/TidyForge/actions/workflows/ci.yml/badge.svg)](https://github.com/Ash7540/TidyForge/actions/workflows/ci.yml)
[![Python Version](https://img.shields.io/badge/python-3.9%20%7C%203.10%20%7C%203.11%20%7C%203.12-blue)](https://github.com/Ash7540/TidyForge)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)

**TidyForge** is a modular, extensible, high-performance Python library that simplifies and automates data cleaning for Pandas DataFrames through an intuitive, chainable API and a robust Command-Line Interface.

---

## 🚀 Key Features

- **Chainable API**: Expressive, fluent interface for combining multiple data cleaning operations cleanly.
- **Modular Architecture**: Separate dedicated modules for missing values, duplicates, type conversions, text, dates, numbers, outliers, contact validation, and range/constraint validation.
- **Declarative Pipelines**: Build reusable cleaning recipes, serialize them to JSON, and reload them on-demand.
- **Parallel Batch Processing**: Clean directories of CSV/JSON files concurrently using multiple CPU cores.
- **Memory Optimization**: Intelligently downcast numerical variables and convert low-cardinality string fields to categories.
- **Diagnostic Reports**: Generate interactive HTML quality dashboards and structural JSON logs.
- **Extensible Plugin System**: Easily register custom plugins subclassing `BasePlugin` to extend the Cleaner method namespace.

---

## 📦 Installation

```bash
pip install tidyforge
```

---

## 📖 How It Works & Core API Guide

TidyForge wraps a Pandas DataFrame in a `Cleaner` instance. Methods return the `Cleaner` instance itself, allowing you to chain operations sequentially. When complete, calling `.to_dataframe()` returns the underlying cleaned DataFrame.

### 1. Basic Cleaning Chaining

```python
import pandas as pd
import tidyforge as tf

df = pd.DataFrame({
    "name": ["  Alice ", "BOB", "Charlie  "],
    "email": ["alice@domain.com", "invalid-email", None],
    "income": ["$50,000", "$60,000", None],
    "age": [25.0, -5.0, 30.0]
})

# Chain clean operations
cleaner = (
    tf.Cleaner(df)
    # Parse currency text to float
    .clean_currency(columns=["income"])
    # Clean text casing and remove whitespaces
    .clean_text(columns=["name"], case="title", strip=True)
    # Impute missing income values with median
    .fill_missing(columns=["income"], strategy="median")
    # Coerce out-of-range age violations to NaN
    .validate_ranges(columns=["age"], min_val=0, errors="coerce")
)

# Export cleaned DataFrame
cleaned_df = cleaner.to_dataframe()
```

### 2. Missing Value Handling
TidyForge handles missing data using directional filling, interpolation, indicator variables, and classic imputation strategies:

```python
cleaner = (
    tf.Cleaner(df)
    # Drop rows with >50% nulls
    .drop_missing_rows(row_threshold=0.5)
    # Forward-fill column missing values
    .fill_missing_directional(columns=["status"], method="ffill")
    # Linear interpolation for timeseries numeric columns
    .interpolate_missing(columns=["temperature"], method="linear")
    # Replace custom nan markers (e.g. "N/A", "null") with actual NaN
    .replace_custom_markers_with_na(markers=["N/A", "missing"])
)
```

### 3. Outliers Winsorization & Trimming
Detect and handle numerical outliers using standard IQR or Z-score boundaries:

```python
cleaner = (
    tf.Cleaner(df)
    # Handle outliers in "sales" using IQR scale-factor 1.5 by clipping values
    .handle_outliers(columns=["sales"], method="iqr", strategy="clip", factor=1.5)
)
```

### 4. Custom Constraints & Validation
Assert boundary rules, enum categories, or check for mixed data types:

```python
# Custom lambda predicate constraint check
cleaner = (
    tf.Cleaner(df)
    # Validate constraints: value must be positive
    .validate_constraints(
        columns=["score"],
        checks=[lambda x: x >= 0],
        errors="coerce"
    )
    # Resolve object columns with mixed data types by converting to string
    .resolve_mixed_types(columns=["notes"], strategy="string")
)
```

---

## 🛠️ Reusable Pipelines

Create declarative pipelines for consistent cleaning configurations across training and inference workflows:

```python
from tidyforge import CleaningPipeline

# Construct pipeline recipe
pipeline = (
    CleaningPipeline()
    .add_step("drop_empty_columns")
    .add_step("clean_currency", columns=["salary"])
    .add_step("optimize_memory")
)

# Apply to a dataframe
clean_df = pipeline.run(df)

# Serialize to file
pipeline.save("recipe.json")

# Reload from file
loaded_pipeline = CleaningPipeline.load("recipe.json")
```

---

## ⚡ Concurrency Batch Processing

Clean directories containing multiple files simultaneously using parallel CPU processes:

```python
from tidyforge import clean_batch_files, CleaningPipeline

pipeline = CleaningPipeline().add_step("optimize_memory")

# Process files concurrently using 4 workers
clean_batch_files(
    pipeline=pipeline,
    input_files=["data1.csv", "data2.csv", "data3.json"],
    output_dir="./cleaned_exports",
    suffix="_cleaned",
    concurrency=4
)
```

---

## 📈 Quality Scoring & Diagnostic Reports

TidyForge calculates an overall dataset quality score based on completeness, uniqueness, validity, and type consistency metrics:

```python
cleaner = tf.Cleaner(df).drop_empty_columns()

# Get dictionary metadata summary
summary = cleaner.cleaning_summary()

# Export a stunningOutfit/Inter-based HTML diagnostic dashboard
cleaner.generate_html_report("dashboard.html")

# Export structured JSON diagnostics
cleaner.generate_json_report("report.json")
```

---

## 💻 Command-Line Interface (CLI)

Install TidyForge to run cleaning operations directly in your terminal:

```bash
# 1. Evaluate dataset quality score
tidyforge score -i raw.csv

# 2. Clean a dataset using a pipeline recipe
tidyforge clean -i raw.csv -o clean.csv -p recipe.json --optimize-memory --report-html report.html

# 3. Clean multiple datasets concurrently
tidyforge clean -i raw1.csv raw2.csv -o ./cleaned_out -p recipe.json -c 4
```

---

## 🔌 Custom Plugins

Extend the `Cleaner` method namespace with domain-specific plugins:

```python
from tidyforge.plugins.base import BasePlugin
from tidyforge import Cleaner

class CustomTrimmer(BasePlugin):
    """Custom plugin to trim specific columns."""
    
    def transform(self, df, columns=None, **kwargs):
        cols = columns or df.columns
        df_copy = df.copy()
        for col in cols:
            df_copy[col] = df_copy[col].astype(str).str.strip()
        return df_copy

# Register the plugin under a custom method name
Cleaner.register_plugin("custom_trim", CustomTrimmer)

# Use it directly in Cleaner method chaining!
clean_df = Cleaner(df).custom_trim(columns=["name"]).to_dataframe()
```

---

## 📄 License

This project is licensed under the MIT License - see the [LICENSE](LICENSE) file for details.
