Metadata-Version: 2.5
Name: darsh
Version: 2.0.1
Summary: A skin and convenience layer over pandas and matplotlib.
Author: Satyam Rana
License: MIT
Keywords: AI Practitioner,Dashboard,Data Analytics,Data Science,Interactive,Python,Satyam Rana,UI,Web Development
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Classifier: Topic :: Software Development :: User Interfaces
Requires-Python: >=3.9
Requires-Dist: fastapi>=0.100.0
Requires-Dist: matplotlib>=3.7.0
Requires-Dist: nbformat>=4.2.0
Requires-Dist: numpy>=1.24.0
Requires-Dist: pandas>=2.0.0
Requires-Dist: plotly>=5.0.0
Requires-Dist: uvicorn>=0.23.0
Provides-Extra: dev
Requires-Dist: build>=1.0.0; extra == 'dev'
Requires-Dist: mypy>=1.0.0; extra == 'dev'
Requires-Dist: pytest>=7.0.0; extra == 'dev'
Description-Content-Type: text/markdown

# ✨ Darsh

**Built on pandas and matplotlib, not instead of them.**

[![PyPI - Version](https://img.shields.io/pypi/v/darsh.svg)](https://pypi.org/project/darsh/)
[![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](https://opensource.org/licenses/MIT)
[![Python Versions](https://img.shields.io/pypi/pyversions/darsh.svg)](https://pypi.org/project/darsh/)

---

> **North Star**: Darsh is a skin and convenience layer over pandas and matplotlib — not a replacement for them. Every Darsh call returns a real `pandas.DataFrame`, `matplotlib.axes.Axes`, or `plotly.graph_objects.Figure` so you never lose access to the ecosystem you already know and trust.

No proprietary DataFrame wrappers. No black-box chart objects. Just clean ergonomics registered via pandas' sanctioned `df.darsh` accessor.

```python
import pandas as pd
import darsh

# 1. Load data with standard pandas — 100% native
df = pd.read_csv("sales.csv")

# 2. Clean effortlessly with chainable, explainable methods
df = (
    df.darsh.clean_names()
    .darsh.fill_missing(strategy="smart")
    .darsh.handle_outliers(columns=["price"], action="clip")
)

# 3. Use standard pandas methods anywhere in the chain
top_regions = df.groupby("region")["revenue"].sum().reset_index()

# 4. Generate styled matplotlib charts — returns real matplotlib.axes.Axes!
ax = top_regions.darsh.bar(x="region", y="revenue", title="Regional Revenue")

# 5. Native matplotlib methods work seamlessly
ax.figure.savefig("revenue.png", dpi=300)
```

---

## ⚡ Installation

```bash
pip install darsh
```

Dependencies: `pandas>=2.0`, `matplotlib>=3.7`, `plotly>=5.0`, `fastapi>=0.100`, `uvicorn>=0.23`.

---

## 🧼 Transparent Data Cleaning

Every cleaning method returns a real `pandas.DataFrame` and adheres to transparent, documented behavior:

```python
df = (
    df.darsh.clean_names()                          # standardizes to snake_case identifiers
      .darsh.drop_duplicates()                      # removes exact duplicate rows
      .darsh.fill_missing(strategy="smart")         # median for numeric, mode for text; columns >30% missing untouched
      .darsh.handle_outliers(columns=["revenue"], action="clip") # IQR clipping
      .darsh.clean_text(columns=["notes"])          # strips whitespace, lowercases
      .darsh.add_date_features(columns=["order_date"]) # extracts year, month, day_name, is_weekend
      .darsh.infer_types()                          # casts dates, downcasts numerics, categorical optimization
)
```

### Explainable Defaults
No silent, unexplainable mutations. For instance, `fill_missing(strategy="smart")` rules are plain English:
* **Numeric columns**: Imputed using column median.
* **Text / Categorical columns**: Imputed using column mode (or `"Unknown"` if no mode exists).
* **Columns with >30% missing data**: Left untouched and flagged in `.darsh.profile()` rather than distorted.

---

## 🔍 Explainable Profiling & Quality Scoring

```python
# Detailed tabular profile: dtype, % missing, cardinality, skew, and outlier count
profile_df = df.darsh.profile()
print(profile_df)

# Inspectable quality score (0 - 100)
score = df.darsh.quality_score()
print(f"Score: {score:.1f}/100")

# Plain-English diagnostic breakdown of score deductions
print(score.explain())
```

Example `.explain()` output:
```text
Darsh Data Quality Score: 92.4/100.0
============================================
Identified 2 issue(s) affecting score:
  1. [Missing Data] Column 'customer_notes': -3.2 pts — 14 missing value(s) (4.2% missing)
  2. [Outliers] Column 'revenue': -4.4 pts — 18 outlier(s) (5.4% outside 1.5*IQR)
--------------------------------------------
Base: 100.0 | Total Deductions: -7.6 | Final: 92.4
```

---

## 📊 Visualizations: Matplotlib 2D & Plotly 3D

### 2D Charts (Matplotlib)
All 2D chart methods return real `matplotlib.axes.Axes`, accept optional `ax=`, and apply modern typography, palettes, and whitespace automatically:

```python
import matplotlib.pyplot as plt

# Standalone or embedded in subplots
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(14, 5))

df.darsh.line(x="date", y="sales", title="Daily Trend", ax=ax1)
df.darsh.bar(x="category", y="sales", title="Category Sales", ax=ax2)

plt.tight_layout()
plt.show()
```

Available 2D chart methods:
* `df.darsh.line(x=..., y=..., title=..., ax=...)`
* `df.darsh.bar(x=..., y=..., title=..., horizontal=False, ax=...)`
* `df.darsh.scatter(x=..., y=..., size=..., color=..., ax=...)`
* `df.darsh.donut(values=..., names=..., hole=0.55, ax=...)`

### 3D Charts (Plotly)
Clearly labeled, Plotly-backed 3D visualizations returning real `plotly.graph_objects.Figure`:

```python
fig = df.darsh.scatter_3d(x="price", y="units", z="margin", color="category")
fig.show()
```

Available 3D chart methods:
* `df.darsh.scatter_3d(x, y, z, color=None, size=None)`
* `df.darsh.line_3d(x, y, z, color=None)`
* `df.darsh.bar_3d(x, y, z)`

---

## 🚀 One Unified Dashboard API

No competing OOP vs. procedural paradigms. Create dashboards directly using the charts and DataFrame you've already built:

```python
import darsh

# 1. Create your charts
ax1 = df.darsh.line(x="order_date", y="revenue", title="Revenue Trend")
ax2 = df.darsh.bar(x="region", y="revenue", title="Revenue by Region")

# 2. Build dashboard with KPIs and reactive filter controls
app = darsh.dashboard(
    title="Sales Intelligence Hub",
    charts=[ax1, ax2],
    data=df,
    kpis=[
        darsh.kpi(label="Total Revenue", value=df["revenue"].sum(), delta="+12.4%"),
        darsh.kpi(label="Avg Order Value", value=df["revenue"].mean()),
        darsh.kpi(label="Orders", value=len(df)),
    ],
    filters=["region", "category"],
    theme="light",
)

# 3. Run the interactive local server
app.run(port=8501, open_browser=True)

# Or export standalone HTML
app.save_html("dashboard.html")
```

When users select filter values in the web UI, Darsh automatically filters the backing DataFrame and re-renders the charts in real-time.

---

## 🛡️ License

MIT License &copy; 2026 Satyam Rana.
