Metadata-Version: 2.4
Name: noweda
Version: 0.2.0
Summary: Automated EDA with insights, scoring, and security-aware detection — built as a native pandas extension.
Author-email: Daniel Peng <danielpeng95@gmail.com>
License: MIT
Project-URL: Homepage, https://github.com/codewithdaniel1/NowEDA
Project-URL: Documentation, https://codewithdaniel1.github.io/NowEDA/
Project-URL: Repository, https://github.com/codewithdaniel1/NowEDA
Project-URL: Bug Tracker, https://github.com/codewithdaniel1/NowEDA/issues
Keywords: eda,data-analysis,pandas,data-quality,pii-detection,automated-eda,profiling
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.8
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Requires-Python: >=3.8
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pandas>=1.3
Requires-Dist: numpy>=1.21
Requires-Dist: pyspark>=3.4
Requires-Dist: openpyxl>=3.1
Requires-Dist: lxml>=4.6
Provides-Extra: excel
Requires-Dist: xlrd>=2.0.1; extra == "excel"
Requires-Dist: pyxlsb>=1.0.10; extra == "excel"
Requires-Dist: odfpy>=1.4.1; extra == "excel"
Provides-Extra: viz
Requires-Dist: matplotlib>=3.5; extra == "viz"
Requires-Dist: scipy>=1.7; extra == "viz"
Provides-Extra: ml
Requires-Dist: statsmodels>=0.13; extra == "ml"
Requires-Dist: scikit-learn>=1.0; extra == "ml"
Provides-Extra: test
Requires-Dist: pytest>=7; extra == "test"
Requires-Dist: pyarrow>=8.0; extra == "test"
Requires-Dist: odfpy>=1.4.1; extra == "test"
Requires-Dist: matplotlib>=3.5; extra == "test"
Requires-Dist: scipy>=1.7; extra == "test"
Provides-Extra: parquet
Requires-Dist: pyarrow>=8.0; extra == "parquet"
Provides-Extra: hdf
Requires-Dist: tables>=3.7; extra == "hdf"
Provides-Extra: spss
Requires-Dist: pyreadstat>=1.1; extra == "spss"
Provides-Extra: full
Requires-Dist: xlrd>=2.0.1; extra == "full"
Requires-Dist: pyxlsb>=1.0.10; extra == "full"
Requires-Dist: odfpy>=1.4.1; extra == "full"
Requires-Dist: matplotlib>=3.5; extra == "full"
Requires-Dist: scipy>=1.7; extra == "full"
Requires-Dist: statsmodels>=0.13; extra == "full"
Requires-Dist: scikit-learn>=1.0; extra == "full"
Requires-Dist: pyarrow>=8.0; extra == "full"
Requires-Dist: tables>=3.7; extra == "full"
Requires-Dist: pyreadstat>=1.1; extra == "full"
Dynamic: license-file

<div align="left">
  <img src="https://raw.githubusercontent.com/codewithdaniel1/NowEDA/main/assets/noweda-wordmark-logo.svg" alt="NowEDA" width="700" />
</div>

[![PyPI version](https://img.shields.io/pypi/v/noweda?cacheSeconds=300)](https://pypi.org/project/noweda/)

# NowEDA

Exploratory data analysis through a native pandas accessor. Profile columns,
find missing data and duplicates, flag potential PII, explore relationships,
and generate reports with `df.eda` or its equivalent alias `df.noweda`.

[Documentation](https://codewithdaniel1.github.io/NowEDA/) ·
[API reference](https://codewithdaniel1.github.io/NowEDA/api-reference/) ·
[Changelog](https://codewithdaniel1.github.io/NowEDA/changelog/) ·
[Report an issue](https://github.com/codewithdaniel1/NowEDA/issues)

## Install

Requires Python 3.8 or later. pip selects dependency versions compatible with
your Python version.

```bash
pip install noweda
pip install "noweda[viz]"  # Add charts and density overlays
```

## Quick start

This example runs without downloading a dataset:

```python
import pandas as pd
import noweda as eda

# Or load your own file: df = eda.read("data.csv")
df = pd.DataFrame({
    "age": [25, 31, None, 42],
    "income": [42000, 55000, 47000, 68000],
    "segment": ["A", "B", "A", "B"],
})

print(df.eda.scores_df())
print(df.eda.missing_df())
df.eda.statsall()
```

## Analysis methods

### 1. `df.eda.statsall()` — Statistical profile

Prints scores, column roles, numeric and categorical statistics, missingness,
outliers, and preprocessing suggestions. VIF uses multivariate regression
in the standard install. Optional `noweda[ml]` dependencies add time-series
diagnostics. Constant columns and insufficient observations yield unavailable VIF.

### 2. `df.eda.mlall()` — Task-aware ML guidance

With no arguments, NowEDA assesses whether supervised or unsupervised analysis is
plausible. It presents possible target columns for review, excludes likely IDs,
and ranks appropriate unsupervised directions. It never silently selects a target.

```python
# Assess a dataset when you do not yet know the ML objective.
df.eda.mlall()

# Infer binary/multiclass classification or regression from the named target.
df.eda.mlall(target="segment")

# State the objective explicitly when you already know it.
df.eda.mlall(target="segment", problem_type="classification")
df.eda.mlall(problem_type="clustering", features=["age", "income"])

# Get the printed guidance and the structured result for use in code.
plan = df.eda.mlall(target="segment", plan=True)
```

**Supported problem types**

`classification`, `regression`, `clustering`, `anomaly_detection`, and
`dimensionality_reduction`. Binary and multiclass are classification subtypes.
Forecasting is planned separately because it needs a time column, horizon, and
time-aware validation.

**Supervised tasks**

Classification and regression require `target=`. If `problem_type` is omitted,
NowEDA infers one from the target dtype and cardinality, shows the reason, and lets
you override it. The target is excluded from features. It reports label coverage,
partial labels, small samples, imbalance, likely IDs, and potential leakage before
recommending a supervised workflow.

**Unsupervised tasks**

Clustering, anomaly detection, and dimensionality reduction do not accept a target.
Use `features=` to limit the analysis to columns available for that objective.
When no usable labels are available for a selected target, NowEDA explains that
supervised training cannot begin and presents unsupervised directions instead.

**Honest guidance**

Recommendations are task-specific candidates with preprocessing, validation,
metrics, and cautions. Stars and `/5` values are estimated dataset-fit ratings
within the selected task, not measured accuracy or expected performance. NowEDA
does not train models in this step. See the [ML guidance documentation](https://codewithdaniel1.github.io/NowEDA/ml-guidance/).

### 3. `df.eda.vizall()` — Automatic charts

Install `noweda[viz]` to draw distributions, correlations, categorical associations,
missingness, and applicable temporal plots. Unsupported or undefined categorical
associations appear as `N/A`, not zero association.

```python
df.eda.vizall()              # Samples 10,000 rows when the input exceeds 50,000
# df.eda.vizall(sample=5000)  # Choose a sample size
# df.eda.vizall(sample=False) # Plot the complete dataset
```

Statistical reports use the complete DataFrame; visualization sampling affects
only charts.

### 4. `df.eda.profile_column("age")` — One-column profile

Inspect a column's distribution, missingness, outliers, and suggested transforms.

### 5. `df.eda.compare(other_df)` — Compare reports

Compare dimensions, inferred column roles, scores, and detected risks between
two DataFrames. This is not a formal statistical test for distribution drift.

## Tables and programmatic reports

| Method | Output |
|---|---|
| `scores_df()` | `Value` column indexed by score name |
| `insights_df()` | `Insight` column |
| `schema_df()` | `Column`, `dtype`, `role`, `confidence`, `unique`, `uniqueness_ratio` |
| `stats_df()` | Per-column statistics; numeric fields include `mean`, `std`, `q25`, `median`, `q75`, `skewness`, `kurtosis` |
| `missing_df()` | Missing percentages; use `format="number"` for counts |
| `duplicates_df()` | Duplicate rows and constant columns |
| `correlation_df()` | Numeric Pearson correlation matrix |
| `outliers_df()` | IQR outlier counts; use `format="percentage"` for percentages |
| `pii_df()` | `Column`, `PII_Type`, `Count` |
| `encoding_df()` | `Column`, `Encoding_Type`; `include_confidence=True` adds sample evidence |
| `summary()` | Dictionary of raw plugin outputs |
| `report()` | `results`, `scores`, `insights`, `score_breakdown`, and `encoding_details` |

In 0.1.4, outlier penalties use the fraction of observed numeric values flagged
by the IQR rule: above 1% deducts 5 points; above 5% deducts 10 from quality and
readiness. `report()["score_breakdown"]` explains the score contributions.
Base64 detection requires at least six matches and an 80% sample match rate,
with extra evidence beyond simply being decodable. Its confidence field is the
sample match fraction, not a probability that data is encoded or malicious.

Column labels may be integers, strings, or tuples but must be unique. Empty
DataFrames produce an explanatory message; their scores are not informative.

Reports are cached. Each request fingerprints the DataFrame's values and schema,
recomputing analysis when they change. This check scans the data; cached access is
not constant-time. Use `df.eda.refresh()` to explicitly force a new report.

PII detection uses patterns and can miss sensitive values or flag false positives.
A risk score of zero means no configured pattern was detected, not that data is
safe to share. Quality and readiness scores are heuristics.

## File formats and optional dependencies

```python
df = eda.read("data.csv", dtype={"customer_id": str})
# eda.read("data.jsonl") defaults to lines=True
```

| Formats / feature | Installation |
|---|---|
| CSV, TSV, TXT, JSON/JSONL, XLSX/XLSM, XML, HTML, Stata, SAS, Pickle | `pip install noweda` |
| XLS, XLSB, ODS/ODF/ODT | `pip install "noweda[excel]"` |
| Parquet, Feather, ORC | `pip install "noweda[parquet]"` |
| HDF5 | `pip install "noweda[hdf]"` |
| SPSS | `pip install "noweda[spss]"` |
| Charts and KDE overlays | `pip install "noweda[viz]"` |
| Stationarity and seasonality dependencies | `pip install "noweda[ml]"` |
| All optional analysis and format dependencies | `pip install "noweda[full]"` |

Only read Pickle files from trusted sources, since loading them can execute code.

## Large files

CSV, JSON, and all chunked reads use pandas consistently regardless of file size.
PySpark is a standard dependency. Large Parquet/ORC files (at least
128 MB) without reader options may use Spark, with a pandas fallback on failure.
Spark requires a compatible Java installation. Reads with options use pandas.
Spark loading still collects the final pandas DataFrame into local memory and
is not guaranteed to be faster.

For data that does not fit in memory, iterate over CSV or line-delimited JSON:

```python
for chunk in eda.read_chunked("large.csv", chunksize=10000, concat=False):
    print(chunk.eda.missing_df())
```

The default `concat=True` retains all chunks and allocates a combined DataFrame;
it is appropriate only when the data and concatenation overhead fit in RAM.
Per-chunk scores describe each chunk, not the whole dataset.

If processing may stop early, wrap the generator in `contextlib.closing()`:

```python
from contextlib import closing

with closing(eda.read_chunked("large.csv", concat=False)) as chunks:
    for chunk in chunks:
        print(chunk.eda.missing_df())
        break  # Optional: the reader closes and the indicator reports stopped.
```

Running indicators show elapsed time; 100% appears only after completion.
Closing a stream early displays "Stopped before completion".

## Export and CLI

```python
from noweda.report.html import generate_html_report
from noweda.report.json import generate_json_report

generate_html_report(df.eda.report(), "report.html")
generate_json_report(df.eda.report(), "report.json")
```

```bash
noweda data.csv --html report.html --json report.json
```

JSON exports replace undefined/nonfinite statistics with `null`. Column labels
become strings; colliding labels raise an error before overwriting a file.
The in-memory report retains its original labels and numeric values.

## Development

```bash
git clone https://github.com/codewithdaniel1/NowEDA.git
cd NowEDA
pip install -e ".[test]"
python -m pytest tests/
```

NowEDA is alpha software, licensed under MIT. See the
[documentation](https://codewithdaniel1.github.io/NowEDA/) for plugins and examples.
