Metadata-Version: 2.4
Name: himanshumehraanalysis
Version: 0.1.0
Summary: A simple, one-line EDA library for pandas DataFrames
Author: Himanshu Mehra
License-Expression: MIT
Project-URL: Homepage, https://pypi.org/project/himanshumehraanalysis/
Project-URL: Issues, https://pypi.org/project/himanshumehraanalysis/
Keywords: eda,pandas,data analysis,data science,databricks
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Operating System :: OS Independent
Classifier: Intended Audience :: Developers
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pandas>=1.3
Requires-Dist: numpy>=1.20
Provides-Extra: plots
Requires-Dist: matplotlib>=3.4; extra == "plots"
Provides-Extra: dev
Requires-Dist: pytest; extra == "dev"
Requires-Dist: build; extra == "dev"
Requires-Dist: twine; extra == "dev"
Dynamic: license-file

# himanshumehraanalysis

A simple, one-line exploratory data analysis (EDA) library for pandas DataFrames.

It detects column types, missing values, duplicates, skewness, IQR outliers, correlations, and numbers stored as text — then returns the results as a dictionary.

## Installation

```bash
pip install himanshumehraanalysis
```

Optional charts:

```bash
pip install himanshumehraanalysis[plots]
```

For local development from the source tree:

```bash
pip install -e ".[dev,plots]"
```

## Basic usage

```python
import pandas as pd
import himanshumehraanalysis as hma

df = pd.read_csv("data.csv")

result = hma.analyze(df)
```

`analyze()` prints a readable report (unless `verbose=False`) and **returns a dictionary** with keys such as `shape`, `columns`, `dtypes`, `null_counts`, `duplicate_rows`, `column_types`, `numbers_stored_as_text`, `skewness`, `outliers`, and `high_correlations`.

```python
import pandas as pd
import himanshumehraanalysis as hma

df = pd.read_csv("data.csv")

result = hma.analyze(df)
types = hma.detect_column_types(df)
table = hma.summary_table(df)
```

- `hma.analyze(df)` — full EDA report (print + dict)
- `hma.detect_column_types(df)` — classify columns as numerical, categorical, datetime, boolean, or text
- `hma.summary_table(df)` — per-column overview as a pandas DataFrame

`df` must be a **pandas DataFrame**. Spark DataFrames are not converted automatically.

## Features

- Dataset overview (shape, columns, memory)
- Column type detection
- Missing-value analysis
- Duplicate-row detection
- Numerical statistics and skewness
- Categorical value counts
- IQR outlier detection
- High-correlation detection
- Numbers stored as text (for example `"23.5"`)

## Databricks

In a Databricks notebook:

```python
%pip install himanshumehraanalysis
```

Then:

```python
import himanshumehraanalysis as hma

result = hma.analyze(df)
```

`df` must be a pandas DataFrame. If you have a Spark DataFrame, convert it first:

```python
pdf = spark_df.toPandas()
result = hma.analyze(pdf)
```

Pin a version if you need a specific release:

```python
%pip install himanshumehraanalysis==0.1.0
```

Notebook `%pip` installs into the current notebook session. Restart the Python session if Databricks prompts you to, so the new package is picked up cleanly. Cluster-wide library installs typically require a cluster restart.

## License

MIT
