Metadata-Version: 2.4
Name: peekframe
Version: 0.2.0
Summary: Tiny, boring exploratory data analysis helpers for pandas DataFrames.
Project-URL: Homepage, https://github.com/haydengalletta/data-visualization
Project-URL: Repository, https://github.com/haydengalletta/data-visualization
Author-email: Hayden Galletta <hdgalletta@dons.usfca.edu>
License: MIT
License-File: LICENSE
Keywords: data-analysis,dataframe,eda,pandas,summary
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Requires-Python: >=3.9
Requires-Dist: pandas>=1.3
Description-Content-Type: text/markdown

# peekframe

Tiny, boring exploratory data analysis helpers for pandas DataFrames.

`peekframe` gives you a handful of plain functions for the first things you
always want to know about a new dataset: how big is it, what's missing, what
types are the columns, what do the numbers look like, and what are the most
common values. Every function **returns a plain object** (a `dict`, a pandas
`Series`, or a pandas `DataFrame`), never prints, and never modifies your data.

## Install

```bash
pip install peekframe
```

## Quick start

```python
import pandas as pd
import peekframe as pf

df = pd.read_csv("data.csv")

pf.summarize(df)      # high-level overview as a dict
pf.missing(df)        # missing counts + percentages per column
pf.dtypes(df)         # column -> dtype string
pf.numeric_summary(df)  # describe() for numeric columns
pf.top_values(df, "city", n=5)  # 5 most common values in a column

print(pf.report(df))  # one pretty, human-readable report of everything
```

## Functions

| Function | Returns | Does |
| --- | --- | --- |
| `summarize(df)` | `dict` | Rows, columns, column names, memory, total missing cells, duplicate rows. |
| `missing(df)` | `DataFrame` | `count` and `percent` missing per column, most-missing first. |
| `dtypes(df)` | `dict` | Maps each column name to its dtype as a string. |
| `numeric_summary(df)` | `DataFrame` | `describe()` restricted to numeric columns (empty if none). |
| `top_values(df, column, n=5)` | `Series` | The `n` most frequent values in one column. |
| `report(df)` | `str` | A formatted, human-readable report of everything above. |

## Pretty report

The functions above return raw data so you can compute with it. When you just
want to *look* at a dataset, `report(df)` renders all of it as one polished
text block (it returns the string — `print` it yourself):

```
>>> print(pf.report(df))
╭────────────────────────────────────────────────────╮
│ peekframe · DataFrame report                       │
╰────────────────────────────────────────────────────╯

▸ Overview
    rows             8
    columns          5
    memory           1.2 KB
    missing cells    3
    duplicate rows   1

▸ Columns
    name     str
    age      float64
    city     str
    spend    float64
    active   bool

▸ Missing values
    age     ██████░░░░░░░░░░░░░░░░░░   25.0%  2
    spend   ███░░░░░░░░░░░░░░░░░░░░░   12.5%  1
    name    ░░░░░░░░░░░░░░░░░░░░░░░░    0.0%  0
    city    ░░░░░░░░░░░░░░░░░░░░░░░░    0.0%  0
    active  ░░░░░░░░░░░░░░░░░░░░░░░░    0.0%  0

▸ Numeric summary
             age   spend
    count      6       7
    mean   35.50  108.29
    ...

▸ Top values
    name    Ana ×2   Ben ×1   Cy ×1
    city    SF ×4   LA ×2   NYC ×2
    active  True ×5   False ×3
```

No extra dependencies — just Unicode text.

### Example output

```python
>>> import pandas as pd, peekframe as pf
>>> df = pd.DataFrame({"a": [1, 2, 2, None], "b": ["x", "x", "y", "y"]})
>>> pf.summarize(df)
{'rows': 4, 'columns': 2, 'column_names': ['a', 'b'],
 'memory_bytes': 396, 'missing_cells': 1, 'duplicate_rows': 0}
# (memory_bytes depends on your pandas version)

>>> pf.missing(df)
   count  percent
a      1     25.0
b      0      0.0

>>> pf.top_values(df, "b", n=2)
b
x    2
y    2
Name: count, dtype: int64
```

## Design

`peekframe` is deliberately small (~250 lines of code) and boring:

- **pandas DataFrames first** — that's the only input type.
- **Plain objects only** — dicts, Series, DataFrames. No custom classes to learn.
- **One job per function** — small signatures, minimal optional arguments.
- **Pure** — nothing is printed and your DataFrame is never changed in place.

The only dependency is `pandas`.

## License

MIT
