Metadata-Version: 2.4
Name: dsvx
Version: 1.0.0
Summary: Dataset Versioning System — a git-style CLI for versioning, comparing, and training on datasets
License: MIT
Project-URL: Homepage, https://github.com/RameshwariS/Dataset-Versioning-System
Project-URL: Repository, https://github.com/RameshwariS/Dataset-Versioning-System
Project-URL: Issues, https://github.com/RameshwariS/Dataset-Versioning-System/issues
Keywords: dataset,versioning,machine-learning,cli,mlops
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Provides-Extra: ai
Requires-Dist: google-genai>=1.0; extra == "ai"
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"
Requires-Dist: pytest-cov; extra == "dev"
Dynamic: license-file

# DSVX — Dataset Versioning System

A **git-style CLI tool** for versioning, comparing, and training on datasets.
Think of it as `git` for your ML training data.

---

## Installation

```bash
# Clone / unzip the project, then install in editable mode:
pip install -e .

# Verify the command is available:
dsvx --help
```

For Gemini-powered insights, install the optional dependency:

```bash
pip install -e ".[ai]"
```

---

## Python Library API

DSVX can be used directly from Python as well as through its CLI:

```python
from dsvx import DSVXRepository

repo = DSVXRepository("dataset_repo").init()
version = repo.create_version(
    data_path="data/reviews.txt",
    name="v1",
    message="Initial import",
    config={"lowercase": True, "remove_duplicates": True},
)

print(version.name, version.metrics)
comparison = repo.compare("v1", "v2")
```

The library returns `DatasetVersion` and `ComparisonResult` objects and raises
DSVX-specific exceptions instead of printing or exiting the Python process.

---

## Quick Start

```bash
# 1. Initialise a repository in the current directory
dsvx init

# 2. Place your raw dataset in dataset_repo/raw/
cp my_data.txt dataset_repo/raw/dataset.txt

# 3. Stage it
dsvx add --data dataset_repo/raw/dataset.txt

# 4. Optionally stage a preprocessing config too
dsvx add --config config.json

# 5. Commit as a named version
dsvx commit --name v1 --message "initial raw dataset"

# 6. List all versions
dsvx list

# 7. Show details for a version
dsvx show v1

# 8. Train a classifier on it
dsvx train --version v1
```

---

## Command Reference

| Command | Description |
|---------|-------------|
| `dsvx init` | Initialise a new DSVX repository |
| `dsvx add --data FILE [--config FILE]` | Stage a dataset file and/or config |
| `dsvx commit --name NAME --message MSG` | Create a new version from staged changes |
| `dsvx status` | Show what is staged vs. committed |
| `dsvx list` | List all versions |
| `dsvx show VERSION` | Show full details of a version |
| `dsvx compare VERSION_A VERSION_B` | Diff metrics and config between two versions |
| `dsvx lineage [VERSION]` | Visualise the version lineage DAG |
| `dsvx dashboard [--port PORT]` | Open the interactive browser dashboard |
| `dsvx train --version NAME` | Train a Naive Bayes classifier on a version |
| `dsvx train --compare A B` | Train on two versions and compare reports |
| `dsvx train --all` | Train on every version and print a summary |
| `dsvx watch [--data FILE]` | Auto-version on every file save |
| `dsvx rollback VERSION` | Restore a past version's config |
| `dsvx create --data FILE --name NAME --message MSG` | *(Legacy)* Direct version creation |

### Global flags

```
--repo PATH   Path to the dataset repository (default: dataset_repo)
--version     Print DSVX version and exit
```

---

## Git-style Workflow

```
dsvx init
dsvx add --data dataset_repo/raw/dataset.txt
dsvx commit --name v1-raw --message "raw data"

# edit dataset or config...

dsvx add --data dataset_repo/raw/dataset.txt --config config.json
dsvx commit --name v2-clean --message "removed duplicates and stopwords"

dsvx compare v1-raw v2-clean
dsvx lineage
```

---

## Preprocessing Config

Edit `config.json` to toggle preprocessing steps:

```json
{
  "lowercase":          true,
  "remove_punctuation": true,
  "remove_duplicates":  true,
  "strip_whitespace":   true,
  "tokenize":           false,
  "remove_stopwords":   false
}
```

Stage it with `dsvx add --config config.json` before the next commit.

---

## Training

Dataset rows must follow the format `<text>|<label>`:

```
Deep learning is powerful.|positive
This data is noisy and bad.|negative
```

Then:

```bash
dsvx train --version v1-raw
dsvx train --compare v1-raw v2-clean
dsvx train --all
```

---

## Debugging

Set `DSVX_DEBUG=1` to get full Python tracebacks on errors:

```bash
DSVX_DEBUG=1 dsvx commit --name v1 --message "test"
```

---

## AI Insights

Create a Gemini API key in Google AI Studio, then provide it for an insights
command by one of these methods (the key is never saved by DSVX):

```bash
# Recommended for local development
export GEMINI_API_KEY="your-api-key"
dsvx insights version v1

# Prompt without showing the key in the terminal or shell history
dsvx insights --prompt-api-key version v1

# One command only; avoid this on shared machines because shells can retain history
dsvx insights --api-key "your-api-key" version v1
```

Choose a different available Gemini model when needed:

```bash
dsvx insights --model gemini-3.8-flash version v1
```

DSVX's insight behavior is adapted to this project through a dataset-versioning
system prompt and the repository's actual metrics, preprocessing configuration,
lineage, and training reports. Gemini API/AI Studio does not currently support
direct model fine-tuning; use Vertex AI supervised tuning if genuine training of
a hosted model is required.

---

## Project Structure

```
dsvx/
├── __init__.py       Package metadata and version
├── cli.py            Entry point — main() registered as `dsvx` console script
├── dsvx.py            Core commands (init, add, commit, list, show, compare, …)
├── train.py          Training and evaluation logic
├── watch.py          File-watcher for auto-versioning
├── storage.py        Filesystem storage layer
├── versioning.py     Version creation, lineage, DAG building
├── preprocessing.py  Text preprocessing pipeline
├── metrics.py        Dataset metrics calculation
└── model.py          TF-IDF vectoriser + Naive Bayes classifier
pyproject.toml        Build metadata and console_scripts entry point
setup.py              Legacy editable-install shim
README.md             This file
```
