Metadata-Version: 2.4
Name: satva
Version: 0.1.0
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Programming Language :: Rust
Classifier: Topic :: Software Development :: Libraries
Summary: Run Satva data pipelines from Python
Keywords: etl,pipeline,csv,parquet,json
Home-Page: https://github.com/harbhim/satvars
Author-email: Hardik Bhimani <hardikbhimani@rocketmail.com>
License-Expression: MIT
Requires-Python: >=3.9
Description-Content-Type: text/markdown; charset=UTF-8; variant=GFM
Project-URL: Homepage, https://github.com/harbhim/satvars
Project-URL: Issues, https://github.com/harbhim/satvars/issues
Project-URL: Repository, https://github.com/harbhim/satvars

# satva

[Satva](https://github.com/harbhim/satvars) is a data pipeline engine. This package runs the same YAML pipelines as the `satva` command-line tool.

Requires Python 3.9 or newer. Licensed under the MIT License.

## Install

```bash
pip install satva
```

## Usage

```python
import satva

summary = satva.run("pipeline.yaml")
print(summary["processed"], summary["succeeded"], summary["skipped"], summary["failed"])

# Raises RuntimeError on the first record failure.
satva.run("pipeline.yaml", stop_on_error=True)
```

`satva.run` returns a dict with `processed`, `succeeded`, `skipped`, `failed`, and `logs`. `logs` is a list of strings. `stop_on_error` defaults to `False`: failed records are counted and the call returns. File paths inside the YAML config are relative to the process working directory.

```yaml
source:
  type: json
  path: input.jsonl

sink:
  type: json
  path: output.jsonl

stages:
  - type: filter
    expression: 'active == true && salary >= 70000'
  - type: set_field
    field: bonus
    expression: 'salary * 0.15'
```

Sources and sinks: CSV, TSV, JSONL (`json`), JSON array (`json_array`), Parquet, and Excel. Pipeline and expression details are in the [repository docs](https://github.com/harbhim/satvars/tree/master/docs).

