Metadata-Version: 2.5
Name: rulebricks-spark
Version: 0.2.0
Summary: A Rulebricks decision-based data processing client for Apache Spark.
Project-URL: Homepage, https://rulebricks.com
Project-URL: Documentation, https://github.com/rulebricks/spark#readme
Project-URL: Repository, https://github.com/rulebricks/spark
Project-URL: Issues, https://github.com/rulebricks/spark/issues
Author-email: Rulebricks <support@rulebricks.com>
License: MIT
License-File: LICENSE
Keywords: business-rules,databricks,decision-tables,pyspark,rulebricks,rules-engine,spark
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Typing :: Typed
Requires-Python: >=3.9
Requires-Dist: httpx<1,>=0.28
Requires-Dist: pandas<3,>=2
Requires-Dist: pyarrow<26,>=18
Requires-Dist: pyspark<3.6,>=3.5
Requires-Dist: rulebricks<3,>=2.6.8
Requires-Dist: setuptools>=68; python_version >= '3.12'
Provides-Extra: dev
Requires-Dist: build>=1.0; extra == 'dev'
Requires-Dist: mypy>=1.0; extra == 'dev'
Requires-Dist: nbformat>=5.10; extra == 'dev'
Requires-Dist: pytest-cov>=4.0; extra == 'dev'
Requires-Dist: pytest>=7.0; extra == 'dev'
Requires-Dist: ruff>=0.4; extra == 'dev'
Requires-Dist: twine>=7.0; (python_version >= '3.10') and extra == 'dev'
Requires-Dist: validate-pyproject>=0.23; extra == 'dev'
Description-Content-Type: text/markdown

![Rulebricks Spark](/assets/cover.png)

# Rulebricks Spark

[![PyPI](https://img.shields.io/pypi/v/rulebricks-spark.svg)](https://pypi.org/project/rulebricks-spark/)

Apply versioned Rulebricks rules and flows to PySpark DataFrames. The library
sends records in bounded requests and keeps each decision, error, and execution
ID alongside its source data, so the pipeline can scale without losing the ability
to explain a run.

```bash
pip install rulebricks-spark
```

```python
import os

from rulebricks_spark import apply_flow

results = apply_flow(
    source,
    "order-routing-flow",
    version="1",
    api_key=os.environ["RULEBRICKS_API_KEY"],
    base_url="https://rulebricks.acme.com/api/v1",
    correlation_id_path="request_id",
    execution_partitions=4,
)
results.write.format("delta").mode("errorifexists").save(result_path)
saved = spark.read.format("delta").load(result_path)
```

Flow results remain JSON by default, ready for your own Spark expressions.
`apply_rule()` uses the same execution model and infers a typed result from the
selected rule version. Context operations provide optional input staging and
ordered stateful processing, while Structured Streaming uses Spark's ordinary
`foreachBatch` interface.

Start with the [documentation](docs/index.md) or the [notebook walkthroughs](examples/README.md).
The guides explain how the concepts fit together; the reference covers parameters,
result fields, failure handling, and recovery. Notebook settings are defined in
the notebooks, with only credentials kept in a secret manager or environment.

The package targets classic PySpark 3.5 and requires Python 3.9+ and PyArrow 18+.
Use a Python version supported by your Spark runtime.
For a documentation preview, follow [the Sphinx build instructions](docs/contributing.md).
