Metadata-Version: 2.4
Name: provium
Version: 0.4.0
Summary: Typed binary artifacts with automatic provenance
Project-URL: Homepage, https://github.com/SirDavidLudwig/provium
Project-URL: Issues, https://github.com/SirDavidLudwig/provium/issues
Project-URL: Repository, https://github.com/SirDavidLudwig/provium
Classifier: Development Status :: 2 - Pre-Alpha
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.12
Requires-Python: >=3.12
Description-Content-Type: text/markdown
Provides-Extra: visualization
Requires-Dist: graphviz>=0.20; extra == "visualization"
Provides-Extra: test
Requires-Dist: pydantic>=2.10; extra == "test"
Requires-Dist: pytest>=8.3; extra == "test"
Requires-Dist: pytest-cov>=6.0; extra == "test"
Requires-Dist: ruff>=0.16; extra == "test"

# Provium

Provium helps you build processing workflows whose results explain where they
came from. Store a result as an artifact, use that artifact as input to another
step, and save the new outputs as artifacts of their own. Provium records those
relationships automatically as your workflow runs.

Each processing step is represented by a versioned procedure. When a procedure
reads existing artifacts and creates new ones, Provium links the outputs to the
procedure and its inputs. That lineage travels with every result, including its
full upstream history, so a final artifact can be traced back through every
intermediate result and the procedures that produced them.

This keeps provenance out of your application logic: you work with inputs,
perform the computation, and write outputs inside a procedure execution. Provium
handles the dependency graph, integrity metadata, and lifecycle of those
artifacts for you.

## Features

- Typed readers and writers for application-specific binary formats
- Automatic input, output, and procedure lineage
- SHA-256 payload integrity checks
- Streaming, body-relative binary I/O
- Runtime artifact discovery through Python entry points
- Optional configuration snapshots, including Pydantic v2 models
- No required runtime dependencies

## Installation

Provium requires Python 3.12 or newer.

```bash
python -m pip install provium
```

## Quick start

Provium includes `JsonArtifact` for storing JSON-compatible values. This example
records a collection of measurements and produces a summary:

```python
from provium import JsonArtifact, Procedure

COLLECT = Procedure(name="collect", version="1")
SUMMARIZE = Procedure(name="summarize", version="1")

with COLLECT.execute():
    measurements = JsonArtifact.create("measurements.pa")
    measurements.write({"measurements": [12.5, 14.0, 13.5]})

with SUMMARIZE.execute():
    measurements = JsonArtifact.open("measurements.pa")
    payload = measurements.read()
    assert isinstance(payload, dict)
    values = payload["measurements"]
    assert isinstance(values, list)
    readings = [float(value) for value in values]

    summary = JsonArtifact.create("summary.pa")
    summary.write(
        {
            "count": len(readings),
            "minimum": min(readings),
            "maximum": max(readings),
            "average": round(sum(readings) / len(readings), 2),
        }
    )
```

`summary.pa` contains `count`, `minimum`, `maximum`, and `average`, together with
the lineage of the measurements and the procedures that collected and
summarized them. A rendered graph looks like this, with identities shortened for
readability:

```mermaid
flowchart LR
    collect(["collect 1<br/>collect-execution"])
    measurements["provium.artifact.prefab.json.JsonArtifact<br/>measurements-id"]
    summarize(["summarize 1<br/>summarize-execution"])
    summary["provium.artifact.prefab.json.JsonArtifact<br/>summary-id"]

    collect --> measurements
    measurements --> summarize
    summarize --> summary
```

When each context exits successfully, Provium closes its handles and finalizes
its output files. If a context exits with an exception, its pending outputs are
not committed. Readers and writers are bound to their execution and cannot be
used after its context exits.

`JsonArtifact` uses deterministic UTF-8 JSON encoding and supports null,
booleans, finite numbers, strings, arrays, and objects with string keys.

## Custom artifact types

For an application-specific binary format, define reader, writer, and artifact
classes. Here is the same number workflow using signed 64-bit integers:

```python
import struct

from provium import Artifact, ArtifactReader, ArtifactWriter

INTEGER = struct.Struct(">q")


class IntegerReader(ArtifactReader):
    def read(self) -> int:
        return INTEGER.unpack(self.body.read(INTEGER.size))[0]


class IntegerWriter(ArtifactWriter):
    def write(self, value: int) -> None:
        self.body.write(INTEGER.pack(value))


class IntegerArtifact(Artifact[IntegerReader, IntegerWriter]):
    reader = IntegerReader
    writer = IntegerWriter
```

Use the custom type just like the prefab JSON artifact:

```python
from provium import session

from your_package.artifacts import IntegerArtifact

SOURCE = Procedure(name="source", version="1")
ADD = Procedure(name="add", version="1")

with SOURCE.execute():
    left = IntegerArtifact.create("left.pa")
    left.write(2)

    right = IntegerArtifact.create("right.pa")
    right.write(3)

with ADD.execute():
    left = IntegerArtifact.open("left.pa")
    right = IntegerArtifact.open("right.pa")
    total = IntegerArtifact.create("sum.pa")
    total.write(left.read() + right.read())
```

The workflow has the same lineage, now with application-specific integer
artifacts. Identities are again shortened in the diagram:

```mermaid
flowchart LR
    source(["source 1<br/>source-execution"])
    left["your_package.artifacts.IntegerArtifact<br/>left-id"]
    right["your_package.artifacts.IntegerArtifact<br/>right-id"]
    add(["add 1<br/>add-execution"])
    total["your_package.artifacts.IntegerArtifact<br/>sum-id"]

    source --> left
    source --> right
    left --> add
    right --> add
    add --> total
```

Registration is optional. Without it, Provium stores the artifact class's full
path, such as `your_package.artifacts.IntegerArtifact`, as its identifier. Typed
calls such as `IntegerArtifact.open()` can read these artifacts directly.

Register the artifact when you want a stable custom identifier, aliases, or
dynamic loading through `provium.open_artifact()`:

```python
from provium import ArtifactCatalog

from .artifacts import IntegerArtifact

catalog = ArtifactCatalog()
catalog.register("example.IntegerV1", IntegerArtifact)
```

Expose that catalog from `pyproject.toml` so Provium can discover it:

```toml
[project.entry-points."provium.catalogs"]
example = "your_package.catalog:catalog"
```

## Inspecting provenance

Every reader exposes the artifact header and lineage:

```python
from provium import Procedure

from your_package.artifacts import IntegerArtifact

with session():
    artifact = IntegerArtifact.open("sum.pa")
    print(artifact.read())
    print(artifact.identity)
    print(artifact.artifact_identifier)
    print(artifact.lineage.to_json())
```

Use `provium.open_artifact()` when the concrete type should be resolved from the
identifier stored in the file rather than selected in advance.

## Reusing artifacts across procedures

A session records every artifact opened within it, even after its reader is
closed. Procedure executions inherit those recorded inputs and create a nested
session for artifacts used only by that execution:

```python
from provium import Procedure, session

PREDICT = Procedure(name="predict", version="1")

with session():
    model_reader = ModelArtifact.open("model.pa")
    model = load_model(model_reader)
    model_reader.close()

    for input_path, output_path in jobs:
        with PREDICT:
            data = DataArtifact.open(input_path)
            result = model.predict(data.read())
            ResultArtifact.create(output_path).write(result)
```

Each result depends on the shared model and its own data artifact. Nested
generic sessions similarly inherit artifacts recorded by their ancestors.

Calling a procedure is shorthand for `execute()`, including configured
executions: `with PREDICT(config=settings): ...`.

## Command-line tools

Inspect an artifact's generic metadata without loading its concrete artifact
type:

```bash
provium inspect result.pa
```

Generate Mermaid or Graphviz source for an artifact's complete lineage:

```bash
provium graph --renderer mermaid result.pa lineage.mmd
provium graph --renderer graphviz result.pa lineage.dot
```

Image output supports SVG, PNG, and PDF and defaults to the Mermaid renderer:

```bash
provium graph result.pa lineage.svg
provium graph --renderer graphviz result.pa lineage.png
```

Mermaid image rendering requires the official `mmdc` executable. Graphviz
rendering requires the optional Python package and Graphviz system package:

```bash
npm install --global @mermaid-js/mermaid-cli
python -m pip install 'provium[visualization]'
```

The output type is inferred from its extension. Library callers can use the
functions in `provium.tool` to produce Mermaid or DOT source and to receive
rendered images as bytes.

## Development

Create a virtual environment and install the project with its test dependencies:

```bash
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[test]'
```

Run the test suite:

```bash
pytest
```

This also runs Ruff linting and `ruff format --check` over `src` and `test`.
The project requires 100% statement and branch coverage for the `provium`
package.
