Metadata-Version: 2.4
Name: makeprov
Version: 0.8.0
Summary: A PROV/JSON-LD provenance tracking library for Python scripts, with optional Snakemake and ReproZip bridges
Author-email: Benno Kruit <b.b.kruit@amsterdamumc.nl>
License-Expression: MIT
Project-URL: Homepage, https://github.com/bennokr/makeprov
Project-URL: Documentation, https://bennokr.github.io/makeprov
Project-URL: Issue Tracker, https://github.com/bennokr/makeprov/issues
Keywords: provenance,prov,rdf,json-ld,snakemake,reprozip
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Operating System :: OS Independent
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: packaging>=21.3
Requires-Dist: parse>=1.20
Provides-Extra: rdf
Requires-Dist: rdflib>=6.0; extra == "rdf"
Requires-Dist: pyshacl>=0.20; extra == "rdf"
Provides-Extra: snakemake
Requires-Dist: snakemake; extra == "snakemake"
Provides-Extra: cli
Requires-Dist: defopt>=6; extra == "cli"
Provides-Extra: dev
Requires-Dist: pytest; extra == "dev"
Requires-Dist: makeprov[cli,rdf]; extra == "dev"
Provides-Extra: docs
Requires-Dist: sphinx>=7; extra == "docs"
Requires-Dist: myst-parser[linkify]; extra == "docs"
Requires-Dist: sphinx-rtd-theme; extra == "docs"
Requires-Dist: sphinx-autodoc-typehints; extra == "docs"
Requires-Dist: makeprov[cli]; extra == "docs"
Dynamic: license-file

# makeprov: Pythonic Provenance Tracking

`makeprov` is a small library for recording W3C PROV/JSON-LD provenance
around Python functions that read and write files: which inputs produced
which outputs, when, with what code and environment. A decorator wraps a
function, tracks the files it declares as inputs/outputs, and writes a
provenance record after each call. A minimal `make`-style dependency
resolver and optional bridges from Snakemake and ReproZip are included, but
the core contract of the library is the provenance record — not workflow
orchestration, which tools like Snakemake already do well.

## Features

- Decorator-based rules that infer dependencies from `InPath`/`OutPath`
  parameters and write a PROV/JSON-LD record after every call.
- A clean `Plan → Run → Artifact` model: the script at a commit is a
  `prov:Plan`, the runtime and the user are agents, and the two are tied
  together by `prov:qualifiedAssociation`/`prov:hadPlan`.
- `ArtifactRef` lets a run cite external entities — a dataset IRI, an
  object-store key, a model checkpoint — without makeprov copying their metadata.
- Provenance write failures are fatal by default (`ProvenanceConfig(strict=True)`),
  so a rule can't silently "succeed" with no record of what it did.
- Resolve templated targets (``results/{sample}.txt``) via ``parse``-style patterns,
  and a small dependency resolver (`build`/`build_all`) for chaining rules.
- Serialize provenance as JSON-LD, or as RDF/TriG when `rdflib` is installed
  (`pip install "makeprov[rdf]"`).
- Optional Snakemake bridge that turns `--d3dag` and `--detailed-summary`
  output into PROV JSON-LD artifacts ready for inclusion in Snakemake HTML reports.
- Converts an existing ReproZip trace (`reprounzip graph --json`) into the
  same PROV model via `makeprov-reprozip`.

## Installation

You can install the module directly from PyPI:

```bash
pip install makeprov
```

Optional extras add RDF/TriG export, CLI subcommand support, or the Snakemake bridge:

```bash
pip install "makeprov[rdf]"        # rdflib + pyshacl for RDF/TriG export
pip install "makeprov[cli]"        # defopt, needed for makeprov.main()
pip install "makeprov[snakemake]"  # the makeprov-snakemake bridge
```

## Usage

Here's an example of how to use this package in your Python scripts:

```python
from makeprov import rule, InPath, OutPath, build

@rule()
def process_data(
    sample: int | None = None,
    input_file: InPath = InPath('data/{sample:d}.txt'),
    output_file: OutPath = OutPath('results/{sample:d}.txt')
):
    with input_file.open('r') as infile, output_file.open('w') as outfile:
        data = infile.read()
        outfile.write(data.upper())

if __name__ == '__main__':
    # Build a specific templated target and its prerequisites
    from makeprov import build
    build('results/1.txt')

    # Or expose rules via a command line interface
    import defopt
    defopt.run(process_data)
```

You can execute `examples/example.py` via the CLI like so:

```bash
python examples/example.py build-all

# Or set configuration through the CLI
python examples/example.py build-all --conf='{"base_iri": "http://mybaseiri.org/", "prov_dir": "my_prov_directory"}' --force --input_file input.txt --output_file final_output.txt

# Or set configuration through a TOML file
python examples/example.py build-all -c @my_config.toml

# Inspect dependency resolution without executing rules
python examples/example.py --explain results/1.txt
python examples/example.py --to-dot results/1.txt
```

For directory outputs, nested/merged provenance, streaming, opt-in rule
metadata, Snakemake integration, and other advanced topics, see the full
[usage guide](https://bennokr.github.io/makeprov/usage.html) and
[configuration reference](https://bennokr.github.io/makeprov/configuration.html).

## The provenance model

makeprov keeps PROV's distinction between the *plan* (the recipe) and the
*agent* (whoever carried it out):

```text
run.py @ git SHA          a prov:Plan, schema:SoftwareSourceCode
CPython 3.11              a prov:Agent, prov:SoftwareAgent
you (opt-in)              a prov:Agent, schema:Person

train-20260910T…-c2f6dc7b a prov:Activity
    prov:used                  dataset-X, the Python environment
    prov:wasAssociatedWith     runtime, person
    prov:qualifiedAssociation  [ prov:agent person ; prov:hadPlan run.py ]

results/model.txt         a prov:Entity
    prov:wasGeneratedBy        train-20260910T…-c2f6dc7b
    dct:identifier             sha256:…
```

Keeping the plan and the agent apart is what makes the graph mappable onto
[Workflow Run RO-Crate](https://www.researchobject.org/workflow-run-crate/),
whose `instrument` (the software that was run) and `agent` (a Person or
Organization) are separate slots:

| makeprov / PROV-O        | Process Run Crate |
| ------------------------ | ----------------- |
| `prov:Plan`              | `instrument`      |
| `prov:Activity`          | `CreateAction`    |
| `prov:used`              | `object`          |
| `prov:wasGeneratedBy`    | `result`          |
| `schema:Person` agent    | `agent`           |
| `startedAtTime`/`endedAtTime` | `startTime`/`endTime` |

The `schema:Person` agent is **off by default**: provenance documents are
routinely committed and published, and a name and email address are personal
data you should choose to publish rather than emit by accident. Turn it on with
`ProvenanceConfig(record_user=True)`, or `--record-user` on the Snakemake
bridge. Without it, the qualified association names the runtime as the
responsible agent.

Note that WRROC is Schema.org-native and defines no normative PROV-O mapping;
the table above is a practical alignment, not an OWL equivalence. makeprov's
own vocabulary stays `prov:`/`schema:` — RO-Crate and OpenLineage are intended
as adapters over this model rather than changes to it.

### Planned structure vs. observed execution

By default a document is purely *retrospective*: it records the activities that
ran. A rule that was already up to date contributes nothing, because asserting
an execution that did not happen would be worse than saying nothing.

Set `emit_plan_graph = true` (or pass `--plan-graph`) to additionally emit the
*prospective* structure — each rule as a `prov:Plan` in its own right, linked to
the plans it depends on by `dct:requires`:

```text
run.py#rule-transform  a prov:Plan
    dct:requires   run.py#rule-extract
    dct:source     run.py
```

The activity's `prov:hadPlan` then points at the specific rule rather than the
whole script. The Snakemake bridge does the same, collapsing the job DAG's
edges to rule-level `dct:requires` edges.

### Forge profiles

When no `base_iri` is set, makeprov derives one from the git remote. The
supported hosts are declared in [`forges.toml`](src/makeprov/forges.toml) —
GitHub, GitLab, Bitbucket, Forgejo/Gitea/Codeberg and SourceHut — each giving
the permalink layout for that host:

```toml
[[forge]]
name = "gitlab"
hosts = ["gitlab.com"]
blob = "{repo}/-/blob/{revision}/"
```

Point `forge_profiles` at your own TOML file to add self-hosted instances;
entries there are matched first, so they can also override a built-in host.
SSH and scp-style remotes (`git@host:owner/repo.git`) are understood, and any
credentials embedded in a remote URL are stripped before it reaches a document.

## Referencing things that aren't local files

`ArtifactRef` describes an entity a run consumed or produced. It is either
*local* (makeprov stats and hashes it) or *external* (makeprov records the IRI
and never touches the filesystem):

```python
from makeprov import ArtifactRef, OutPath, rule

@rule()
def train(
    dataset: ArtifactRef = ArtifactRef.external(
        "https://example.org/datasets/train-v17",
        types=("prov:Entity", "schema:Dataset"),
        digest="sha256:...",
    ),
    model: OutPath = OutPath("models/m.pkl"),
):
    ...
```

The external object keeps its own detailed metadata; makeprov only records that
this run used its stable IRI. External refs take no part in staleness checks,
since they have no local mtime to compare.

## More examples

- [`examples/complex_example.py`](examples/complex_example.py) — a CSV-to-RDF
  workflow that aggregates multiple inputs and embeds an `rdflib.Graph` result
  directly into the provenance dataset.
- [`examples/merge_outdir_example.py`](examples/merge_outdir_example.py) —
  bundling nested provenance and directory outputs with `merge=True` and
  `OutDir`/`InDir`.
- [`examples/context_demo_example.py`](examples/context_demo_example.py) —
  pinning a base IRI and isolating rules and buffers in their own `Session`.

Walkthroughs of these, plus streaming/recovery mode and opt-in rule metadata,
are in the [usage guide](https://bennokr.github.io/makeprov/usage.html).

### Snakemake workflows

Install the `snakemake` extra (`pip install "makeprov[snakemake]"`) to get the
`makeprov-snakemake` command, which shells out to Snakemake and converts its
job DAG and `--detailed-summary` metadata into a PROV document. See the
[Snakemake integration guide](https://bennokr.github.io/makeprov/snakemake.html)
for the CLI flags and an example `report()` wiring.

```bash
makeprov-snakemake --prov-path prov/snakemake -- --snakefile Snakefile --nolock
```

### ReproZip traces

`makeprov-reprozip` converts an existing `reprounzip graph --json` file into
PROV/JSON-LD or TriG — observed file accesses and process relationships only,
with no ReproZip runtime dependency. See the
[ReproZip conversion guide](https://bennokr.github.io/makeprov/reprozip.html).

```bash
makeprov-reprozip graph.json --output prov/command --base-iri https://example.org/my-experiment/
```

### Configuration

You can customize the provenance tracking with the following options:

 - `base_iri` (str): Base IRI for new resources
 - `prov_dir` (str): Directory for writing PROV `.json-ld` or `.trig` files
 - `force` (bool): Force running of dependencies
 - `dry_run` (bool): Only check workflow, don't run anything
 - `strict` (bool, default `True`): Raise `makeprov.ProvenanceWriteError` if a
   rule's provenance record fails to write, instead of only logging a
   warning. A rule that produces a result but no provenance record is
   treated as a failure by default; set `strict=False` to opt out per-rule
   or globally.
 - `run_id` (str | None): Adopt an externally supplied run identity, such as a
   CI job id. When unset, each run gets a fresh unique id.
 - `record_user` (bool, default `False`): Record the invoking user, taken from
   `git config user.name`/`user.email`, as a `schema:Person` agent. Off by
   default so personal data isn't published by accident.
   CLI: `--record-user`.
 - `emit_plan_graph` (bool, default `False`): Also emit prospective structure —
   the rule dependency graph as `prov:Plan` nodes linked by `dct:requires`. Off
   by default, so a document describes only what actually ran. CLI:
   `--plan-graph`.
 - `forge_profiles` (str | None): TOML file of extra forge profiles, for
   self-hosted git hosts. CLI: `--forge-profiles`.

See [`CHANGELOG.md`](CHANGELOG.md) for what changed in past releases, including
the 0.7 provenance-model rework.

### Scoped spans and cached downloads

Use `makeprov.span(label, prov_path=None, frame=None, context=None)` as a
context manager or decorator to bracket a chunk of work in its own provenance
buffer, and `makeprov.CachedDownload` to lazily fetch and record provenance for
a remote resource cached locally. See the
[usage guide](https://bennokr.github.io/makeprov/usage.html) for examples of
both.

## Documentation

Build the Sphinx docs locally (including autosummary API stubs) with the docs
extra so that the CLI dependencies needed for imports are available:

```bash
pip install -e ".[docs]"
python docs/build.py
```

## Contributing

Contributions are welcome! Please open an issue or submit a pull request.

## License

This project is licensed under the MIT License - see the [LICENSE](LICENSE) file for details.
