Metadata-Version: 2.4
Name: modelzone-sdk
Version: 4.0.0
Summary: Modelzone SDK - experiment tracking and AzureML integration
License-Expression: Apache-2.0
Author: Team Enigma
Author-email: enigma@energinet.dk
Requires-Python: >=3.13
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3.13
Provides-Extra: azureml
Requires-Dist: azure-ai-ml (>=1.32.0) ; extra == "azureml"
Requires-Dist: azure-identity (>=1.25.3) ; extra == "azureml"
Requires-Dist: azure-storage-blob (>=12.28.0) ; extra == "azureml"
Requires-Dist: mlflow (>=3) ; extra == "azureml"
Requires-Dist: requests (>=2.32.0) ; extra == "azureml"
Description-Content-Type: text/markdown

# ModelZone SDK

A Python package for local experiment tracking, model artifacts, and optional
Azure ML integration. The package consists of:

- A **Python SDK** for tracking project-owned training code and saving or loading model artifacts.
- A **CLI** for uploading runs, registering model versions, and fetching models from Azure ML.

## Installation

Requires Python 3.13 or newer. The base package has no third-party dependencies
and covers local experiment tracking and artifact loading:

```bash
pip install modelzone-sdk
```

Cli commands upload, register, and fetch operations additionally need the `azureml` extra:

```bash
pip install 'modelzone-sdk[azureml]'
```

Install your model's own training, plotting, and inference dependencies in its
project. Azure operations use
[DefaultAzureCredential](https://learn.microsoft.com/en-us/python/api/azure-identity/azure.identity.defaultazurecredential);
configure an identity with access to the target workspace before using them.

## Overview

Taking a look at the various Python machine learning frameworks out there, you will come across many different workflows used for training and prediction. ModelZone SDK uses the following definition:

> A **model project** is a Python package with a **training function**, which is responsible for (1) **tracking** metrics and graphs used for comparing different model experiments and (2) generating model **artifacts** (such as model parameters and hyper-parameters) that need to be shipped with the package.

ModelZone SDK assumes your machine learning workflow looks more or less like the following:

1. Experimentation by training different models (and hyper-parameters).
2. Release of the best suited model for production usage.
3. Using the released model for inference in operations.

---

## Quick guide

We have made an example of how a training and inference project should look like:

- Training project: [examples/training_example](./examples/training_example/)
- Inference project: [examples/predict_example](./examples/predict_example/)
- Script running the full workflow: [examples/happy_path.sh](./examples/happy_path.sh)

From the repository root, in a dedicated Python environment, install the SDK and
training example and run it locally:

```bash
python -m pip install -e .
cd examples/training_example
python -m pip install -e .
train
```

The project's `train` console script uses `mz.default_parser()` and
`mz.run(__file__, ...)`. `python -m example_forecaster.experiment` remains available.
For uploads, install `python -m pip install -e '.[train]'` in the example directory
and run `train --azureml`. The example's `train` extra installs the SDK's
`azureml` extra; these are different project dependency groups.

The full workflow script recreates or cleans its named Python environments,
removes previous example outputs, and uploads and registers a model in the
configured Azure ML workspace. Use it only when those side effects are intended.

---

## Detailed guide

### Project configuration

Local training uses the nearest `pyproject.toml` above the training script to
locate the project root, but does not require a `[tool.modelzone]` section.

For uploads and registration, a **training project** needs:

```toml
[tool.modelzone]
subscription_id = "<azure-subscription-id>"
resource_group = "<resource-group>"
workspace_name = "<azure-ml-workspace>"
experiment_name = "my_model"
```

An **inference project** needs the workspace settings plus the model to fetch:

```toml
[tool.modelzone]
subscription_id = "<azure-subscription-id>"
resource_group = "<resource-group>"
workspace_name = "<azure-ml-workspace>"
model_name = "my_model"
model_version = 32
```

The CLI reads configuration from the nearest `pyproject.toml` at or above the
current directory. Run `modelzone fetch` from the inference project root: it
places `.model/` in the current directory, replacing any existing contents.

### Model project structure

**Training entry point**

A training project owns its entry point and command-line arguments. For a
no-argument training function, `mz.default_cli(train)` parses `--azureml` and
calls the function inside a tracked run. This minimal example logs a metric;
replace the function body with your training code.

In `mypackage/experiment.py`:

```python
import modelzone as mz

def train():
    mz.tracker.log_metric("MAE", 1.23)

def main():
    mz.default_cli(train)
```

Expose a uniform `train` command in the project's `pyproject.toml`:

```toml
[project.scripts]
train = "mypackage.experiment:main"
```

Install or reinstall the project after adding the console script:

```bash
python -m pip install -e .
train
```

A project needing custom options uses `mz.default_parser()` and
`mz.run(__file__, ...)` directly instead:

```python
import modelzone as mz

def train(quick: bool):
    mz.tracker.log_tag("quick", str(quick))

def main():
    parser = mz.default_parser()
    parser.add_argument("--quick", action="store_true")
    args = parser.parse_args()

    with mz.run(__file__, azureml=args.azureml):
        train(quick=args.quick)
```

The default parser supplies only `--azureml` (plus argparse's `--help`).
`--quick` is project-defined; implement any reduced-data behavior in `train`.

The nearest `pyproject.toml` above the script defines the project root for run
output, code snapshots, and upload configuration, independent of the working
directory. `mz.default_cli` uses the training function's source file automatically.

Code snapshots use Git's ignore rules. Keep `.runs/` and your virtual environment
in `.gitignore` so repeated runs do not snapshot previous runs or dependencies.

**Tracking**

`mz.tracker` forwards calls to the currently active local tracker. Calling it
outside a tracked run raises `RuntimeError`. `mz.run(__file__)` manages the
tracker and code snapshot; `mz.LocalTracker(output_dir)` is available when you
want only tracking, without project discovery, snapshots, or automatic upload.

See [Logging](#logging) for metrics, files, and messages.

**Model artifacts**

Inherit from `mz.Artifact` to save and load a model using pickle. Define the class
in an importable module such as `mypackage/forecaster.py`, not in `__main__`, so
inference code can import the same class when unpickling:

```python
import modelzone as mz

class Forecaster(mz.Artifact):
    def __init__(self, model):
        self.model = model

    def predict(self, features):
        return self.model.predict(features)
```

Pass an instance to `log_artifact` inside your training function. This example
requires scikit-learn:

```python
from sklearn.linear_model import LinearRegression
import modelzone as mz

from mypackage.forecaster import Forecaster

def train():
    features = [[0.0], [1.0], [2.0]]
    targets = [0.0, 2.0, 4.0]
    model = LinearRegression().fit(features, targets)
    mz.tracker.log_artifact("forecaster", Forecaster(model))

def main():
    mz.default_cli(train)
```

The saved file is `artifacts/forecaster.pkl` under the run directory.

In the inference project, fetch the registered version and install its code and
dependencies:

```bash
modelzone fetch
python -m pip install .model/package
```

`modelzone fetch` puts the code in `.model/package` and artifacts in
`.model/artifacts`. For an inference script located at that project root:

```python
from pathlib import Path

from mypackage.forecaster import Forecaster

def get_model_artifact() -> Forecaster:
    artifacts_dir = Path(__file__).parent / ".model/artifacts"
    return Forecaster.load("forecaster", artifacts_dir)
```

Only load artifacts from trusted sources: unpickling can execute code. The
model's Python package and required libraries must be installed when loading it.

### Logging

Call the tracking methods inside `mz.run(__file__)` or a
`mz.LocalTracker(output_dir)` context.

**Metrics**

Log a scalar metric:

```python
mz.tracker.log_metric("MAE", 1.23)
```

Use dimensions to record the same metric in different contexts:

```python
mz.tracker.log_metric("MAE", 1.23, dimensions={"country": "DK"})
mz.tracker.log_metric("MAE", 4.56, dimensions={"country": "SE"})
```

All metrics are appended to `results/metrics.jsonl` and the run log. When the run
ends, dimensioned metrics are also written to `results/metrics/<name>.txt`:

```text
country  MAE
-------  ----
DK       1.23
SE       4.56
```

Use the same dimension keys for all entries of a given metric name.

**Figures and files**

Both figures and files handled in the same way just as files.

```python
mz.tracker.log_file("notes.txt", "Finished evaluating the model")
mz.tracker.log_file("payload.bin", b"example data")
mz.tracker.log_file(fig.to_html())
```

**Tags and messages**

```python
mz.tracker.log_tag("model_type", "LinearRegression")
mz.tracker.log_message("Finished fitting the model")
```

Tags are stored in `results/tags.jsonl`. Messages are printed to stdout and
appended to `results/logs.log`. On upload, tags appear in Azure ML's _Overview_,
and the saved files and logs appear under _Output + logs_.

### CLI commands

| Command                       | Description                                           |
| ----------------------------- | ----------------------------------------------------- |
| `modelzone upload <run_dir>`  | Upload an existing local run directory to Azure ML.   |
| `modelzone register <run_id>` | Register an Azure ML run as a new model version.      |
| `modelzone fetch`             | Download the configured model version into `.model/`. |

The SDK CLI does not have a training command. The project's `train` script is
separate and uses the entry-point helpers above.

The full workflow is: run the training entry point with `--azureml` → `register`
(in the training project) → `fetch` (in the inference project).

### How training and AzureML tracking work

`mz.run(__file__)` tracks your code locally in `.runs/<timestamp>` and stores a
code snapshot under `package/`. At exit, the directory is renamed to
`.runs/<timestamp>__FINISHED` or `.runs/<timestamp>__FAILED`.

A completed run can contain:

```text
.runs/<timestamp>__FINISHED/
    package/
    artifacts/forecaster.pkl
    results/
        run.json
        tags.jsonl
        metrics.jsonl
        metrics/MAE.txt
        files/forecast_comparison.html
        logs.log
```

Files are created as the corresponding data is logged.

`azureml=True` performs the _same_ local run and then uploads it to Azure ML
once it finishes:

```bash
train --azureml
```

Upload happens only after the tracked block completes successfully. A run that
raises is retained locally as `__FAILED` and is not uploaded. Manual upload also
rejects runs whose saved status is not `FINISHED`.

Local runtime is measured with a monotonic timer and stored in `results/run.json`
alongside the run status:

```json
{ "duration_seconds": 0.425, "status": "FINISHED" }
```

The SDK does not store or override start/end timestamps. Azure ML assigns its
own timestamps during upload. Use the `duration_seconds` metric for the local
runtime; the same value is kept in `output/results/run.json` under _Output + logs_.

Upload prints a portal URL of the form
`https://ml.azure.com/runs/<run_id>?wsid=...`; use the ID after `/runs/` with
`modelzone register <run_id>`.

### Uploading a local run

If you ran the training entry point without `--azureml` and later decide a local
run is worth keeping, you can persist it to Azure ML without re-running the
training:

```bash
modelzone upload .runs/2026-07-16T12:11:55__FINISHED
```

This creates a new Azure ML run in the configured experiment, replays the run's
tags and metrics into Azure ML (so they show up in the _Overview_ and _Metrics_
sections), and uploads the full run directory (code snapshot, artifacts, figures,
logs) to the run's _Output + logs_. It is exactly what `mz.run(__file__, azureml=True)`
does after a run finishes.

Upload requires `results/run.json` with `status` and `duration_seconds`, and the
status must be `FINISHED`. Older runs without a stored duration cannot be
uploaded as-is. Metrics are replayed from `results/metrics.jsonl`; plain text
tables and HTML files are uploaded as files, not parsed for metrics.

