Metadata-Version: 2.5
Name: jointfm-client
Version: 0.9.0
Summary: A Python SDK for the JointFM REST API
Project-URL: Repository, https://github.com/datarobot/joint-client-python
Author-email: Stefan Hackmann <stefan.hackmann@datarobot.com>
License-Expression: Apache-2.0
License-File: AUTHORS
License-File: LICENSE
Keywords: datarobot,forecasting,jointfm,sdk
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Typing :: Typed
Requires-Python: >=3.11
Requires-Dist: pydantic>=2.13.4
Requires-Dist: python-dotenv>=1.2.2
Requires-Dist: pyyaml>=6.0.3
Requires-Dist: requests>=2.33.1
Requires-Dist: tenacity>=9.0.0
Provides-Extra: notebooks
Requires-Dist: pandas>=3.0.2; extra == 'notebooks'
Description-Content-Type: text/markdown

# JointFM Python SDK

`jointfm-client` is the Python SDK package for callers of the JointFM REST API. The import namespace is `jointfm_client`, and the first supported Python version is Python 3.11.

The SDK targets the DataRobot-hosted unstructured prediction route and the same direct local service contract used by the JointFM inference container. Public code for the contract lives in `jointfm_client.contract`; the README mirrors it for package users.

## Package Contract

- Distribution package: `jointfm-client`
- Import namespace: `jointfm_client`
- Supported Python: `>=3.11`
- Current SDK package version: `0.9.0`
- Current JointFM service schema: `schema_version="v5"`

The public API shape is a synchronous low-level `JointFMClient` with `health()`, `health_instances()`, and `predict(payload)` methods plus high-level `forecast(...)`, `forecast_mean(...)`, `forecast_samples(...)`, and `forecast_quantiles(...)` helpers. The SDK is not a proxy service; callers use it as a local Python library that talks to the hosted or local JointFM endpoint.

SDK package versions are standard Python distribution versions: `[project].version` in `pyproject.toml` is the single declared value, and `jointfm_client.__version__` reports it back from the installed distribution metadata. JointFM `schema_version`, `image_version`, `model_version`, and `checkpoint_version` are service compatibility identifiers carried in configuration, health metadata, requests, and responses. They are not SDK package versions, and changing a deployment pin does not by itself require changing the SDK package version.

See [docs/api-reference.md](docs/api-reference.md) for the checked-in API reference covering public classes, functions, exceptions, environment variables, and payload fields.

## Service Contract

The DataRobot-hosted prediction URL is built from the DataRobot API v2 endpoint and deployment ID:

```python
from urllib.parse import urljoin

service_base_url = DATAROBOT_ENDPOINT.rstrip("/") + "/"
predict_url = urljoin(
		service_base_url,
		f"deployments/{deployment_id}/predictionsUnstructured",
)
```

The direct local service exposes `GET /healthz` and `POST /predict`.

## Configuration And Authentication

Structured SDK defaults live in `jointfm_client.configuration.JointFMConfig` and are mirrored in the checked-in `config.sample.yaml`. Copy `config.sample.yaml` to `config.yaml` and change only the fields needed for your deployment or transport defaults. `JointFMClient.from_env()` and `load_settings()` read `config.yaml` by default, then layer `.env` values over it, then layer process environment variables or the supplied `env` mapping over both. Explicit Python arguments such as `timeout=` and `retry_config=` still override YAML transport defaults.

`JointFMClient.from_env()` and `load_settings()` resolve `JOINTFM_SCHEMA_VERSION` and exactly one service selector from that layered configuration. Hosted options include `JOINTFM_DEPLOYMENT_ID` or load-balanced `JOINTFM_DEPLOYMENT_IDS` (comma-separated same-checkpoint peers; mutually exclusive with other selectors). `JOINTFM_MODEL_VERSION` is optional: when unset the SDK discovers the model version from `/healthz` on first use, and when set the SDK validates it against `/healthz` as a drift-detection guard. Hosted selectors also require `DATAROBOT_ENDPOINT` and `DATAROBOT_API_TOKEN`; the direct local selector does not use DataRobot credentials. Missing credentials, missing schema version, malformed credentials, unsupported schema versions, missing selectors, and multiple selectors raise `JointFMConfigurationError`.

`DATAROBOT_ENDPOINT` must be a normalized HTTPS DataRobot API v2 URL ending in `/api/v2`; the SDK stores it without a trailing slash. `DATAROBOT_API_TOKEN` must be non-empty and whitespace-free. The token is excluded from `JointFMSettings` repr output.

Required `.env` entries for hosted SDK calls are `DATAROBOT_ENDPOINT`, `DATAROBOT_API_TOKEN`, `JOINTFM_SCHEMA_VERSION`, and exactly one hosted selector from the list below. Required `.env` entries for local REST calls are `JOINTFM_LOCAL_BASE_URL` and `JOINTFM_SCHEMA_VERSION`. `JOINTFM_MODEL_VERSION` may be set in `.env` to pin a specific deployment artifact (the SDK then hard-errors on mismatch with `/healthz`); leave it unset to let the SDK use whatever version the deployment currently advertises. `.env` is the right place for these pins when using `from_env()` because they describe the selected JointFM service rather than a package-wide default. They are not secrets, and callers can still override them with process environment variables. Optional live DataRobot smoke tests additionally read `DATAROBOT_DEPLOYMENT_ID` from `.env` and use it as the hosted deployment ID for the `deployments/{deployment_id}/predictionsUnstructured` route.

Example deployment configuration:

```yaml
deployment:
	datarobot_endpoint: https://app.datarobot.com/api/v2
	datarobot_api_token: <token>
	schema_version: v5
	deployment_id: <deployment-id>
	# Optional model-version pin; the SDK discovers it from /healthz when unset:
	# model_version: jointfm-inference:0.3.0+ckpt.fin-2026-05-22
transport:
	timeout:
		connect_seconds: 10.0
		read_seconds: 120.0
	retry:
		max_attempts: 20
		backoff_seconds: 2
```

Equivalent `.env` deployment configuration:

```dotenv
DATAROBOT_ENDPOINT=https://app.datarobot.com/api/v2
DATAROBOT_API_TOKEN=<token>
JOINTFM_SCHEMA_VERSION=v5
JOINTFM_DEPLOYMENT_ID=<deployment-id>
# Optional drift-detection pin; the SDK discovers the model version from /healthz when unset:
# JOINTFM_MODEL_VERSION=jointfm-inference:0.3.0+ckpt.fin-2026-05-22
```

Equivalent local REST configuration for a service started from the `joint` repository with `task service:start CONFIG=nvidia-studentt-m4cr2`:

```dotenv
JOINTFM_LOCAL_BASE_URL=http://127.0.0.1:8080
JOINTFM_SCHEMA_VERSION=v5
# Optional drift-detection pin; the SDK discovers the model version from /healthz when unset:
# JOINTFM_MODEL_VERSION=jointfm-inference:0.3.0+ckpt.fin_i504_o63_f0_t10_h16l16_mam7_af_t3r1_cnn_k3l4_hpst_h16l2_studentt_m4cr2df8skew
```

Choose exactly one service selector:

- `JOINTFM_DEPLOYMENT_ID`: builds `DATAROBOT_ENDPOINT.rstrip("/") + "/"` plus `deployments/{deployment_id}/predictionsUnstructured`
- `JOINTFM_DEPLOYMENT_IDS`: comma-separated hosted deployment IDs (≥2 unique, same checkpoint) for round-robin load balancing; mutually exclusive with other selectors
- `JOINTFM_DEPLOYMENT_URL`: appends `/predictionsUnstructured` to a hosted deployment URL
- `JOINTFM_PREDICT_URL`: uses a full hosted prediction URL ending in `/predictionsUnstructured`
- `JOINTFM_DEPLOYMENT_TARGET` with `JOINTFM_PULUMI_OUTPUTS_PATH`: resolves a named target from saved Pulumi outputs JSON, preferring `deployment_id`, then `deployment_url`, then `predict_url`
- `JOINTFM_LOCAL_BASE_URL`: builds direct local `GET /healthz` and `POST /predict` URLs without DataRobot authentication

Pulumi deployment discovery is explicit and file-backed. Export stack outputs to a JSON object keyed by target name, then set `JOINTFM_DEPLOYMENT_TARGET` to the key and `JOINTFM_PULUMI_OUTPUTS_PATH` to that JSON file:

```json
{
	"fin-studentt": {
		"deployment_id": "<deployment-id>"
	}
}
```

For each target, the SDK accepts exactly one of these output fields: `deployment_id`, `deployment_url`, or `predict_url`. `deployment_id` is preferred because it lets the SDK build both the hosted health URL and the `deployments/{deployment_id}/predictionsUnstructured` URL from the configured `DATAROBOT_ENDPOINT`. `deployment_url` is normalized and extended with `/healthz` and `/predictionsUnstructured`. `predict_url` is accepted when the full hosted prediction URL has already been discovered; the SDK derives the owning deployment URL from it for health checks.

Hosted prediction calls use the same authorization scheme as the notebook helper:

```python
{
	"Authorization": f"Bearer {DATAROBOT_API_TOKEN}",
	"Accept": "*/*",
	"Content-Type": "application/json;charset=UTF-8",
}
```

The SDK still decodes hosted prediction responses as JSON; the broad `Accept` value avoids hosted unstructured prediction content negotiation failures before the deployment body is returned.

Direct local URL helpers are used by the local service selector: `build_local_health_url("http://localhost:8080")` returns `/healthz`, and `build_local_predict_url("http://localhost:8080")` returns `/predict`.

Hosted settings also derive `health_url` from the resolved deployment URL as `deployments/{deployment_id}/healthz`. `JointFMClient.health(cache=True)` stores typed `HealthMetadata` only when the caller asks for caching, and `JointFMClient.refresh_health()` fetches a fresh copy.

Each endpoint's health payload describes only that endpoint. With `JOINTFM_DEPLOYMENT_IDS`, the client probes every configured peer and aggregates locally: `health()` returns consensus metadata whose `max_sample_count` is the **minimum** reachable cap (the sample-batch size), while `health_instances()` returns one `InstanceHealth` per configured ID plus the **sum** of reachable caps as overall parallel capacity, a compact `topology` / `topology_label` (for example `2x5000` or `1x7000, 1x3000`), and errors for unavailable peers.

## CLI Workflows

The package installs a `jointfm-client` command. It reads `.env` by default, accepts `--dotenv <path>` for another file, and accepts `--no-dotenv` when the process environment should be the only source.

Validate credentials, resolve the deployment, probe health, and print non-secret service metadata plus per-instance availability and sample topology:

```bash
uv run jointfm-client health
```

The health command includes consensus `service` metadata, an `instances` list (available/unavailable, per-instance sample cap, errors), `topology` (for example `1x7000, 1x3000`), overall `max_sample_count` (sum of reachable caps), and non-secret `deployment` settings when configured.

Submit one low-level JSON request file and write the JSON response file:

```bash
uv run jointfm-client predict request.json response.json
```

Forecast from CSV history and write tidy forecast rows as CSV:

```bash
uv run jointfm-client forecast-csv history.csv forecast.csv \
	--query-times 2,3,4 \
	--target-column target \
	--return-mode mean
```

`forecast-csv` supports `--time-index-mode ordinal|continuous_float|absolute_datetime`, `--time-column`, repeated `--target-column`, repeated `--requested-column`, `--return-mode mean|samples|quantiles`, `--n-samples`, `--quantiles`, and `--seed`. The output is the same tidy shape returned by the Python result helpers.

## Notebook Workflows

Checked-in example notebooks live under `notebooks/`. Every example starts with:

```python
from jointfm_client import bootstrap_notebook

bootstrap_notebook(add_src_root=True)
```

Run `task setup` first so VS Code can select the registered `Python (joint-client-python)` notebook kernel backed by this repository's `.venv`.

The bootstrap helper resolves the nearest src-layout Python project root, switches the working directory there, and prepends that project's local `src` tree during development. The examples cover hosted health checks, low-level JSON prediction, mean forecasts, sample forecasts, quantile forecasts, conditional forecasts (one conditional read as a mean, as draws, as quantiles inside a bounded band, and as the log density of observed values, plus the ranking of candidate conditions), pandas/NumPy result conversion, and CSV forecast workflows. They use `.env.sample` placeholders and checked-in fixture payloads; no real tokens or deployment IDs are stored in notebooks.

The current forecast request contract is:

- `schema_version`: exactly `"v5"`, configured as `JOINTFM_SCHEMA_VERSION` for `from_env()` clients
- `model_version`: exact model version advertised by `/healthz` or otherwise selected by the caller. Optional for `from_env()` clients: when `JOINTFM_MODEL_VERSION` is unset the SDK reads it from `/healthz` on first use; when set it acts as a drift-detection pin
- `query_mode`: `"forecast"` for the unconditional forecast, or `"condition"` for the forecast given conditions on some columns; the high-level helpers set it from whether a `condition` was passed
- `return_mode`: one of `"mean"`, `"samples"`, or `"quantiles"`
- `time_index_mode`: one of `"ordinal"`, `"continuous_float"`, or `"absolute_datetime"`
- `time_column`: required for `"absolute_datetime"`, and used for ordered ordinal or continuous histories when supplied
- `query_times`: non-empty future forecast times only
- `requested_columns`: optional column names or integer column indices, with duplicates rejected
- `n_samples`: positive sample count for sampled forecasts and quantile estimation. When `return_mode="samples"` exceeds the `max_sample_count` advertised by the deployment's health metadata, `forecast_samples(...)` splits the request into capped prediction batches up front and returns one merged `SampleForecastResult`.
- `condition`: required with `query_mode="condition"` and forbidden otherwise. One `EqualityCondition` or `IntervalCondition`, or a list of them, each covering the future positions its `query_time_indices` names; sent on the wire as the `conditions` list, see below.

### Conditional Queries

The `condition` query mode asks for the forecast *given* something about some of its columns. Each condition names the future positions it covers by index into `query_times` through `query_time_indices`, and covers every position when that is `None`. Two conditions on the same column must cover disjoint positions; conditions on different columns may share positions, which is how one position mixes both kinds:

- `EqualityCondition(column, value, query_time_indices=None)` pins the column to a finite value. It stays nameable in `requested_columns` and reads back the value the request supplied, so a scenario answer lines up column for column with an unconditioned one.
- `IntervalCondition(column, lower=None, upper=None, query_time_indices=None)` confines the column to a range; `None` leaves that side open, and at least one side must be bounded. The column stays readable, and what comes back is its distribution inside the range.

The response answers every entry of `query_times`. A position no condition covers carries the unconditioned forecast, and that is exact: the model draws each future position from its own joint, independently of the others, so a condition relates columns to each other within a position and never reaches another one. At every covered position at least one column must stay unconditioned, because what the position reads out is the conditional distribution of those columns. `requested_columns` chooses what the response carries independently of that, defaulting to every declared column in declared order. Pass one condition or a list of them to `forecast(...)`, `forecast_mean(...)`, `forecast_samples(...)`, `forecast_quantiles(...)`, or `forecast_log_prob(...)`:

```python
from jointfm_client import EqualityCondition, IntervalCondition

result = client.forecast_mean(
	history,
	query_times=query_times,
	requested_columns=["portfolio_nav", "realized_volatility"],
	columns=plan.columns,
	condition=[
		EqualityCondition(column="equity_index_level", value=4780.0),
		IntervalCondition(
			column="treasury_10y_yield", lower=0.041, upper=0.045, query_time_indices=[0]
		),
	],
)
print(result.plausibility)
```

Whether a deployment can condition depends on the mounted checkpoint's head. `/healthz` advertises `condition` in `supported_query_modes` and the kinds it answers in `supported_condition_kinds` (empty when the mode is absent). The client checks that advertisement before sending, so a deployment that cannot condition is refused with `UnsupportedServiceContractError` rather than after a paid round trip; `require_condition_support(metadata, condition)` exposes the same check.

A condition response carries a `plausibility` block: `equality_log_density` is the log density the model assigns to the pinned values and `region_log_probability` the log probability it gives the interval region *given the pinned values at the same position*, so the two add to the plausibility of the whole condition set, each `None` when the request carried no condition of that kind. Over several covered positions each is the sum of the per-position values, exact because positions are independent. They separate a confident answer from one conditioned on something the model finds implausible; the service reports them and never refuses on them. `diagnostics.condition_draws` counts the draws behind sampled outputs, and `diagnostics.interval_estimator` reports the numerical accounting (`points`, `effective_sample_size`) when some position bounds more than one column and its region probability had to be estimated; across several estimated positions it describes the one with the smallest effective sample size.

Column descriptors support the server fields `name`, `modality`, `role`, `nullable`, `vocabulary_size`, `level_count`, `mapping`, `lower_bound`, `upper_bound`, `time_value_kind`, `time_value_scale_seconds`, `time_value_use_local_normalized_time`, `time_value_calendar_id`, and `time_value_timezone`.

DataFrame helpers and the notebook examples are available through one optional extra that pulls in `pandas`:

```bash
uv add "jointfm-client[notebooks]"
```

Use `build_forecast_payload_from_dataframe(...)` when history is already in a pandas DataFrame. It can accept explicit `ColumnSpec` objects or infer basic numeric, categorical, ordinal, count, binary, and time-valued columns from the DataFrame plus role, mapping, nullable, and bounds hints. The helper emits `history_rows` in the same order as the service frame builder: `time_column` first when present, followed by the ordered modeled columns. `build_forecast_payload_from_arrays(...)` provides the same request path for two-dimensional NumPy-like arrays when callers already have array values and column metadata. `build_datetime_query_times(...)`, `build_ordinal_query_times(...)`, `build_continuous_query_times(...)`, and `validate_forecast_horizon(...)` perform local future-horizon validation before the SDK sends the request.

Successful forecast responses preserve `schema_version`, `image_version`, `model_version`, `checkpoint_version`, `head`, `query_mode`, `return_mode`, `outputs`, `plausibility`, and `diagnostics`. Structured service errors use this shape:

```json
{
	"schema_version": "v5",
	"errors": [
		{
			"code": "VALIDATION_ERROR",
			"message": "request field explanation",
			"field": "request_field"
		}
	]
}
```

Known error codes are `VALIDATION_ERROR`, `UNSUPPORTED_HEAD_QUERY_COMBINATION`, `UNSUPPORTED_RETURN_MODE`, `SCHEMA_VERSION_MISMATCH`, `MODEL_VERSION_MISMATCH`, `INPUT_SIZE_EXCEEDED`, and `INTERNAL_ERROR`.

## Compatibility Policy

The SDK supports only `schema_version="v5"`. `validate_service_metadata()` checks `/healthz` metadata and raises typed compatibility errors before prediction if the service advertises a different schema, an unexpected model version, mode capabilities outside the recorded service contract, or an unsupported `decoding_strategy`. Return modes and time-index modes must match the SDK's lists exactly. Query modes and condition kinds are derived by the service from the mounted head, so a deployment may advertise fewer of them than the SDK knows; it must advertise at least one query mode and nothing the SDK does not know.

Callers should pass an expected `model_version` when they already know which deployment artifact they intend to use. A mismatch is treated as a hard compatibility error rather than silently downgrading, guessing, or retrying another model.

High-level forecast helpers build the same validated request payloads as `build_forecast_payload(...)` and return `ForecastResponse`. Row-list inputs require an explicit `DataFrameSchema` or `ColumnSpec` sequence, while pandas DataFrame inputs can use the DataFrame adapter inference options. `predict(payload)` remains the low-level JSON method and requires the payload to include `model_version`.

## Quick Start

### Install `uv` and `task`

Use the same curl-based bootstrap flow on Ubuntu, macOS, and AWS Linux:

```bash
curl -LsSf https://astral.sh/uv/install.sh | sh
uv python install 3.13

sh -c "$(curl --location https://taskfile.dev/install.sh)" -- -d -b ~/.local/bin
```

Ensure `~/.local/bin` is on your `PATH` in new shells.

### Set Up The Project

```bash
task setup
```

`task setup` creates or reuses `.venv` pinned to Python 3.13.3, synchronizes the project and development dependencies, installs the local Git pre-commit hooks, installs or verifies `typos`, and prints the shell activation hint for `.venv`.

Verify the installed package import:

```bash
uv run python -c "import jointfm_client; print(jointfm_client.__version__)"
```

### Configure A Deployment

Create `.env` from `.env.sample` or set the same values in your shell. A hosted forecast needs the DataRobot API v2 endpoint, token, schema pin, and exactly one deployment selector:

```dotenv
DATAROBOT_ENDPOINT=https://app.datarobot.com/api/v2
DATAROBOT_API_TOKEN=<token>
JOINTFM_SCHEMA_VERSION=v5
JOINTFM_DEPLOYMENT_ID=<deployment-id>
# Or: JOINTFM_DEPLOYMENT_IDS=chevron-id,research-id
# Optional drift-detection pin; the SDK discovers the model version from /healthz when unset:
# JOINTFM_MODEL_VERSION=jointfm-inference:0.3.0+ckpt.fin-2026-05-22
```

Check that the SDK can resolve the deployment and that the service metadata matches the configured schema and model pins:

```bash
uv run jointfm-client health
```

### Run The First Forecast

This minimal example uses row dictionaries plus explicit column metadata, so it does not require pandas:

```python
from jointfm_client import (
	ColumnSpec,
	DataFrameSchema,
	JointFMClient,
	build_ordinal_query_times,
)

history_rows = [
	{"t": 0, "sales": 10.0},
	{"t": 1, "sales": 12.0},
	{"t": 2, "sales": 13.5},
	{"t": 3, "sales": 15.0},
]
schema = DataFrameSchema(
	columns=(ColumnSpec(name="sales", modality="numeric", role="target"),),
	time_index_mode="ordinal",
	time_column="t",
)

client = JointFMClient.from_env()
result = client.forecast_mean(
	history_rows,
	schema=schema,
	query_times=build_ordinal_query_times([row["t"] for row in history_rows], periods=3),
	requested_columns=("sales",),
)

print(result.to_pandas_tidy())
```

Use `forecast_samples(...)` for sampled trajectories or `forecast_quantiles(...)` with `quantiles=(0.1, 0.5, 0.9)` for quantile surfaces. For pandas inputs, install the optional dataframe extra and pass a `DataFrame` to `forecast(...)` with `target_columns` and `requested_columns` set explicitly.

## Development Commands

- `task lint`: run Ruff lint checks (read-only)
- `task format`: run Ruff formatter (rewrites files in place)
- `task license-check`: verify every Python source file has the required copyright/SPDX header
- `task typecheck`: run `ty` static type checks
- `task test`: run unit tests
- `task coverage`: run tests with coverage enforcement above 90%
- `task build`: build the source distribution and wheel, then validate artifact metadata and contents
- `task check`: run the static code quality gate (typos, lint, format check, type checks)
- `task release:dry`: preview the next SemVer bump without changing any files
- `task release`: cut a SemVer release with Commitizen (writes `CHANGELOG.md`, bumps versions, creates tag)
- `task release:publish`: push the release commit and its tag to `origin`, which triggers the PyPI publish workflow
- `task pre-commit`: run every configured pre-commit hook

Contributors do not need to add copyright or license headers manually. The `insert-license` pre-commit hook runs [skywalking-eyes](https://github.com/apache/skywalking-eyes) (via the `apache/skywalking-eyes` Docker image, so a running Docker daemon is required) to stamp the standard Apache-2.0 header (`Copyright 2026 DataRobot, Inc. and its affiliates.` followed by the standard "Licensed under the Apache License, Version 2.0" notice) into every `.py` file, and the companion `insert-license-notebooks` hook stamps the same notice into a leading markdown cell of every notebook the first time you run `task pre-commit`. Verify the headers are present at any time with `task license-check`.

Your user must belong to the `docker` group so the hook can reach the daemon:

```shell
sudo usermod -aG docker "$USER"
```

Group membership is captured when a process starts, so it must be in effect **before** the editor or its language server launches — start a fresh login session (or reboot) after adding yourself to the group, otherwise the commit hook inherits the old groups and fails with a Docker permission error.

## Versioning & Commits

The package follows strict [Semantic Versioning](https://semver.org/spec/v2.0.0.html). Releases are cut with [Commitizen](https://commitizen-tools.github.io/commitizen/), driven by [Conventional Commits](https://www.conventionalcommits.org/), so the commit log is the source of truth for what a release contains.

### Commit message format

Every commit subject must follow:

```text
<type>(<optional scope>): <imperative summary>
```

Examples:

```text
feat(adapters): add quantile forecast helper
fix(transport): retry on idempotent 5xx responses
perf(adapters): cache schema validation across forecast batches
refactor(configuration): split URL resolution into helpers
docs(readme): document the release workflow
test(transport): cover retry on 502/503/504
build(deps): bump pydantic to 2.13.4
ci(pre-commit): pin commitizen to v3.31
chore(repo): move fixtures to tests/fixtures/v1
```

#### Commit types

The full set of accepted types comes from the `cz_conventional_commits` rule set referenced above. Pick the type that matches the **primary intent** of the commit; split commits that mix concerns. Only `feat`, `fix`, `refactor`, and `perf` (plus breaking-change markers) drive a SemVer bump — every other type is recorded in `git log` but does not, on its own, cause `cz bump` to cut a new version.

| Type | When to use | SemVer bump |
| --- | --- | --- |
| `feat` | A user-visible feature: new public API, new CLI flag, new helper exposed to callers. | **minor** |
| `fix` | A bug fix in shipped behavior — the symptom is observable to callers or operators. | **patch** |
| `perf` | A change that improves performance without changing observable behavior. | **patch** |
| `refactor` | An internal restructure that neither adds a feature nor fixes a bug (renames, moves, extractions, internal type changes). | **patch** |
| `docs` | Documentation-only changes (READMEs, prose, docstrings, comments, example payloads, API reference). | none |
| `test` | Adding, fixing, or restructuring tests, fixtures, or test helpers with no production-code change. | none |
| `build` | Build system, packaging, or dependency-pinning changes (`pyproject.toml`, lockfile, wheel build hooks). | none |
| `ci` | CI configuration changes (GitHub Actions workflows, `.pre-commit-config.yaml`, release hooks). | none |
| `chore` | Repository maintenance not covered above (tooling tweaks, repo-level renames, housekeeping, fixture moves). | none |
| `style` | Pure formatting or whitespace, no behavior change — rare here because Ruff format runs in pre-commit. | none |
| `revert` | Reverts a previous commit; the body should include `Refs: <sha>`. Re-add `!` or a `BREAKING CHANGE:` footer if the reverted commit was a breaking change. | none |
| any type with `!` or a `BREAKING CHANGE:` footer | A breaking change to public API, configuration schema, environment variables, or the service wire contract. | **major**¹ |

¹ This project sets `major_version_zero = true` in `pyproject.toml`, so breaking changes are downgraded to a **minor** bump while the SDK is on `0.x`. They will become major bumps once the SDK ships `1.0.0`.

A release window that contains only `none`-bump commits is not releasable on its own: `cz bump` exits without writing a new version. Either land a `feat`/`fix`/`refactor`/`perf` first, or force the bump explicitly with `task release -- --increment PATCH`.

#### Breaking changes

Declare a breaking change with `!` after the type, e.g. `feat(adapters)!: rename forecast() to predict()`, or — preferred when the change needs explanation — with an explicit footer:

```text
feat(adapters): rename forecast() to predict()

BREAKING CHANGE: forecast() is gone; callers must use predict().
```

The `commit-msg` pre-commit hook runs `cz check` and rejects malformed messages locally before the commit is created.

### First-time tag seed (one-off)

Commitizen bumps **from** the tag matching the current `version =` in `pyproject.toml`. A brand-new repo has no such tag, so the first `task release` or `task release:dry` will fail with a clear message. Seed it once:

```bash
git tag -a v0.0.1 -m 'Seed initial release tag for Commitizen'
git push --tags
```

After that, every future `task release` finds its base tag automatically.

### Cutting a release

```bash
task release:dry      # preview the next version + CHANGELOG entries
task release          # bump, write CHANGELOG.md, create the annotated tag
task release:publish  # push the bump commit and the tag
```

`task release` first runs `task release:check` (clean tree, on `main`, in sync with `origin/main`), then calls `cz bump` which:

- reads commits since the last `v*` tag,
- picks the SemVer bump from the types it sees,
- updates `CHANGELOG.md`,
- bumps `version =` in `pyproject.toml` and the "Current SDK package version" line in this README,
- commits the bump and creates the annotated tag.

Publishing is a separate task on purpose, so you can inspect the inferred bump before anything leaves your machine — Commitizen derives the version from commit messages, and a stray `feat:` where you meant `fix:` is only fixable while the release is still local. Override the inferred bump level when needed: `task release -- --increment minor`.

`task release:publish` re-checks that the working tree is clean, that you are on `main`, that the tag points at `HEAD`, and that `origin` does not already have the tag, prints the commits about to be pushed, asks for confirmation on a terminal, and then pushes `main` and that single tag. It pushes one explicit tag rather than `git push --tags`, so unrelated local tags are never published.

Pushing the `v*` tag triggers the [`Publish to PyPI`](.github/workflows/publish.yml) workflow, which rebuilds and validates the distribution with `task build` and uploads it to PyPI via [`pypa/gh-action-pypi-publish`](https://github.com/pypa/gh-action-pypi-publish) — keeping the git tag, the wheel filename, `jointfm_client.__version__`, and the PyPI version all in lockstep.

The workflow authenticates with PyPI through [trusted publishing](https://docs.pypi.org/trusted-publishers/) (OIDC), so no API-token secret is stored in the repository. Before the first release, configure it once on PyPI:

- register a [pending trusted publisher](https://docs.pypi.org/trusted-publishers/creating-a-project-through-oidc/) for the `jointfm-client` project pointing at owner `datarobot`, repository `joint-client-python`, workflow `publish.yml`, and environment `pypi`;
- create a GitHub Actions environment named `pypi` on the repository (optionally gated with required reviewers) so the publish job can run.

After pulling these changes for the first time, run `uv run pre-commit install` (or `task setup`) once so the new `commit-msg` hook is registered with git.

### Hand-editing the version (don't)

The `version =` line in [`pyproject.toml`](pyproject.toml) and the "Current SDK package version" line in this README are **owned by `cz bump`** — treat them the way you'd treat a lockfile. Each release rewrites both of them atomically.

Nothing in the package source carries a version literal: `jointfm_client.contract.PACKAGE_VERSION` reads the installed distribution metadata that the build backend generates from `[project].version`, and `jointfm_client.__version__` is that same value. They follow a bump automatically, so there is nothing to hand-edit and nothing that can drift.

If you hand-edit them, the next `task release` will catch you:

- if you bumped only some lines, `cz bump` aborts with "Configured files cannot be updated, check consistency";
- if you bumped them all but never created the matching `vX.Y.Z` tag, `task release:require-tag` refuses to run and tells you to seed the tag.

If you genuinely need to set a specific version outside the normal release flow (bootstrap, recovery from a botched state), use the escape hatch:

```bash
uv run cz bump --files-only X.Y.Z
```

That atomically rewrites every `version_files` line to `X.Y.Z` without committing or tagging. After it returns, commit and tag manually.
