Metadata-Version: 2.4
Name: fourtheorem-tempo-worker
Version: 0.7.0
Summary: Tempo worker SDK: failure reporting, logging, and custom metrics for worker tasks
Author: fourTheorem
License-Expression: Apache-2.0
License-File: LICENSE
Requires-Dist: opentelemetry-api>=1.25
Requires-Python: >=3.12
Project-URL: Homepage, https://fourtheorem.com/high-performance-compute/
Description-Content-Type: text/markdown

# Tempo Worker SDK (Python)

A Tempo **worker** is a process that Tempo runs on its compute backends to
execute tasks. A **task** is your own custom business logic, packaged as a
function with this signature:

```python
from pathlib import Path
from tempo_worker import WorkerContext

def run(working_dir: Path, params: dict, ctx: WorkerContext) -> None: ...
```

Tempo calls `run` with a working directory (inputs under
`working_dir/inputs/`, outputs written to `working_dir/outputs/`), the
job parameters, and a `WorkerContext` that carries execution-environment
identifiers for the current attempt. Tempo creates the directory only when
the job declares inputs or outputs. A job that declares neither gets a path
that does not exist, and must create it before use. Task code needs no
platform imports beyond the SDK, so you do not need anything else to write a
task. The SDK gives you easy, standardised access to the job run context,
structured logging, custom metrics, and failure reporting.
`fourtheorem-tempo-worker` (import name `tempo_worker`) provides all of this.

```bash
pip install fourtheorem-tempo-worker
```

When you run task code yourself (locally, or in unit tests) instead of on a
Tempo worker, this is a **standalone run**: the SDK detects it and falls
back to the safe defaults noted in each section below.

## Features

Here is the list of features you find in the SDK.

### Run context

The harness passes a `WorkerContext` directly as the third argument to `run`.
It carries execution-environment identifiers for the current attempt.

```python
from pathlib import Path
from tempo_worker import WorkerContext

def run(working_dir: Path, params: dict, ctx: WorkerContext) -> None:
    print(ctx.workload_id, ctx.job_id, ctx.attempt, ctx.worker_id, ctx.resource_class)
```

The context object carries five properties:

- `workload_id`: the workload the job belongs to.
- `job_id`: the job the worker is executing.
- `attempt`: the zero-based retry attempt number.
- `worker_id`: the worker process running the job.
- `resource_class`: the resource class serving this job (e.g. `"default"`, `"gpu"`).

`current_context()` provides the same object for helper code that cannot
receive `ctx` as a parameter. In a standalone run, `current_context()` returns
`None`; `ctx` is not available outside the harness.

### Logging

`get_logger` returns a standard library logger that emits structured,
single-line JSON records, ready for observability and monitoring tools. On a
Tempo worker each record automatically carries `workload_id`, `job_id`,
`attempt`, and `worker_id`, so you can filter and correlate task logs with
the rest of the platform logs.

```python
from tempo_worker import get_logger

logger = get_logger(__name__)

def run(working_dir, params, ctx):
    logger.info("Pricing started", extra={"scenario_count": len(params["scenarios"])})
```

In a standalone run the SDK installs a JSON handler with the same field
shape, so log lines look identical in both environments. Standalone runs
read the service name from `TEMPO_SERVICE_NAME` (default `tempo-task`) and
the level from `LOG_LEVEL` (default `INFO`).

### Custom metrics

The `tempo_worker.metrics` module provides `counter`, `histogram`, and
`up_down_counter`. They create instruments that export through the worker's
metrics pipeline. Tempo prefixes each name with `tr_custom_` and attaches
the `workload_id` label from the run context. The series appear in the
Tempo Grafana dashboard under "Custom task metrics".

```python
from tempo_worker import metrics

rows = metrics.counter("rows_processed", description="Rows processed per job.")
latency = metrics.histogram("pricing_seconds", unit="s")

def run(working_dir, params, ctx):
    rows.add(1_000)
    latency.record(2.31)
```

The attribute keys `job_id` and `worker_id` are rejected: their cardinality
is unbounded and would overload the time-series store.

A standalone run has no configured metrics provider, so every call is a
no-op.

> **A note on naming.** You might have noticed that logging has a factory
> function, `get_logger`, while metrics is a module namespace. This is
> deliberate: each follows the strongest convention in its own domain.
> `get_logger(__name__)` mirrors the stdlib `logging.getLogger(__name__)`
> idiom, and `metrics.counter(...)` mirrors the OpenTelemetry module-style
> API. A logger is a stateful object you keep; the metrics functions are
> factories for named instruments.

### Explicit failure reporting

Raise `TaskFailure` to report a failure with structured context. Set
`retriable=False` to tell the scheduler not to retry the job, even when retry
budget remains. Use it when a retry cannot change the outcome, for example
corrupt input data or invalid parameters: every attempt would fail the same
way, so skipping the retries saves time and compute, and the workload
reaches its final state sooner.

The `detail` dict is stored in the job attempt record, so
keep its values JSON-serialisable.

```python
from tempo_worker import TaskFailure

def run(working_dir, params, ctx):
    if "rate_table" not in params:
        raise TaskFailure(
            "missing rate_table parameter",
            retriable=False,
            detail={"missing_field": "rate_table"},
        )
```

By default, any exception other than `TaskFailure` is reported as a
retriable failure.

In a standalone run, `TaskFailure` is a plain exception and propagates
normally.

## Versioning

The SDK follows the Tempo platform release version.
