Metadata-Version: 2.4
Name: fourtheorem-tempo-worker
Version: 0.4.0
Summary: Tempo worker SDK: failure reporting, logging, and custom metrics for worker tasks
Author: fourTheorem
License-Expression: Apache-2.0
License-File: LICENSE
Requires-Dist: opentelemetry-api>=1.25
Requires-Python: >=3.11
Project-URL: Homepage, https://fourtheorem.com/high-performance-compute/
Description-Content-Type: text/markdown

# Tempo Worker SDK (Python)

A Tempo **worker** is a process that Tempo runs on its compute backends to
execute tasks. A **task** is your own custom business logic, packaged as a
function with this signature:

```python
from pathlib import Path

def run(working_dir: Path, params: dict) -> None: ...
```

Tempo calls `run` with a working directory (inputs under
`working_dir/inputs/`, outputs written to `working_dir/outputs/`) and the
job parameters. Task code needs no platform imports, so you do not need an
SDK to write a task. An SDK is still convenient: it gives you easy,
standardised access to the job run context, structured logging, custom
metrics, and failure reporting. `fourtheorem-tempo-worker` (import name
`tempo_worker`) provides all of this.

```bash
pip install fourtheorem-tempo-worker
```

When you run task code yourself (locally, or in unit tests) instead of on a
Tempo worker, this is a **standalone run**: the SDK detects it and falls
back to the safe defaults noted in each section below.

## Features

Here is the list of features you find in the SDK.

### Run context

`current_context()` returns the identifiers of the running job attempt.

```python
from tempo_worker import current_context

def run(working_dir, params):
    ctx = current_context()
    if ctx is not None:
        print(ctx.workload_id, ctx.job_id, ctx.attempt, ctx.worker_id)
```

The context object carries four properties:

- `workload_id`: the workload the job belongs to.
- `job_id`: the job the worker is executing.
- `attempt`: the zero-based retry attempt number.
- `worker_id`: the worker process running the job.

In a standalone run, `current_context()` returns `None`.

### Logging

`get_logger` returns a standard library logger that emits structured,
single-line JSON records, ready for observability and monitoring tools. On a
Tempo worker each record automatically carries `workload_id`, `job_id`,
`attempt`, and `worker_id`, so you can filter and correlate task logs with
the rest of the platform logs.

```python
from tempo_worker import get_logger

logger = get_logger(__name__)

def run(working_dir, params):
    logger.info("Pricing started", extra={"scenario_count": len(params["scenarios"])})
```

In a standalone run the SDK installs a JSON handler with the same field
shape, so log lines look identical in both environments. Standalone runs
read the service name from `TEMPO_SERVICE_NAME` (default `tempo-task`) and
the level from `LOG_LEVEL` (default `INFO`).

### Custom metrics

The `tempo_worker.metrics` module provides `counter`, `histogram`, and
`up_down_counter`. They create instruments that export through the worker's
metrics pipeline. Tempo prefixes each name with `tr_custom_` and attaches
the `workload_id` label from the run context. The series appear in the
Tempo Grafana dashboard under "Custom task metrics".

```python
from tempo_worker import metrics

rows = metrics.counter("rows_processed", description="Rows processed per job.")
latency = metrics.histogram("pricing_seconds", unit="s")

def run(working_dir, params):
    rows.add(1_000)
    latency.record(2.31)
```

The attribute keys `job_id` and `worker_id` are rejected: their cardinality
is unbounded and would overload the time-series store.

A standalone run has no configured metrics provider, so every call is a
no-op.

> **A note on naming.** You might have noticed that logging has a factory
> function, `get_logger`, while metrics is a module namespace. This is
> deliberate: each follows the strongest convention in its own domain.
> `get_logger(__name__)` mirrors the stdlib `logging.getLogger(__name__)`
> idiom, and `metrics.counter(...)` mirrors the OpenTelemetry module-style
> API. A logger is a stateful object you keep; the metrics functions are
> factories for named instruments.

### Explicit failure reporting

Raise `TaskFailure` to report a failure with structured context. Set
`retriable=False` to tell the scheduler not to retry the job, even when retry
budget remains. Use it when a retry cannot change the outcome, for example
corrupt input data or invalid parameters: every attempt would fail the same
way, so skipping the retries saves time and compute, and the workload
reaches its final state sooner.

The `detail` dict is stored in the job attempt record, so
keep its values JSON-serialisable.

```python
from tempo_worker import TaskFailure

def run(working_dir, params):
    if "rate_table" not in params:
        raise TaskFailure(
            "missing rate_table parameter",
            retriable=False,
            detail={"missing_field": "rate_table"},
        )
```

By default, any exception other than `TaskFailure` is reported as a
retriable failure.

In a standalone run, `TaskFailure` is a plain exception and propagates
normally.

## Versioning

The SDK follows the Tempo platform release version.
