Metadata-Version: 2.5
Name: robotensor-horizon-runtime-zerowam
Version: 0.2.0
Summary: The Zero-WAM model family: the served policy, its family definition, and the weight and knob checks a submission must pass.
License: Apache-2.0
License-File: LICENSE
Requires-Python: >=3.10
Requires-Dist: numpy>=1.23
Requires-Dist: pyyaml>=6.0
Requires-Dist: robotensor-horizon-protocol
Provides-Extra: dev
Requires-Dist: hatchling>=1.25; extra == 'dev'
Requires-Dist: imageio[ffmpeg]>=2.30; extra == 'dev'
Requires-Dist: pyarrow>=14; extra == 'dev'
Requires-Dist: pytest>=8; extra == 'dev'
Requires-Dist: ruff==0.16.7; extra == 'dev'
Requires-Dist: scipy>=1.10; extra == 'dev'
Provides-Extra: model
Requires-Dist: safetensors>=0.4; extra == 'model'
Requires-Dist: torch>=2.9; extra == 'model'
Provides-Extra: recipes
Requires-Dist: imageio[ffmpeg]>=2.30; extra == 'recipes'
Requires-Dist: pyarrow>=14; extra == 'recipes'
Description-Content-Type: text/markdown

# horizon-runtime-zerowam

The Zero-WAM model family for the [Robotensor Horizon competition](https://github.com/robotensor/horizon-competition):
the served policy, the family definition the competition schedules from, and the checks a submission
must pass before it is given a GPU.

A submission is **weights only**. Every model in an epoch runs on this code, so evaluation time is
predictable, a submission cannot run code on our machines, and two models differ only by what they
learned.

```bash
pip install horizon-runtime-zerowam                 # checks and conversions: no torch needed
pip install "horizon-runtime-zerowam[model]"        # and the model itself, on a machine with the weights
```

## Checking a submission

```bash
horizon-runtime-zerowam check --weights /models/submission
horizon-runtime-zerowam check --base --weights /models/zero-wam-posttrain-robotwin
horizon-runtime-zerowam fingerprint --weights /models/zero-wam-posttrain-robotwin
horizon-runtime-zerowam family --knobs icl_guidance_scale=7
```

`check` reads **safetensors headers only** — never a tensor — so a 22 GB checkpoint is checked in
milliseconds on any laptop: the architecture must match the base model's, the repository must be
within the size limit, the required files must be there, and nothing the base provides (`vae/`,
`text_encoder/`, `tokenizer/`) may be.

It covers exactly what diffusers' loader reads from `transformer/`:

- the shards `diffusion_pytorch_model.safetensors.index.json` names in its `weight_map`, all of them
  present and nothing else - a shard the index does not name is never loaded, so it is refused - or,
  with no index, the one file `diffusion_pytorch_model.safetensors`;
- `config.json`, which must be the base model's: `attn_window`, `rope_max_seq_len` and the like
  change what the model does without changing a shape. It is compared as the loader reads it (keys
  starting with `_`, such as `_diffusers_version`, do not count), so a config Zero-WAM's trainer
  saves passes, and a changed value does not.

A submission adds two files to the weights: `zerowam.yaml`, its knobs (below), and its own action
statistics, `norm_stats/<name>.json` for each set the family's `submission_requires` lists -
`norm_stats/robotwin.json` today, which is what the embodiment table's `aloha-agilex` entry is
served on - in Zero-WAM's released format: `method: abs`, and the finite
`q01` and `q99` of `action.hand.position` (14 values) and `action.effector.position` (2), no q01
above its q99. `check` refuses a submission without them, and `serve` never falls back to the base
model's statistics.

The table also holds `PandaOmron`, RoboCasa's robot, whose statistics are a 7-wide hand and one
gripper (`norm_stats/robocasa.json`). That set is **not** required of a submission: no checkpoint
has been fine-tuned for RoboCasa, so nobody could write one, and decision Q2 keeps that axis off
any epoch scoring a real model. A submission that ships one anyway is checked against it, and
`serve` refuses to serve that embodiment without it. So a submission is three things and nothing
else:

```
<submission>/
  transformer/              the shards the index names, the index, and the base model's config.json
  norm_stats/robotwin.json  its own action statistics, in Zero-WAM's released format
  zerowam.yaml              its knobs
```

`check` also reports `weights_sha256`, the submission's content hash, and `weights_files`, the
files it covers: the shards and the index the loader reads, `transformer/config.json`,
`zerowam.yaml` and the `norm_stats/<name>.json` files - everything of the submission's own
that is served, and nothing else in the repository. The components a submission does not train
(`vae/`, `text_encoder/`, `tokenizer/`) are the base's, and `serve --weights <submission> --base
<base>` serves a submission's transformer against them, so they are not in the hash; the
pinned-revision check of `check --base` is handed that same list and every file of those three
directories. `fingerprint_sha256` sees names, dtypes and shapes, so
every fine-tune of this model shares it; `weights_sha256` sees the values, so only an exact copy
shares it. It is `sha256sum` over those files, in byte order of their paths (the order `check`
lists them in), hashed again, so intake can recompute it without this package:

```bash
cd /models/submission && sha256sum <weights_files, in the order listed> | sha256sum
```

It is the one part of `check` that reads every byte (as bytes: no tensor is parsed, no torch):
1.7 s for the base's 22.9 GB from the page cache, 8 files at a time, and as long as the disk takes
when cold. `serve` does not compute it.

`check --base` checks the base checkpoint itself, which has no files of a submission's own. The
family pins it to one commit (`base_model.revision`), and the check refuses a checkpoint whose
download records (`.cache/huggingface/download/`, written by `snapshot_download(...,
local_dir=...)`) name another commit, are missing, or are older than their file. Download it at
the pinned commit:

```bash
python -c "
from huggingface_hub import snapshot_download
from horizon_runtime_zerowam import family
base = family.load().base_model
snapshot_download(base['repo'], revision=base['revision'],
                  local_dir='/models/zero-wam-posttrain-robotwin')"
```

## Serving it

```bash
export HORIZON_AUTHKEY=$(python -c 'import secrets; print(secrets.token_bytes(32).hex())')
horizon-runtime-zerowam serve --weights /models/submission --base /models/zero-wam-posttrain-robotwin \
    --address 127.0.0.1:7100 --authkey-env HORIZON_AUTHKEY --max-sessions 0 \
    --knobs num_inference_steps=50
```

A submission is served with its own `transformer/` and `norm_stats/`, and with the pinned base's
`vae/`, `text_encoder/` and `tokenizer/` (`weights.from_base` in the family file): those three never
come from a submission, and `check` refuses one that ships them. `--base DIR` alone serves the bare
base.

A robot is served on the entry the family's `embodiments` table has for it, and on no other: the
released server's configuration, its camera map, the statistics a submission is served on, and the
whole action space the benchmark must declare — every field of Q3's `action_spec`. An
`info.action_spec` that differs from the entry in any field (`base_poses` by more than 1e-5) is
refused, naming the field, rather than converted: a two-armed robot is never served through the
one-arm path, a `robot_base` unit is never composed as if it were `world`, and another tool,
execution or gripper reading is never served as RoboTwin's. The pinned block is the fork's own spec
function's output, never typed: `tools/action_spec_from_fork.py` builds one scene in the fork's
environment and prints it (`base_poses` is read live from the scene), and with
`ZEROWAM_TEST_ROBOTWIN` naming the fork a test reruns it and fails on a stale pin (C-R1).

A one-armed robot observes and executes 8 numbers, and the model works in 16 — its single arm goes
in Zero-WAM's slot 0 (channels 0-6 and 28, the rest masked), because that is where its own loader
puts one arm. The wire calls that arm `right`; the name picks nothing (decision Q3).

One server is loaded once per GPU and serves every unit of that submission: the weights take
minutes to load and a unit takes a few. It tells the benchmark what it serves - the family file's
sha256 and version, the resolved knobs, the weights' fingerprint and `weights_sha256`, the hash of
the weight bytes - and the benchmark records that in every `result.json`, so a result says what
produced it. `weights_sha256` is the server's own reading of the files it loads, taken once when
it starts, over the files and in the order `check` hashes: the value intake recorded when the bytes
are the ones it checked, and another when a checkpoint was changed after intake.

A unit is one client session, and `close` ends the session, not the model. The policy releases what
belonged to that unit - the demonstration, the episode, and the caches the model filled for them -
and keeps the checkpoint loaded for the next client, which arrives with its own `hello` and its own
demonstration. So `--max-sessions 0` loads the weights once however many units it serves, and the
process's footprint on the card does not grow with the unit count.

What the model keeps across those sessions, besides the weights, is the caption cache: a bounded
LRU of 16 T5 embeddings, ~4 MiB each. An axis reuses one HumanGen caption over many units, so a
kept server encodes that caption once rather than once per unit, and at its limit the cache holds
~64 MiB. It is dropped when the model is unloaded, with the T5 that produced it.

A submission is served on the knobs its own `zerowam.yaml` sets. `--knobs num_inference_steps=10`
is an operator's override of those - a reduced-knob smoke run, say - and is recorded: `serve`
prints the knobs it serves and what was overridden, and the policy logs the override again when
a benchmark connects. `serve --base` serves the bare base checkpoint on the family's defaults.

## The family definition

`families/zerowam.yml` is everything model-specific in one file: the base model, the weight
fingerprint, the knobs a participant may turn and their bounds, the GPU it needs, what it accepts as
a demonstration, and how a benchmark's native-rate video becomes the frames it was trained on:
`native`, one frame every 1/12 s over the whole video, however long it is. That is the window on
both axes: the three candidates are the same window on a 5 s HumanGen clip, and on the simulation
axis's three-times-longer expert videos the file records what each of them does to one and why
`native` is the one Zero-WAM's own code would pick, and what argues against it (#17). The
competition core knows only the family's name, so **another model is another repository like this
one** — the benchmarks and the epoch machinery do not change.

Its `inputs` block says what the runtime reads of a demonstration, in the protocol's own vocabulary
(decision Q4, `horizon-protocol` `docs/demonstrations.md` §9), and the runtime reads nothing else:

| Setting | Value here | What it means |
|---|---|---|
| `demonstration` | `[video, caption]` | the video, through the channel `info.demo_cameras` names first, and the video's own caption where a benchmark carries one |
| `prompt_language` | `demonstration_caption` | the text the demonstration is conditioned on: that caption, or the benchmark's generic instruction where there is none |

Change either and what is served changes with it. `proprio` and `actions` are refused: the
demonstrator's own state and actions reach no policy, on any axis.

Knobs a participant may set in `zerowam.yaml`, which holds `knobs` and nothing else:

```yaml
knobs:
  num_inference_steps: 25
  icl_guidance_scale: 6.5
```

| Knob | Range | Default |
|---|---|---|
| `icl_guidance_scale` | 1.0 – 10.0 | 5.0 |
| `num_inference_steps` | 1 – 50, whole | 50 |
| `action_num_inference_steps` | 1 – 50, whole | 50 |

A knob left out takes its default. `check` refuses a knob not in this table, a value outside its
range, a step count that is not a whole number, and a file it cannot read unambiguously (not YAML,
another shape, a key given twice).

`action_guidance_scale` is deliberately absent: with a video prompt, Zero-WAM's action branch runs a
single pass with no guidance, so the knob would do nothing.

## Training on the competition's own data

`recipes/` turns closed epochs' bundles, or your own `demo make` output, into the dataset Zero-WAM's
own trainer loads: the demonstrator's state and actions read only from `private/expert.npz`
(decision Q4), `observation.state` and `action` per step, one `action_config` span per episode, the
action statistics in the model's channel groups, the three views at the evaluation's size, and the
in-context pairing. A second pass encodes the latents on a GPU, as the runtime encodes when it
serves; [recipes/README.md](recipes/README.md) has both, training, and packaging the result.

```bash
pip install "horizon-runtime-zerowam[recipes]"
python recipes/bundles_to_lerobot.py --pool <bundles> --out data/robotwin_sim --axis robotwin_sim
```

## What is tested and what is not

Everything that does not need the weights is tested and runs anywhere: the family definition, the
weight check, the demonstration window, the action conversion.

`model.py` drives Zero-WAM's released server (`VA_Server.infer`) in the order its own RoboTwin
client does, and changes one thing: the demonstration arrives as frames and a caption instead of a
latent file, and is encoded the way the shipped latents were made (`families/zerowam.yml`, `video`).
It loads and answers on the real checkpoint (`tests/test_model_gpu.py`), the demonstration it
builds matches the seven shipped latents, and the epoch's policy seed is what decides an episode's
sampling: on the real checkpoint one seed answered the same chunks bit for bit in two fresh sessions
of a served base, and another seed answered different ones. **That is not yet a reproduction of
Zero-WAM's rates.** The acceptance test (G1, robotensor/horizon-competition#49) passed in fast mode,
as a right-versus-broken test on one task: this path solved a seed Zero-WAM's own client solved, on
the same harness. Whether the two score the same on the 7 held-out tasks is the rate comparison,
robotensor/horizon-competition#78, deferred out of fast mode. Until it runs, a timing or memory
figure from this runtime is a measurement only for the scope it was taken in, and no success rate
from it is a measurement.

The Zero-WAM checkout is pinned in the family file (`code.commit`) and checked before the server is
loaded: a process that finds another commit on its path refuses to serve, because everything that
decides an action is that code.

`tests/test_sessions_gpu.py` covers the other thing a served process must get right: three sessions
one after another build the model once and leave nothing of themselves on the card.

A **weights-only submission** - the shape an epoch actually serves, rather than the bare base - is
loaded and answered from on the real checkpoint too (`tests/test_submission_gpu.py`): the
released server is pointed at the submission's own `transformer/` and the pinned base's VAE, T5 and
tokenizer, runs on the submission's own statistics and `zerowam.yaml`, and answers an episode's acts
as absolute end-effector targets, moved and turned by the pose the episode started at. Its
submission is built by hardlinking the base's `transformer/` file by file, so it needs no second
checkpoint and is still a directory of its own, which is what lets the test tell the two apart;
until a fine-tune exists, the values it serves are the base's own.

```bash
# the model's environment: Zero-WAM 08e2c4a and its requirements, torch for this GPU, setuptools
# (torch.compile builds flex_attention with it), and this runtime; then
ZEROWAM_TEST_WEIGHTS=/models/zero-wam-posttrain-robotwin ZEROWAM_TEST_HUMANGEN=/data/HumanGen \
    pytest -m gpu tests/test_model_gpu.py tests/test_sessions_gpu.py \
        tests/test_submission_gpu.py
horizon-runtime-zerowam serve --base /models/zero-wam-posttrain-robotwin \
    --address 127.0.0.1:7100 --authkey-env HORIZON_AUTHKEY --max-sessions 0
```

## T2: the model smoke

```bash
python scripts/t2.py --axis robotwin_sim --out runs/t2/robotwin_sim --benchmark ../RoboTwin-Horizon \
    --weights /models/zero-wam-posttrain-robotwin --base --runtime-python ../Zero-WAM/.venv/bin/python
```

One axis, one task, one episode (plan §12.1): the fork's `demo make` builds the axis's smoke unit,
`horizon-runtime-zerowam serve` serves this checkout with reduced knobs, and the fork's `eval run` drives it.
`t2.json` reports the outcome, the seconds per act, the load, prompt and reset times, and the served
process's peak GPU and host memory, against the 10-minute budget. `--control no_icl` and
`--control empty_icl_text` serve decision Q4's two controls, for operators only. A T2 run is a smoke
run: its outcome tells right from broken, never a success rate.

A rerun of the same `--out` keeps the unit and rebuilds nothing, so the unit in place must be the one
this run asks for: a `--axis`, `--task`, `--task-config` or `--seed` that does not match the one it
was built with is a usage error, refused before anything is served and before the last run's report
is touched. `--dry-run` prints the three commands; its `--policy-seed` reads `<unit scene_seed>`,
because the seed the episode runs at is the unit's own and no plan knows it before `demo make` picks
one.

Recorded on 2026-09-18 (#18). G1, the acceptance test, has since passed in fast mode
(robotensor/horizon-competition#49), so these are measurements for their fast-mode scope, and only
for it: one task (`click_bell`), two scene seeds per served axis (the first of blocks 0 and 1), the
reduced knobs below, one card. They say nothing of a success rate, of the released knobs (50/50,
whose timings are in horizon-competition's plan §10.1), of other tasks or of a cold load.

Everything ran with the base checkpoint `zero-wam-posttrain-robotwin` on one RTX PRO 6000 Blackwell
(97,887 MiB, driver 580.178.04). The knobs were `num_inference_steps=10` and
`action_num_inference_steps=10` (defaults 50), with `icl_guidance_scale=5` left at its default. The
commits were runtime 74a4691 (main 1eaa6c6 plus this script), RoboTwin fork 7a5d8a6 and
horizon-protocol bdd94a1.

| axis | scene seed | outcome | acts | s per act (mean, max) | load s | serve + eval s | total s |
|---|---|---|---|---|---|---|---|
| robotwin_sim | 100000 | success, green | 3 | 3.80, 4.39 | 6.5 | 58.0 | 118.8 |
| robotwin_sim | 200000 | success, green | 3 | 4.26, 5.06 | 7.8 | 60.2 | 95.0 |
| robotwin_humangen | 100000 | success, green | 3 | 4.24, 4.88 | 6.3 | 54.4 | 82.9 |
| robotwin_humangen | 200000 | success, green | 5 | 4.07, 5.02 | 6.4 | 68.3 | 96.1 |

On every run the served process peaked at 47.2 GB allocated and 47.3 GB reserved on the GPU, with
6.1 to 6.3 GB of host RSS. The whole card, with the server and the simulator together, peaked at
54,948 MiB (`nvidia-smi` sampled every 2 s). The weights were already in the page cache, so a cold
load takes longer. `total` includes `demo make` (28 to 61 s), which the 600 s budget leaves out.
robocasa_sim was not run: decision Q2 keeps it off real-model epochs until a checkpoint is
fine-tuned for `PandaOmron` and passes T2 there. Since G7 the runtime can *serve* that embodiment -
the `robocasa` configuration, the true scalar-last reading, its own channels, statistics and
cadence are all in the table - but the competition's base checkpoint was post-trained on RoboTwin,
so serving it on RoboCasa would score 0 on every unit. Serving a second embodiment needs
`--policy-arg embodiment=PandaOmron`: `hello` declares one `observe_every` before any demonstration
names a robot, and RoboCasa's is 2 where RoboTwin's is 4.

## Development

```bash
uv venv --python 3.10 .venv && uv pip install -e ".[dev]" -e ../horizon-protocol
ruff check . && ruff format --check .
pytest -q
```

Licensed under [Apache-2.0](LICENSE). Zero-WAM is Apache-2.0 from
[robbyant-research](https://github.com/robbyant-research/Zero-WAM).
