TRACKMANIARL / SYSTEM GUIDE / 01
How off-policy training runs
Actors collect experience. Durable ingestion feeds a learner that publishes the next policy.
01 Collect
02 Persist
03 Learn
activate next
episode
Actor + game adapter
Episode policy
Features + rollouts
Authenticated ingest
Commit before ACK
SQLite WAL
Replay + learner
Sample + update
Advance WAL frontier
Policy snapshot
Portable model tensors
Checkpoint + logs
Replay + optimizer
RNG + WAL frontier
BEFORE COLLECTION
RunSpec 2.0 resolves components, seeds and contracts into an immutable run manifest.
DURABILITY
The actor spools unaccepted rollouts for retry.
WAL pruning waits until a durable checkpoint
covers the applied frontier.
SCOPE
This is the off-policy path. PPO uses a separate
local, on-policy lifecycle.
Download editable diagram