TRACKMANIARL / SYSTEM GUIDE / 03
From demonstrations to a driving policy
Prepare trustworthy data, fit offline, then require closed-loop evidence before an RL warm-start.
01 Prepare data
02 Fit offline
03 Prove in game
pass
Record + validate
Complete laps + timing
Dataset provenance
Split by lap
Train / validation
Then augment training
Behavior cloning
Encoder + temporal
Categorical action head
BC checkpoint
Held-out selection
Full trainer state
Closed-loop
benchmark
Measure finish rate,
pace and interventions.
Promotion gate
Warm-start RL
Encoder + temporal
New value head
IF THE POLICY FAILS
Collect DAgger recovery labels at student states, then repeat
preparation and offline training. Promotion criteria must be
explicit.
EXACT BC RESUME
Restore trainer state, RNG and the same
dataset split. Offline demonstrations bypass RL
replay and actor WAL.
TRANSFER BOUNDARY
A warm-start does not transfer the categorical
action head, optimizer state or episodic hidden
state.
Download editable diagram