TRACKMANIARL / SYSTEM GUIDE / 03 From demonstrations to a driving policy Prepare trustworthy data, fit offline, then require closed-loop evidence before an RL warm-start. 01 Prepare data 02 Fit offline 03 Prove in game pass Record + validate Complete laps + timingDataset provenance Split by lap Train / validationThen augment training Behavior cloning Encoder + temporalCategorical action head BC checkpoint Held-out selectionFull trainer state Closed-loopbenchmark Measure finish rate,pace and interventions. Promotion gate Warm-start RL Encoder + temporalNew value head IF THE POLICY FAILSCollect DAgger recovery labels at student states, then repeatpreparation and offline training. Promotion criteria must beexplicit. EXACT BC RESUMERestore trainer state, RNG and the samedataset split. Offline demonstrations bypass RLreplay and actor WAL. TRANSFER BOUNDARYA warm-start does not transfer the categoricalaction head, optimizer state or episodic hiddenstate.