TRACKMANIARL / SYSTEM GUIDE / 06
Learn from a complete replay history
Episode-local windows preserve real context. Masks and n-step horizons determine valid learning.
01 One contiguous history, t0 ... tL-1
02 Construct targets and train
Burn-in [0, B)
Rebuild state.
No gradient.
Learning positions
B ... L-n-1
Train only valid positions.
Final anchor
tL-1
Own n-step target.
Eligible sample
One episode + IDs
Complete n-step horizon
N-step target
Rewards + bootstrap
Double-Q evaluation
Masked value loss
Valid positions only
Absolute TD errors
Sequence priority
Assign to the final anchor ID.
Next sampling distribution
Use priorities and optional mixtures.
Correct loss with normalized IS weights.
BOOTSTRAP BOUNDARY
A true terminal uses zero bootstrap. Truncation
follows the transition discount contract.
SEQUENCE CONTRACT
n_step < L and burn_in < L. Do not combine
replay sequences with stacked feature history.
Download editable diagram