TRACKMANIARL / SYSTEM GUIDE / 06 Learn from a complete replay history Episode-local windows preserve real context. Masks and n-step horizons determine valid learning. 01 One contiguous history, t0 ... tL-1 02 Construct targets and train Burn-in [0, B) Rebuild state.No gradient. Learning positions B ... L-n-1Train only valid positions. Final anchor tL-1Own n-step target. Eligible sample One episode + IDsComplete n-step horizon N-step target Rewards + bootstrapDouble-Q evaluation Masked value loss Valid positions onlyAbsolute TD errors Sequence priority Assign to the final anchor ID. Next sampling distribution Use priorities and optional mixtures.Correct loss with normalized IS weights. BOOTSTRAP BOUNDARYA true terminal uses zero bootstrap. Truncationfollows the transition discount contract. SEQUENCE CONTRACTn_step < L and burn_in < L. Do not combinereplay sequences with stacked feature history.