Metadata-Version: 2.5
Name: pichak
Version: 0.3.0
Summary: Derive fine-tuning hyperparameters from measurements of your actual machine, and print the measurement behind every number.
Project-URL: Homepage, https://github.com/olesxg/pichak
Project-URL: Findings, https://github.com/olesxg/flap-findings
Author-email: Oleksandr Pichak <stasselust@gmail.com>
License: MIT
License-File: LICENSE
Keywords: batch-size,benchmark,fine-tuning,gpu,hyperparameters,learning-rate,llm,lora,vram
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.9
Provides-Extra: dev
Requires-Dist: pytest>=7; extra == 'dev'
Provides-Extra: hf
Requires-Dist: safetensors>=0.4; extra == 'hf'
Requires-Dist: torch>=2.0; extra == 'hf'
Requires-Dist: transformers>=4.40; extra == 'hf'
Provides-Extra: torch
Requires-Dist: torch>=2.0; extra == 'torch'
Description-Content-Type: text/markdown

# pichak

**Fine-tuning hyperparameters derived from your machine and your task, with the
measurement printed next to every number.**

```bash
pip install pichak
```

```python
from pichak import derive
from pichak.tune import (LookaheadSchedule, ModelGradFn, TrainingStep, forbid_spill,
                         lookahead_from)

forbid_spill("cuda:0")          # past the room this process has: OOM, never paging —
                                # and a watch on every forward and optimizer step
opt = torch.optim.AdamW(trained_params, lr=1e-4)      # fresh; pichak sets the lr

plan = derive(
    data="train.jsonl", tokenizer=tok,      # or rows=[...] already tokenized, in memory
    # the loss of b rows at the length the run trains at
    step_fn=TrainingStep(lambda b, length: model(**rows(b, pad_to=length)).loss, opt),
    grad_fn=ModelGradFn(model, lambda i: fresh_micro_batch(i)),
    named_weights=trained_weights,
    trained_as="factors",                   # LoRA/PiSSA; "weights" for anything trained
)                                           # as a weight: full FT, grown neurons, a head
print(plan.report())

# THE RUN: the plan's rate, checked every look-ahead horizon on held-out rows and
# rolled back when it gets worse; the ending decided at the end. Do not skip it —
# on AG News the plan alone made the model worse than untrained (E11 below).
sched = LookaheadSchedule(trained_params, opt, train_step, held_out_losses_per_row,
                          lookahead_from(plan), total_steps=steps)
for step in range(steps):
    sched.before(step)
    opt.zero_grad(set_to_none=True)
    train_step(step)                        # plan.grad_accum micro-batches of plan.micro_batch
    opt.step()
sched.finish()
```

What it did, on the task that broke the plan alone — `validation/e11_accuracy.py`:
AG News topic classification, SmolLM2-135M, LoRA rank 16, 16,000 rows, GTX 1060
6GB, accuracy on 2,000 test articles:

```
  untrained model                                        0.6565
  0.2.0's plan                                           0.3065
  0.3.0's plan, alone                                    0.2500   answers "Sports" to everything
  0.3.0's plan + LookaheadSchedule (3 rollbacks)         0.9255   vs untrained z = +21.4
```

The plan alone learns fast — 0.88 at step 200 — then a loss spike at 5 rows a
step collapses the run to one class and the rate is too high to climb out
(`validation/e11d_diagnose.py`). The schedule saw each spike at its next check,
put the best check back and halved the rate. Shared GPU memory stayed at 58 MB,
where any idle CUDA process sits, through every run on this page.

---

## Why not just use a good default

Because you cannot tell a good default from a bad one when it fails.

`micro_batch=4` tells you nothing about whether 4 came from a measurement on your
card, a heuristic in a blog post, or a number someone picked in 2023 for a
different model. So when it OOMs at step 40 you bisect instead of reading.

Every number `pichak` returns carries the sentence that produced it. Anything it
could **not** measure is collected under `plan.constants()` and printed together,
so a constant can never quietly pass for a derivation.

```python
plan.micro_batch          # 3
plan.why("micro_batch")   # the ramp, rung by rung, and which one failed
plan.constants()          # {} — nothing unmeasured, unless you pass your own number
```

## Tokens per step: measured, not 65,536

0.1.0 aimed every optimizer step at 65,536 tokens and said honestly that it was a
guess at "the scale where gradient noise stops dominating". That scale has a name
and a formula — the gradient noise scale `B_noise = tr(Σ)/|G|²` (McCandlish et al.,
2018) — and it depends on the task. Measured on SmolLM2-135M, two independent draws
of rows each:

| task | loss scores | per row | B_noise draw 1 | draw 2 |
|---|---|---|---|---|
| MetaMathQA | the completion | 166 | 4,171 | 4,390 |
| Python code | every token | 383 | 1,557 | 1,543 |
| English web text | every token | 383 | 2,125 | — |
| Ukrainian Wikipedia | every token | 383 | 865 | 814 |
| True/False deduction | the completion | 2 | 1 | 1 |

Draws agree within 1–7%. Tasks differ by **~4,700x**. At 65,536 raw tokens a step,
0.1.0 spent 11x the noise scale per step on math and ~1,300x on the deduction task.

**A coarse estimate is not a license.** When the probe's largest batch (`K × b`)
is far below the scale it reports, the estimate leans on a difference barely
above its own noise. On a Ukrainian-Wikipedia stage the probe saw 1% of its
answer, and 0.2.0 cut the learning rate 15x on it. Now the probe extends itself
while it is coarse, and whatever is left coarse cuts the rate only as far as the
probe reached (`NoiseScale.confirmed_tokens`).

**It drifts, so it is also measured inside the loop.** `NoiseMonitor` hooks the
trained parameters and reads the scale from the micro-batches an accumulating
loop computes anyway — no extra pass, no extra copy — and `suggest()` returns the
`grad_accum` and learning rate for the scale as it is now.

**How many micro-batches the probe takes is not a setting either.** It draws them
until it has seen the scale it reports — |G|² above zero by one standard error,
and its largest batch covering B_noise at the upper edge of that error — or until
the corpus runs out. A task with a knee at 4,000 tokens settles after about one
knee's worth; a noisier one takes as many as it needs. 0.2.0 took 16.

What makes it universal rather than tuned: **scored tokens, not tokens read.** A
row of 900 prompt tokens and a one-token label is one scored token.
`profile_corpus` counts what the loss scores at your `seq_len`, including what
truncation cuts from the end of the completion.

**B_noise is the knee, not the target.** At a batch B, reaching a given loss takes
`S_min(1 + B_noise/B)` steps and `E_min(1 + B/B_noise)` tokens (McCandlish et
al.); at B_noise both are twice their minimum. Which to spend depends on their
price, and on one card that accumulates, a token costs the same in any batch
while a step adds a fixed cost — the optimizer update, a per-step penalty. The
least time to a given loss is then at

```
B* = sqrt(B_noise × tokens per micro-batch × t_step / t_micro)
```

four measured numbers: the noise scale from the probe, the micro-batch from the
ramp, and the two times, which `TrainingStep` takes apart during the ramp. When
the optimizer step is cheap, B* is well below the knee, and aiming at the knee
would spend twice the tokens for a cleaner gradient.

## During the run

**Measured on a real model, this lost — read this first.** E6 on SmolLM2-135M, MetaMathQA, one 540k-token budget each
(validation/e6_online_vs_plan.py): batch at the knee 0.5248 in 699s; the
time-optimal B* 0.5233 in 632s; this Controller 0.5706 in 647s; cosine over the
known length 0.5436. The noise scale did grow as measured (~4k -> ~30k tokens),
and cutting the rate by sqrt(B/B_noise) with it took the rate to ~3e-5 — too low:
E5 found the best constant rate the same at 30, 60 and 90 steps, so there was
nothing to decay. The rule that held on toys does not hold here. The monitor cost
7% per micro-batch, not the <1% estimated. By the rule written before the run,
the Controller stays opt-in; the plan (B* and a constant anchor) is the default.

What follows is what was built and why it looked right on toys.

The noise scale grew ~5x in 60 steps on the 135M, and the best constant rate
for one batch moved 4x between a 20-step and a 200-step toy stage. Both are the
same fact: as the loss falls the gradient's signal shrinks against its noise. A
schedule written in advance guesses how fast; `Controller` measures it:

```python
ctl = Controller(trained_params, opt, anchor_lr=anchor, tokens_per_micro=m,
                 grad_accum=plan_accum, t_micro=t_micro, t_step=t_step)
for step in range(steps):
    for _ in range(ctl.grad_accum):
        (loss_of(next_micro_batch()) / ctl.grad_accum).backward()
        ctl.micro(scored_tokens)
    ctl.step()                          # before opt.step(): it sets the rate
    opt.step(); opt.zero_grad(set_to_none=True)
```

From the micro-batches the loop computes anyway — a hook reads each gradient's
norm as the backward writes it, one synchronisation per step — it pools B_noise
over the last ~128 micro-batches, moves `grad_accum` toward B* when the move is
larger than the estimate's own error, and sets the rate to
`anchor × min(1, sqrt(B / B_noise))` with the noise as it is now. On toy
regressions, one anchor that never learns the stage length stayed within 1.19x /
1.00x / 1.32x of the best constant rate swept for 20 / 60 / 200 steps; the rate
best at 20 steps, kept, was up to 10x worse at 200. A cosine swept per length was
better at 200 steps (0.66x): this replaces the guess, not a sweep. The real-model
test is `validation/e6_online_vs_plan.py`.

## The learning rate

An Adam update moves a parameter by roughly `lr`, whatever the gradient's scale. So
there is a learning rate at which one step rewrites a weight matrix instead of
nudging it: where `lr * sqrt(numel(W))` reaches `||W||_F`. Every sane learning rate
is a fraction of that.

**Adapters** (LoRA/PiSSA factors), measured at rank 32. One thing to know first:
for classic LoRA, B starts at zero and has no size to measure, so the rate comes
from A — and A is PEFT's random init, `1/sqrt(3 × fan_in)` for the widest input.
That makes the LoRA rate a function of the architecture's width and the init
scheme, not of what the model learned; a different init (gaussian, olora, loftq)
moves it, and the report says so when it happens. PiSSA's factors are cut from
the weights and do not have this gap. E5 tests a rule that sizes the first step's
change to W against W itself.

```
rewrite/2    2.26e-3   destroyed the model twice (CE 1.74 -> 10.4, 1.11 -> 19.3)
rewrite/32   1.41e-4   the first scale that actually learned
rewrite/45             where a hand-tuned LoRA run at 2e-4 landed
```

**A full fine-tune** — the model's own weights — measured on SmolLM2-135M,
MetaMathQA, 60 steps at the noise-scale batch, held-out CE from 0.793:

```
rewrite/32    1.39e-3   -> 0.733    the adapter anchor: 3.7x less learned
rewrite/320   1.39e-4   -> 0.568    best
rewrite/1483  3.00e-5   -> 0.605
rewrite/3200  1.39e-5   -> 0.639
```

The adapter anchor does not carry over: a LoRA step moves the factors, and W only
through their product, which starts at zero. Train W itself and the same rate is a
several-times bigger step. The question is not "is it a full fine-tune" but "is
it trained as a weight": grown neurons and a new head are, too. Pass
`trained_as="weights"` and pichak opens at `rewrite/320`.

**A smaller batch and the cut do not stack.** When the batch is below the noise
scale and the rate is cut by `sqrt(B/B_noise)`, that is not a double penalty:
below the knee a batch 1/16 the size at 1/4 the rate is the same run per token. Measured on Adam over noisy regressions at a fixed budget, the
small batch's own best rate was exactly the cut one on a short stage, and without
the cut it lost on short and long stages alike (`validation/toys.md`).

**What a divisor cannot know: the horizon.** In the same toy, the best rate for one
batch fell 4x between a 20-step and a 200-step stage. `/320` was measured at 60
steps. So the rate is measured instead, in two steps:

1. **alpha1** — `measure_first_step()` walks the trained tensors along Adam's own
   first update and finds where the held-out loss stops falling. It is the first
   safe step, and where the search starts: E5 put the best rate at 1.23x alpha1
   for a full fine-tune and 1.3-2.6x for LoRA.
2. **Running ahead** — `lookahead_lr()` runs the REAL optimizer from the same
   weights on the same rows for n steps at alpha1 x 2^j, reads the held-out loss,
   and doubles n until the best rate stops depending on it (or follows a steady
   trend, carried to the run's length). `derive()` does this at the plan's own
   batch, on the rows after the probe's, and the plan's rate is its answer.

E8, against E5's full sweep of six constant rates over 90 steps on the same model,
data and batch (`validation/e8_lookahead_vs_sweep.py`, rule set before the run:
within 2x of the sweep's best, and a 90-step CE no worse than its second best):

```
              look-ahead    sweep's best     CE at 90: look-ahead / best / 2nd   cost
full FT       1.25e-4       1.39e-4          0.5479 / 0.5474 / 0.5541             194 steps
LoRA r32      1.13e-3       1.59e-3          0.5494 / 0.5527 / 0.5565             194 steps
```

Both settled on their own — the same grid point at 8, 16 and 32 steps — for 194
steps each, where the sweep ran 540. The short horizons alone would have
misled: at 2 and 4 steps LoRA's best was 2.5e-3, twice the rate that won.

The divisors remain only where nothing could be run, listed under
`plan.constants()`. Where they do: the anchor is right at the noise scale, and
below it the step is cut by Adam's square-root rule (Malladi et al., 2022),
`lr = rewrite / divisor * min(1, sqrt(B / B_noise))`, never raised above the
anchor. The look-ahead needs no such rule: it runs at the batch the plan takes.

**The rate during the run — decided at the end.** A lower rate always wins over
a short horizon: it settles the noise it no longer adds and pays in progress only
later. So `LookaheadSchedule` asks at the horizon that matters: at step T - 2h it
copies the weights and the optimizer's state, the run goes on to the end at its
rate, and `finish()` reruns the same 2h steps from the copy with the rate falling
linearly to zero — then keeps the ending with the lower held-out loss. E9, on a
256-step full fine-tune, rule set before each arm ran, verdict on 512 held-out
rows no decision used (`validation/e9_schedule_and_warmup.py`):

```
                                                          final CE vs constant
constant look-ahead rate (with warmup)                    0.4972
greedy: halve when h steps at half beat h at the rate     +0.0053 ± 0.0022  FAIL
the ending, decided at the end (+64 steps)                -0.0084 ± 0.0019  helps
weight decay 1/(lr x steps) = 31 instead of 0             +2.96             (0 wins)
weight decay 10x weaker (a 10-run-length time constant)   +0.0086 ± 0.0009  (0 wins)
betas measured by running ahead, (0.8, 0.968) vs defaults -0.0011 ± 0.0008  (no effect)
warmup from alpha1 vs none                                -0.0001 ± 0.0001
```

(the last three from E13 and E9; on rank-32 LoRA the measured betas stayed the
defaults, and decay at a 10-run-length time constant lost +0.0020 ± 0.0002)

The greedy version halved twice, the second time on a gap of 0.0007 on 64 rows,
and led by 0.0055 at step 128 before ending behind. The Controller's
noise-driven decay lost the same way in E6. **Warmup** is derived — from alpha1
up to the rate over 1/(1 - beta1) steps. It changed nothing measurable on this
full fine-tune (rate 1.1x alpha1) or on rank-32 LoRA (rate 1.84x alpha1, E9b:
-0.0001 ± 0.0003): harmless, 10 steps, and it keeps the first step where alpha1
said it is safe.
**Weight decay** is 0 by the argument in `derive()`'s report: a decay that acts
within the run pulls fine-tuned weights toward zero — measured above — and one
that does not is 0.

## It catches the failure that does not raise

On Windows, a batch that exceeds VRAM does not throw. The driver pages the excess
to system memory and the step simply crawls. Measured on a GTX 1060 6GB with a
6-layer model at sequence 512:

| | micro_batch | peak | seconds/step |
|---|---|---|---|
| watching only for OOM | 14 | **18.24GB** on a 6GB card | 47.5 |
| watching per-sample time too | 3 | 5.05GB | **2.0** |

No margin, reserve or threshold is guessed:

- **`forbid_spill()` — a cap, not a check.** It limits the process to the dedicated
  VRAM it can have — free plus what it already holds — so an over-allocation raises
  at the line that caused it. The ramp sets it itself if nothing has. Measured: an
  uncapped probe on this card went to 5,999 of 6,144 MiB and put **1.1 GB** in
  shared memory, silently; capped, every run since held shared memory at **58 MB**,
  the 56 MB any idle CUDA process shows. With a cap, paging cannot happen, so the
  old timing test for it and its 1.6x threshold are gone.
- **Fits, or only squeezes in.** PyTorch's allocator, refused, empties its cache and
  retries before it raises. A batch that needed that retry is not kept, and the
  kept batch must run twice in a row with the cache not emptied between — the
  steady state every later step sees. That is what 0.2.0's 12% margin guessed at,
  read off the allocator's own counter.
- **The memory outside the allocator is measured.** The CUDA context grows as the
  first steps load kernels; the ramp reads how much free-plus-held shrank while it
  ran and lowers the cap by it. 0.2.0 held back a fixed 512 MB for this.
- **`SpillWatch` moves the cap while the run lasts — and `forbid_spill()` installs
  it on every optimizer step by itself.** A cap set once is not enough: 0.3.0's
  first look-ahead run called `forbid_spill()` and nothing else, was capped at all
  5.07 GB that were free, and lost 178 MB to shared memory when the allocator's
  cache reached the cap and the CUDA context loaded ~170 MB of kernels on top.
  Other programs grow too — 130 MB in 20 minutes one day, 560 MB during one ramp
  another — and on the machine these numbers come from, the desktop window
  manager (dwm) alone swung by ~400 MB, to 1.3 GB, while runs went on. So the
  watch reads the card before every outermost forward (once per micro-batch) and
  after every optimizer step, of any model and optimizer in the process (global
  hooks), and inside pichak's own loops, and moves the cap to what is free plus
  what the process holds. Reading only after the step left the cap a whole step
  stale: dwm gave memory back mid-step, the card had 284 MB free, and a
  micro-batch that fit the card raised OOM at a cap read while dwm was large.
- **It reads the spill itself.** `pichak.measure.wddm` reads Windows' own
  counters (0.02 ms a reading). Measured on the 1060: an allocation past the
  card's edge raises the process's shared usage and lowers its dedicated usage
  minus the allocator's reserved bytes by the same amount (1,204 MB spilled,
  1,206 MB more shared); pinned memory raises shared alone. The watch counts only
  what moves both — and, on a card at its edge, the CARD's shared memory growing
  by more than the process's own, which is this process pushing another program
  out (what Task Manager's "Shared GPU memory" shows). Either way the cap comes
  down by what was read and the allocator's cache is emptied, so the next step
  lands in dedicated memory. `watched("cuda:0").peak_spilled_bytes` and
  `.peak_pushed_bytes` are the numbers a run can quote.
- **No margin.** A card read at 0 MB free did not push others into shared
  memory on this machine — Windows trimmed them instead (the desktop window
  manager gave back 55 MB) and the card's shared memory stayed put. A margin of
  "the most memory outside the allocator grew between two readings" was tried
  and cost a run: the first step's one-off kernels made it ~400 MB, and it held
  the cap 400 MB under a card with 522 MB free until a step raised OOM.

## The batch a training run fits, not a single step

A ramp is only as right as the step it runs. On a 135M full fine-tune with a
4.57 GB cap, three hand-written steps each cleared 3 rows that then OOMed at 3, for
two reasons:

| the step left out | ramp said | the run needed |
|---|---|---|
| rows at the run's longest sequence (ramped at 268 tokens, trained at 384) | 3 fit | OOM |
| the gradient buffers an accumulating backward runs on top of | 3 at 3.95 GB | OOM |
| — (`TrainingStep`: all of it) | 3 at **4.57 of 4.57 GB**, kept 2 | 60 steps, no OOM |

`TrainingStep(loss_of, optimizer)` pads nothing for you — `loss_of(b)` must build
`b` rows at your longest sequence — but it puts the optimizer's state and the
gradient buffers in the peak, steps at learning rate 0 so the weights do not move,
and hands the optimizer back fresh.

Put in `loss_of` only what the real step backpropagates in the same pass. An EWC
penalty that the run backpropagates once per step, after the micro-batches, went
into `loss_of` in one run and the ramp said not even one row fits — that peak
never happens in training. Pass it as `per_step_loss=`, and the tensors that only
sit on the card (the Fisher, the anchor) as `resident=`.

## What it derives, and from what

| | measurement |
|---|---|
| `seq_len` | the longest row, read from every row — nothing cut (a caller may choose `quantile=` / `buckets=`, and the report says what that costs) |
| `loss_policy` | whether the rows have a separable completion to mask against |
| `micro_batch` | real training steps, doubling then bisecting to the exact edge; kept only if it runs twice without the allocator emptying its cache |
| `seconds_per_step` | the wall time of that rung |
| `noise_scale_tokens` | the gradient noise scale of your task, from `grad_fn` |
| `grad_accum` | rows per step to reach the time-optimal batch B* in *scored* tokens (the noise scale without timings), capped by the corpus |
| `micro_tokens` | micro_batch x seq_len: the padded-token budget a token-packing loader (`pack_steps`) fills |
| `lora_rank` | the largest rank whose adapter and optimizer state fit, from the optimizer's own per-element copies — no ladder. (`rank_probe=` adds `signal_rank`, the directions of the first gradient above its noise, as a description only: E10 trained ranks 2x and 4x above it better) |
| `first_step` | alpha1: where the held-out loss stops falling along Adam's first update |
| `learning_rate` | running ahead from alpha1 with the real optimizer at the plan's batch (`lookahead_horizon`); without a `grad_fn`, the tensors' norms over a divisor, listed as a constant |
| `warmup_steps` | from alpha1 up to the rate over 1/(1 - beta1) steps, beta1 as measured |
| `betas` | running ahead: each beta's averaging window and its neighbours at x2 and /2, the real optimizer from the same start; moved only on a win of more than 2 SE, paired over the held-out micro-batches |
| `weight_decay` | 0, and why: a decay that acts within the run pulls fine-tuned weights toward zero |

Everything is optional. It will not invent a number it could not measure — a plan
with three derived values and an honest gap is more useful than one with ten you
cannot tell apart.

## What it does not know yet

Said here so nobody finds it the hard way:

- **The noise scale is measured once, at the start, and it grows fast.** Same
  135M full fine-tune on MetaMathQA, three runs, re-measured on held-out rows with
  64 micro-batches of the run's own size:

  | step | 3 rows a micro-batch | 2 rows | 1 row |
  |---|---|---|---|
  | 0 | 4,280 | 4,280 | 4,280 |
  | 30 | 13,258 (3.1x) | 20,212 (4.7x) | 20,009 (4.7x) |
  | 60 | 22,527 (5.3x) | 18,945 (4.4x) | 22,479 (5.3x) |

  About 5x in 60 steps, in all three runs. So a batch planned at step 0 is a fifth
  of the noise scale by step 60, and the gradient drifts back into the
  noise-dominated regime the plan was meant to leave. `NoiseMonitor` and
  `Controller` follow it from inside the loop at no extra cost; the plan at step 0
  does not.

  These are recomputed (`validation/recompute_uneven_tokens.py`). 0.2.0 published
  3.4x/5.9x, 5.4x/5.0x and 8.4x/10.1x, and blamed the last on a coarse probe. A
  review found the real cause: the estimator assumed every micro-batch scores the
  same number of tokens, and answers differ in length — most with one row per
  micro-batch. Rebuilt from the same rows in the same order, with the corrected
  estimator, the three runs agree.
- **The noise scale is Adam-blind.** It is measured on the raw gradient, as
  McCandlish et al. define it for SGD; for Adam they suggest measuring the
  gradient preconditioned by the optimizer's second moment. The measurements here
  are consistent with the raw one, but a model whose gradient scales differ wildly
  between layers could read differently under preconditioning.
- **B_noise at step 0 favours the easiest direction.** On the True/False task it is
  1 token: every row agrees on the answer format and the end-of-sequence token
  before any reasoning is learned. A batch sized from that is right only until the
  format is learned.
- **The look-ahead has passed on one model.** E8 matched a full sweep for a full
  fine-tune and for LoRA on SmolLM2-135M / MetaMathQA, at 36% of the sweep's cost.
  A bigger model, another task, a longer run are each untested. It costs real
  steps — 194 here, about two run-lengths of 90 — which on a short run is most of
  the run; on a long one it is a rounding error. The divisors are still single
  measurements (one 135M, one task, 60 steps) and remain only where nothing could
  be run, listed under `plan.constants()`.
- **The time-optimal batch held on a real model; the Controller did not.** E6:
  B* matched the knee's loss (0.5233 vs 0.5248) in 10% less time; the Controller's
  noise-driven decay ended at 0.5706. One model, one task, one budget.
- **The cap defends this process, not the card.** SpillWatch follows the room other
  programs leave from one micro-batch to the next, but memory taken inside one
  micro-batch can still make Windows demote something. The watch reads that from
  Windows' own counters at the next reading and pulls the process back into
  dedicated memory. On the runs above it read 0 MB, over 19,367 readings in one of
  them; one earlier run, reading only after each optimizer step, had its own
  memory demoted by up to 570 MB during evaluation loops, which led to the
  per-forward reading.
- **The batch comes from the noise at the START.** On AG News the first gradient
  is all one direction — learn the answer's format — and the probe read a noise
  scale of 2 tokens, so steps of 5 rows. At the same rate, 16x the rows ran
  without a spike and reached 0.892 (`validation/e11d_diagnose.py`). The schedule's
  rollbacks rescue the small-batch run (0.9255); the batch itself is not re-derived
  as the noise grows.
- **The held-out rows for the look-ahead are as many as the probe took.** Where
  the probe stops early (2 micro-batches on AG News, 16 rows), the rate and the
  betas are chosen on few rows. The betas' walk moves only on a 2-SE win, so it
  stays put there; the rate's walk has no such test.
- **A plan on the card's edge is fragile.** A full fine-tune of the 135M fit one
  row of ~890 tokens; LoRA one row of 1,292 — at 4.81 of 4.81 GB. Minutes later,
  other memory had moved or the allocator's cache had fragmented over changing
  lengths, and a step raised OOM. `derive()` now holds the kept batch to the least
  room the card offered while it ran; a run can still meet less room later.
- **`signal_rank` is not the rank to train at** (E10): ranks 2x and 4x above it
  trained better. The rank is what the memory holds.
- **Padding moves the gradient.** Packing is exact to 1e-8 as arithmetic, but a
  row padded in a batch gives a gradient 1e-5 to 1e-4 of the norm away from the
  same row alone (E12) — the attention kernel a padded batch takes. Any loader
  that pads does this. Packing (`pack_steps`) trained the same rows 1.85x faster
  than one row per micro-batch on MetaMathQA LoRA, with no spill.

## Measurement tools, usable on their own

```bash
pichak gpu                    # virtualisation, launch latency, transfer bandwidth
pichak disk "models/*.safetensors"   # queue-depth sweep, page cache bypassed
pichak corpus train.jsonl mistralai/Mistral-Small-24B-Instruct-2501
```

```python
from pichak.measure import disk_queue_depth, transfer_bandwidth, virtualisation
```

These exist because the numbers people quote are rarely the numbers their machine
gives:

- A buffered disk benchmark reported **4079 MB/s on a drive rated 2100** — that was
  the page cache. Unbuffered, the same drive peaks at **queue depth 3** and gets
  *slower* past it.
- A rented A6000 measured **D2H 457 MB/s against H2D 1502**, and pinned memory —
  the standard fix — bought nothing. Nothing about PCIe explains that; it was a
  fabric, and it decided where hidden states could live.
- One rented machine had **84.8us kernel launches** against a normal 2-5. A
  workload issuing 400,000 launches per step would have spent nine hours on
  overhead and looked slower than a 2016 card.

## Install

```bash
pip install pichak            # the plan, the disk sweep — no dependencies
pip install pichak[torch]     # the ramp, the noise scale, the learning rate, the GPU
pip install pichak[hf]        # + transformers, for the corpus CLI
```

Python 3.9+. The core has **no dependencies at all**; torch is only needed for the
parts that touch a GPU.

## Where the numbers come from

The 0.1.0 measurements were made while fine-tuning a 24B model on a 6GB GTX 1060 and
on a rented A6000 held down to the same 6GB; the raw logs, including the runs that
failed, are at [flap-findings](https://github.com/olesxg/flap-findings). The 0.2.0
measurements — noise scale across five tasks, the full-fine-tune divisor, the
drift, and the spill cap — are reproducible from `validation/` in this repository,
with their JSON output beside each script.

## Licence

MIT. By [Oleksandr Pichak](https://github.com/olesxg).
