Network
Circles are neurons, lines are weights. Column k shows layer k's outputs. The line from neuron i to neuron j is the weight wij: blue is positive, orange is negative, and thickness is |w|. Forward pass for neuron j: zj = Σi wij ai + bj, aj = f(zj).
gradients shows ∂L/∂wij = mean over the batch of δj·ai, the direction backprop pushes each weight (SGD moves it the opposite way). updates shows how much each weight changed since the previous snapshot. Node colour is the neuron's mean activation on the logged batch, or its % active (a ReLU that is never active is dead, ringed in red). Click a neuron to inspect it.
Large layers draw at most 32 evenly spaced neurons; all statistics still use every neuron.
Loss & metrics
Every scalar you log (loss=, metrics={...}) over training steps. The vertical line marks the step
selected on the timeline.
What to look for: training loss should fall steadily. Bumps and spikes mean the learning rate is too high. If the test/validation loss turns upward while the training loss keeps falling, the network is overfitting: it memorises the training set instead of generalising.
Insights
Automatic, plain-language diagnoses of what the numbers mean: vanishing or exploding gradients, dead or saturated units, a learning rate that is too high or too low, overfitting. Each fires once, at the step where it first held. Click one to jump there.
Gradient flow
Mean |∂L/∂W| for each layer at the selected step (log scale), from the input side to the output side. Backprop computes these from the output backwards: δℓ = (Wℓ+1 δℓ+1) ⊙ f′(zℓ), so each layer multiplies the signal by its weights and by f′.
Healthy: bars of similar height. Vanishing: bars shrink towards the input (a sigmoid's f′ ≤ 0.25 per layer). Exploding: bars grow towards the input.
Layer health
Per layer over time: the weight norm ‖W‖, the gradient norm ‖∇W‖ and the update ratio ‖ΔW‖/‖W‖ (how much the weights changed since the previous log, relative to their size). Below: histograms of W and ∇W at the selected step.
Rule of thumb: an update ratio around 10−3 is healthy. Around 10−1 the learning rate is too high; below 10−6 nothing is learning. A histogram of W that keeps widening means the weights are growing (consider weight decay).
Activations
For each hidden layer, one cell per drawn neuron: the top strip is its mean activation on the logged batch, the bottom strip the fraction of samples where it is active (ReLU: output > 0; tanh/sigmoid: not saturated).
A neuron with 0 % active passes no gradient back, so it has stopped learning (dead ReLU or
saturated unit). Activation stats need activations: they are read automatically from layers that cache
their input (e.g. Linear.x), or pass activations=[a1, a2, …].
Latent space
The features of the layer you pass as latent=Z, one dot per sample, coloured by labels=. With more than 2
dimensions they are projected onto their first two principal components (PCA); the axis titles show how much variance each keeps.
Scrub the timeline to watch training reshape the representation: the classes usually start mixed and end up (linearly) separable, because the output layer is linear in these features.
latent=Z, labels=y to scope.log(...) to see this view.