alg-stable-log-softmax
llm-logsumexp-logsoftmax-from-scratch
llm-pagedattention-block-table-gather
llm-per-channel-vs-per-tensor-error
llm-round-trip-attention-stages-s-p-o
llm-scale-invariance-weights-s-activations-s
llm-sliding-window-mask-width-w-mistral
llm-weight-tying-identity-check
num-attainable-flop-s-from-roofline
num-blocked-tiled-matmul-correctness
num-catch-a-broken-analytic-gradient
num-error-growth-naive-vs-kahan-vs-pairwise
num-logsigmoid-gradient
num-sequential-inclusive-scan
num-shared-memory-tiled-cuda-matmul
num-vjp-of-matmul
rwa-classify-per-step-prefill-vs-decode-token-allocation
rwa-measure-blocks-slot-mapping-and-internal-fragmentation
rwa-measure-the-o-n-memory-of-the-mem-efficient-path
rwc-effective-bits-weight-including-scale-bias-overhead
rwc-fix-a-too-narrow-dynamic-shape-range-that-fails-at-runtime
rwq-math-equivalent-smoothing-transform
rws-select-keep-sets-for-a-target-width-depth
sys-apply-causal-mask-to-dense-scores
sys-emulated-triton-tiled-matmul
sys-grid-block-coverage-of-n
sys-ring-all-gather-correctness
sys-tree-attention-mask-construction
