Metadata-Version: 2.4
Name: experts4bit-qlora
Version: 0.47.0
Summary: Train and serve Mixture-of-Experts models that do not fit in VRAM: fused 4-bit experts, QLoRA, CPU/NVMe offload, and fast inference on consumer NVIDIA GPUs.
Author-email: Jordan Anderson <paul.jordan.anderson@gmail.com>
License-Expression: MIT
Project-URL: Homepage, https://cerinamroth.com/ml/experts4bit-qlora/
Project-URL: Documentation, https://cerinamroth.com/ml/experts4bit-qlora/
Project-URL: Source, https://github.com/pjordanandrsn/experts4bit-qlora
Project-URL: Issues, https://github.com/pjordanandrsn/experts4bit-qlora/issues
Project-URL: Changelog, https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/CHANGELOG.md
Project-URL: Release Notes, https://github.com/pjordanandrsn/experts4bit-qlora/releases/latest
Project-URL: Status, https://cerinamroth.com/ml/status/
Project-URL: Capabilities, https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/capabilities.json
Project-URL: Solutions, https://cerinamroth.com/ml/solutions/
Project-URL: Benchmarks, https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/claims.json
Project-URL: Kernel: grouped-nf4-gemm, https://pypi.org/project/grouped-nf4-gemm/
Project-URL: Upstream (bitsandbytes#1965), https://github.com/bitsandbytes-foundation/bitsandbytes/pull/1965
Project-URL: Tracking (bitsandbytes#1849), https://github.com/bitsandbytes-foundation/bitsandbytes/issues/1849
Keywords: mixture-of-experts,moe,qlora,lora,bitsandbytes,transformers,pytorch,quantization,nf4,mxfp4,int4,fp8,cuda,triton,gpu-offload,nvme,llm-inference,llm-training,fine-tuning,consumer-gpu
Classifier: Environment :: GPU :: NVIDIA CUDA
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: POSIX :: Linux
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Typing :: Typed
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
License-File: THIRD_PARTY_NOTICES.md
Requires-Dist: torch>=2.2
Requires-Dist: bitsandbytes>=0.43
Provides-Extra: fast
Requires-Dist: grouped-nf4-gemm>=0.30.0; extra == "fast"
Provides-Extra: train
Requires-Dist: transformers>=5.0; extra == "train"
Requires-Dist: datasets>=2.14; extra == "train"
Requires-Dist: accelerate>=0.30; extra == "train"
Requires-Dist: safetensors>=0.4; extra == "train"
Requires-Dist: huggingface_hub>=0.23; extra == "train"
Provides-Extra: serve
Requires-Dist: fastapi>=0.110; extra == "serve"
Requires-Dist: uvicorn>=0.29; extra == "serve"
Requires-Dist: pydantic>=2; extra == "serve"
Requires-Dist: transformers>=5.0; extra == "serve"
Requires-Dist: accelerate>=0.30; extra == "serve"
Requires-Dist: safetensors>=0.4; extra == "serve"
Requires-Dist: huggingface_hub>=0.23; extra == "serve"
Provides-Extra: test
Requires-Dist: pytest>=7; extra == "test"
Requires-Dist: transformers>=5.0; extra == "test"
Requires-Dist: accelerate>=0.30; extra == "test"
Requires-Dist: safetensors>=0.4; extra == "test"
Requires-Dist: fastapi>=0.110; extra == "test"
Requires-Dist: httpx>=0.27; extra == "test"
Requires-Dist: grouped-nf4-gemm>=0.30.0; extra == "test"
Requires-Dist: jsonschema>=4.18; extra == "test"
Requires-Dist: compressed-tensors==0.18.0; extra == "test"
Dynamic: license-file

# experts4bit-qlora

[![CI](https://github.com/pjordanandrsn/experts4bit-qlora/actions/workflows/ci.yml/badge.svg)](https://github.com/pjordanandrsn/experts4bit-qlora/actions/workflows/ci.yml)
[![PyPI](https://img.shields.io/pypi/v/experts4bit-qlora)](https://pypi.org/project/experts4bit-qlora/)

<!-- release-block:start -->
<!-- generated by `python scripts/check_readme_claims.py --write-release-block` from CHANGELOG.md's latest release heading; do not edit by hand -->
**Latest released package:** [`experts4bit-qlora` 0.47.0](https://pypi.org/project/experts4bit-qlora/0.47.0/) · **Current development status:** [`docs/STATUS.md`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/STATUS.md) on `main` (this README describes `main`) · **Released documentation for 0.47.0:** [`docs/`](https://github.com/pjordanandrsn/experts4bit-qlora/tree/v0.47.0/docs) · [`README.md`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/v0.47.0/README.md) · [`CHANGELOG.md`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/v0.47.0/CHANGELOG.md)
<!-- release-block:end -->

Train and serve **fused Mixture-of-Experts** models — MoEs whose experts
transformers v5 stores as one 3-D parameter per layer rather than as
per-expert `nn.Linear` — in 4-bit on hardware that cannot hold them in bf16.

**The problem in one line:** `load_in_4bit=True` leaves a fused MoE's
expert weights in bf16, so the model still OOMs; this package quantises
exactly those experts, fine-tunes them with QLoRA, keeps them in host RAM
or on NVMe when they do not fit, and serves them on one consumer NVIDIA
GPU. **Canonical package:** `experts4bit-qlora` on PyPI (`import
experts4bit_qlora`); the lookup aliases are listed under Install.
**Two repositories:** this one owns loading, quantisation orchestration,
adapters, training, residency (which tier each expert's bytes live in
while the model runs: VRAM, pinned host RAM or an NVMe arena) and serving;
the kernels it calls through the `[fast]` extra live in
[`grouped-nf4-gemm`](https://github.com/pjordanandrsn/grouped-nf4-gemm).
**Environment:** Linux, a CUDA GPU, torch ≥ 2.2 and bitsandbytes ≥ 0.43,
with transformers ≥ 5.0 for the streaming loader and trainer via `[train]`
(the floors are `pyproject.toml`'s; Python 3.11 is what CI tests; the
kernels need Triton on an sm_80+ GPU). **The material limitation:** on a
model that already fits in bf16, 4-bit here is a memory trade, not a
speed-up, and on the measured comparator it cost energy
(`e4b.train.energy-honest.scoped-a2000`) — this is for models that do not
fit. Machine-readable capabilities and evidence:
[`docs/capabilities.json`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/capabilities.json)
and [`docs/claims.json`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/claims.json).

transformers v5 stores a MoE's experts as one fused 3-D parameter per
layer. bitsandbytes' 4-bit walker only replaces `nn.Linear`, so it
**silently skips the experts** — the overwhelming majority of the weights
([bitsandbytes#1849](https://github.com/bitsandbytes-foundation/bitsandbytes/issues/1849)).
This package quantises exactly that fused stack (`Experts4bit`, the 4-bit
face of `ExpertsNbit`: nf4 / fp4 / int8 / fp8 / bf16 / fp16 storage, with
a test-pinned fidelity ordering), pairs it with a streaming loader and
per-expert LoRA so you can fine-tune, and serves the result through a
paged decode engine that is measured against each model's own attention.

**Current position, one page:** [`docs/STATUS.md`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/STATUS.md).
**Every number, with its evidence and status:** [`docs/claims.json`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/claims.json) — the claims register ('the register' below).
This README describes `main`: every link is to `main`, every number is the
current value of the claim it names, and the documentation as released with
a version is reached from the release block at the top.

## Use this when

- `load_in_4bit=True` / `BitsAndBytesConfig` loads your MoE but the
  expert tensors (`gate_up_proj`, `down_proj`) stay bf16 and the model
  still OOMs — the experts are fused 3-D parameters, not `nn.Linear`.
- You need QLoRA or LoRA on the experts themselves, and PEFT or the
  bitsandbytes walker never sees them.
- The quantised experts fit in host RAM but not VRAM (stream per layer),
  or fit on NVMe but not host RAM (serve or train from an *arena* — a baked,
  expert-row-addressable file on NVMe that `grouped-nf4-gemm` produces;
  [`docs/solutions/offload-moe-experts-to-cpu-or-nvme.md`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/solutions/offload-moe-experts-to-cpu-or-nvme.md)).
- You want to train a 30B-class MoE on a 12–24 GB consumer GPU (expert
  offload; `e4b.offload.fits-30b-class`), or serve one on an RTX 5090 —
  the only card the serving claims are measured on
  (`e4b.serve.census.bo7.*` is the current position; `e4b.serve.buildout.*`
  are the build-out lanes behind it).
- You are choosing between the reference per-expert path (the default
  `ExpertsNbit` forward — one expert at a time, no flag needed; "the
  reference loop" and "the per-expert loop" below name the same path), the
  batched path, the fused kernel path, host-streamed residency and the NVMe
  tier —
  [`docs/CHOOSING.md`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/CHOOSING.md) is the decision page.
- Your experts are NF4, native MXFP4 (gpt-oss, DeepSeek-V4), int4-b32 for
  serving, fp8 for the KV cache, or a mix across storage and residency
  tiers.

## Do not use this when

- The model is dense (no experts): bitsandbytes' own 4-bit path already
  covers every `nn.Linear`.
- The model already fits in bf16 with headroom: 4-bit is a memory trade
  there, and on the measured comparator it was slower and used more
  energy (`e4b.train.energy-honest.scoped-a2000`; the scope and the
  withdrawal note are under "What is measured").
- You expect a general-purpose serving engine or a vLLM replacement: on
  the same box, with identical prompt ids, vLLM 0.30.0 decodes 1.087× faster
  than this package's current int4 stack at B=1 and 1.396× at B=16
  (`e4b.serve.h2h.vllm-0.30.0.p58.qwen3.b1.5090.2026-09-22`,
  `e4b.serve.h2h.vllm-0.30.0.p58.qwen3.b16.5090.2026-09-22`; **bounded** to
  graph decode at B=1 and B=16 on one RTX 5090 box with one prompt set —
  never a general position; vLLM's number includes its serving loop and
  this package's does not, so the engine advantage is understated; quality
  is quoted, never equated; footprint and prefill/TTFT were not compared);
  this is a measured 4-bit path for models that otherwise do not run at
  all.
- You need Windows, macOS, ROCm or a non-CUDA accelerator.
- The model family or expert layout is not in
  [`docs/ARCHITECTURE_SUPPORT.md`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/ARCHITECTURE_SUPPORT.md)
  — unsupported architectures fail fast with a named error; the
  accelerated paths fall back to the reference loop, so assert every
  `enable_*` count (a `0` looks identical to the per-expert loop).

## Start here

| | |
|---|---|
| [`docs/SOLUTIONS.md`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/SOLUTIONS.md) | one page per problem: symptoms, cause, install, smallest example, verification, limits |
| [`docs/capabilities.json`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/capabilities.json) | the machine-readable capability contract (entry points, environments, limitations, claim IDs) |
| [`docs/STATUS.md`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/STATUS.md) | the current position — claims tiered in the public register: confirmed, measured, measured-private, open, superseded, retired |
| [`docs/claims.json`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/claims.json) | every number with its evidence and status |
| [`docs/INDEX.md`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/INDEX.md) | what each document is and whether it is current |
| [`grouped-nf4-gemm`](https://github.com/pjordanandrsn/grouped-nf4-gemm) | the kernel package this one drives (`pip install "experts4bit-qlora[fast]"`) |
| [PyPI: experts4bit-qlora](https://pypi.org/project/experts4bit-qlora/) | the canonical distribution |
| [`llms.txt`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/llms.txt) · [`AGENTS.md`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/AGENTS.md) | orientation for language models and coding agents |
| [cerinamroth.com/ml/experts4bit-qlora/](https://cerinamroth.com/ml/experts4bit-qlora/) | the routing page for this project (problem-first index, status, compatibility) |

## Install

```bash
pip install experts4bit-qlora           # primitive + adapters (torch + bitsandbytes)
pip install "experts4bit-qlora[train]"  # + the streaming MoE trainer
pip install "experts4bit-qlora[fast]"   # + the fused grouped-GEMM path (grouped-nf4-gemm)
```

`e4b`, `e4b-qlora`, `experts4bit`, `expertsnbit` and `experts-mxfp4` are
lookup aliases that install this package; always install and cite
`experts4bit-qlora`. Trusted publishing; every wheel carries a PEP 740
attestation. Runs on stock bitsandbytes; every feature has a
reference path. Building from source — `pip install --no-build-isolation`,
or any build outside pip's isolated build environment — needs setuptools ≥ 77
for the PEP 639 license metadata in `pyproject.toml`; an ordinary
`pip install` gets it automatically through build isolation. The `[fast]`
extra's floor on `grouped-nf4-gemm` is `pyproject.toml`'s and is not
repeated here; which version of this package needs which kernel release,
and why, is the `compatibility` record in
[`docs/system-manifest.json`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/system-manifest.json),
validated in CI against `pyproject.toml`. `[fast]` also brings the kernel
package's own torch floor, torch 2.8 or newer (grouped-nf4-gemm's
`pyproject.toml`; pre-releases accepted) — above this package's 2.2, so
check `python -c "import torch; print(torch.__version__)"` first on a CUDA
image whose torch you want to keep.

## Which door? Start from what does not fit

| what ran out | call | needs |
|---|---|---|
| you do not know yet — ask before loading | `estimate_qlora_footprint(describe_moe(id), QLoRASetup(...), tokens_per_microbatch=n)`, then `prepare_qlora_training(id, setup)` | `[train]` |
| you do not know yet — serving, ask before loading | `estimate_serve_footprint(describe_moe(id), ServeSetup(max_seqs=n, max_tokens_per_seq=t))` | `[fast]` |
| nothing — just train a fused MoE | `load_moe_4bit_streaming(...)` | `[train]` |
| each step is slow | `enable_fast_train(model, dgrad=True)` | `[fast]` |
| …and there is spare VRAM to trade | `E4B_MOE_KEEP_LAYERS=n` + `NF4_QLORA_COMPACT_DELTA=1`, then `enable_fast_train` | `[fast]` + grad ckpt |
| …and `[fast]` will not build | `enable_batched_train(model)` | — |
| the experts do not fit VRAM | `load_moe_4bit_streaming(..., offload=True)` | — |
| the experts do not fit host RAM, serving | `enable_nvme_residency(...)` | `[fast]` + arena |
| …and they are native MXFP4 | `enable_mxfp4_nvme_residency(...)` | `[fast]` + arena |
| the experts do not fit host RAM, training | `enable_nvme_train_residency(...)` | `[fast]` + arena + grad ckpt |
| the dense side does not fit | `enable_dense_offload(model, "cuda")` | — |
| serving, want it faster | `enable_fast(model)` | `[fast]` |
| serving, spare VRAM to trade | `enable_pipelined_residency(model, hot_sets, k_slots=k)` | `[fast]` |

An arena is baked by `grouped-nf4-gemm`, not by this package; the bake
tools, and how to bind an arena to a model, are on
[`docs/solutions/offload-moe-experts-to-cpu-or-nvme.md`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/solutions/offload-moe-experts-to-cpu-or-nvme.md).

Reasoning and caveats for each: [`docs/CHOOSING.md`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/CHOOSING.md).
**Assert the return value** of every `enable_*`: `0` and "silently still
on the per-expert loop" look identical from the caller's side.

## Quickstart

Needs the `[train]` extra: the streaming loader imports transformers ≥ 5.0,
so on the base install the import below raises `ModuleNotFoundError`. It
also needs a CUDA GPU, network access and the checkpoint download.
`verify_moe_4bit(..., strict=True)` returning silently is the success signal.

```python
import torch
from experts4bit_qlora import Experts4bit, ExpertsLoRA, load_moe_4bit_streaming, verify_moe_4bit

# A real fused-MoE checkpoint, quantised on the way to the GPU (never bf16-resident).
# As written the 4-bit experts stay resident, so the card must hold them; on a 12 GB
# card add offload=True (experts in pinned host RAM, one layer on the GPU at a time;
# `e4b.offload.fits-30b-class`: Qwen3-30B-A3B QLoRA-trains at a 7.16 GB peak).
model, config = load_moe_4bit_streaming(
    "Qwen/Qwen3-30B-A3B", "cuda", torch.bfloat16, r=8, alpha=16, quant_type="nf4",
)
verify_moe_4bit(model, strict=True)   # raises if any expert stack is still high precision
```

The shipped trainer and `infer` read their settings from environment
variables (`python -m experts4bit_qlora.train --help` lists them). Without
`MODEL=` both default to `allenai/OLMoE-1B-7B-0924` (the trainer on
`tatsu-lab/alpaca`); `R`/`ALPHA` must match the adapter and `QUANT_TYPE`
the training run. To train and serve the model loaded above:

```bash
MODEL=Qwen/Qwen3-30B-A3B OFFLOAD_EXPERTS=1 STEPS=150 R=8 TRAIN_EXPERTS=1 OUT=./out python -m experts4bit_qlora.train   # QLoRA fine-tune; writes ./out/adapter_best.pt
MODEL=Qwen/Qwen3-30B-A3B OFFLOAD_EXPERTS=1 ADAPTER=./out/adapter_best.pt python -m experts4bit_qlora.infer            # serve it
```

Do **not** load these models with stock `from_pretrained(...,
load_in_4bit=True)`: it quantises the `nn.Linear` layers, leaves the
experts in bf16, and OOMs.

## What is measured

Each row names its entries in `docs/claims.json`, which carry the value,
the conditions and the receipt path (the entry's `evidence`, a location
in this repository; `evidence_private` where the receipt is not public);
the last column is the status there.
**measured** means the receipt is in this repository; **measured-private**
means the run happened but the receipt lives in a private audit tree and
you cannot check it from here. Every number in the result column is the
named claim's *current* value: `scripts/check_readme_claims.py` holds this
table to the register in CI and fails on drift, on a superseded or retired
id, and on a private receipt presented as public — a number that moves in
the register moves here or the build goes red.

| | result | status |
|---|---|---|
| OLMoE-1B-7B fits a 12 GB card and trains (`e4b.train.olmoe-fits`, `e4b.train.olmoe-converges`) | 4.70 GB load; held-out eval 1.4813 → 1.0290 | measured |
| Expert offload trains 30B-class MoEs on 12 GB (`e4b.offload.fits-30b-class`) | Qwen3-30B-A3B peaks 7.16 GB, Gemma-4-26B-A4B 8.47 GB | measured |
| Fused training path, two 30B MoEs × five datasets (`e4b.train.flagship-matrix`) | 1.52–1.81× per step at 0.75–0.81× VRAM, loss parity, frozen stack bit-identical (12.85 GB hashed per Gemma-4 cell; 16.31 GB on Qwen3, by a separate gate) | measured |
| Training on real weights, per family, under the shipped code — the fused path vs the per-expert loop on one rented RTX 5090, verdicts in the registered units (`e4b.train.parity.tp1.granite.fused.2026-09-05`, `e4b.train.parity.tp1.olmoe.fused.2026-09-05`, `e4b.train.parity.tp1.qwen3.fused.2026-09-05`, `e4b.train.parity.tp1.gemma4.fused.2026-09-05`, `e4b.train.parity.tp1.mixtral.fused.2026-09-05`, `e4b.train.parity.tp1.granite.batched.2026-09-05`, `e4b.train.parity.tp1.mixtral.batched.2026-09-05`) | `enable_fast_train(dgrad=True)` PASS on every family that has one, \|Δ final train loss\| 0.01329 on Granite, 0.01327 on OLMoE, 0.01315 on Qwen3 (resident), 0.02385 on Gemma-4 (the `-it` checkpoint), 0.00953 on Mixtral (offload); `enable_batched_train` PASS 0.01553 on Granite and 0.00766 on Mixtral, VOID on OLMoE, Qwen3 and Gemma-4 (the kernel not reached on every layer); gpt-oss expert-LoRA REFUSED, its attention-only arm trains, its MXFP4 route experimental | measured |
| Against Unsloth's 4-bit MoE QLoRA path, end-to-end, one identical training problem on one rented RTX 5090 — Qwen3-30B-A3B, the fused `dgrad` path + NF4 attention vs Unsloth 2026.9.2 (`e4b.train.h2h.unsloth.qwen3.5090.2026-09-05`, `e4b.train.h2h.unsloth.qwen3.5090.2026-09-05.quality-n60`, `e4b.train.h2h.unsloth.qwen3.5090.2026-09-05.curve-n200`, `e4b.train.h2h.unsloth.qwen3.5090.2026-09-05.e4b-internal-parity`) | at 60 steps, s/step Unsloth/e4b 1.413 (2.151 vs 1.522 s — e4b faster per step at this workload), peak 21.371 vs 23.141 GB, 157.1 vs 224.7 J/step, time to a held-out loss of 0.32 92.5 vs 130.3 s, held-out comparable (0.2923 vs 0.2975, \|Δ\| 0.0052 ≤ 0.05); **at 200 steps Unsloth's held-out loss is lower, 0.2713 vs 0.2881 (Δ −0.017)** — quoted beside the position, causes not established; e4b fused vs its own reference PASS (0.00131 / 0.01138), 2.92× per step | measured |
| Arena vs pinned host RAM, at a descending cap (`e4b.offload.arena-vs-host-ram`) | 2.56× / 3.80× / 6.40× less host RAM (OLMoE / Gemma-4 / Qwen3-30B) | measured |
| Paged decode vs the model's own attention (`e4b.parity.*.paged-vs-own-attention`, `e4b.parity.gemma4.no-reference`, `e4b.parity.gemma4.fp8-share`) | indistinguishable on Granite (0.00229 nats), gpt-oss (0.00288) and Qwen3 (0.00173) against a chunk-free reference, each below its own floor; Gemma-4 has no reference at this resolution — its own cached forward swings −0.107 … +0.271 nats across windows — and the paged path's one measured cost there is the fp8 cache, 0.046 nats ([#359](https://github.com/pjordanandrsn/experts4bit-qlora/issues/359)) | measured-private |
| Serving: the licensed best per family under the shipped code, one rented RTX 5090, every ratio vs e4b's own NF4 control on the same box — never a field-engine speedup; Granite, OLMoE, gpt-oss, Gemma-4 and Mixtral have no field comparator measured; Qwen3's field comparator is named on the vLLM row below (`e4b.serve.census.bo7.*.b1.5090.2026-09-05`, `e4b.serve.census.bo7.*.b16.5090.2026-09-05`) | Qwen3-30B-A3B ×2.067 at B=1 (238.1 tok/s on that box; anchor-class projection 159.2 × 2.067 ≈ 329 tok/s, a projection) and ×2.602 at B=16 (1327.5 tok/s); Granite-3.1-3B ×1.341 (304.9) / ×1.160 (1836.8); Gemma-4-26B ×1.281 (103.6) / ×1.106 (675.8), exact arithmetic on NF4 because that family has no K8 instrument; OLMoE (282.5 / 1347.5), Mixtral (50.3 / 191.4) and gpt-oss (144.5 / 761.6) sit at ×1.000, their NF4 or reference arm — nothing above it is licensed; the measured-but-unlicensed arms are in [`SERVING-THROUGHPUT.md`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/SERVING-THROUGHPUT.md) | measured |
| Qwen3-30B-A3B's licensed serving stack passes the registered K8 gate on both texts: streamed 64k-token GPTQ-calibrated int4 experts + C4-calibrated int4 attention + round-1/2 folds + router epilogue + decode glue (`e4b.serve.buildout.bo6c.qwen3.all-calibexp-streamed-64k.k8.2026-09-05`) | −0.0528 ppl on wikitext and −0.0662 on C4 validation against the same-cut NF4, both inside the family's 0.0095-nat floor — at parity or better, licensed under the unchanged gate, no improvement claimed by a number | measured |
| Same box, same session, identical prompt token ids, against vLLM 0.30.0 (Qwen's GPTQ-Int4 via Marlin, CUDA graphs) on one rented RTX 5090 (`e4b.serve.h2h.vllm-0.30.0.p58.qwen3.b1.5090.2026-09-22`, `e4b.serve.h2h.vllm-0.30.0.p58.qwen3.b16.5090.2026-09-22`) | vLLM 260.3 tok/s at B=1 and 1925.6 aggregate at B=16 against this package's current int4 stack 239.4 / 1379.2 — vLLM / e4b-int4 1.087 and 1.396, vLLM ahead, bounded to graph decode on that box and prompt set (vLLM's number includes its serving loop and this package's does not; quality quoted, never equated; footprint and prefill/TTFT not compared); against the same box's NF4 control vLLM is 2.562 / 3.980 ahead. The 2026-09-05 comparison against vLLM 0.28.0 (`e4b.serve.h2h.vllm-0.28.0.qwen3.5090.2026-09-05`: 2.52 / 4.06 against that box's NF4 control) stays as history | measured |
| DeepSeek-V4-Flash (284B, 147 GB of experts on disk) (`e4b.serve.deepseek-v4`) | loads in ~10 s at 8.74 GiB peak VRAM and generates | measured |
| Informed hot sets vs by-index, identical VRAM (`e4b.serve.informed-hot-sets`) | +37.1% on DeepSeek-V4-Flash; the gain is a property of the host | measured |

Four things to read beside that table, because they change what it
means:

- **A parity delta is read against a per-model noise floor, never
  against zero.** Two arithmetically equivalent forwards of an MoE
  disagree, because rounding flips which experts the router picks; on
  gpt-oss 4.5% of layer-token choices flip and those tokens carry the
  whole disagreement. "Below the floor" means indistinguishable.
  [`docs/METHODOLOGY.md` §13.1](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/METHODOLOGY.md).
- **4-bit on a card that already fits the model was a 1.2–2.3× energy
  penalty on the measured comparator**, not a saving: one OLMoE-dims
  expert projection on an RTX A2000, dequantize-then-`linear` and a
  bitsandbytes 0.50-dev fork build's `matmul_4bit` routing against native
  bf16 (`e4b.train.energy-honest.scoped-a2000`). It inverts when memory
  binds. *Note, 2026-09-04:* the earlier wording "NF4 is storage-only and
  the GEMM runs in bf16 either way" was a universal mechanism statement
  and is withdrawn as such — bitsandbytes ≥ 0.50.0 can run supported
  ordinary 2-D 4-bit inference cells on the packed weights directly, while
  routed grouped MoE execution and training's input gradient are separate
  contracts ([`docs/BITSANDBYTES.md`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/BITSANDBYTES.md)).
  The measurement stands as its receipt made it.
- **A head-to-head is one workload on one box.** The Unsloth row above is
  ≈86 tokens per step at batch 1 on a resident 30B MoE; its 200-step curve
  favours Unsloth and is quoted beside the 60-step position wherever that
  position is quoted (`bench/h2h-20260905/p38/`, the pre-registration and
  every amendment in the bundle). The vLLM row is lane P58 (`bench/p58/`,
  2026-09-22): vLLM 0.30.0 against this package's current int4 stack —
  round-to-nearest int4 experts + uncalibrated int4 attention, not the
  licensed stack on the row above it — 1.087 at B=1 and 1.396 at B=16,
  and against the same box's NF4 control as the secondary ratio (2.562 /
  3.980); no lane quotes vLLM against the licensed stack. The 2026-09-05
  lane (`bench/h2h-20260905/p37/`, vLLM 0.28.0) could quote vLLM only
  against its box's NF4 control, the slowest, licence-free configuration
  this package ships (2.52 / 4.06): its licensed arms were void on that box
  (a pack fingerprint that did not reproduce, and the registered gate run
  on that pack failed its second text). That pack has since been rebuilt,
  gated on its bytes and hash-pinned, and P37's divergence is read as an
  outlier (`docs/STATUS.md`). The 2026-09-05 lane and the 2026-09-03
  comparison (×1.47 / ×1.55) it superseded stay as history.
- **Ratios travel; absolutes do not.** The 5090 class carries ~8.5%
  inter-box dispersion; the same config on two 4090s moved 8.6% in
  s/step. Quote the card, or quote a ratio.

## What was retired

Claims this project published and then withdrew, each with the
measurement that withdrew it, are listed in
[`docs/STATUS.md`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/STATUS.md#what-changed--retired-superseded-corrected)
and kept as `retired` entries in `docs/claims.json` so they stay
findable. The most recent: the "+0.047 ppl fp8 KV cost" on Qwen3 (below
the model's own floor), and the "+0.078 nats gpt-oss sinks/windows
defect" (the chunked oracle was the drifting arm, not the serving path).

## Scope

The primitives are model-agnostic. The streaming loader admits 14
fused-MoE `model_type`s (`SUPPORTED_ARCHITECTURES` in
`experts4bit_qlora/loader.py`; the list below groups aliases and
variants under one name — `qwen3_moe`/`qwen3_5_moe`,
`gemma4`/`gemma4_text`, `granitemoe`/`granitemoeshared`/`granitemoehybrid`),
stored per-expert or pre-fused: **OLMoE**, **Qwen3-MoE / Qwen3.5-MoE**,
**Gemma-4** (text tower), **gpt-oss** (MXFP4 experts with per-expert
biases and a clamped GLU, dequantised bit-identically), **GraniteMoe**
with its shared-expert and Mamba-hybrid variants (granite-4.0-h),
**Nemotron-H** (non-gated ReLU² experts), **LFM2-MoE**, **Jamba**,
**Kimi K3** and **DeepSeek-V4** (Flash / Pro). Mixtral is not in that
list: it loads through the loader's second admission route, the
read-compatible per-expert conventions (`READ_COMPATIBLE_CONVENTIONS`),
and Mixtral-8x7B-Instruct-v0.1 has a real-weight training receipt (the
tp1 rows in the table above). Ten of the 14 have a reference-tier
passing row on a real published checkpoint; which ones, and which load,
run and CUDA-graph-capture, are in
[`docs/ARCHITECTURE_SUPPORT.md`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/ARCHITECTURE_SUPPORT.md).
The five model_types (four conventions) that have a convention but no
loader path yet — `axk1`, `qwen3_vl_moe`, `qwen3_vl_moe_text`, `jetmoe`,
`dbrx`; `moe_conventions.STAGED_NOT_WIRED` is the authority — each have
the blocker they are waiting on recorded as a test in
[`tests/test_staged_blockers.py`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/tests/test_staged_blockers.py),
summarised in
[`experts4bit_qlora/README-LAYOUT.md`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/experts4bit_qlora/README-LAYOUT.md).
Unsupported architectures fail fast with a clear error. Which of those
families have a *training* receipt on real weights under the shipped
code, per path (quantize, reference, fused, batched, NVMe, native MXFP4),
and which enablers refuse: the tp1 section of the same document and
`training_support` in `docs/capabilities.json`. The capability's
`model_families` (`olmoe`, `qwen3_moe`, `gemma4_text`, `mixtral`,
`granitemoe`) is exactly the families whose fused path passes with a
receipt here. gpt-oss is refused; its experts train only through the
kernel package's experimental MXFP4 route.

Known open: Gemma-4-26B-A4B needs a parity instrument that survives its
batch-shape variance before any verdict is quoted for the family
([#359](https://github.com/pjordanandrsn/experts4bit-qlora/issues/359);
its other half — finer fp8 K-cache groups on the 512-dim heads — shipped
in 0.32.0 with grouped-nf4-gemm 0.26.0); the September load fault on
2 of 6 rented hosts is unreproduced
([#344](https://github.com/pjordanandrsn/experts4bit-qlora/issues/344));
reproducing the TR2 training receipt from published artifacts is still
open in the register (`e4b.open.tr2-repro-gap`), although grouped-nf4-gemm's
`nvme_bake_nf4` bakes the NF4 arena its trainer lane reads.

## Docs

`docs/STATUS.md`, `docs/claims.json` and `docs/INDEX.md` are in *Start
here* above. The reference documents:

| | |
|---|---|
| [`docs/CHOOSING.md`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/CHOOSING.md) | which mode, and why |
| [`docs/METHODOLOGY.md`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/METHODOLOGY.md) | hosts, protocols, every measurement's provenance |
| [`docs/SERVING-PARITY.md`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/SERVING-PARITY.md) | paged decode vs each model's own attention |
| [`docs/SERVING-THROUGHPUT.md`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/SERVING-THROUGHPUT.md) | per-family decode throughput under one protocol, with the refusal list |
| [`docs/STORAGE-MODES.md`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/STORAGE-MODES.md) | the six storage modes and what each promises |
| [`docs/RESIDENCY-ENGINES.md`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/RESIDENCY-ENGINES.md) | residency engines, hot-set selection, host-regime laws |
| [`docs/SERVING.md`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/SERVING.md) | the HTTP shim and Docker deployment |
| [`docs/DEEPSEEK-V4.md`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/DEEPSEEK-V4.md) | V4's storage split, epilogue, arena bake |
| [`docs/BITSANDBYTES.md`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/BITSANDBYTES.md) | relationship to bitsandbytes, prior art |

## The package family

- **`experts4bit-qlora`** (this repo) owns everything *around* the expert
  GEMM: the fused-stack primitives and per-expert LoRA, the streaming
  loaders, offload, training, the paged serving engine, hot-expert
  residency.
- **[`grouped-nf4-gemm`](https://pypi.org/project/grouped-nf4-gemm/)**
  owns the GEMM itself: one launch over 4-bit-packed expert stacks with
  in-register decode and fp32 accumulation, plus the fp8 paged decode
  attention and the decode glue kernels. `[fast]` is the seam.

The kernel makes one expert-stack matmul cheap; this package decides
which bytes are where.

## Provenance

Every number traces to a committed script and a named host, with
receipts under `bench/` and `docs/` — or, where the receipt is private,
the register says so. `PROVENANCE.md` is the OpenTimestamps-anchored
record for the v0.2.0 convergence result; anchored documents are never
edited in place (see `docs/INDEX.md`). Falsification work lives under
`audits/`.

## License

MIT ([`LICENSE`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/LICENSE)). `experts4bit_qlora/_vendor/experts.py`
is vendored from bitsandbytes (also MIT) pending upstream merge; its
notice is in [`THIRD_PARTY_NOTICES.md`](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/THIRD_PARTY_NOTICES.md).
