Metadata-Version: 2.4
Name: tt-bio
Version: 0.11.0
Summary: Boltz-2, ESMFold2, Protenix-v2, OpenFold3, and OpenDDE structure prediction, BoltzGen/RFdiffusion3 design, and ESMC/SaProt protein embeddings for inference on Tenstorrent Blackhole and Wormhole
Requires-Python: !=3.11.*,<3.13,>=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
License-File: NOTICE
Requires-Dist: torch
Requires-Dist: numpy
Requires-Dist: rdkit
Requires-Dist: requests
Requires-Dist: pandas
Requires-Dist: einops
Requires-Dist: mashumaro
Requires-Dist: modelcif
Requires-Dist: ihm
Requires-Dist: click
Requires-Dist: pyyaml
Requires-Dist: omegaconf
Requires-Dist: biopython
Requires-Dist: biotite<1.7
Requires-Dist: matplotlib
Requires-Dist: scipy
Requires-Dist: numba
Requires-Dist: gemmi
Requires-Dist: scikit-learn
Requires-Dist: chembl_structure_pipeline
Requires-Dist: hydride
Requires-Dist: pydssp
Requires-Dist: tqdm
Requires-Dist: rich
Requires-Dist: huggingface_hub<2.0,>=1.5.0
Requires-Dist: safetensors
Requires-Dist: transformers<6.0,>=5.5.0
Requires-Dist: zstd
Requires-Dist: msgpack-numpy
Requires-Dist: brotli
Requires-Dist: cloudpathlib
Requires-Dist: pydantic
Requires-Dist: pdbeccdutils
Requires-Dist: func_timeout
Requires-Dist: networkx
Requires-Dist: packaging
Requires-Dist: toolz
Requires-Dist: cytoolz
Requires-Dist: typing_extensions
Requires-Dist: zstandard
Requires-Dist: beartype
Requires-Dist: jaxtyping
Requires-Dist: pyarrow
Requires-Dist: opt_einsum
Provides-Extra: test
Requires-Dist: pytest; extra == "test"
Provides-Extra: reference
Requires-Dist: ml_collections; extra == "reference"
Provides-Extra: tenstorrent
Requires-Dist: ttnn==0.68.0; extra == "tenstorrent"
Provides-Extra: train
Requires-Dist: pytorch_lightning; extra == "train"
Requires-Dist: ml_collections; extra == "train"
Dynamic: license-file

```text
████████╗████████╗        ██████╗  ██╗  ██████╗
╚══██╔══╝╚══██╔══╝        ██╔══██╗ ██║ ██╔═══██╗
   ██║      ██║    █████╗ ██████╔╝ ██║ ██║   ██║
   ██║      ██║    ╚════╝ ██╔══██╗ ██║ ██║   ██║
   ██║      ██║           ██████╔╝ ██║ ╚██████╔╝
   ╚═╝      ╚═╝           ╚═════╝  ╚═╝  ╚═════╝
```

> [!IMPORTANT]
> **TT-Boltz is now TT-Bio**

TT-Bio runs biomolecular structure prediction, protein design and protein language models on
Tenstorrent Blackhole and Wormhole hardware, on one card or many (a QuietBox has 4 cards, a Galaxy
server 32).

Every model is validated against its official reference implementation on the same input and
reproduces it within that reference's own run-to-run noise. The methodology, per-target results and
reproduction commands are in [`docs/implementation-parity.md`](docs/implementation-parity.md).
Predictions and designs per hour per server, and throughput per dollar of purchase price and of
total cost of ownership against NVIDIA DGX H200, B200 and A100, are on
[tt-bio.com](https://tt-bio.com). The NVIDIA figures run each model's own upstream code, not
BioNeMo or Anthropic's life-sciences kit; an optimised serving stack might be faster and is not
measured. The [benchmark page](https://tt-bio.com/benchmarks/) has the measured seconds behind
every figure, the fixtures, the run conditions and the cost model.

Each runs as `tt-bio <command> --model <name>`:

| Model | What it does | Command | `--model` |
|---|---|---|---|
| [Boltz-2](https://github.com/jwohlwend/boltz) | Folds protein, DNA, RNA and ligand complexes; [binding affinity](#binding-affinity-prediction-boltz-2) | [`predict`](#structure-prediction) | `boltz2` |
| [ESMFold2](https://github.com/Biohub/esm) | Folds complexes without an MSA | [`predict`](#structure-prediction) | `esmfold2`, `esmfold2-fast` |
| [Protenix-v1, Protenix-v2](https://github.com/bytedance/Protenix) | AlphaFold3-family folding, with PAE/PDE output | [`predict`](#structure-prediction) | `protenix-v1`, `protenix-v2` |
| [OpenFold3](https://github.com/aqlaboratory/openfold-3) | AlphaFold3-family folding of protein, RNA and DNA | [`predict`](#structure-prediction) | `openfold3` |
| OpenBind-0 | OpenFold3 tuned for protein-ligand co-folding | [`predict`](#structure-prediction) | `openbind` |
| OpenDDE | Antibody-antigen co-folding | [`predict`](#structure-prediction) | `opendde`, `opendde-abag` |
| [RoseTTAFold3](https://github.com/RosettaCommons/foundry) | AlphaFold3-family folding, or refining a start structure | [`predict`](#structure-prediction) | `rf3` |
| AlphaFold2 initial guess | Scores a designed binder complex | [`predict`](#structure-prediction) | `af2ig` |
| [BoltzGen](https://github.com/HannesStark/boltzgen) | Protein, peptide, nanobody and antibody binder design | [`design`](#design) | `boltzgen` |
| [RFdiffusion3](https://www.biorxiv.org/content/10.1101/2025.09.18.676967) | All-atom design: binders, motif scaffolding, nucleic-acid binders | [`design`](#design) | `rfd3` |
| [PXDesign](https://github.com/bytedance/PXDesign) | Binder backbones against a target structure | [`design`](#design) | `pxdesign` |
| ESMC | Protein language model embeddings | [`embed`](#protein-embeddings-esmc) | `esmc-300m`, `esmc-600m`, `esmc-6b` |
| SaProt | Structure-aware embeddings, variant-effect scoring | [`saprot`](#structure-aware-protein-embeddings-saprot) | `saprot-35m`, `saprot-650m`, `saprot-1.3b` |
| Nesso-1 | Protein-ligand binding affinity without a structure | [`affinity`](#binding-affinity-without-a-structure-nesso-1) | `nesso1` |

Several machines each run their own controller, and a scheduler of your choice spreads the work
between them: [Running tt-bio on many machines](#running-tt-bio-on-many-machines).

## Installation

Create a Python virtual environment with Python 3.10 or 3.12, install with the Tenstorrent extra, then install the matching Tenstorrent system dependencies.

We test on Ubuntu 24.04 (glibc 2.39, Python 3.12). The `ttnn` wheel is tagged
`manylinux_2_34` and imports `GLIBC_2.34` symbols, so glibc 2.34 or newer is required:
RHEL 8 and its rebuilds ship glibc 2.28 and cannot install it. See
[Troubleshooting](#troubleshooting) for the CPU frequency driver check, which matters more
than anything else about the host.

```bash
python3.10 -m venv env
source env/bin/activate
pip install 'tt-bio[tenstorrent]'
tt-bio install-deps
```

`tt-bio install-deps` installs the Tenstorrent system dependencies that match this release. It may ask for your sudo password.

On a host without a Tenstorrent card, plain `pip install tt-bio` is enough: the Boltz-2 CPU/GPU path (`--accelerator cpu` / `--accelerator gpu`) and the CLI work without the Tenstorrent SDK. The other models run on Tenstorrent only.

### From GitHub / source
Pin to a tagged release, track nightly `main` (may be untested), or work from an editable clone:
```bash
pip install "tt-bio[tenstorrent] @ git+https://github.com/moritztng/tt-bio.git@v0.11.0"   # pinned release, see Releases for the latest
pip install "tt-bio[tenstorrent] @ git+https://github.com/moritztng/tt-bio.git@main"     # nightly
# or
git clone https://github.com/moritztng/tt-bio.git
cd tt-bio
pip install -e '.[tenstorrent]'
tt-bio install-deps
```
Drop the `[tenstorrent]` extra on a host without a Tenstorrent card.

### Optional: Build TT-Metal / TT-NN from Source
If you need to build from source, follow the [Tenstorrent Installation Guide](https://github.com/tenstorrent/tt-metal/blob/main/INSTALLING.md).

### Verify Installation
```bash
tt-bio --version   # or -V; prints the installed version
tt-bio --help
tt-bio predict --help
tt-bio msa --help
```

Then [fold a structure](#structure-prediction), [predict affinity](#binding-affinity),
[design binders](#design) or [embed sequences](#embeddings). For reference: the
[input format](#input-format), [output files and scores](#understanding-results),
[weights and MSA databases](#weights-and-msa), [every command-line option](#command-line-options),
[tuning flags](#tuning-flags) and [training](#training).

### Troubleshooting

**Check the CPU frequency driver first.** On an AMD host:

```bash
cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_driver   # want: amd-pstate-epp
```

If it reads `acpi-cpufreq`, every ttnn op pays a fixed extra host cost, measured at roughly
5 to 12 microseconds, and throughput drops across every model. The symptom is distinctive:
per-op device time is unchanged and the loss tracks a model's op count rather than its
size, so a fold that issues 25k ops loses far more than an embedding that issues 1k. One
user measured Boltz-2 at 1.29 structures/s under a 5.4 kernel with `acpi-cpufreq` and
2.14 structures/s on the same box, same binaries, after booting a kernel that gives
`amd-pstate-epp`.

`amd-pstate-epp` needs Linux 6.3 or newer. A kernel configured with
`CONFIG_X86_AMD_PSTATE_DEFAULT_MODE=3` selects it with no boot flag; Ubuntu 24.04 does
this. On an older kernel that has the driver but does not default to it, add
`amd_pstate=active` to the kernel command line. Kernels before 6.3 (RHEL 8, Oracle UEK 6,
anything 4.18 or 5.4) have no amd-pstate at all.

Single-host prediction needs no MPI setup of yours. tt-metal ships the OpenMPI it wants, and a
system OpenMPI on the environment breaks that build instead of replacing it: with `OMPI_MCA_*` or
`OPAL_PREFIX` set, or a foreign `libmpi` on `LD_LIBRARY_PATH`, `MPI_Init` aborts before the fold
starts. tt-bio warns when it sees one and changes nothing for you. To clear it for the current
shell:

```bash
unset OMPI_MCA_pml OMPI_MCA_plm OPAL_PREFIX LD_LIBRARY_PATH
```

## Structure Prediction

```bash
tt-bio predict examples/prot.yaml --model boltz2 --override
```

Every command names its model with `--model`:

- **`boltz2`**: folds complexes of proteins, DNA, RNA, and ligands and predicts binding affinity. MSA-dependent (uses an MSA by default).
- **`esmfold2`** / **`esmfold2-fast`**: fold complexes of proteins, DNA, RNA, and ligands (CCD code or SMILES) on-device, no MSA required (`esmfold2-fast` is the lighter, faster checkpoint). Modified residues are folded as the modified chemistry, and covalent `bond` constraints and cyclic peptides are supported. Templates and pocket constraints are refused.
- **`protenix-v1`** / **`protenix-v2`**: fold complexes of proteins, RNA, DNA, and ligands (an AlphaFold3-family model, the [Protenix](https://github.com/bytedance/Protenix) reproduction); MSA-dependent for proteins (uses an MSA by default), and also emit a PAE/PDE matrix with `--write_pae`. `protenix-v1` is upstream's own v0.5.0 base checkpoint: half the pair width and 4 trunk recycles against `protenix-v2`'s 10, so it is the cheaper of the two. Covalent `bond` constraints and modified residues are supported on both; cyclic peptides and templates, as a structure file or a precomputed alignment (see [Templates](#templates)), on `protenix-v2` only. `protenix-v1`'s checkpoint ships no template blocks, and it does not close a cyclic peptide's ring.
- **`openfold3`**: folds proteins, RNA and DNA (an AlphaFold3-family model, the [OpenFold3](https://github.com/aqlaboratory/openfold-3) reproduction); MSA-dependent (uses an MSA by default), with optional templates, as a structure file or a precomputed alignment (see [Templates](#templates)). Polymer chains only, so ligands are refused. Cyclic peptides, modified residues and a covalent bond to a modified residue are supported; a bond between two standard residues (a disulfide) is refused. Weights come from the OpenFold consortium; point `OF3_CKPT` at them.
- **`openbind`**: OpenBind-0, the same OpenFold3 stack on upstream's v0.5.0 checkpoint, tuned for protein-ligand co-folding. Takes ligands by SMILES or CCD code alongside protein, RNA and DNA chains; MSA-dependent (uses an MSA by default), with optional templates, as a structure file or a precomputed alignment (see [Templates](#templates)). Cyclic peptides, modified residues and covalent bonds to a ligand or modified residue are supported; a bond between two standard residues (a disulfide) is refused. Weights are a separate file from `openfold3` and are not downloaded; point `TT_BIO_OPENBIND` at them (see [`docs/weights.md`](docs/weights.md)).
- **`saprot`**: structure-aware protein embeddings, an ESM-2 encoder over a fused amino-acid + Foldseek-3Di vocabulary (446 tokens). Needs a structure for the 3Di structural tokens (`--structure`); runs sequence-only without it. Use for variant-effect / mutation-fitness scoring and function prediction.
- **`nesso1`** (`tt-bio affinity`): protein-ligand binding affinity without a structure. Predicts a soft distogram and reads the affinity off that, so it is much cheaper than folding and it returns no coordinates. Proteins and ligands only.
- **`opendde`** / **`opendde-abag`**: antibody-antigen co-folding built on the Protenix-v2 stack plus a structural-token expander; `opendde-abag` selects the antibody-antigen checkpoint. Protein, RNA, DNA and ligand chains, with covalent `bond` constraints, cyclic peptides, modified residues and templates in either form (see [Templates](#templates)). Proteins are MSA-dependent (uses an MSA by default, like Protenix-v2). `opendde-abag` matches the upstream checkpoint on the standard 1AHW antibody-antigen target; both implementations perform poorly on 9DSG.
- **`af2ig`**: AlphaFold2 initial-guess, the filter binder-design pipelines use to tell a real design from a plausible one. It takes a design you already have (a structure carrying the target chain and the binder backbone, plus the binder's sequence), re-predicts the complex starting from those coordinates, and reports pLDDT, pTM, ipTM, pAE and interface pAE. Single-sequence, no diffusion, no seed: the same input gives the same answer. Input format and an example are in [`examples/af2_designed_complex.yaml`](examples/af2_designed_complex.yaml); the file takes `target:` and `binder:` and nothing else, and any other key is refused rather than ignored. Weights are DeepMind's AlphaFold2 monomer pTM parameters; `tt-bio weights --download af2ig` fetches them (4 GB, one file kept).
- **`rf3`**: folds complexes of proteins, RNA, DNA, and ligands (an AlphaFold3-family model, [RoseTTAFold3](https://github.com/RosettaCommons/foundry) from the Institute for Protein Design); MSA-dependent for proteins (uses an MSA by default). Writes AlphaFold3-style `<name>_summary_confidences.json` (pTM, ipTM, chain-pair PAE/PDE, ranking score) next to each structure. Modified residues and covalent bonds to a ligand or modified residue are supported; cyclic chains, a bond between two standard residues (a disulfide) and pocket constraints are refused. Templates work in either form (see [Templates](#templates)). Weights download from the IPD on first use.

```bash
tt-bio predict examples/prot.fasta --model esmfold2-fast --fast
tt-bio predict examples/fkg_ligand.yaml --model esmfold2   # protein + ligand co-fold, no MSA
tt-bio predict examples/prot.yaml --model protenix-v2   # MSA on by default; NA/ligand chains are single-sequence
tt-bio predict examples/prot.yaml --model protenix-v1   # upstream's v0.5.0 base checkpoint, 4 recycles
tt-bio predict examples/prot.fasta --model openfold3    # MSA on by default; set OF3_CKPT to the weights file
tt-bio predict examples/affinity.yaml --model openbind  # protein + ligand co-fold; set TT_BIO_OPENBIND
tt-bio predict examples/9dsg_abag.yaml --model opendde-abag   # antibody-antigen co-fold, MSA on by default
tt-bio predict examples/prot.yaml --model rf3            # MSA on by default; weights fetch from the IPD
tt-bio predict examples/prot.yaml --model rf3 \
    --partial_t 150 --partial_structure start.cif       # refine start.cif instead of folding from scratch
tt-bio predict targets.yaml --model rf3 --early_stop_plddt 0.5   # skip the rollout on hopeless targets
tt-bio predict examples/af2_designed_complex.yaml --model af2ig  # score a designed complex: ipTM + interface pAE
```

| Feature | Boltz-2 | ESMFold2 | Protenix-v1 | Protenix-v2 | OpenFold3 | OpenBind-0 | OpenDDE | RF3 |
|---|---|---|---|---|---|---|---|---|
| Input | protein/DNA/RNA/ligand complex | protein/DNA/RNA/ligand complex | protein/DNA/RNA/ligand complex | protein/DNA/RNA/ligand complex | protein/RNA/DNA (polymer-only) | protein/DNA/RNA/ligand complex | protein/DNA/RNA/ligand complex | protein/DNA/RNA/ligand complex |
| MSA | MSA-dependent (on by default) | single-sequence | proteins MSA-dependent (on by default), NA/ligand single-sequence | proteins MSA-dependent (on by default), NA/ligand single-sequence | proteins MSA-dependent (on by default) | proteins MSA-dependent (on by default) | proteins MSA-dependent (on by default), NA/ligand single-sequence | proteins MSA-dependent (on by default) |
| PAE/PDE output (`--write_pae`) | no | no | yes | yes | no | no | no | in `_summary_confidences.json` |

Ligands, nucleic acids, modified residues, cyclic chains, covalent and pocket constraints,
templates and affinity are one generated matrix:
[`docs/model-capabilities.md`](docs/model-capabilities.md). Anything a model cannot honour is
refused by name before the fold starts, never accepted and dropped.

### Size Limits

Every structure model folds at least 1536 residues on a single 12 GiB Wormhole chip. The limit
below is the largest size that folded, under the first measured failure where one was found,
walked with the settings the platform sends:

| model | Wormhole limit | first measured failure |
|---|---:|---:|
| `boltz2` | 1920 | 2048 |
| `opendde`, `opendde-abag` | 1536 | 1664 (it runs for hours rather than crashing) |
| `openfold3` | 1664 | 1792 |
| `openbind` | 1664 (residues; a ligand adds tokens) | 1792 |
| `pxdesign` | 1536 (target residues; the binder is on top) | none found; top of the ladder |
| `protenix-v2` | 2048 (residues; a ligand adds tokens) | none found; top of the ladder |
| `esmfold2` | 1664 (residues; a ligand adds tokens) | 1792 |
| `esmfold2-fast` | 1664 (residues; a ligand adds tokens) | 1792 |
| `rf3` | 1600 | 1664 |
| `protenix-v1` | 2048 | none found; top of the ladder |
| `rfd3` | 1536 (motif + designed) | none found; top of the ladder |
| `boltzgen` | 14786 (atoms in the target) | none found; top of the ladder |
| `esmc-6b` (embed) | 1968; 8192 with `--fast` | 1984 |

For every model here that reads an alignment, the limit was measured with 16384 alignment rows,
the most any of them reads, so a deeper a3m does not lower it.

Ask for more than a model's limit and tt-bio refuses before it opens a device, naming the
model, the limit and any model that does take the input. `nesso1` and `af2ig` have no measured
limit and are never refused: nothing above the top of either ladder has been run, so there is no
failing size to refuse on. AF2-IG's top rung is 1024 tokens (944 target residues plus an
80-residue binder) in 488 s on one Wormhole chip, and 1536 tokens (1008 plus a 528-residue
binder) in 665 s on a Blackhole p150a. What sets each wall is in
[docs/large-targets.md](docs/large-targets.md#what-stops-each-model-above-1024-on-a-galaxy-chip).

`boltzgen` is the one model sized on atoms rather than residues, because its wall follows the
target's atom count and atoms per residue vary with what the target is made of: the 14786-atom
rung is 1831 residues of a deposited protein, and a lighter target reaches more residues in the
same number of atoms. The designed binder is outside the number, which was 80 residues at every
rung. tt-bio counts the atoms out of the structure file your `file:` entity points at, so the
refusal still lands before a device opens.

Blackhole has more memory per chip, so it is a separate table and the Wormhole numbers do not
carry over. Six models have been walked to a failure on a p150a:

| model | Blackhole limit | first measured failure |
|---|---:|---:|
| `saprot-35m` (embed) | 126976 (longest sequence) | 131072 |
| `esmc-300m`, `esmc-600m`, `saprot-650m`, `saprot-1.3b` (embed) | 114688 (longest sequence) | 126976 |
| `pxdesign` | 2500 (target residues) | 3000 |

The embedding wall is one allocation: the model asks for the whole sequence-by-sequence attention
matrix as a single buffer that grows with the square of the sequence length, 34.4 GB at 131072
residues against a 31.9 GiB card. `saprot-35m` gets one rung further than the rest because its
weights are the smallest and leave more room for it. Everything else is unmeasured on Blackhole
and refused nothing, including `boltzgen`, which designs against a 16948-atom target there.

The Blackhole ladders have been walked separately. This is the largest rung each structure model
is on record folding, from `docs/size_ladder_baseline.d/`:

| model | p150a | p300c |
|---|---:|---:|
| `boltz2` | 1536 | 1024 |
| `esmfold2` | 1536 | 1024 |
| `nesso1` | 1536 | 1024 |
| `openbind` | 1536 | 1024 |
| `opendde` | 1024 | 1024 |
| `openfold3` | 1536 | 1024 |
| `protenix-v1` | 1536 | 1024 |
| `protenix-v2` | 1536 | 1024 |
| `rf3` | 1536 | 1088 |

Only `opendde` and `opendde-abag` are enforced on Blackhole, at 1024; every other model here
accepts an oversized request rather than refusing it, so the table is what you have and not a
guard. A blank rung above a model's top is unwalked, not a measured failure. 1536 is the top of
the ladder itself, so the models sitting there have no failure on record at all.

The limits were measured with an MSA, which is the default for the models that take one, and at the
deepest alignment the MSA pipeline actually produces. Folding single-sequence is roomier, so if you
know your run is lighter than the ladder that set the limit, `TT_BIO_SIZE_LIMIT=0` turns the refusal
into a warning and runs it anyway.

A ligand counts against these limits. Its heavy atoms are tokens the model pays for exactly like
residues, and on `esmfold2`, `esmfold2-fast`, `openbind` and `protenix-v2` the wall is on tokens,
so a cocrystal is checked on residues plus ligand atoms rather than on the residue count alone.
`esmfold2` folds 1664 residues, which leaves no room at all: 1664 residues plus any ligand is
refused, and 1640 residues with a 24-atom ligand is admitted. `openbind` is the same: its 1664 was
walked apo, so a ligand counts against it atom for atom. Either way the
refusal names the token count and the wall, and it arrives before a device is opened instead of
as an out-of-memory error part way through the fold.

`esmfold2-fast` is the same architecture at half the trunk depth. It has its own row because it
was walked on its own; today both fold 1664 residues and fail at 1792.

The pair track switches to row-blocked execution at a size threshold smaller targets never reach,
so their speed and numerics are untouched. See [docs/large-targets.md](docs/large-targets.md).
Perf levers are gated at several sequence lengths, not just one; the release gate re-checks the
ladder against a recorded baseline. See [docs/size-generality.md](docs/size-generality.md).

### Options, Weights and MSA

All structure models support the sampling, output-format, and scheduling options.
MSA, affinity, constraint, and auxiliary-output options apply only where listed
below. Each model downloads its weights automatically on first use, except
OpenFold3: fetch the consortium checkpoint yourself from
`https://openfold3-data.s3.amazonaws.com/openfold3-parameters/of3-p2-155k.pt`
and put it at `~/.boltz/of3-p2-155k.pt`, or point `OF3_CKPT` at it. `tt-bio weights` lists every artifact with
its status, size and path; `--download` prefetches, `--prune` reclaims disk. Set
`TT_BIO_CACHE` to move all of it (both `~/.boltz` and the Hugging Face cache, about
65 GiB) somewhere with room. See [docs/weights.md](docs/weights.md).

Boltz-2, Protenix-v1, Protenix-v2, OpenFold3, OpenDDE, and RF3 are MSA-dependent and use an MSA **by default**, a local
ColabFold DB (`~/.boltz/msa_db`) if one is set up (see [Offline MSA](#offline-msa-optional)),
otherwise the online ColabFold server. Sending sequences to the online server (`api.colabfold.com`)
leaves your machine; a one-line notice is printed when that fallback is used. Pass
`--msa_db_path` for a private offline database, or `--single_sequence` to deliberately fold
without an MSA (lower accuracy; for batch-screening orphan sequences). A complex with two or
more different protein sequences also gets a species-paired MSA, searched once per complex, the
way each model's upstream pairs; a homodimer is not paired. Two models fold unpaired because
their upstream does: RF3 pairs by taxonomy IDs that ColabFold alignments do not carry, and
OpenFold3's preview2 checkpoint runs on an upstream release that drops the paired rows
(OpenBind pairs).
ESMFold2 needs no MSA and uses one when a source is given.

`--fast` makes some operations use a lower-precision numeric format that runs faster. Accuracy is typically very close; the Boltz-2 measurement is in [`docs/boltz2-fast-parity.md`](docs/boltz2-fast-parity.md).

### Many Inputs and Cards

`predict` accepts either a single YAML/FASTA file or a directory containing many input files.
An input the model refuses is reported, recorded as failed in `results.json` and skipped, and
the rest of the directory folds; the exit status is 2 when some inputs failed and 1 when all did.
Every chain comes back under the id you gave it.

A live display shows the progress of each target. Prediction uses up to one card
per pending target, labelled in the display (`quietbox:tt0`, `quietbox:tt1`, ...).
Models load once per active card and stay resident:

```bash
tt-bio predict proteins/ --model boltz2 --out_dir results --fast
```

Pass `--devices 0,1,2,3` to pick or limit the available cards. A single target
remains a single-card fold; additional cards increase throughput only when
multiple targets are queued.

Once enough folds run at once that each worker is down to a couple of CPU threads,
tt-bio also has their idle thread pools sleep between device syncs instead of
spinning on them. The output is identical either way, and a single fold is
unaffected. See [Tuning flags](docs/tuning-flags.md) for the measurement, or set
`OMP_WAIT_POLICY` yourself to take the decision back.

To spread work across several machines, see [Running tt-bio on many machines](#running-tt-bio-on-many-machines).

## Binding Affinity

### Binding Affinity Prediction (Boltz-2)

```bash
tt-bio predict examples/affinity.yaml --model boltz2 --use_msa_server --override --affinity_mw_correction
```

The `--affinity_mw_correction` flag applies molecular weight correction for more accurate predictions.

An affinity run folds the complex and then runs a second model that has its own
64-block trunk, so it costs more than a structure-only fold. All of it runs on the
card. FKBP12+SB3 at the default affinity protocol (200 sampling steps, 5 affinity
samples, single sequence) takes about 206 s per ligand on one Blackhole p150a,
measured as a whole `tt-bio predict` invocation with model load included.
`--sampling_steps_affinity` and `--diffusion_samples_affinity` are the two flags that
move that wall most.

The affinity trunk runs in fp32 because the predicted log10(IC50) is sensitive to
activation precision, and that is not configurable. Earlier releases ran it in fp32
on the host CPU instead, which is why affinity used to take minutes per ligand and
looked CPU-bound.

### Binding Affinity Without a Structure (Nesso-1)

`tt-bio affinity` predicts protein-ligand affinity without folding anything. Nesso-1 has no
structure module, so it returns a number, not coordinates:

```bash
tt-bio affinity examples/affinity.yaml                  # one complex
tt-bio affinity ligands/ --out_dir screen               # a directory is a screen
```

A directory keeps the model resident across inputs, so a ligand series against one target pays the
weight load and the kernel compile once. Output is one `<id>_affinity.json` per input plus an
`affinity.csv` for the whole run: the affinity value (mean of a two-member ensemble, and each
member), a binary binder probability, and six distogram entropies.

On DAVIS it reaches 0.662 mean within-target Pearson against measured Kd (0.175 for a
molecular-weight-only control), matching the 0.636 the upstream implementation gets on an H200.

It is far cheaper than folding for the same question. One 512 aa prediction takes 8.3 s of model
time on one Blackhole card, 33 s for the whole command including featurisation; Boltz-2 affinity
takes 386 s for the same command on the same input, which is what tt-bio shipped for this before.
Against a GPU it is 7.9x off an H200 at that size, so choose it for what the answer costs on this
hardware rather than expecting it to beat a GPU.

Use it to rank a series; use `predict --model boltz2` when you need the pose. Proteins and ligands
only, one ligand scored per input. One Wormhole chip scores a 3072-residue target with any
ligand up to cobalamin's size in about 15 minutes. 3584 residues also completes but takes about
two hours, because the chip's memory spills to the host; 4096 is refused. The trunk runs bf16 by default: it is about 6x faster than fp32
and no less accurate from 276 tokens up, and fp32 runs out of DRAM around 1000 tokens. On inputs
under ~150 tokens fp32 is the more faithful arm, and `--trunk fp32` switches back. See
[`docs/nesso1.md`](docs/nesso1.md) for the input schema, the four upstream limits, and what to watch
when comparing numbers against another implementation.

## Design

Design new binders and protein structures from a target or motif specification, with one command:

```bash
tt-bio design examples/binder.yaml --model boltzgen --num_designs 10
tt-bio design specs.json --model rfd3 --from_pdb --out_dir designs/
```

| Model | Designs | Input |
|-------|---------|-------|
| `boltzgen` (default) | protein / peptide / nanobody / antibody binders against a target | design YAML, same entity grammar as `predict` |
| `rfd3` | all-atom structures: binders, motif scaffolding, nucleic-acid binders | JSON spec with contig strings |
| `pxdesign` | binder backbones against a target structure | target YAML: structure file, chains to condition on, binder length |

**[BoltzGen](https://github.com/HannesStark/boltzgen)** designs binders against a target structure. The pipeline runs design → inverse folding → folding → analysis → filtering and writes the top-ranked binders to `<out_dir>/final_ranked_designs/`. Pass `--seed N` to make a design reproducible; without it every run draws fresh. Input grammar, protocols, pipeline subsets, and options: [`docs/boltzgen-design.md`](docs/boltzgen-design.md). Designability (scRMSD) QA: [`docs/boltzgen-designability.md`](docs/boltzgen-designability.md).

**[RFdiffusion3](https://www.biorxiv.org/content/10.1101/2025.09.18.676967)** (RFD3) is an all-atom generative model that designs new protein structures and sequences from a specification, rather than folding an existing one. Design modes, the contig-string input grammar, and which conditioning fields a spec can and cannot ask for: [`docs/rfd3-design.md`](docs/rfd3-design.md).

**[PXDesign](https://github.com/bytedance/PXDesign)** generates binder backbones against a target structure, conditioned on a distogram of the target rather than its coordinates. Input is a target YAML naming a structure file, the chains to condition on (with optional per-chain crop and hotspots) and a `binder_length`; each design is written as a CIF in the target structure's own frame, so it opens alongside your input file. A `designs.json` lands beside them with each design's numbers: fit RMSD against the target, binder residue and atom counts, and how many target tokens it was conditioned on. The binder is written as GLY because PXDesign generates a backbone with no sequence. Hotspot residues are `label_seq` numbers, not the author numbering a viewer shows, and a number that names no residue is refused rather than dropped. `--num_designs` is also the batch axis for this model: every requested design comes from one batched diffusion trajectory, and the gain per design grows with the batch and shrinks with the target: 2.7x at 8 designs against a 256-residue target, 1.5x against a 512-residue one, and flat from 16 up rather than turning back. A given `--seed` and `--num_designs` always reproduce the same designs, but `--num_designs 1` and `--num_designs 2` do not share their design 0: asking for more designs currently changes which ones you get, so pin both values when you want a run back. Selecting designs, which upstream does with a Protenix and an AF2-IG filter, is not on the CLI yet.

Each model downloads its weights automatically on first use. BoltzGen and RFdiffusion3 fan out across every available card (`--devices 0,2` restricts); PXDesign runs on one card locally, or one design per card across a host's controller with `--controller http://127.0.0.1:8765`. `tt-bio gen` still works as a deprecated alias for `tt-bio design --model boltzgen`.

How many designs a card returns per hour, how `--num_designs` and `--devices` move it, and how to size a campaign: [`docs/design-throughput.md`](docs/design-throughput.md).

**[BindCraft 2](https://github.com/PacesaLab/BindCraft2)** is not a tt-bio model and has no CLI entry; it is a third-party design loop you install yourself, and `tt_bio.bindcraft2` gives it an AlphaFold 2 Evoformer that runs on a card, interleaving as many design trajectories over that one chip as the box and card have memory for. On a roomy box that is three for a complex up to 352 tokens and fewer above it; at 288 tokens three take a completed trajectory from 927 s to 774 s and a round from 7.41 s to 6.00 s, 8.6x an H200 rather than 10.6x, with the same designs accepted. Its gradient loop runs on card, with the validation ensemble on BindCraft 2's own trunk so it stays the reference's. On the shipped PD-L1 example it accepts binders at 7 per 31 trajectories against BindCraft 2's own JAX at 1 per 5, which Fisher exact does not separate (p = 1.00). One p150a carries a complex up to 864 tokens and one Wormhole Galaxy chip up to 512, counting the fused complex rounded to 32 rather than target plus binder, and a larger one is refused with a message naming the axis and the memory; an input it cannot design against, from a hotspot in an unresolved loop to a nucleotide sequence pasted in where the protein goes, is refused before a card is opened rather than quietly worked around. What you need, how to point a campaign at a chip, what fits, and how to read that comparison: [`docs/bindcraft2.md`](docs/bindcraft2.md).

## Embeddings

### Protein Embeddings (ESMC)

Turn protein sequences into ESMC language-model embeddings on-device (no
folding, no MSA). `DATA` is a FASTA file, a directory of them, a YAML
`{id: sequence}` mapping, or a bare sequence string:

```bash
tt-bio embed proteins.fasta --model esmc-600m --out_dir embeddings
tt-bio embed "MQIFVKTLTGKTITLEV..." --model esmc-600m   # one-off sequence
```

`--model` selects the ESMC variant (`esmc-300m`, `esmc-600m`, `esmc-6b`). For
each sequence you get its **per-residue** embeddings (`[length, d_model]`
float32, one row per amino acid, row order == input order) and a **pooled**
whole-sequence vector (`[d_model]` float32, `--pool mean`/`max`/`cls`).
`--out_dir` (default `./embeddings`) gets:

- `<id>.npz` per sequence: `per_residue`, `pooled` (+ `logits` with `--logits`); `--format npz`, default
- `embeddings.parquet`: pooled vectors, one row per sequence; `--format parquet`
- `manifest.json`: model/pool/shapes/dtype and which file holds each sequence

Add `--logits` for the per-residue amino-acid predictions (300M/600M only),
and `--fast` for the lower-precision weight path. Weights download automatically on
first use.

Sequences batch automatically on 300M/600M (`--batch_size`, default 8): a
padded, length-bucketed device forward per batch, masked so results are
identical to running each sequence alone. Single-sequence calls
(`--batch_size 1`, e.g. serving one sequence at a time) replay through a
captured device trace once a length bucket repeats, up to ~1.5x faster per
call on QuietBox-class hosts, bit-identical, no flags needed.

To embed a large batch faster, shard it across several cards with
`--devices 0,1,2,3`: one worker per card, results reassembled in input order
and identical to a single-card run:

```bash
tt-bio embed proteins.fasta --model esmc-600m --devices 0,1,2,3
```

Fanout only pays off when there's enough work per shard to amortize each worker's model-load and device-init cost. On small batches it can be flat or worse than a single card. `esmc-6b` scales to 4 cards on suitably large batches. Reach for `--devices` on large batches, not small ad-hoc jobs; use `--controller` (below) for repeated/production embedding.

For repeated/production embedding, submit to a persistent pool instead: a worker
loads its model once and keeps it resident across every call, so the reload cost
above is paid once per worker, not once per invocation:

```bash
tt-bio controller --port 8765            # starts + keeps a worker per local card
tt-bio embed proteins.fasta --model esmc-6b --controller http://127.0.0.1:8765
```

The same capability is available from Python:

```python
from tt_bio import esmc

emb = esmc.embed("MQIFVKTLTGKTITLEV...", model="esmc-600m")[0]
emb.per_residue   # [L, d_model] float32
emb.pooled        # [d_model] float32

# Shard a large set across cards (data-parallel, order preserved):
embs = esmc.embed(sequences, model="esmc-600m", devices=[0, 1, 2, 3])
```

### Structure-Aware Protein Embeddings (SaProt)

SaProt is a structure-aware protein language model, an ESM-2 encoder over a fused
amino-acid + Foldseek 3Di vocabulary (446 tokens). Where ESMC is sequence-only, SaProt
also encodes local structure, so its embeddings and MLM logits reflect both sequence
and shape. Use it for variant-effect / mutation-fitness scoring and function prediction
when you have a structure (predicted or experimental: fold it with `tt-bio predict`
first, then score it with SaProt).

```bash
tt-bio saprot proteins.fasta --model saprot-650m --structure structs/ --out_dir embeddings
tt-bio saprot proteins.fasta --model saprot-650m                # sequence-only (3Di = '#')
tt-bio saprot proteins.fasta --model saprot-650m --devices 0,1    # data-parallel across 2 cards
```

`--structure` is a PDB/cif file (single sequence) or a directory of `<id>.pdb`/`<id>.cif`
files, one per FASTA id. The 3Di structural tokens are computed on host with
[Foldseek](https://github.com/steineggerlab/foldseek) (`conda install -c bioconda foldseek`;
`--foldseek PATH` or `FOLDSEEK_BIN` if it is not on PATH); it runs off-device. Residues
the structure does not resolve get the `#` unknown-structure token, and a structure that
is not of the sequence you passed is refused rather than lined up by length. Omit
`--structure` for sequence-only mode (lower accuracy for 35M/650M; the 1.3B works
sequence-only).

For each sequence you get **per-residue** structure-aware embeddings (`[length, d_model]`
float32) and a **pooled** vector, plus per-residue MLM logits (`[length, 446]` with
`--logits`) over the fused vocabulary, the log-likelihoods used for zero-shot mutation
scoring. Output layout matches `tt-bio embed` (`<id>.npz` / `embeddings.parquet` /
`manifest.json`).

`--model` selects the variant (`saprot-35m`, `saprot-650m`, `saprot-1.3b`). `--devices 0,1,2,3`
shards the input across cards data-parallel (one pinned subprocess each, results reassembled in
input order), bit-exact vs single-card with `--batch_size 1`. Parity vs the reference HuggingFace
checkpoint, the multi-card bit-exactness check, and warm throughput are in
[`docs/saprot-parity.md`](docs/saprot-parity.md).

Python:

```python
from tt_bio import saprot

emb = saprot.embed(("MQIFVKTLTGKTITLEV...", "dweweaepvrdidi..."), model="saprot-650m")[0]
emb.per_residue   # [L, d_model] float32, structure-aware
emb.logits        # [L, 446] float32 (with return_logits=True)
```

## Input Format

ESMFold2, Protenix-v1 and Protenix-v2 accept proteins, DNA, RNA, ligands and covalent `bond`
constraints. OpenFold3 accepts proteins, DNA and RNA plus per-chain templates, and refuses
ligands. OpenDDE accepts proteins and ligands with `bond` constraints. Boltz-2 additionally
supports affinity, pocket/contact constraints, potentials, and user-supplied templates. Which
model takes cyclic chains and which kind of bond is in
[`docs/model-capabilities.md`](docs/model-capabilities.md).

Create a YAML file describing your complex:

```yaml
version: 1
sequences:
  - protein:
      id: A
      sequence: MVTPEGNVSLVDESLLVGVTDEDRAVRSAHQFYERLIGLWAPAVMEAAHELGVFAALAEAPADSGELARRLDCDARAMRVLLDALYAYDVIDRIHDTNGFRYLLSAEARECLLPGTLFSLVGKFMHDINVAWPAWRNLAEVVRHGARDTSGAESPNGIAQEDYESLVGGINFWAPPIVTTLSRKLRASGRSGDATASVLDVGCGTGLYSQLLLREFPRWTATGLDVERIATLANAQALRLGVEERFATRAGDFWRGGWGTGYDLVLFANIFHLQTPASAVRLMRHAAACLAPDGLVAVVDQIVDADREPKTPQDRFALLFAASMTNTGGGDAYTFQEYEEWFTAAGLQRIETLDTPMHRILLARRATEPSAVPEGQASENLYFQ
  - ligand:
      id: B
      smiles: 'N[C@@H](Cc1ccc(O)cc1)C(=O)O'
properties:
  - affinity:
      binder: B
```

**Entity Types:**
- **Polymers** (`protein`, `dna`, `rna`): provide `sequence`
- **Ligands** (`ligand`): provide `smiles` or `ccd` code

**Multiple Identical Chains:**
```yaml
- protein:
    id: [A, B]  # Two identical chains
    sequence: ...
```

### Proteins with Custom MSA
```yaml
- protein:
    id: A
    sequence: MVTPEGNVSLVDES...
    msa: ./path/to/msa.a3m
```

### Proteins with Modifications
```yaml
- protein:
    id: A
    sequence: MVTPEGNVSLVDES...
    modifications:
      - position: 5
        ccd: PTR  # Modified residue code
```

### Ligands
```yaml
- ligand:
    id: B
    smiles: 'CC1=CC=CC=C1'  # SMILES string
    # OR
    ccd: ATP                # CCD code
```

### Constraints

Pocket and contact constraints are **Boltz-2 only** (they need a trained constraint embedder). A covalent `bond` to a ligand or a modified residue works on every structure model that takes the ligand. A bond between two standard residues (a disulfide) works on Boltz-2, ESMFold2, Protenix and OpenDDE; RF3 and the OpenFold3 family refuse it by name.

**Pocket Constraints** (binding site):
```yaml
constraints:
  - pocket:
      binder: B              # Ligand chain
      contacts: [[A, 10], [A, 11], [A, 12]]  # Binding site residues
      max_distance: 6.0      # Angstroms (4-20A, default 6A)
      force: false           # Use potential to enforce (default: false)
```

**Contact Constraints:**
```yaml
constraints:
  - contact:
      token1: [A, 10]
      token2: [A, 50]
      max_distance: 8.0
      force: false
```

**Bond Constraints** (covalent link, e.g. a covalent inhibitor, glycosylation, or disulfide):
```yaml
constraints:
  - bond:
      atom1: [A, 10, SG]     # [chain, residue, atom]
      atom2: [B, 1, C1]      # SMILES ligand: element + count in SMILES order
```

> **OpenDDE + covalent bonds:** OpenDDE honors a `bond` constraint between a protein
> residue and a ligand atom (the covalent-inhibitor case) or between two protein
> residues (a disulfide or crosslink). Both ride the same `token_bonds` machinery as
> Protenix-v2 and are honored in the output (device-verified against upstream OpenDDE
> within the reference's own seed noise floor); see `examples/opendde_covalent_ligand.yaml`
> and `examples/opendde_covalent_bond.yaml`.

### Templates

Use experimental structures as templates:

```yaml
templates:
  - cif: ./template.cif
    chain_id: A
    template_id: A
    force: true              # Enforce template alignment
    threshold: 2.0           # Max deviation in Angstroms
```

`chain_id` names the chains to template (default: every protein chain) and
`template_id` the template's chains by mmCIF `label_asym_id` (default: the best
match). Each chain is aligned to the template's sequence for you. This block
works on `boltz2`, `protenix-v2`, `opendde`, `opendde-abag`, `openfold3`,
`openbind` and `rf3`. `force` (with its `threshold`) and pdb files are Boltz-2
only; the other models refuse them, and RF3 takes one template per chain.

The same models except `boltz2` also take a precomputed alignment `.npz` per
protein chain (the format the upstream benchmark cache ships); Boltz-2 refuses it
and takes the template as a cif instead. Its structures are fetched from
RCSB, and a missing one is a hard error rather than a silently dropped
template. See `examples/7xi5_tmpl.yaml`. There is no template search.

```yaml
sequences:
  - protein:
      id: A
      sequence: MSSATPDPAEILT...
      templates: ./templates.npz
```

## Understanding Results

### Output Structure

```text
<model>_results_prot/   # e.g. protenix_results_prot, boltz2_results_prot
├── structures/
│   ├── prot.cif                      # Best-ranked predicted structure
│   └── prot_model_1.cif              # Additional samples (if diffusion_samples > 1)
├── results.json                      # One entry per target with confidence/affinity metrics
├── power_profile.csv                 # (optional, --report-energy)
├── power_profile.png                 # (optional, --report-energy)
├── prot_pae.npz                      # (optional, --write_pae)
├── prot_plddt.npz                    # (optional, --write_pae, Boltz-2)
├── prot_pde.npz                      # (optional, --write_pde)
└── prot_embeddings.npz               # (optional, --write_embeddings)
```

MSA results are cached in `<out_dir>/msa/` (default `./msa/`), keyed by sequence hash. The same protein sequence is never searched twice, even across different input files or runs. The MSA search uses all available CPU threads and keeps the database index memory-mapped for maximum speed.

### Confidence Scores

Each target entry in `results.json` contains confidence metrics. The fields below are Boltz-2's; Protenix-v2 and OpenFold3 report the same `confidence_score` / `ptm` / `iptm` / `plddt` (and `all_runs` when `--diffusion_samples` > 1, ranked best-first), while an ESMFold2 entry instead carries `plddt` (mean, 0-1), `ptm` when available, and `n_residues` / `n_chains`. Every model reports its complex mean pLDDT under `plddt`. Boltz-2, Protenix-v2 and OpenDDE also report the two per-chain fields below on a multi-chain target, in `all_runs` as well as for the best sample.

```json
{
    "id": "prot",
    "status": "ok",
    "confidence_score": 0.84,
    "ptm": 0.84,
    "iptm": 0.82,
    "complex_plddt": 0.84,
    "plddt": 0.84,
    "chains_ptm": {
        "0": 0.85,
        "1": 0.83
    },
    "pair_chains_iptm": {
        "0": {"0": 0.85, "1": 0.72},
        "1": {"0": 0.82, "1": 0.83}
    }
}
```

- `confidence_score`: Overall confidence (0-1, higher is better). Models are ranked by it and the best sample is written as `<name>.cif`. Boltz-2 and BoltzGen use their own published rule, 0.8 × `complex_plddt` + 0.2 × `iptm`, or 0.2 × `ptm` on a single chain where there is no interface and `iptm` is 0; it is not a pLDDT and can sit either side of one, on CDK2 0.008 above `complex_plddt` at 298 aa and 0.050 below it at 512 aa. The AlphaFold3-family models (OpenFold3, OpenBind-0, RF3, Protenix-v1/v2, OpenDDE) share one rule: AlphaFold3's 0.8 × `iptm` + 0.2 × `ptm` + 0.5 × disorder − 100 × clash, with `plddt` taking ipTM's weight on a single chain, where ipTM is zero by construction. OpenFold3 and OpenBind-0 are the only two that compute the disorder and clash terms; the others pass 0 for them. RF3 reports the same number as `ranking_score`. So the family's scores are comparable to each other but not to Boltz-2's
- `ptm`: Predicted TM-score for complex (0-1)
- `iptm`: Interface TM-score (0-1)
- `complex_plddt`, `plddt`: Mean confidence (0-1), the same value under both names. It is the mean of the B-factor column of the structure file the same fold wrote, so averaging that column reproduces it. Boltz-2 writes one pLDDT per residue, so average over one atom per residue (CA); Protenix-v2, OpenFold3, OpenBind-0 and OpenDDE write one per atom, so average over all of them
- `chains_ptm`: Per-chain TM-scores (0-1)
- `pair_chains_iptm`: Per-chain-pair interface TM-scores (0-1), with each chain's own `chains_ptm` on the diagonal. Read `pair_chains_iptm[binder][target]` to score one named interface of a complex; the global `iptm` is the whole-interface number and on a two-chain target the two agree. Like every other confidence value, these are comparable between targets of the same model, not between models

### Interface Scores

With `--write_pae`, a multi-chain Boltz-2 entry also carries `interface_scores`: ipSAE (both
directions, their max and their min), ipTM, interface pAE, pDockQ, pDockQ2 and LIS for every chain
pair. The definitions match the script Adaptyv scored its Nipah binder competition with, at the
same 15 A cutoffs. `interface_score_distribution` gives the same scores for every diffusion
sample, with mean, standard deviation and range, so `--diffusion_samples 5` shows how stable a
design's score is as well as its value. To score a fold you already have, from tt-bio or from
upstream Boltz, without a device:

```bash
tt-bio score prot_model_0.cif pae_prot_model_0.npz --plddt plddt_prot_model_0.npz
```

Tenstorrent arithmetic differs from a GPU's, so the same inputs fold to slightly different
numbers than they do on CUDA. See [docs/interface-scores.md](docs/interface-scores.md) for the
definitions, the choices made where the reference is ambiguous, and how this relates to the ipSAE
BoltzGen reports while designing.

### Affinity Predictions

For affinity targets, the same `results.json` entry also contains:

```json
{
    "affinity_pred_value": 2.47,
    "affinity_probability_binary": 0.41,
    "affinity_pred_value1": 2.55,
    "affinity_pred_value2": 2.19,
    "affinity_probability_binary1": 0.50,
    "affinity_probability_binary2": 0.42
}
```

- `affinity_probability_binary`: Probability of binding (0-1). Use for hit discovery (higher = more likely to bind)
- `affinity_pred_value`: Predicted binding affinity as log10(IC50) in μM. Use for ligand optimization (lower = stronger binding). Only compare between known active molecules
- `affinity_pred_value1`, `affinity_pred_value2`: Individual model predictions for binding affinity
- `affinity_probability_binary1`, `affinity_probability_binary2`: Individual model predictions for binding probability
- `runtime_s`: Wall-clock seconds for the whole target, structure and affinity together. Affinity
  targets also carry `structure_runtime_s` and `affinity_runtime_s`; the affinity leg is normally the
  larger of the two by several times, so read the split before pricing a screen

## Weights and MSA

### Weights

Weights download on first use, so nothing here is required. `tt-bio weights` is for when
you want to see or move them:

```bash
tt-bio preflight protenix-v1         # can this machine run it right now?
tt-bio weights                       # every artifact: status, size, resolved path
tt-bio weights --download            # prefetch everything (e.g. before going offline)
tt-bio weights --download boltz2     # or just one model's set
tt-bio weights --prune               # reclaim superseded revisions and leftovers
```

`tt-bio preflight` answers before you submit a job, and exits non-zero when something is
missing, so it works in a script. When weights are missing it also measures whether the
hosts they come from can be reached from this machine.

A full set is about 65 GiB. It lands in `~/.boltz` and the Hugging Face cache; set
`TT_BIO_CACHE` to put both somewhere with more room. Each artifact also takes its own
override, so `TT_BIO_BOLTZ2_CONF=/mnt/weights/boltz2_conf.ckpt` loads that file instead of
downloading. Rows show as `corrupt` if a download was interrupted, and are re-fetched rather
than loaded. No download waits forever: a source that sends nothing is dropped for the next
one, and the error names every host tried. See [docs/weights.md](docs/weights.md).

### Offline MSA (Optional)

A local database avoids the online MSA server and is faster for repeated runs, if you have the disk and RAM for it.

```bash
tt-bio msa
tt-bio predict examples/prot.yaml --model boltz2 --override
```

`tt-bio msa` downloads UniRef30 to `~/.boltz/msa_db` (~100GB download, ~500GB on disk after indexing). `predict` auto-detects this path.

EnvDB can improve MSA coverage when UniRef30 hits are weak, at higher disk and RAM cost. To add it and use it in prediction:

```bash
tt-bio msa --db all
tt-bio predict examples/prot.yaml --model boltz2 --use_envdb --override
```


### Shared MSA Server (Optional)

Host the database on one machine and let others fetch MSAs from it over HTTP, so each prediction machine need not keep its own ~500GB copy.

```bash
# On the machine with the database:
tt-bio msa-server --listen 0.0.0.0:8765

# On any other machine (no local database needed):
tt-bio predict examples/prot.yaml --model protenix-v2 --msa_endpoint http://HOST:8765
```

The server runs the same offline `colabfold_search` and serves unpaired `{hash}.a3m`, with a shared cache and a search-concurrency cap (`--max_concurrent`). Add `--token` to require `Authorization: Bearer <token>`. `--msa_endpoint` applies to `--model esmfold2`, `protenix-v1`, `protenix-v2`, `openfold3`, `opendde`, and `rf3`.

### MSA Server Authentication

For `--use_msa_server`:

**Basic Authentication:**
```bash
export BOLTZ_MSA_USERNAME=myuser
export BOLTZ_MSA_PASSWORD=mypassword
tt-bio predict ... --model boltz2 --use_msa_server
```

**API Key Authentication:**
```bash
export MSA_API_KEY_VALUE=your-api-key
tt-bio predict ... --model boltz2 --use_msa_server
```

## Command-Line Options

Model-specific options are labelled below.

**Common Options:**

| Option | Default | Description |
|--------|---------|-------------|
| `--model` | `boltz2` | `boltz2`, `esmfold2`, `esmfold2-fast` (single-sequence ESMFold2; protein / DNA / RNA / ligand complexes), `protenix-v1` / `protenix-v2` (AlphaFold3-family folder; protein / RNA / DNA / ligand complexes; `protenix-v1` is upstream's v0.5.0 base checkpoint at 4 trunk recycles, `protenix-v2` the wider one at 10), `openfold3` (AlphaFold3-family folder; protein / RNA / DNA polymers, optional templates, `OF3_CKPT` weights), `openbind` (OpenBind-0; the OpenFold3 stack on upstream v0.5.0 weights, protein-ligand co-folding, `TT_BIO_OPENBIND` weights), `opendde` / `opendde-abag` (antibody-antigen co-folding on the Protenix-v2 stack plus a structural-token expander; `opendde-abag` selects the antibody-antigen checkpoint; protein-only for now), or `rf3` (RoseTTAFold3, AlphaFold3-family folder; protein / RNA / DNA / ligand complexes) |
| `--out_dir` | `./` | Output directory |
| `--cache` | `~/.boltz` | Weight cache directory. Whole-repo models (ESMFold2, ESMC, SaProt, OpenDDE) use the Hugging Face cache; `TT_BIO_CACHE` moves both, see [docs/weights.md](docs/weights.md) |
| `--accelerator` | `tenstorrent` | **(Boltz-2)** `tenstorrent`, `cpu`, or `gpu`; other models run on Tenstorrent |
| `--recycling_steps` | model-specific | 3 for Boltz-2 and OpenFold3; 4 for Protenix-v1 (its checkpoint's own `N_cycle`); 10 for Protenix-v2/OpenDDE/ESMFold2/RF3 (the ESMFold2 paper's benchmark setting). Boltz-2, ESMFold2 and the OpenFold3 family run one more trunk cycle than asked and accept 0; Protenix, OpenDDE and RF3 count cycles, so their smallest value is 1 |
| `--sampling_steps` | model-specific | Requested diffusion sampling steps: 200 for Boltz-2/Protenix-v1/Protenix-v2/OpenFold3/OpenDDE; 100 for ESMFold2 (executes 68 after the sigma-schedule clip, the paper's protocol) |
| `--diffusion_samples` | `1` | Number of structure samples. Device memory stays flat past the chunk width and time grows linearly, see [docs/sample-scaling.md](docs/sample-scaling.md) |
| `--partial_t` | `0` | rf3 only. Schedule index the diffusion rollout starts at, so it refines `--partial_structure` instead of folding from scratch. Higher stays closer to that structure |
| `--partial_structure` | none | rf3 only. The `.cif`/`.pdb`/`.json` structure `--partial_t` refines. It supplies the sequences too, so no MSA is attached |
| `--early_stop_plddt` | none | rf3 only. Abandon a target after the first trunk recycle if its mean pLDDT is below this. Writes no structure; the results entry carries `early_stopped` |
| `--max_parallel_samples` | `5` | **(Boltz-2/Protenix/OpenDDE)** Diffusion samples denoised in one batched forward. Device memory grows linearly in it; when the chip refuses a batch the fold halves it on its own, down to one sample, instead of failing. ESMFold2 sizes its own chunk to free memory; OpenFold3, OpenBind and RF3 denoise one sample at a time |
| `--output_format` | `cif` | `cif` or `pdb`. A PDB has one column for the chain id, so a longer name is rewritten `A`, `B`, `C`... and the originals go into a `REMARK 999` block; `cif` keeps them as submitted. See [docs/model-capabilities.md](docs/model-capabilities.md#outputs) |
| `--seed` | `0` | Random seed for the diffusion sampler |
| `--trace` | `False` | **(Protenix-v1/Protenix-v2/OpenDDE)** Replay a captured trace of the per-step diffusion device stream instead of dispatching it from the host every step. The output is identical to a run without it. On Wormhole at 512 tokens it did not change the end-to-end time, and it reserves 0.2-0.3 GB of device memory |
| `--diffusion_trace` | `False` | **(Boltz-2)** The same for Boltz-2's diffusion DiT stream; `tt-bio design --model boltzgen` takes the same flag |
| `--write_pde` | `False` | **(Boltz-2)** Write the PDE matrix to its own `<name>_pde.npz`. The Protenix family and OpenDDE put PDE next to PAE in one file under `--write_pae` instead |
| `--write_embeddings` | `False` | **(Boltz-2)** Write the `s`/`z` embeddings per target |
| `--override` | `False` | Re-run from scratch, ignoring cached files |
| `--debug` | `False` | Show all raw output from the hardware and libraries instead of the progress display |
| `--log` | `False` | With `--debug`, also print what each device is currently working on |
| `--use_msa_server` | auto | Use the online ColabFold API; auto-enabled for Boltz-2/Protenix-v1/Protenix-v2/OpenFold3/OpenBind-0/OpenDDE/RF3 when no local DB is found |
| `--single_sequence` | `False` | **(Boltz-2/Protenix-v1/Protenix-v2/OpenFold3/OpenDDE)** Skip all MSA requests; lower accuracy |
| `--msa_endpoint` | none | Fetch unpaired MSAs from a `tt-bio msa-server`. A complex is not paired through it unless its paired MSA is already in `--msa_dir` |
| `--write_pae` | `False` | **(Protenix-v1/Protenix-v2/OpenDDE)** Write the token-token PAE/PDE matrices to `<name>_pae.npz` |
| `--use_potentials` | `False` | **(Boltz-2)** Apply physical constraints |
| `--affinity_mw_correction` | `False` | **(Boltz-2)** Apply MW correction to affinity |
| `--num_devices` | `0` | Number of TT devices (0=all available) |
| `--device_ids`, `--devices` | all | Comma-separated TT device IDs (e.g. `0,2`); `--devices` is the shorter alias (matches `tt-bio embed`) |
| `--host_threads` | all cores | Total CPU threads this process may use, split across its cards. Set it when you run several single-card predicts side by side on one host: each one otherwise sizes its thread pools to every core and they fight for the CPU. Use cores ÷ concurrent predicts. At two threads per card or fewer the pools also stop spinning through device syncs ([Tuning flags](docs/tuning-flags.md)) |
| `--fast` | `False` | Makes some operations use a lower-precision numeric format that runs faster; accuracy is typically very close |
| `--report-energy` | `False` | **(Boltz-2)** Enables optional energy profiling for one TT device (requires `tt-mgmt` add-on); writes `power_profile.csv` and `power_profile.png` |
| `--energy-metric` | `both` | **(Boltz-2)** Choose power channel(s): `tdp`, `input`, or `both` |
| `--energy-sample-hz` | `20.0` | **(Boltz-2)** Sampling rate in Hz for both `power_w` and `input_power_w` channels |

**Affinity-Specific Options (Boltz-2):**

| Option | Default | Description |
|--------|---------|-------------|
| `--sampling_steps_affinity` | `200` | Sampling steps for affinity |
| `--diffusion_samples_affinity` | `5` | Number of affinity samples |

**MSA Options** (Boltz-2, Protenix-v1, Protenix-v2, OpenFold3, OpenBind-0, OpenDDE and RF3 use an MSA by default; ESMFold2 only when requested):

| Option | Default | Description |
|--------|---------|-------------|
| `--msa_db_path` | auto-detect | Path to local ColabFold database (`~/.boltz/msa_db` if present), e.g. `--msa_db_path /data/colabfold_db` |
| `--msa_dir` | `<out_dir>/msa` | MSA cache directory. Point it at a shared persistent path to reuse `{seq_hash}.a3m` across runs |
| `--msa_cache_only` | `False` | Treat `--msa_dir` as the only MSA source: never search, and fail rather than quietly fold a chain single-sequence |
| `--use_envdb` | `False` | Also search environmental database |
| `--use_msa_server` | auto | Use ColabFold API for MSA (auto-enabled when no local DB is found) |
| `--single_sequence` | `False` | Fold without an MSA (Boltz-2/Protenix-v1/Protenix-v2/OpenFold3/OpenDDE) |
| `--msa_server_url` | `https://api.colabfold.com` | MSA server URL |
| `--msa_pairing_strategy` | `greedy` | `greedy` or `complete` |
| `--max_msa_seqs` | `8192` | Maximum MSA depth. The default applies to Boltz-2 only. Unless you set it, the other models read what their upstream reads: ESMFold-2, Protenix, OpenDDE, OpenFold3 and OpenBind up to 16384 rows, RF3 1024 rows drawn per recycle. Each fold reports the depth it used as `msa_depth` |
| `--subsample_msa` | `False` | Subsample MSA |
| `--num_subsampled_msa` | `1024` | Number of subsampled sequences |

**MSA Database Setup Options:**

| Option | Default | Description |
|--------|---------|-------------|
| `--db` | `uniref30` | `uniref30` (~500GB), `envdb` (~800GB), or `all` |
| `--path` | `~/.boltz/msa_db` | Where to store the databases |
| `--install-tools` | `True` | Auto-install missing `mmseqs`/`colabfold_search` |

## Tuning Flags

The engine ships its device optimizations on. Each one is an environment variable you can set to
`0` to fall back to the path it replaced, which is what you want if you are bisecting a result.
"Bit for bit" means the same structure, byte for byte; a flag that is not bit-exact says how far it
moves a structure, next to the seed-to-seed spread.

| Flag | Default | What it does |
|------|---------|--------------|
| `BOLTZ2_TOKEN_DIT_SDPA` | on | Boltz-2 only: token-level DiT attention as one fused SDPA. Not bit-exact: the kernel holds its exponentiated scores in bf16. |
| `TT_BIO_ATOM_AXIS_BUCKET` | on | Sizes the atom axis on the real atom count instead of assuming every token is a tryptophan. Byte-identical at 298 residues; at 512 it reassociates one matmul's contraction. |
| `TT_BIO_ATOM_SHIFT_GATHER` | on | Builds each atom's attention key window by slicing instead of a matrix multiply. Bit for bit. |
| `TT_BIO_DEVICE_CONDITIONING` | on | Boltz-2 only: runs the diffusion conditioning's pair track on the card. Not bit-exact (device bf16 where the host path was fp32); scored against 1HCL it is flat to slightly closer. |
| `TT_BIO_DEVICE_CONFIDENCE` | on | Boltz-2 only: assembles the confidence head's pair input on the card. Coordinates are bit-identical; pLDDT moves by at most 0.362 of 100. |
| `TT_BIO_DEVICE_CONF_HEADS` | on | Boltz-2 only: runs the confidence head's pae/pde projections on the card, so 67.1 MB of bin logits at 512 residues becomes 2.1 MB on Wormhole and 16.8 MB on Blackhole. Coordinates and per-residue pLDDT are bit-identical; pTM shifts by 0.0031 at 512 residues against a 0.0758 four-seed spread. |
| `TT_BIO_DEVICE_ZINIT` | on | Boltz-2 only: builds the trunk's `z_init` pair tensor on the card. Not bit-exact: moves a 298 aa structure 0.264 Å all-atom inside its 0.35 Å bar; CA-lDDT against 1HCL is flat over eight seeds. |
| `TT_BIO_DIT_COND_HOIST` | on | Hoists the token diffusion transformer's conditioning out of the layer loop; RF3's token DiT inherits it. Not bit-exact: one bf16 rounding changes order. |
| `TT_BIO_FUSE_BIAS_STACKS` | on | Boltz-2 only: builds the diffusion conditioning's per-layer bias stack in one pass. Not bit-exact: moves a 298 aa structure 0.218 Å all-atom, inside its 0.35 Å bar. |
| `TT_BIO_FUSE_MASK_ADD` | on | Gated-residual write-back as one `ttnn.addcmul`. Bit for bit. |
| `TT_BIO_FUSE_NORM_RESIDUAL` | on | Passes an add whose only consumer is a layer norm to the norm as its residual input. Bit for bit. |
| `TT_BIO_FUSE_SCALE_ADD` | on | Attention's scale-then-bias as one `ttnn.addalpha`, fp32 operands only. Bit-identical at every call shape. |
| `TT_BIO_GATE_GRANULARITY` | 2 | Tiles per DST acquire in the reblock-permute gate kernel. Bit for bit at every value. |
| `TT_BIO_HOST_LEVERS` | on | Master switch for the host-side Boltz-2 levers (`TT_BIO_FUSE_BIAS_STACKS` and `TT_BIO_HOST_BLOCK_PAIRWISE`); `0` takes the host path for all of them. |
| `TT_BIO_MSA_LADDER` | on | Boltz-2 and BoltzGen: pads the MSA depth to the smallest of 64, 128, 256, 512, 1024 that holds the alignment instead of always 1024. Not bit-exact; scored against 1HCL it is as accurate or closer. |
| `TT_BIO_OPM_LEGACY_LAYOUT` | off | Restores `OuterProductMean`'s old output stage in every model that builds it (Boltz-2, BoltzGen, Protenix, OpenFold3, RF3, AF2). The default differs by one bf16 step: 0.29 to 1.59 A worst pseudo-domain on the hinged 512-residue fixture, against 1.09 to 1.42 A for the old path against its own seeds. |
| `TT_BIO_PAIR_FFN_L1_FC1` | on | ESMFold2 only: keeps the pair transition's first matmul in L1 up to 512 residues. Bit for bit. |
| `TT_BIO_PAIR_INPLACE` | on | Writes a too-big pair tensor's blocks back on the card instead of assembling on the host. Bit for bit. Only OpenDDE's structural-token refiner reaches that size at 1536 residues or below. |
| `TT_BIO_PWA_BATCH_HEAD_WEIGHTS` | on | All heads' MSA row weights from one projection of the pair tensor. Bit for bit. |
| `TT_BIO_REBLOCK_PERMUTE_GATED` | on | Folds a triangle multiplication's chunk and sigmoid gates into its channel move. Bit for bit. |
| `TT_BIO_RESIDUAL_L1` | on | Two Pairformer residual updates go to L1 instead of DRAM, up to 512 residues. Bit for bit. |
| `TT_BIO_SDPA_ADD_GRANULARITY` | auto | Batches the fused SDPA kernel's running-sum/max and mask adds. Bit for bit at every granularity. |
| `TT_BIO_SDPA_FUSED_LARGE_S` | on | Triangle attention through the fused mask kernel above 1024 tokens. Not bit-exact above 1024: moves a 1536-residue structure 1.007 Å where a different seed moves it 36.6 Å. Untouched at and below 1024. |
| `TT_BIO_SDPA_GRID_Q_CHUNK` | on | Sizes each attention's query chunk to the card's compute grid. Bit for bit. |
| `TT_BIO_SDPA_WIDE_K` | on | Wider SDPA key chunk at twenty padded token lengths (288 to 1504) the shipped chunk does not divide. Not bit-exact there: moves a 686-residue structure 0.060-0.146 Å where a different seed moves it 3.69-7.28 Å. Every other length is unchanged. The fold-level speedup is unquoted because its one arm recorded no clock: [`docs/sdpa-wide-k-parity.md`](docs/sdpa-wide-k-parity.md). |
| `TT_BIO_TOKEN_BUCKET` | on | The one global off switch for token bucketing; `0` folds the exact token count instead of a padded one, slower. The legacy per-model flags are ANDed with it. |
| `TT_BIO_TRANSITION_L1_ROWS` | on | Blackhole only: sizes each transition's row block from the card's L1 budget (Boltz-2, BoltzGen, OpenFold3). Bit-identical on the reference fixture at 298, 512, 768 and 1024 residues; 0.165 Å on the no-MSA prot leg. |
| `TT_BIO_TRIATT_FUSED_QKVG` | on | Projects a triangle attention's query, key, value and gate in one pass. Bit for bit. |
| `TT_BIO_TRIATT_FUSED_QKVGB` | on | Adds the pair-bias projection to that pass; chains of 32 residues or fewer keep it separate. Bit for bit. |
| `TT_BIO_TRIATT_HIFI_PAD_UP` | 2 | Pads a triangle attention axis the fused HiFi kernel cannot tile (32 x a prime, such as 544 or 608) up to one it can, by at most this many 32-wide tiles, and masks the added keys; `0` turns it off. Same attention over the same keys. Lets BindCraft 2 on a p150a run 608 to 864 tokens where 608 was refused; moves an OpenFold3 structure at 585 residues 0.03 A where a different seed moves it 0.83 A. |
| `TT_BIO_TRIATT_SDPA_HIFI_AB` | on for the OpenFold3 trunk, off elsewhere | Triangle attention through the fused SDPA at HiFi4. Not bit-exact below 1088 residues: at 298 it moves a structure 2.58-8.01 A CA where a different seed moves it 5.23-9.78 A, 0.396 A closer to the deposited structure. Per construction site: `openfold3.trunk` forces it on, `-openfold3.trunk` off, `all` / `-all` every site. |
| `TT_BIO_TRIMUL_FUSED_GOUT` | on | A triangle multiplication's output gate as a second output of its input projection. Bit for bit. |
| `TT_BIO_TRIMUL_GP_BANK_SPLIT` | on | Interleaves the fused input projection's gate and value columns across DRAM banks. Bit for bit. |
| `TT_BIO_TRIMUL_INPROJ_ROWBLOCK_NORM` | on | Row-blocked input projection for a pair tensor over 3 GiB; OpenDDE's refiner only. Bit for bit. |
| `TT_BIO_TRIMUL_MASK_AFTER_MOVE` | on | Applies a triangle multiplication's pair mask after the channel move. Bit for bit. |
| `TT_BIO_TRIMUL_MASK_L1` | on | Keeps a triangle multiplication's pair mask in L1. Bit for bit. |
| `TT_BIO_TRIMUL_MM_TRANSPOSE` | on | The matmul takes a triangle multiplication's operand transpose. Bit for bit; `--fast` keeps the separate op. |
| `TT_BIO_TRIMUL_TAIL_F1` | on | Output projection, gate projection and gate multiply as one kernel. Bit for bit. |
| `TT_BIO_TRIMUL_TAIL_F1_L1_OUT` | on | Lands that fused tail's product in L1. Bit for bit. |
| `TT_PROTENIX_CONF_DEVICE` | off | Protenix-v2 and OpenDDE: runs the confidence head on the card. Off because pLDDT is precision-sensitive here; coordinates are unaffected. |

What each one is worth in seconds, at what clock, and how it was measured: [`docs/tuning-flags.md`](docs/tuning-flags.md).

## Running tt-bio on many machines

Each machine runs its own `tt-bio controller`, which listens on 127.0.0.1 and keeps a
worker on every chip. Work reaches it through `--controller`, and whatever spreads jobs
across machines sits on top: a platform, a cluster scheduler, or the fifty-line
[`examples/many_hosts.py`](examples/many_hosts.py), which folds a directory across
several hosts over ssh. [docs/multi-host.md](docs/multi-host.md) is the contract a
scheduler builds against: the endpoints, what a host advertises, the lease and how a
result settles once.

## Optional: Energy Measurement (Boltz-2)

Use `--report-energy` to profile energy during prediction:

```bash
tt-bio predict examples/686.yaml --model boltz2 --override --device_ids 0 --report-energy --energy-metric both --energy-sample-hz 5
```

Behavior:
- Select metric channel(s) with `--energy-metric` (`tdp`, `input`, `both`)
- Uses one sampling rate (`--energy-sample-hz`, default 20 Hz)
- Supports only Tenstorrent runs with one selected device
- Records two power channels when available:
  - `power_w`: `tt-mgmt` UMD telemetry power (TDP channel)
  - `input_power_w`: `tt-mgmt` UMD telemetry input power
- Requires optional `tt-mgmt` installation:
  - `git clone --recursive https://github.com/aperezvicente-TT/tt-mgmt.git`
  - `pip install -e ./tt-mgmt`
- Prints energy summary metrics for selected channels
- Always writes:
  - `power_profile.csv`
  - `power_profile.png`

## Training

Train OpenFold3 with the same forward the inference path uses:

```bash
pip install 'tt-bio[tenstorrent,train]'
curl --create-dirs -o ~/.boltz/of3-p2-155k.pt \
  https://openfold3-data.s3.amazonaws.com/openfold3-parameters/of3-p2-155k.pt
tt-bio train --model openfold3
```

The `curl` line fetches the OpenFold3 weights (2.29 GB, no login). tt-bio does not download them
for you because the consortium publishes no parameter licence; see
[`docs/weights.md`](docs/weights.md). With no data given, `tt-bio train` fetches upstream's
8-structure training sample (73 MB, from OpenFold3's public bucket, checked file by file against
a shipped sha256 list), trains the model's own weights at a 384-token crop for one pass over it,
and writes to `runs/openfold3`. A step takes about 42 s on one p300c chip. Run the same command again and it
resumes from the last checkpoint; run it with different settings on the same `--out` and it
refuses rather than mixing two runs.

Your own data is OpenFold3's training-set layout, a `pdb_training_set/` directory beside one
`training_cache*.json`, featurised on the fly by upstream's own pipeline:

```bash
tt-bio train data/ --model openfold3 --steps 2000
tt-bio train data/ --model openfold3 --steps 2000 --chips 1,2 --global-batch 2
tt-bio train data/ --model openfold3 --dry-run      # does it fit, and how long a step takes
```

A run ends by writing `<out>/weights.pt`, the trained weights in the same format as the shipped
checkpoint, whether it ran on one chip or several. Fold with them by passing that file to
`predict`:

```bash
tt-bio predict target.yaml --model openfold3 --checkpoint runs/openfold3/weights.pt
```

Without `--checkpoint`, `predict` folds with the shipped weights exactly as before.

`--help` lists the next layer: `--steps`, `--chips`, `--global-batch`, `--lr`,
`--warmup-steps`, `--checkpoint-every`, `--tokens`, `--seed`. `--help-all` adds the expert
ones: precision (`--exact`), the objective, the recipe, and `--train` with the LoRA shape.
OpenFold3 trains its own weights only; `--train adapters` is refused for it with what to do
instead. [`docs/training.md`](docs/training.md#data) has the data layout and the rest.

**For an agent**, everything a run says it also writes, in JapanFold's vocabulary:

| file | what is in it |
|---|---|
| `<out>/status.json` | `status` (`running`, `succeeded`, `failed`), `step`, `steps`, `loss`, the chips, the clock sampled during the run, the checkpoint, and on failure `error: {title, detail}` |
| `<out>/progress.jsonl` | one row per step: `step`, `loss`, `lr`, `grad_norm`, `s` (wall seconds, the first row includes model load), `healthy` |
| `<out>/weights.pt` | the trained weights, for `tt-bio predict --checkpoint`; `status.json` names it as `weights` |
| `<out>/run.json` | the full record after success: history, provenance, the per-rank data-parallel record |
| `<out>/traceback.txt` | the traceback, when the run failed |

The exit code is 0 on success and 1 on failure, with a one-line reason on stderr. A step whose
loss or gradient norm is not finite stops the run rather than training on it. After a resume
the progress rows after the last checkpoint appear twice, once from each process.

Underneath the command are three more levels, and you pick the one that matches what you want
to write:

| Level | You write | You own |
|---|---|---|
| `tt-bio train ...` | a command line | the config |
| `train.finetune(...)` | one call | the objective |
| `plan`, `batches`, `objectives`, `AdamW`, `Checkpointer`, `Mesh`, `trainable` | the loop | the `for` statement |
| `tt_bio.autograd` + `train.gradcheck` | an op and its backward | the gradient |

Dropping a level is not a rewrite. `tt-bio train --show-recipe` prints the loop the command
runs, the body of `train.recipes.source("default")`, written only in names the level below
exports, and a test keeps it that way: if the recipe ever needed a private hook, the test fails
and the hook becomes public. More chips is `mesh=train.Mesh({"dp": [1, 2]})` one level down,
and adapters instead of weights is `train="adapters"`; one loop body serves both, which the
escape-hatch test checks instruction for instruction.

**What works today:** OpenFold3 end to end from the command line (data, featurisation, the
weights, training its weights, checkpoints, resume, the agent files, folding with the result), the dry run, gradient
checking, and data parallelism across the chips in one box. **What does not:** OpenFold3 is the
only model that ships a training featuriser; `--model` offers only it, and another model
registers its own with `tt_bio.train.catalogue.register`. LoRA adapters are not offered for
OpenFold3: they have never been measured on it, and `predict` loads full weights only.

`tt-bio train` follows OpenFold3's optimizer setup rather than Adam's library defaults, which
differ in three places that no loss curve shows: `betas=(0.9, 0.95)`, no weight decay, and the
AlphaFold 2 learning-rate schedule. Each is an argument, and the loop clips every sample
separately, so a batch of 8 is 8 forwards per step. See
[`docs/training.md`](docs/training.md) for what each one costs if you get it wrong.

A training step runs entirely on the device, with the softmax backward in fp32. On OpenFold3
that keeps the gradient inside its accuracy bar against upstream. `--exact` computes softmax and
layer norm in float64 on the host instead, a diagnostic reference that makes a step many times
slower. Inference is unaffected. See
[`docs/training.md`](docs/training.md#training-runs-on-the-device-float64-is-a-diagnostic).

ABodyBuilder3 wants its data staged first: `data.tar.gz` from Zenodo `10.5281/zenodo.11354577`,
extracted so that `structures/structures/*.pt` sits under the path you pass. Fine-tuning also
wants their checkpoint beside it; without one the command refuses rather than adapting a random
initialisation. To train from scratch instead, `scripts/abb3_port/repro.py` is the reproduction's
own entry point.

Four things the API enforces rather than documents, because each is a bug we hit:

- `plan()` answers from measured numbers or returns `UNMEASURED`. It refuses a crop size whose
  forward is measured to run out of memory instead of estimating one, and it will not report a
  4-chip step time from a 2-chip measurement. Protenix's own 384-token crop is one of the
  refusals.
- The optimizer refuses a bfloat16 master copy of the weights. An update accumulated at
  bfloat16 stops moving the weight while the gradient still looks healthy.
- `opt.step()` raises if you spread training over several chips and never gave it a way to
  combine their gradients. Otherwise you train one model per chip and see one loss curve.
- Every run records the clock it actually ran at, sampled during the work, plus the seed and
  the commit. A time without its clock is not a measurement on this hardware.

Global batch is always yours to set and is never derived from how many chips you have, so a
recipe means the same thing on a bigger box.

Training on several chips is one flag, `--chips 2`, or one argument,
`mesh=train.Mesh({"dp": [0, 1]})`. Under it a launcher runs your program once per chip and sums
the gradients between them, the way `torchrun` does, so the program has to be re-runnable and
must not open a card before the `finetune` call. Both are checked before anything starts. Two
p150a chips on a QuietBox measured **1.96x** at a 0.33 MB adapter gradient and **1.70x** at 5.24
MB, both at 1350 MHz; the gap is host-side Adam contending between the two processes, not the
exchange, which costs 2.3 % of the step. Training OpenFold3's weights at a 384-token crop, two p300c chips on a QuietBox take 46.5 s per two-sample step against 77.4 s on one, **1.66x**, both at 1350 MHz; each step exchanges a 2.1 GB gradient through shared memory. ABodyBuilder3 on one QuietBox's four chips measures **3.69x at 92 %**, 7.65 s a step against 28.19 s on one, with one master hash across the four ranks. It needs the host's cores split between the ranks, which the launcher does by default; left at torch's own width every rank claims all 16 cores and the same step takes 931 s.
Two boxes work too and the link is what it costs: the same two chips split across two QuietBoxes give **1.554x** where one box gives 1.887x, all of the difference being the 28.4 MB gradient exchange over the second box's WiFi. That box has no cable, so 77.7 % efficiency is close to the floor for this; a wired link would cut the exchange to about 2 % of the step. Name the other box on the mesh, `train.Mesh({"dp": [0, 1]}, hosts=["ttuser@tt-quietbox2"])`, and it composes the rendezvous the transport takes. Starting the ranks on both boxes is still your own launcher rather than `--chips`.

Every run carries the check that makes a multi-chip number mean something. The ranks' weights
must stay identical, so the launcher compares every rank's master weights at the end and refuses
a run where they differ, and it reads each rank's chip off that rank's own open file descriptors
rather than trusting `TT_VISIBLE_DEVICES`, which names a different number than the device node.
Both failures look like a healthy run otherwise.

Tiers, cut lines, the escape-hatch test and where `plan()` gets its numbers:
[`docs/training.md`](docs/training.md). What one gradient step through a pairformer stack costs,
how it scales with depth and width, and where the headroom is:
[`docs/gradient-step-cost.md`](docs/gradient-step-cost.md).

## Cite

If you use this code or the models in your research, please cite the following papers:

```bibtex
@article{passaro2025boltz2,
  author = {Passaro, Saro and Corso, Gabriele and Wohlwend, Jeremy and Reveiz, Mateo and Thaler, Stephan and Somnath, Vignesh Ram and Getz, Noah and Portnoi, Tally and Roy, Julien and Stark, Hannes and Kwabi-Addo, David and Beaini, Dominique and Jaakkola, Tommi and Barzilay, Regina},
  title = {Boltz-2: Towards Accurate and Efficient Binding Affinity Prediction},
  year = {2025},
  doi = {10.1101/2025.06.14.659707},
  journal = {bioRxiv}
}

@article{stark2025boltzgen,
  author = {Stark, Hannes and Faltings, Felix and Choi, MinGyu and Xie, Yuxin and Hur, Eunsu and O'Donnell, Timothy John and Bushuiev, Anton and U{\c c}ar, Talip and Passaro, Saro and Mao, Weian and Reveiz, Mateo and Bushuiev, Roman and Pluskal, Tom{\'a}{\v s} and Sivic, Josef and Kreis, Karsten and Vahdat, Arash and Ray, Shamayeeta and Goldstein, Jonathan T. and Savinov, Andrew and Hambalek, Jacob A. and Gupta, Anshika and Taquiri-Diaz, Diego A. and Zhang, Yaotian and Hatstat, A. Katherine and Arada, Angelika and Kim, Nam Hyeong and Tackie-Yarboi, Ethel and Boselli, Dylan and Schnaider, Lee and Liu, Chang C. and Li, Gene-Wei and Hnisz, Denes and Sabatini, David M. and DeGrado, William F. and Wohlwend, Jeremy and Corso, Gabriele and Barzilay, Regina and Jaakkola, Tommi},
  title = {BoltzGen: Toward Universal Binder Design},
  year = {2025},
  doi = {10.1101/2025.11.20.689494},
  journal = {bioRxiv}
}

@article{wohlwend2024boltz1,
  author = {Wohlwend, Jeremy and Corso, Gabriele and Passaro, Saro and Getz, Noah and Reveiz, Mateo and Leidal, Ken and Swiderski, Wojtek and Atkinson, Liam and Portnoi, Tally and Chinn, Itamar and Silterra, Jacob and Jaakkola, Tommi and Barzilay, Regina},
  title = {Boltz-1: Democratizing Biomolecular Interaction Modeling},
  year = {2024},
  doi = {10.1101/2024.11.19.624167},
  journal = {bioRxiv}
}

@misc{candido2026language,
  author = {Candido, Salvatore and Hayes, Thomas and Derry, Alexander and Rao, Roshan and Lin, Zeming and Verkuil, Robert and others},
  title = {Language Modeling Materializes a World Model of Protein Biology},
  year = {2026},
  url = {https://biohub.ai/papers/esm_protein.pdf},
  note = {Preprint; ESMC / ESMFold2}
}

@misc{protenix2025,
  author = {{ByteDance AML AI4Science Team}},
  title = {Protenix: An AlphaFold3 Reproduction for Biomolecular Structure Prediction},
  year = {2025},
  url = {https://github.com/bytedance/Protenix}
}

@misc{openfold3,
  author = {{OpenFold Consortium}},
  title = {OpenFold3: An Open-Source Reproduction of AlphaFold3},
  year = {2026},
  url = {https://github.com/aqlaboratory/openfold-3}
}

@article{butcher2025rfdiffusion3,
  author = {Butcher, Jasper and Krishna, Rohith and Mitra, Raktim and Brent, Rafael Isaac and Li, Yanjing and Corley, Nathaniel and Kim, Paul T and Funk, Jonathan and Mathis, Simon Valentin and Salike, Saman and Muraishi, Aiko and Eisenach, Helen and Thompson, Tuscan Rock and Chen, Jie and Politanska, Yuliya and Sehgal, Enisha and Coventry, Brian and Zhang, Odin and Qiang, Bo and Didi, Kieran and Kazman, Maxwell and DiMaio, Frank and Baker, David},
  title = {De novo Design of All-atom Biomolecular Interactions with RFdiffusion3},
  year = {2025},
  doi = {10.1101/2025.09.18.676967},
  journal = {bioRxiv}
}
```

In addition if you use the automatic MSA generation, please cite:

```bibtex
@article{mirdita2022colabfold,
  title={ColabFold: making protein folding accessible to all},
  author={Mirdita, Milot and Sch{\"u}tze, Konstantin and Moriwaki, Yoshitaka and Heo, Lim and Ovchinnikov, Sergey and Steinegger, Martin},
  journal={Nature methods},
  year={2022}
}
```

## License

tt-bio is released under the MIT License (see [`LICENSE`](LICENSE)) and is built on the MIT-licensed Boltz-2 / Boltz-1 code. It bundles third-party code, each under its upstream license: the ESMFold2 host-side reference under `tt_bio/_vendor/` (the `esm` pipeline, MIT, © Chan Zuckerberg Biohub; and the HuggingFace ESMFold2 model definition, Apache-2.0), the OpenFold3 host-side data pipeline under `tt_bio/_vendor/openfold3/` (Apache-2.0, OpenFold Consortium), the BoltzGen binder-design source under `tt_bio/boltzgen/` (MIT, © Hannes Stärk), and RF3's host featurizer under `tt_bio/_vendor/rf3/`, `tt_bio/_vendor/foundry/` and `tt_bio/_vendor/atomworks/` (BSD-3-Clause, University of Washington / Institute for Protein Design). Protenix-v2, OpenFold3's on-device model, RFdiffusion3, RF3's on-device model, and PXDesign's are independent ttnn reimplementations (no upstream compute code is vendored); Protenix-v2's weights download from ByteDance's Hugging Face mirror under Apache-2.0, RFdiffusion3's and RF3's checkpoints download directly from the Institute for Protein Design (BSD-3-Clause), PXDesign's generator checkpoint downloads from ByteDance's release host under Apache-2.0 and its AF2-IG selection stage reads DeepMind's AlphaFold2 monomer pTM parameters (CC BY 4.0, from `storage.googleapis.com/alphafold/`), and OpenFold3's `of3-p2-155k.pt` and OpenBind-0's `of3-ob-2025-06-30-174k.pt` are the consortium's ungated public parameter releases, which you fetch yourself (the project is Apache-2.0, stated by upstream as free for academic and commercial use; the consortium publishes no separate parameter license). See [`NOTICE`](NOTICE) for sources, versions, and modifications.
