Metadata-Version: 2.5
Name: attr-eomt
Version: 1.1.2
Summary: Standalone EoMT (Encoder-only Mask Transformer) for instance segmentation, with DINOv2 init, COCO training/validation and inference.
Project-URL: Homepage, https://github.com/imagra93/attr-eomt
Project-URL: Documentation, https://imagra93.github.io/attr-eomt
Project-URL: Repository, https://github.com/imagra93/attr-eomt
Project-URL: Issues, https://github.com/imagra93/attr-eomt/issues
Author: attr-eomt contributors
License-Expression: Apache-2.0
License-File: LICENSE
Keywords: coco,dinov2,eomt,instance-segmentation,transformers
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.10
Requires-Dist: huggingface-hub>=0.23.0
Requires-Dist: matplotlib>=3.5.0
Requires-Dist: numpy>=1.19.0
Requires-Dist: opencv-python>=4.8.0
Requires-Dist: pillow>=9.1.0
Requires-Dist: pycocotools>=2.0.0
Requires-Dist: pyyaml>=6.0
Requires-Dist: requests>=2.25.0
Requires-Dist: scipy>=1.7.0
Requires-Dist: supervision<0.30,>=0.22.0
Requires-Dist: torch>=2.4.0
Requires-Dist: torchao>=0.7.0
Requires-Dist: torchvision>=0.19.0
Requires-Dist: tqdm>=4.65.0
Requires-Dist: transformers>=5.1.0
Provides-Extra: dev
Requires-Dist: build>=1.0; extra == 'dev'
Requires-Dist: pytest>=7.0; extra == 'dev'
Requires-Dist: twine>=5.0; extra == 'dev'
Provides-Extra: logging
Requires-Dist: tensorboard>=2.10; extra == 'logging'
Requires-Dist: wandb>=0.15; extra == 'logging'
Description-Content-Type: text/markdown

<div align="center">

<img src="https://raw.githubusercontent.com/imagra93/attr-eomt/main/docs/assets/00-hero-banner.png" alt="attr-eomt — one DINOv2 encoder predicts instances plus independent per-instance attribute heads in a single pass, contrasted with flat combinatorial labels and a detector-plus-second-model pipeline" width="100%">

<p>
  <a href="https://pypi.org/project/attr-eomt/"><img src="https://img.shields.io/pypi/v/attr-eomt.svg?color=4ec9b0" alt="PyPI version"></a>
  <a href="https://pypi.org/project/attr-eomt/"><img src="https://img.shields.io/pypi/pyversions/attr-eomt.svg" alt="Python versions"></a>
  <a href="LICENSE"><img src="https://img.shields.io/badge/license-Apache--2.0-blue.svg" alt="License"></a>
  <a href="https://imagra93.github.io/attr-eomt"><img src="https://img.shields.io/badge/docs-annotated%20explainer-e2b341.svg" alt="Annotated explainer"></a>
</p>

**One query embedding, many independent labels.**

📖 **[Read the annotated explainer →](https://imagra93.github.io/attr-eomt)**

</div>

---

## What it is

**attr-eomt** is a standalone **EoMT** (Encoder-only Mask Transformer) for **instance
segmentation** and **object detection**, with one feature that sets it apart: **independent
per-instance attribute heads**. Alongside the mask/box + class output, it predicts one or
several orthogonal attributes for *every* detected instance — read straight off the **same**
per-query embedding the detector already computes. No second model, no second pass, and the
primary detection metric is untouched (the figure above tells the whole story).

It's a clean-room, Apache-2.0 reimplementation: the weights you train are yours to release.

```python
from eomt import EoMT

model = EoMT("l")                            # fresh large model (DINOv2 backbone)
model.train(data="coco", epochs=50)          # COCO 2017 auto-downloads if missing

model = EoMT("runs/train/eomt-l")            # reload a run — size/classes/heads auto-detected
model.predict("images/", plot=True)          # render masks/boxes + per-instance attributes
```

---

## Architecture

EoMT is a **DINOv2-with-registers ViT** whose last few transformer blocks are augmented
with a fixed set of **learnable queries** (the Mask2Former idea) — each query is one
"slot" that latches onto one object instance. After the encoder runs, every query emits
a single vector, the **per-query embedding** of shape `[B, Q, hidden]`. The whole model
is then just "turn that embedding into predictions": a **class head** for the primary
label and a **mask/box head** for geometry. It is **NMS-free**, so two overlapping
garments stay two distinct queries instead of being merged — the property that lets
attributes stay attached to the right instance.

The attribute heads add nothing to this picture except themselves: they tap the **exact
same embedding** (captured non-invasively with a forward hook), each a small classifier
on top.

This collapses what is classically a *two-stage* pipeline — detect, crop each box, run a
second classifier per crop — into a single pass. Attributes therefore cost only a thin head
each, see **full-image context** (not just a cropped box), and never inherit a second
model's cropping errors — the modern, single-stage formulation of the DETR / Mask2Former
lineage (see the figure at the top).

### Two model families: segmentation & detection

Both families share the same DINOv2 encoder, query mechanism, NMS-free matching and
auxiliary heads — they differ only in the head on top and what they output:

| family | `--task` | output | metric driving `best.pt` |
|--------|----------|--------|--------------------------|
| **instance** (default) | `instance` | per-instance **masks** + boxes + class | `segm/mAP` |
| **detect** | `detect`  | per-instance **boxes** + class (DETR-style box head, no masks) | `bbox/mAP` |

```python
EoMT("l").train(data="coco", family="instance")   # masks (default)
EoMT("l").train(data="coco", family="detect")     # boxes only
```

The family is recorded in the checkpoint, so `val` / `predict` pick the right
post-processing automatically. Everything below applies identically to both.

### Models & sizes

| size | backbone        | hidden | layers | heads | queries |
|------|-----------------|--------|--------|-------|---------|
| `s`  | DINOv2-small    | 384    | 12     | 6     | 100     |
| `b`  | DINOv2-base     | 768    | 12     | 12    | 200     |
| `l`  | DINOv2-large    | 1024   | 24     | 16    | 200     |

Default input is a patch-14-aligned square (`644 = 14 × 46`) so DINOv2 weights load 1:1.

### Compute & inference speed

Measured on a single **NVIDIA GeForce RTX 5090**, `644 × 644` input, batch size 1.
GFLOPs are multiply-accumulates at that resolution (attention included); latency /
throughput are the median over 50 runs after warm-up, under `torch.amp.autocast`
(fp16) — the package's own inference path.

**`instance` family** (masks + boxes + class):

| size | params | GFLOPs | latency (fp16) | throughput (fp16) | throughput (fp32) |
|------|--------|--------|----------------|-------------------|-------------------|
| `s`  | 24.0 M | 128    | 8.4 ms         | 119 img/s         | 70 img/s          |
| `b`  | 93.9 M | 430    | 17.4 ms        | 58 img/s          | 32 img/s          |
| `l`  | 317 M  | 1144   | 30.2 ms        | 33 img/s          | 15 img/s          |

**`detect` family** (boxes + class, no mask head):

| size | params | GFLOPs | latency (fp16) | throughput (fp16) | throughput (fp32) |
|------|--------|--------|----------------|-------------------|-------------------|
| `s`  | 22.7 M | 89     | 2.9 ms         | 348 img/s         | 120 img/s         |
| `b`  | 88.6 M | 276    | 5.3 ms         | 190 img/s         | 60 img/s          |
| `l`  | 308 M  | 881    | 13.6 ms        | 74 img/s          | 21 img/s          |

Dropping the mask-upsampling head makes `detect` substantially lighter and ~1.3–3×
faster. Figures are for the detector itself (backbone + queries + heads); the
attribute heads add a thin linear/MLP per head and are negligible by design.

---

## Factorizing the label space

This is the contribution. Conventional detectors fold every distinction into one flat
label space: an object's `type × viewpoint × occlusion × …` becomes a Cartesian product of
leaf classes that explodes combinatorially, starves each leaf of examples, and multiplies
the Hungarian matcher's targets. **attr-eomt factorizes instead** — a small, general primary
head plus independent attribute heads that **add, not multiply**.

Because the heads are independent, the primary taxonomy stays compact and every class keeps
its full sample count; attributes ride along for near-zero compute; and the model composes
`attribute × class` combinations that **never appear in the training data** — combinations a
flat label space cannot even represent.

### Example: clothing with per-instance attributes

One model segments each garment (primary classes like `vest_dress` / `short_sleeve_top`
/ `long_sleeve_dress` / `skirt` / `trousers` …) and, for **every** detection, reads off
four **independent** attribute heads — `scale` (`small` / `modest` / `large`),
`occlusion` (`no` / `slight` / `medium`), `zoom_in` (`no` / `medium` / `large`) and
`viewpoint` (`frontal` / `side` / `back`). The renderer prints the primary class + score
on the first row and each attribute + its confidence on the rows beneath it.

![Two people in dresses; each instance labelled with its garment class plus scale, occlusion, zoom and viewpoint attributes](https://raw.githubusercontent.com/imagra93/attr-eomt/main/docs/examples/sample.jpg)

The four attributes are *orthogonal* to the garment class — they vary independently —
which is exactly the case that's awkward to fold into the primary class space. The same
pattern fits any "class **plus** per-instance sub-labels" task: **retail shelves →
product + facing**, **documents → element + role**, **cells → type + health**.

> Trained on the public **[DeepFashion2](https://github.com/switchablenorms/DeepFashion2)**
> dataset (13 garment classes + 4 attribute heads) and rendered with the package's own
> renderer ([`eomt.visualize.draw_instances`](eomt/visualize.py)).

---

## Training — it rides on the detector's own match

Attributes never run their own matcher. Detection already solves "which query is
responsible for which ground-truth object" via the **Hungarian matcher**; attributes
simply reuse that same query→GT assignment and read the answer off the matched queries.

- **Embedding source.** Each head reads the per-query embedding — the input to EoMT's
  `class_predictor`, captured with a forward hook (`[B, Q, hidden]`).
- **Matching.** Supervision reuses EoMT's *own* Hungarian matcher
  (`model.eomt.criterion.matcher`), so every attribute is trained on the **same**
  query→GT assignment the detection loss used; the attribute is read *after* matching.
- **Gate.** An optional IoU gate drops barely-overlapping matched pairs (common early in
  training) so attributes only learn from queries that actually localize the object.
- **Loss.** Cross-entropy per head over matched queries, summed across heads and scaled
  by `aux_w` (default `1.0`), added to the detector loss. Empty-match batches contribute
  a graph-preserving zero, and missing labels use `ignore_index` and contribute nothing.
- **Checkpoint selection is unchanged.** The attribute "rides along": its per-head
  matched-query accuracy is shown live and written to `metrics.csv`, but never drives
  `best.pt` (still `segm/mAP` or `bbox/mAP`).
- **Inference.** Each result attaches `aux = {head: {"ids", "probs"}}` for the kept
  detections, and `predict(plot=True)` renders each attribute next to the class label
  using names stored in the checkpoint.

---

## Data format (auto-discovered from the COCO JSON)

Attributes live **inside the COCO annotations** — each annotation is already a
per-instance object, so alignment is automatic and `pycocotools` still parses it. Just
two additions to a standard COCO file; **no YAML changes** — heads (count, classes,
names) are discovered from the JSON, the same as `nc`.

**1. A top-level `attributes` list** — one entry per head, defining its vocabulary:

```jsonc
"attributes": [
  {
    "name": "scale",
    "categories": [
      {"id": 1, "name": "small"},
      {"id": 2, "name": "modest"},
      {"id": 3, "name": "large"}
    ]
  },
  {
    "name": "viewpoint",
    "categories": [
      {"id": 0, "name": "frontal"},
      {"id": 1, "name": "side"},
      {"id": 2, "name": "back"}
    ]
  }
]
```

**2. A per-annotation `attributes` map** — `{head: raw_id}` on each instance:

```jsonc
{
  "id": 1, "image_id": 42, "category_id": 1,
  "segmentation": [...], "bbox": [...], "area": 1234, "iscrowd": 0,
  "attributes": {"scale": 3, "viewpoint": 0}
}
```

Notes:

- Raw ids are remapped to a contiguous `0..n-1` per head (so `scale`'s `1`/`2`/`3` become
  `0`/`1`/`2`); `categories` may be omitted, in which case the id set is inferred.
- A **missing or out-of-vocab** per-annotation value is *ignored* (`-100`), not trained as
  class `0` — so a **partially tagged** dataset is valid: each head learns only from the
  instances that actually carry its value. A JSON with **no** `attributes` ⇒ detection-only,
  exactly as before.

**Class-conditional heads.** Give an attribute definition an optional `applies_to` list of
primary-class names or ids, and that head is only trained on — and only emitted for —
instances of those classes (hard routing on the primary class). Omit it and the head applies
to every class. So different attributes can attach to different classes, each with its own
label set, in one model:

```jsonc
"attributes": [
  {"name": "posture", "categories": [...], "applies_to": ["cat", "dog"]}
]
```

At inference a scoped head reports `ids = -1` ("not applicable") for detections whose class it
does not cover. The scope is stored in the checkpoint, so it survives reload.

**Sidecar format (optional).** You can keep the COCO JSON as plain, standard COCO and put the
attributes beside it instead of inside it: an `attributes.yaml` schema in the dataset root plus
`attributes/<split>.json` values keyed by annotation id (`{ann_id: {head: value}}`). If present
(and the JSON has no embedded `attributes`), it is merged in memory at load — so a plain COCO
dataset always works and the sidecar is picked up automatically when you add it. Embedded
`attributes` in the JSON take precedence.

A tiny, self-contained example (two heads, including a non-contiguous id set) lives in
[sample_data/](sample_data/).

---

## Cross-photo re-identification

Several photos of the same scene from different viewpoints, and some instance appears
in three of them — is that one instance seen three times, or three separate ones?
`infer_match()` answers that at **inference time, with no second model and nothing
retrained**: every detected instance already carries a fingerprint, the per-query
embedding that feeds the class head, and two detections are the same physical instance
when those embeddings agree.

```python
from eomt import EoMT

model = EoMT("runs/train/eomt-l/weights/best.pt")
result = model.infer_match("photos/subject_42/", group_by=("class",))

for rec in result["identities"]:
    print(rec["identity_id"], rec["class_name"], rec["attributes"], rec["num_photos"])
```

One call handles one folder = one subject. Each detection is stamped with an
`identity_ids` entry, `identities` summarizes each identity (dominant class, smoothed
attributes, how many photos it appears in), and `matches` records every candidate pair
— accepted or not, with the reason — so a threshold can be retuned from the JSON
without re-running inference. With `plot` you get a grid image: every photo, instances
colored by identity, plain lines linking the matched pairs (thickness = similarity, on
an absolute scale that does not vary with panel size), and a legend.

```bash
python scripts/match.py weights/best.pt photos/subject_42/
python scripts/match.py weights/best.pt photos/ --each-subdir   # a folder per subject
```

### How it works

It is the DeepSORT recipe — detect, describe, associate — applied *between photos*
instead of between video frames, which is also what [`track()`](eomt/engine/track.py)
does for video. Three EoMT properties make the fingerprint free:

1. the model is **NMS-free**, so two overlapping instances stay two distinct queries;
2. each query owns one instance, so its embedding describes that instance alone;
3. the embedding is already computed on the way to the class head.

Matching runs the **Hungarian** matcher per pair of photos, then links the accepted
pairs into identities with union-find. Two constraints keep that honest: one physical
instance appears **at most once per photo** (union-find would otherwise chain two
instances of the same photo together through a third), and a **merge guard** rejects a
merge whose cross-cut similarity falls below the threshold — without it, A↔B and B↔C
chain into one identity even when A and C are nothing alike.

### Gating and thresholds

`group_by` decides which pairs may match at all. The default `("class",)` means only
same-class instances compete. Adding attribute heads tightens it, which matters
whenever the primary class is coarser than the distinction you care about: if one class
covers instances that sit at different places on the subject, an attribute head that
separates them — `("class", "position")`, say — stops an instance in one location from
ever matching one in another, however alike they look. Any head in the checkpoint can
be named; unknown names raise before the first forward pass. `group_by=None` disables
gating entirely and needs a much higher `sim_thres`, since far more pairs then compete
with no structural prior ruling any of them out.

`sim_thres` (default `0.7`) is the cosine floor for "same instance". Hungarian always
returns a full assignment, so the threshold is what turns *best available partner* into
*no partner — this is a new instance*. Measured on one checkpoint across four real
photo sets, with the gate on:

| | non-candidate pairs | true-match pairs |
|---|---|---|
| mean | 0.11 – 0.16 | 0.85 – 0.98 (the true-match mode) |
| p99 / max | 0.40 – 0.67 / 0.49 – 0.79 | — |

The default was raised from `0.6` to `0.7` after a 30-case run showed non-candidate
pairs reaching `0.73` and the two distributions overlapping in over half the cases —
`0.6` was admitting matches with no margin at all. Every run
writes its own `diagnostics` (within-identity vs across-identity percentiles) into
`<subject>_identities.json`; if those two distributions overlap, no threshold will
save the run.

### Limitations

- **A gate group with one instance per photo gets no benefit from the embedding.**
  Hungarian has no choice to make, and only `sim_thres` can veto the pairing. The
  fingerprint earns its keep where several instances of one group compete in the same
  photo.
- **Small-instance recall caps what can be matched.** An instance that is never
  detected in a view cannot be linked to it.
- **One folder must be one subject.** A folder holding photos of more than one subject
  will happily link generic-looking instances across them; that is a data problem, not
  a matching one.
- **Letterboxing shifts scale with orientation**, so a landscape and a portrait shot of
  the same subject are not on quite equal footing.

---

## Install

```bash
pip install attr-eomt                  # from PyPI
pip install "attr-eomt[logging]"       # + tensorboard/wandb
pip install -e ".[dev]"                # from source (editable; [dev] adds pytest/build/twine)
```

## Usage

Everything goes through one class. Initialize from a **size** (fresh model, pretrained
DINOv2 backbone) or from a **checkpoint / run folder** (family, size, classes, image
size, normalization and any auxiliary heads are auto-detected from the `.pt`):

```python
from eomt import EoMT

# Train on COCO 2017 (auto-downloaded on first run):
EoMT("l").train(data="coco", epochs=50, batch=4)

# ...or any COCO-format dataset (point at its data.yaml):
EoMT("s").train(data="sample_data/data.yaml", epochs=1, batch=1)

# Validate and predict from a trained run:
EoMT("runs/train/eomt-l").val(data="coco")
EoMT("runs/train/eomt-l").predict("images/", plot=True)   # writes annotated images
```

For the full training recipe, every `train()` knob, and int8 compression, see the
**[annotated explainer →](https://imagra93.github.io/attr-eomt)** — it's the deep dive.

---

## Augmentation

Training augmentation is one config, [`AugConfig`](eomt/data/transforms.py), run by one pipeline
(`TrainAugment`) for both model families. The defaults are a **strong general recipe**: the original
EoMT/Mask2Former one (horizontal flip, Large-Scale Jitter 0.1–2.0, random crop, colour jitter) **plus** small
rotation / shear / perspective, gamma, grayscale, Gaussian blur, sensor noise, JPEG recompression, glare and
"safe" random erasing (a rectangle that never overlaps an instance). Masks are resized by area averaging (soft,
mass-preserving), so thin objects are not shredded into dots by nearest-neighbour resampling.

| group | knobs (default probability) |
|---|---|
| geometry | `flip_prob` 0.5 · `vflip_prob` 0 · `rot90_prob` 0 · LSJ `min_scale`/`max_scale` 0.1–2.0 · `rotate_prob` 0.3 (±10°) · `shear_prob` 0.2 (±5°) · `perspective_prob` 0.15 |
| optics / sensor | colour jitter 1.0 · `gamma_prob` 0.3 · `grayscale_prob` 0.05 · `blur_prob` 0.2 · `noise_prob` 0.2 · `jpeg_prob` 0.3 · `glare_prob` 0.15 · `erasing_prob` 0.2 |
| multi-image (instance family) | `mosaic_prob` 0 · `mixup_prob` 0 · `copy_paste_prob` 0 |
| crop | `instance_crop_prob` 0 (instance-aware crop) · `instance_crop_empty_prob` 0 |
| masks | `mask_resize="area"` (or `"nearest"`) |

Override per run, per dataset, or from the CLI — precedence is explicit keywords (`flip_prob`, `min_scale`,
`max_scale`) **>** `aug=` **>** the dataset YAML's `train_aug` block **>** defaults:

```python
model.train(data="coco", aug={"rotate_prob": 0.5, "blur_prob": 0})   # in code
model.train(data="coco", aug={"preset": "legacy"})                    # the original recipe (hard masks, no extras)
model.train(data="coco", train_transform=my_callable)                 # bring your own (image, masks) -> (image, masks)
```

```yaml
# data.yaml — settings for THIS dataset
train_aug:
  min_scale: 0.6                  # thin / tiny objects: keep the effective scale r = imgsz/long_side * s in ~[0.4, 1.0]
  max_scale: 1.6
  instance_crop_prob: 0.8         # optional instance-aware crop: window placed around a (rarity-weighted) instance
  instance_crop_rarity_attr: typology
```

```bash
python scripts/train.py --data my/data.yaml --aug min_scale=0.6 --aug instance_crop_prob=0.8 --aug-preset default
```

Things worth knowing:

- **Orientation.** Flips, 90° turns and rotations change the pixels but not the label. If an attribute encodes
  orientation (e.g. a `viewpoint` head) set `flip_prob=0` and `rotate_prob=0`; `vflip_prob` / `rot90_prob` stay off
  unless your scenes have no canonical "up".
- **Thin or tiny objects.** Large-Scale Jitter at 0.1 shrinks a 5 px crack to under a pixel; keep the effective
  scale within ~[0.4, 1.0] (`min_scale = 0.4·long_side/imgsz`, `max_scale = 1.0·long_side/imgsz`). The instance-aware
  crop pays off when the crop is much smaller than the image (high zoom, small `imgsz`), and
  `instance_crop_rarity_attr` makes rare attribute values the crop's focus more often.
- **Multi-image ops** (`mosaic`, `mixup`, `copy_paste`) add instances from other samples and are *off*: they splice or
  blend long thin structures, so try them deliberately and look at the result.
- `args.yaml` of a run records the resolved `aug` dict; the log prints the active ops at start.
- `build_train_transform(imgsz, aug=...)` returns the same pipeline for use in your own loaders
  (`tf(image_uint8, masks) -> (image, masks)`).

---

## Roadmap / future work

- **Model export.** ONNX / TensorRT (and friends) for deployment — currently out of
  scope; the inference path is being kept export-friendly.
- **Keypoints.** A keypoint/pose head family alongside `instance` and `detect` (the code
  already carries a `family` parameter so new heads slot in without API churn).
- **Pretrained COCO checkpoints.** None are published yet. COCO-trained `s`/`b`/`l`
  weights will be released on the Hugging Face Hub (the `from_pretrained` / `hf://`
  loading plumbing is already in place and waiting for them).
- **Contrastive re-ID training.** Cross-photo re-identification already works at
  inference time (see [above](#cross-photo-re-identification)) on the embedding the
  detector computes anyway. What remains is *training* that embedding for the job: a
  contrastive objective on the matched queries would make each one a purpose-built
  re-identification vector rather than a by-product of the class head, which should
  widen the margin between matches and non-matches and make `sim_thres` transferable
  across datasets. Feeding those embeddings into the video tracker to re-associate
  objects across occlusions is the same lever.
