Metadata-Version: 2.4
Name: medh5
Version: 1.2.1
Summary: Self-describing HDF5 container for one medical imaging sample and all of its ground truth: multi-timepoint, multi-modal images with segmentation, detection, classification and registration annotations, provenance and integrity.
Author-email: Puyang Wang <pauliwang411@gmail.com>
License: MIT License
        
        Copyright (c) 2026 Puyang Wang
        
        Permission is hereby granted, free of charge, to any person obtaining a copy
        of this software and associated documentation files (the "Software"), to deal
        in the Software without restriction, including without limitation the rights
        to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
        copies of the Software, and to permit persons to whom the Software is
        furnished to do so, subject to the following conditions:
        
        The above copyright notice and this permission notice shall be included in all
        copies or substantial portions of the Software.
        
        THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
        IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
        FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
        AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
        LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
        OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
        SOFTWARE.
        
Project-URL: Homepage, https://github.com/XwK-P/medh5
Project-URL: Repository, https://github.com/XwK-P/medh5
Project-URL: Issues, https://github.com/XwK-P/medh5/issues
Project-URL: Changelog, https://github.com/XwK-P/medh5/blob/main/CHANGELOG.md
Project-URL: Documentation, https://medh5.readthedocs.io/
Project-URL: Specification, https://github.com/XwK-P/medh5/blob/main/docs/spec/medh5-1.0.md
Keywords: medical imaging,hdf5,blosc2,machine learning,segmentation,detection,registration,longitudinal,nifti,dicom,rtstruct,nnunet,pytorch,monai
Classifier: Development Status :: 5 - Production/Stable
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering
Classifier: Topic :: Scientific/Engineering :: Medical Science Apps.
Classifier: Typing :: Typed
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: h5py>=3.10
Requires-Dist: hdf5plugin>=4.1
Requires-Dist: numpy>=1.24
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"
Requires-Dist: pytest-cov>=4.0; extra == "dev"
Requires-Dist: ruff==0.16.3; extra == "dev"
Requires-Dist: mypy==1.18.2; extra == "dev"
Requires-Dist: build>=1.0; extra == "dev"
Requires-Dist: twine>=5.0; extra == "dev"
Requires-Dist: jsonschema>=4.18; extra == "dev"
Requires-Dist: hypothesis>=6.100; extra == "dev"
Provides-Extra: schema
Requires-Dist: jsonschema>=4.18; extra == "schema"
Provides-Extra: torch
Requires-Dist: torch>=2.0; extra == "torch"
Provides-Extra: monai
Requires-Dist: monai>=1.3; extra == "monai"
Requires-Dist: torch>=2.0; extra == "monai"
Provides-Extra: nifti
Requires-Dist: nibabel>=5; extra == "nifti"
Provides-Extra: dicom
Requires-Dist: pydicom>=2.4; extra == "dicom"
Provides-Extra: dicomseg
Requires-Dist: highdicom>=0.22; extra == "dicomseg"
Requires-Dist: pydicom>=2.4; extra == "dicomseg"
Provides-Extra: itk
Requires-Dist: SimpleITK>=2.3; extra == "itk"
Provides-Extra: interp
Requires-Dist: scipy>=1.10; extra == "interp"
Dynamic: license-file

# medh5

[![PyPI version](https://img.shields.io/pypi/v/medh5.svg)](https://pypi.org/project/medh5/)
[![Python versions](https://img.shields.io/pypi/pyversions/medh5.svg)](https://pypi.org/project/medh5/)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)
[![CI](https://github.com/XwK-P/medh5/actions/workflows/ci.yml/badge.svg?branch=main)](https://github.com/XwK-P/medh5/actions/workflows/ci.yml)
[![Coverage](https://img.shields.io/badge/coverage-93%25-brightgreen.svg)](#)
[![Typed](https://img.shields.io/badge/typed-mypy%20strict-informational.svg)](medh5/py.typed)
[![Code style: ruff](https://img.shields.io/badge/code%20style-ruff-000000.svg)](https://github.com/astral-sh/ruff)

**One medical imaging sample — a subject, at every timepoint, with all of its
ground truth — in a single self-describing HDF5 file.**

Multi-modality images, segmentation in five encodings, detection boxes,
keypoints, contours, meshes, classification, registration between visits,
provenance and quality records, and per-object integrity digests. Format
version **1.0**, with a [normative specification](https://medh5.readthedocs.io/en/latest/spec/medh5-1.0/) and a
[103-case conformance suite](https://medh5.readthedocs.io/en/latest/spec/conformance/) any implementation can run.

```python
import medh5

with medh5.open("case_0001.medh5") as s:
    s.identity.subject_id                              # "BRATS-GLI-01234"
    s.at("tp1").images["CT_tp1"].read(physical=True)   # HU, not raw counts
    s.annotations["organs"].dense(["liver", "spleen"]) # any encoding, one API
    s.transform_between("tp0", "tp1")                  # resolved via frames
    s.tracks("lesion")                                 # lesions joined across visits
```

## Install

```bash
pip install medh5
pip install "medh5[torch,nifti,dicom]"
```

Reading and writing needs only `h5py`, `hdf5plugin` and `numpy`. Extras:
`torch`, `monai`, `nifti`, `dicom`, `dicomseg`, `itk`, `schema`, `interp`.

## Documentation

**[medh5.readthedocs.io](https://medh5.readthedocs.io/)** — tutorials, how-to
guides, the Python and CLI reference, and the normative specification.

[Write your first sample](https://medh5.readthedocs.io/en/latest/tutorials/first-sample/) ·
[How-to guides](https://medh5.readthedocs.io/en/latest/guides/) ·
[Python API](https://medh5.readthedocs.io/en/latest/reference/python-api/) ·
[CLI](https://medh5.readthedocs.io/en/latest/reference/cli/) ·
[Specification](https://medh5.readthedocs.io/en/latest/spec/medh5-1.0/)

## What the format is for

- **One file per subject, not per scan** — every visit in one place, so
  longitudinal work has a referent and splitting by file cannot leak a patient.
- **Geometry is stated once and never guessed** — declared grids, boxes at voxel
  edges, and converters that refuse rather than invent.
- **Absence is not silence** — a class examined and not found is recorded as
  such, which is a different training signal from one nobody examined.
- **Every claim is checkable** — per-object digests, a Merkle `content_id` that
  survives recompression, a stable diagnostic-code table, and a 103-case
  conformance corpus.
- **Reading a patch is fast** — a 64³ multi-class patch in ~4 ms, and O(1)
  foreground sampling once `build_index()` has run.

[The reasoning behind each](https://medh5.readthedocs.io/en/latest/).

## Write a sample

```python
import numpy as np
import medh5
from medh5 import LabelClass, LabelSet

labels = LabelSet("demo-v1", version="1.0.0", classes=[
    LabelClass(1, "liver", "Liver", category="organ"),
    LabelClass(2, "spleen", "Spleen", category="organ"),
    LabelClass(3, "lesion", "Lesion", parents=[1], category="lesion"),
])

with medh5.create("case_0001.medh5", sample_id="case_0001",
                  subject_id="DEMO-0001") as w:
    w.label_set(labels)
    w.add_timepoint("tp0", label="baseline", days_from_baseline=0)
    w.add_grid("ct", shape=ct.shape, spacing=(2.0, 0.8, 0.8),
               origin=(-64.0, -38.4, -38.4), timepoint="tp0")
    w.add_image("CT", ct, grid="ct", modality="CT",
                value_type="quantitative", value_units="HU")
    w.add_segmentation("organs", grid="ct",
                       masks={"liver": liver, "lesion": lesion},
                       annotated_classes=["liver", "spleen", "lesion"])
    w.build_index()   # optional; foreground sampling is O(1) only with it
```

`annotated_classes` names the spleen although there is no spleen mask: that
records "we looked and found none". The encoding is chosen by measuring the
class overlap graph — liver and lesion overlap, so it picks one that can
represent that — and the write is atomic.

## Train on it

```python
from torch.utils.data import DataLoader
from medh5.torch import PatchDataset, collate, worker_init_fn
from medh5.sampling import PatchSampler

sampler = PatchSampler((96, 96, 96), strategy="balanced",
                       foreground_classes=["liver", "lesion"])
dataset = PatchDataset(paths, sampler, images=["CT"],
                       annotations={"organs": ["liver", "lesion"]},
                       samples_per_volume=8)

loader = DataLoader(dataset, batch_size=2, num_workers=8,
                    worker_init_fn=worker_init_fn, collate_fn=collate)
```

`worker_init_fn` drops handles inherited across a `fork`. It is recommended
rather than required: the handle cache is PID-keyed and re-checks ownership on
every access, so a forked worker abandons the parent's handles on first use
rather than reading through or closing them.

## Command line

```bash
medh5 info case.medh5                  # grids, images, annotations, coverage
medh5 validate case.medh5 --level strict
medh5 verify case.medh5                # digests and content_id
medh5 timeline case.medh5              # visits and intervals
medh5 track case.medh5 --class lesion  # per-lesion volumes across visits

medh5 dataset index studies/ -o cohort.json
medh5 dataset split cohort.json --group-by group_id --stratify-by site_id
medh5 dataset stats cohort.json --partition train --workers 8
medh5 dataset check cohort.json --deep

medh5 convert from-dicom /studies out/     # one sample per patient, all visits
medh5 convert from-nifti case.medh5 --image CT=ct.nii.gz
medh5 convert from-rtstruct plan.dcm case.medh5 --rasterize
medh5 migrate old/*.medh5 -o new/ --group-by subject

medh5 scrub out/*.medh5 --apply --date-shift-days -117
medh5 pack cohort/*.medh5 -o shard.medh5c
medh5 recompress cohort/*.medh5 --profile training
medh5 bench                                # reproduce the performance targets
medh5 conformance publish suite/           # the suite, for another implementation
```

## Interoperability

| Format | |
|---|---|
| **NIfTI** | affine and voxels bit-identical on round trip; RAS↔LPS is a sign flip, never a resample |
| **DICOM** | slices ordered by geometry, spacing measured between origins, modality LUT stored not applied, tags on an explicit allow-list; slices that disagree about orientation, spacing or rescale are refused rather than read off the first one |
| **DICOM SEG** | frames placed by geometry; segments matched by label, not number; import preserves overlap and `FRACTIONAL`, export writes `BINARY` |
| **RTSTRUCT** | contours stay contours; rasterisation is opt-in and recorded in provenance |
| **nnU-Net v2** | class ids kept; region labels become label-set DAG parents; `dataset.json` round-trips |
| **MONAI** | `to_metatensor` gives a `MetaTensor` with the correct affine |
| **0.x** | `medh5 migrate`, reporting every decision and every guess |

Every conversion returns a report distinguishing what it **decided** from the
data and where it **guessed** — the encoding chosen, the class ids minted, a
half-voxel convention changed, a timepoint order inferred rather than read.

COCO is deliberately unsupported: it has no world geometry, spacing or frame of
reference, so importing means inventing a grid and exporting means discarding
the geometry that makes a medical annotation reproducible.

## Reading it without medh5

```python
import h5py, json, hdf5plugin       # hdf5plugin only for blosc2 profiles

with h5py.File("case_0001.medh5") as f:
    doc = json.loads(f["meta"][()])
    doc["identity"]["subject_id"]
    dict(f["grids"]["ct"].attrs)     # spacing, origin, direction
    f["images"]["CT"][10:20]
```

`medh5 recompress --profile portable` writes gzip, readable by any HDF5 build.

## Versioning

The **format** is 1.0. A minor version may add optional objects, profiles,
encodings and diagnostic codes; it may not change what an existing one means
(spec §16). The **package** follows semantic versioning from 1.0.0.

0.x files are not readable by 1.0 and are not meant to be — `medh5 migrate`
converts them once. See [Converters](https://medh5.readthedocs.io/en/latest/guides/migrate-0x/).

## License

MIT
