Metadata-Version: 2.4
Name: auroraomics
Version: 0.1.0.dev1
Summary: Virtual spatial transcriptomics from H&E histology: patch QC, tile packing and the .h5ad result contract.
Author: Kalin Nonchev
License-Expression: PolyForm-Noncommercial-1.0.0
Keywords: histopathology,spatial-transcriptomics,h5ad,anndata,whole-slide-image
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
Classifier: Typing :: Typed
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: numpy>=1.23
Requires-Dist: h5py>=3.9
Requires-Dist: opencv-python-headless>=4.5
Requires-Dist: pillow>=9
Requires-Dist: httpx>=0.27
Requires-Dist: pydantic>=2
Requires-Dist: anndata>=0.10
Requires-Dist: tifffile>=2023.7.10
Requires-Dist: imagecodecs>=2023.3.16
Provides-Extra: embed
Requires-Dist: torch>=2.0; extra == "embed"
Requires-Dist: timm>=1.0; extra == "embed"
Requires-Dist: transformers>=4.40; extra == "embed"
Requires-Dist: huggingface-hub>=0.23; extra == "embed"
Provides-Extra: mcp
Requires-Dist: mcp<3,>=2; extra == "mcp"
Provides-Extra: dev
Requires-Dist: auroraomics[mcp]; extra == "dev"
Requires-Dist: pytest>=7; extra == "dev"
Requires-Dist: pytest-cov>=5; extra == "dev"
Requires-Dist: setuptools>=77; extra == "dev"
Requires-Dist: wheel; extra == "dev"
Requires-Dist: build>=1.0; extra == "dev"
Requires-Dist: numpy<2; extra == "dev"
Requires-Dist: huggingface-hub>=0.23; extra == "dev"
Dynamic: license-file

# auroraomics

Virtual spatial transcriptomics from H&E histology.

This package holds the pieces of that pipeline that are pure Python, so the
same code runs on your laptop, in a GPU container and on the service:

- **`auroraomics.qc`** — the three patch-quality predicates (foreground, blur,
  stained tissue) applied to every candidate tile before a model sees it, and
  the short-circuiting cascade that combines them.
- **`auroraomics.pack`** — build the `patches` archive a prediction takes as
  input (`pack_tiles`), and validate one you have been handed (`read_archive`).
- **`auroraomics.h5ad`** — write the standard `.h5ad` result with `h5py`
  alone, streaming the expression matrix batch by batch so peak memory is one
  batch rather than the whole matrix.
- **`auroraomics.postprocess`** — apply a model card's `postprocess` entries
  to a result file: a registry of entries, one runner, `h5py` alone, streaming
  one chunk at a time. Ships the unmeasured-gene mask, which states per gene
  how many training datasets measured it and blanks the ones none did.
- **`auroraomics.subsample`** — pick the densest contiguous square of spots
  when a slide yields more tiles than a run is allowed to spend.
- **`auroraomics.genes`** — the model family's gene symbols mapped to Ensembl
  stable gene ids, which is what every result's `var` index is keyed by. Ships
  as one committed table, resolved in the same order the service resolves it.
- **`auroraomics.contracts`** — the shared contract values (container layout,
  input-kind caps, result layout) as data, so nothing here re-types a number
  the service also reads.
- **`auroraomics.embed`** *(extra)* — run a pinned image encoder over your
  tiles locally and write the embeddings container, so a prediction can be made
  from numbers instead of pixels.

One module describes work that does not happen here:

- **`auroraomics.runtimes.deepspotm`** — the models the service will run, as
  data: what each one accepts, what it returns, and whether an archive matches
  it. There is no install that makes one of them run locally; the definition is
  not distributed, and `predict_local` refuses and says where prediction
  happens.

## Install

```
pip install auroraomics
```

That is the whole documented path: open a slide, judge and pack its tiles,
submit, and read the result. One extra adds local embedding extraction
(`embed`), because a tensor runtime is gigabytes and specific to the machine it
was built for. No extra brings the model: predicting gene expression runs on
the service, not here.

## Pack tiles, then look at the report

```python
import auroraomics as ao

report = ao.pack_tiles(tiles, "sample.zip", mpp=0.499, thumbnail=thumb)
print(report.written, "tiles kept,", report.rejected, "dropped")
print(report.rejected_by_reason)          # {'foreground_ratio': 12, ...}
```

`tiles` is any iterable of `ao.Tile(image, x, y)`, where `image` is an RGB
`uint8` array and `x`/`y` are the tile's top-left position in full-resolution
pixels. Tiles are quality-checked as they stream past, and only the ones that
pass are written, so an iterable that reads a slide lazily never has to hold
more than one tile in memory.

## Validate an archive before trusting it

```python
archive = ao.read_archive("sample.zip")   # raises ArchiveError on anything odd
for tile in archive.tiles():              # decoded one at a time
    ...
```

`read_archive` checks the member names, the manifest, the tile geometry and the
declared sizes *before* decoding a single pixel, and refuses an archive whose
members do not match the container contract.

## Write a result

```python
ao.write_result(
    "result.h5ad",
    obs=obs,                    # per-spot columns, as plain arrays
    var=var,                    # per-gene columns, indexed by gene id
    spatial=coords,             # (n_spots, 2) array -> obsm["spatial"]
    x=batches,                  # an array, or an iterable of row batches
    uns={"model": {"id": "..."}},
    layers={"image_only": other_batches},
)
```

The file reads back as an ordinary `AnnData` in anndata 0.10 and 0.11. Passing
an iterable for `x` streams it: each batch is compressed into the file as it
arrives and then dropped, which is what makes a matrix larger than memory
writable.

## Embed locally, and send numbers instead of pixels

```
pip install "auroraomics[embed]"
```

```python
import auroraomics as ao
from auroraomics.embed import available_encoders, describe_encoders

print(available_encoders())        # ('deepspot-h', 'dinov2-b14', 'midnight')

with ao.Client() as client:
    # Both are the service's, so the file records the encoder and the
    # thresholds the model it is submitted to was served with.
    encoder = ao.resolve_encoder(client.encoders())
    thresholds = client.qc_thresholds()

report = ao.embed_tiles(
    slide, "sample.npz", mpp=0.499, patch_um=55, encoder=encoder, thresholds=thresholds
)
print(report.rows, "rows of", report.dim, "numbers from", report.encoder.name)
```

`slide` is anything with a `(height, width, 3)` RGB `uint8` shape that can be
sliced — an array, or a memory-mapped or lazily-read one, so a slide larger than
memory works: only one crop exists at a time. Crops are taken at `patch_um`
micrometres using the `mpp` you give, quality-checked with the same three
predicates as `pack_tiles`, and pooled over the encoder's patch tokens.

Every encoder is pinned to a repository **and** a revision sha, with its licence
and access gate recorded beside it. `describe_encoders()` returns the whole
registry, including the encoders this package refuses to run — a refusal names
the licence or the approval you would have to accept, and which encoders remain.

The written file carries one extra vector, `calibration`: the embedding of a
fixed image shipped inside this package, through the same encoder, revision and
preprocessing as your rows. Comparing that one vector against a known reference
tells a reader whether the file was produced by the model it claims, before they
look at a single row.

The weights are downloaded from their publisher on first use, into the ordinary
model cache. Pass `allow_download=False` to guarantee no request is made: for
the duration of that load, name resolution and internet sockets are refused in
this process, so the guarantee does not rest on a model library reading an
environment variable it may already have read.

An encoder runs at one input size: the size the served model's embeddings were
computed at, which the registry pins and which is not always what the encoder's
own repository declares. The DINOv2 releases carry a 518-px position grid; the
model was trained on 224-px crops fed to that grid through an interpolated
position embedding, so that is what this package does too. `resize_px` is
checked against the pin rather than obeyed, because a vector computed at another
size is a different measurement that nobody else's numbers can be compared with.
## Where a prediction runs

Predicting gene expression runs on the service, and `predict_local` refuses —
by design, and the design is the product rather than a limitation.

This is split inference at the encoder. An open-weight pathology foundation
model turns your tiles into patch embeddings **on your machine**, so your slides
never leave it; our gene decoder turns embeddings into expression **on ours**, so
its weights never leave it. What crosses between us is a vector from a model
neither side owns.

```python
from auroraomics.runtimes.deepspotm import available_models, check_archive

available_models()                      # what the service will run
check_archive("sample.zip", model=...)  # are my tiles the right physical size?
```

Then submit with the API client and read the result back — the same `.h5ad`
contract whichever lane ran it.

## Post-process a result

```python
from auroraomics.postprocess import apply_card

apply_card("result.h5ad", card=card)        # -> ('unmeasured_gene_mask',)
```

Or, from a pipeline rule, without writing a script:

```
python -m auroraomics.postprocess result.h5ad --card card.json
```

The card's `postprocess` list decides what runs. Every entry writes only where
the result contract says its product lives — a matrix under `layers/`, a
per-gene statement as a `var` column — and an entry the installed version does
not implement is refused **by name**, before the file is opened, rather than
skipped: a card is the description of a file, and a reader promised a layer
cannot tell a missing one from a model that predicted zeros.

The entry this version ships states what the model measured. `var` gains
`n_train_datasets` — the number of training datasets each gene was measured in,
because a gene seen in one and a gene seen in twenty are not the same claim —
and `measured_in_training`, derived from that count. A gene the model never
measured becomes `nan` in `X` and in every layer, not zero: zero is a
measurement, and an unmeasured gene left as one averages, correlates and
colours a heat map exactly like a prediction.

## Prepare a bulk RNA profile

```python
report = ao.bulk_rna("sample.star_gene_counts.tsv", "sample.bulk.tsv", units="counts")
print(report.rows, "genes;", f"{report.coverage:.0%} of", report.coverage_of)
```

A long or wide table, CSV or TSV, or a GDC STAR gene-counts file exactly as
downloaded. Gene names are resolved to unversioned Ensembl ids through the same
table every result is keyed by; the value column of the written table is named
by the unit you declared, so the file states its own units; and a table that is
already log-transformed is refused rather than logged a second time. Given the
model's card, a profile that covers less of the model's measured genes than the
service requires is refused here, with the service's own code, before anything
is uploaded.

## Typing

The package ships `py.typed`, so annotations are visible to type checkers in
your project.

## Licence

The code in this package is licensed under
[PolyForm Noncommercial 1.0.0](https://polyformproject.org/licenses/noncommercial/1.0.0),
which permits use for any purpose that is not commercial. It is the same licence the
model package this client is built for carries, so installing both puts you under one
rule rather than two.

The model weights are licensed separately by whoever publishes them, and access to them
may be gated. Read those terms before you use a model: they are not this licence, and a
permission granted here is not a permission granted there.

