Metadata-Version: 2.4
Name: auroraomics
Version: 0.1.0.dev3
Summary: Prepare H&E slides on your own machine, submit them to Aurora, and read the spatial gene expression it predicts.
Author: Kalin Nonchev
License-Expression: PolyForm-Noncommercial-1.0.0
Keywords: histopathology,spatial-transcriptomics,h5ad,anndata,whole-slide-image
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
Classifier: Typing :: Typed
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: numpy>=1.23
Requires-Dist: h5py>=3.9
Requires-Dist: opencv-python-headless>=4.5
Requires-Dist: pillow>=9
Requires-Dist: httpx>=0.27
Requires-Dist: pydantic>=2
Requires-Dist: anndata>=0.10
Requires-Dist: tifffile>=2023.7.10
Requires-Dist: imagecodecs>=2023.3.16
Provides-Extra: embed
Requires-Dist: torch>=2.0; extra == "embed"
Requires-Dist: timm>=1.0; extra == "embed"
Requires-Dist: transformers>=4.40; extra == "embed"
Requires-Dist: huggingface-hub>=0.23; extra == "embed"
Provides-Extra: mcp
Requires-Dist: mcp<3,>=2; extra == "mcp"
Provides-Extra: dev
Requires-Dist: auroraomics[mcp]; extra == "dev"
Requires-Dist: pytest>=7; extra == "dev"
Requires-Dist: packaging>=22; extra == "dev"
Requires-Dist: pytest-cov>=5; extra == "dev"
Requires-Dist: setuptools>=77; extra == "dev"
Requires-Dist: wheel; extra == "dev"
Requires-Dist: build>=1.0; extra == "dev"
Requires-Dist: numpy<2; extra == "dev"
Requires-Dist: huggingface-hub>=0.23; extra == "dev"
Dynamic: license-file

# auroraomics

Virtual spatial transcriptomics from H&E images.

This package is the client of the **Aurora API**, the hosted service that runs
the prediction: from Python, from a shell with the `auroraomics` command, or
from an agent through its MCP server.

This package holds the pieces of that pipeline that are pure Python, so the
same code runs on your machine and on the service:

- **`auroraomics.qc`** — the three patch-quality predicates (foreground, blur,
  stained tissue) applied to every candidate tile before a model sees it, and
  the short-circuiting cascade that combines them.
- **`auroraomics.pack`** — build the `patches` archive a prediction takes as
  input (`pack_tiles`), and validate one you have been handed (`read_archive`).
- **`auroraomics.h5ad`** — write the standard `.h5ad` result with `h5py`
  alone, streaming the expression matrix batch by batch so peak memory is one
  batch rather than the whole matrix.
- **`auroraomics.postprocess`** — apply a model card's `postprocess` entries
  to a result file: a registry of entries, one runner, `h5py` alone, streaming
  one chunk at a time. Ships the unmeasured-gene mask, which blanks the genes
  the model never measured.
- **`auroraomics.subsample`** — pick the densest contiguous square of spots
  when a slide yields more tiles than one submission takes.
- **`auroraomics.genes`** — Ensembl stable gene ids, which is what every
  result's `var` index is keyed by, and the shape a gene the service resolved
  comes back in.
- **`auroraomics.contracts`** — the shared contract values (container layout,
  input-kind caps, result layout) as data, so nothing here re-types a number
  the service also reads.
- **`auroraomics.embed`** *(extra)* — run a pinned image encoder over your
  tiles locally and write the embeddings container, so a prediction can be made
  from numbers instead of pixels.

One module describes work that does not happen here:

- **`auroraomics.runtimes.deepspotm`** — the models the service will run, as
  data: what each one accepts, what it returns, and whether an archive matches
  it. There is no install that makes one of them run locally; the definition is
  not distributed, and `predict_local` refuses and says where prediction
  happens.

## Install

```
pip install auroraomics
```

That is the whole documented path: open a slide, judge and pack its tiles,
submit, and read the result. One extra adds local embedding extraction
(`embed`). No extra brings the model: predicting gene expression runs on the
service, not here.

## Pack tiles, then look at the report

```python
import auroraomics as ao

with ao.Client() as client:
    thresholds = client.qc_thresholds()   # the service's quality floors, no key needed

report = ao.pack_tiles(tiles, "sample.zip", mpp=0.499, thresholds=thresholds, thumbnail=thumb)
print(report.written, "tiles kept,", report.rejected, "dropped")
print(report.rejected_by_reason)          # {'foreground_ratio': 12, ...}
```

`tiles` is any iterable of `ao.Tile(image, x, y)`, where `image` is an RGB
`uint8` array and `x`/`y` are the tile's top-left position in full-resolution
pixels. Tiles are quality-checked as they stream past, and only the ones that
pass are written, so an iterable that reads a slide lazily never has to hold
more than one tile in memory.

## Validate an archive before trusting it

```python
archive = ao.read_archive("sample.zip")   # raises ArchiveError on anything odd
for tile in archive.tiles():              # decoded one at a time
    ...
```

`read_archive` checks the member names, the manifest, the tile geometry and the
declared sizes *before* decoding a single pixel, and refuses an archive whose
members do not match the container contract.

## Write a result

```python
ao.write_result(
    "result.h5ad",
    obs=obs,                    # per-spot columns, as plain arrays
    var=var,                    # per-gene columns
    var_index=gene_ids,         # the Ensembl gene id of each column
    spatial=coords,             # (n_spots, 2) array -> obsm["spatial"]
    x=batches,                  # an array, or an iterable of row batches
    uns={"model": {"id": "..."}},
    layers={"image_only": other_batches},
)
```

The file reads back as an ordinary `AnnData` in anndata 0.10 and 0.11. Passing
an iterable for `x` streams it: each batch is compressed into the file as it
arrives and then dropped, which is what makes a matrix larger than memory
writable.

## Embed locally, and send numbers instead of pixels

```
pip install "auroraomics[embed]"
```

```python
import auroraomics as ao
from auroraomics.embed import available_encoders, describe_encoders

with ao.Client() as client:
    # Both are the service's, so the file records the encoder and the
    # thresholds the model it is submitted to was served with. There is no
    # list of encoders inside this package: every name comes from the registry
    # the service publishes, which is why these calls take one.
    registry = client.encoders()
    thresholds = client.qc_thresholds()
    # The tile size the model's card states, for the model the file is for.
    patch_um = client.model_card(model_id)["input_spec"]["patch_um"]

print(available_encoders(registry))   # the names this package will run today
encoder = ao.resolve_encoder(registry)

report = ao.embed_tiles(
    slide, "sample.npz", mpp=0.499, patch_um=patch_um, encoder=encoder, thresholds=thresholds
)
print(report.rows, "rows of", report.dim, "numbers from", report.encoder.name)
```

`slide` is anything with a `(height, width, 3)` RGB `uint8` shape that can be
sliced — an array, or a memory-mapped or lazily-read one, so a slide larger than
memory works: only one crop exists at a time. Crops are taken at the tile size
the model's card states, using the `mpp` you give, and quality-checked with the
same three predicates as `pack_tiles`.

Every encoder is pinned to an exact revision.
`describe_encoders(registry)` returns the whole registry, including the
encoders this package refuses to run: a refusal says why, and which encoders
remain.

The written file carries one extra embedding, `calibration`: computed from a
fixed image shipped inside this package, through the same encoder and revision
as your rows. Comparing that one embedding against a known
reference tells a reader whether the file was produced by the model it claims,
before they look at a single row.

The weights are downloaded on first use. Pass `allow_download=False` to
guarantee no request is made: for the duration of that load, name resolution
and internet sockets are refused in this process.

An encoder runs at the input size the registry pins for it. `resize_px` is
checked against that pin rather than obeyed, because an embedding computed at
another size cannot be compared with the served model's.

## Where a prediction runs

Predicting gene expression runs on the service, and `predict_local` refuses.

**DeepSpot-H**, the foundation model for H&E images, turns each tile of your
slide into one embedding. It can run on your machine or on ours. **DeepSpot-M**
turns those embeddings into spatial gene expression. It only ever runs on ours.

```python
from auroraomics.runtimes.deepspotm import check_archive, describe_models

with ao.Client() as client:
    cards = client.models()                                # the catalogue is the service's
describe_models(cards)                                     # what it will run, and on what
check_archive("sample.zip", model=model_id, cards=cards)   # are my tiles the right size?
```

Then submit with the API client and read the result back — the same `.h5ad`
contract whichever runtime ran it.

## Post-process a result

```python
from auroraomics.postprocess import apply_card

apply_card("result.h5ad", card=card)        # -> ('unmeasured_gene_mask',)
```

Or, from a pipeline rule, without writing a script:

```
python -m auroraomics.postprocess result.h5ad --card card.json
```

The card's `postprocess` list decides what runs. Every entry writes only where
the result contract says its product lives — a matrix under `layers/`, a
per-gene statement as a `var` column — and an entry the installed version does
not implement is refused **by name**, before the file is opened, rather than
skipped: a card is the description of a file, and a reader promised a layer
cannot tell a missing one from a model that predicted zeros.

The entry this version ships states what the model measured. `var` gains
`measured_in_training`, which is `False` for a gene the model never saw
measured. Such a gene becomes `nan` in `X` and in every layer, not zero: zero
is a measurement, and an unmeasured gene left as one averages, correlates and
colours a heat map exactly like a prediction.

## Prepare a bulk RNA profile

```python
with ao.Client() as client:
    report = ao.bulk_rna(
        "sample.star_gene_counts.tsv", "sample.bulk.tsv", units="counts",
        card=client.model_card(model_id), resolve=client.resolve_keys,
    )
print(report.rows, "genes;", f"{report.coverage:.0%} of", report.coverage_of)
```

A long or wide table, CSV or TSV, or a GDC STAR gene-counts file exactly as
downloaded. Gene names are resolved to unversioned Ensembl ids through the same
table every result is keyed by; the value column of the written table is named
by the unit you declared, so the file states its own units; and a table that is
already log-transformed is refused rather than logged a second time. Given the
model's card, a profile that covers less of the model's measured genes than the
service requires is refused here, with the service's own code, before anything
is uploaded.

## Typing

The package ships `py.typed`, so annotations are visible to type checkers in
your project.

## Licence

The code in this package is licensed for non-commercial use; its terms are in the
`LICENSE` file it ships with. Commercial evaluation and use are governed by a written agreement with Aurora.

