Metadata-Version: 2.4
Name: rustscenic
Version: 0.5.0
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Science/Research
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Operating System :: MacOS
Classifier: Operating System :: Microsoft :: Windows
Classifier: Operating System :: POSIX :: Linux
Classifier: Programming Language :: Rust
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Typing :: Typed
Requires-Dist: numpy>=1.21
Requires-Dist: pandas>=1.5
Requires-Dist: pyarrow>=10
Requires-Dist: scipy>=1.8
Requires-Dist: anndata>=0.10
Requires-Dist: anndata>=0.10 ; extra == 'benchmarks'
Requires-Dist: scanpy>=1.10 ; extra == 'benchmarks'
Requires-Dist: scikit-learn>=1.3 ; extra == 'benchmarks'
Requires-Dist: tomotopy>=0.12 ; extra == 'benchmarks'
Requires-Dist: gensim>=4.3 ; extra == 'benchmarks'
Requires-Dist: psutil>=5.9 ; extra == 'benchmarks'
Requires-Dist: anndata>=0.10 ; extra == 'examples'
Requires-Dist: scanpy>=1.10 ; extra == 'examples'
Requires-Dist: igraph>=0.11 ; extra == 'examples'
Requires-Dist: leidenalg>=0.10 ; extra == 'examples'
Requires-Dist: anndata>=0.10 ; extra == 'reference'
Requires-Dist: scanpy>=1.10 ; extra == 'reference'
Requires-Dist: scikit-learn>=1.3 ; extra == 'reference'
Requires-Dist: setuptools<82 ; extra == 'reference'
Requires-Dist: pyscenic>=0.12 ; extra == 'reference'
Requires-Dist: arboreto>=0.1.6 ; extra == 'reference'
Requires-Dist: ctxcore>=0.2 ; extra == 'reference'
Requires-Dist: anndata>=0.10 ; extra == 'validation'
Requires-Dist: scanpy>=1.10 ; extra == 'validation'
Requires-Dist: igraph>=0.11 ; extra == 'validation'
Requires-Dist: leidenalg>=0.10 ; extra == 'validation'
Requires-Dist: scikit-learn>=1.3 ; extra == 'validation'
Provides-Extra: benchmarks
Provides-Extra: examples
Provides-Extra: reference
Provides-Extra: validation
License-File: LICENSE
License-File: NOTICE
Summary: Rust and PyO3 implementation of SCENIC-style regulatory-network analysis. Includes GRN, AUCell, topics, cistarget, peak calling, cell QC, enhancer-gene links, and eRegulon assembly. Installs without dask, Java, or CUDA.
Keywords: single-cell,scRNA-seq,scATAC-seq,SCENIC,SCENIC+,gene regulatory network,GRN,transcription factor,regulon,AUCell,pycisTopic,pycistarget,bioinformatics,computational biology,Rust,PyO3
Author-email: Ekin Kahraman <evk23umu@uea.ac.uk>
License-Expression: Apache-2.0
Requires-Python: >=3.10
Description-Content-Type: text/markdown; charset=UTF-8; variant=GFM
Project-URL: Changelog, https://github.com/Ekin-Kahraman/rustscenic/blob/main/CHANGELOG.md
Project-URL: Documentation, https://ekin-kahraman.github.io/rustscenic/
Project-URL: Homepage, https://github.com/Ekin-Kahraman/rustscenic
Project-URL: Issues, https://github.com/Ekin-Kahraman/rustscenic/issues
Project-URL: Repository, https://github.com/Ekin-Kahraman/rustscenic

<h1 align="center">RustScenic</h1>

<p align="center">
  <strong>Fast, memory-efficient gene-regulation analysis for single-cell data.</strong>
</p>

<p align="center">
  Infer gene networks and score their activity using RNA and chromatin-accessibility data.
  A Python package accelerated with Rust. Runs on CPUs; no GPU required.
</p>

<p align="center">
  Created and maintained by Ekin Kahraman, developed in collaboration with the
  Kuan-Lin Huang Lab at the Icahn School of Medicine at Mount Sinai.
</p>

<p align="center">
  <a href="https://ekin-kahraman.github.io/rustscenic/">Documentation</a> |
  <a href="site_docs/benchmarks.md">Benchmarks</a> |
  <a href="site_docs/validation.md">Validation</a> |
  <a href="CITATION.cff">Citation</a> |
  <a href="https://doi.org/10.5281/zenodo.20246040">Zenodo DOI</a>
</p>

<p align="center">
  <a href="https://github.com/Ekin-Kahraman/rustscenic/actions/workflows/audit.yml"><img alt="CI" src="https://github.com/Ekin-Kahraman/rustscenic/actions/workflows/audit.yml/badge.svg"></a>
  <a href="https://github.com/Ekin-Kahraman/rustscenic/actions/workflows/docs.yml"><img alt="Docs" src="https://github.com/Ekin-Kahraman/rustscenic/actions/workflows/docs.yml/badge.svg"></a>
  <a href="https://github.com/Ekin-Kahraman/rustscenic/actions/workflows/nightly-real-data.yml"><img alt="Nightly real-data validation" src="https://github.com/Ekin-Kahraman/rustscenic/actions/workflows/nightly-real-data.yml/badge.svg"></a>
  <a href="https://pypi.org/project/rustscenic/"><img alt="PyPI" src="https://img.shields.io/pypi/v/rustscenic"></a>
  <br>
  <a href="https://doi.org/10.5281/zenodo.20246040"><img alt="Zenodo DOI" src="https://img.shields.io/badge/DOI-Zenodo-1682d4"></a>
  <a href="LICENSE"><img alt="License: Apache-2.0" src="https://img.shields.io/badge/License-Apache--2.0-blue.svg"></a>
  <a href="https://www.python.org/"><img alt="Python" src="https://img.shields.io/badge/Python-3.10%2B-blue"></a>
  <a href="https://www.rust-lang.org/"><img alt="Rust" src="https://img.shields.io/badge/Rust-stable-orange"></a>
</p>

## Highlights

- Gene-network inference on **1.3 million mouse-brain cells** in under 47 minutes,
  with **4.28 GB** peak memory during analysis on 16 CPU cores
  ([v0.5.0 benchmark](https://github.com/Ekin-Kahraman/rustscenic/blob/0c8eb00539e3860c78e452c8661cc2735c169386/validation/scaling/IFB_REAL_RNA_GRN_2026-08-28.md)).
- **3.3x faster with about 81% less peak physical memory than arboreto** in a
  controlled 20,000-cell gene-network comparison on the same hardware.
- **21.4% lower topic-model peak memory**, with unchanged output files in repeated
  mouse-brain tests ([v0.5.0 memory audit](https://github.com/Ekin-Kahraman/rustscenic/blob/0c8eb00539e3860c78e452c8661cc2735c169386/validation/scaling/IFB_SCALE_2026-08-28.md#compact-gibbs-token-audit)).
- `11x` to `52x` faster than SCENIC+ for selected analysis stages on sampled real-data inputs, measured on one machine.
- Huang Lab collaborator run recovered `16/17` expected brain transcription factors in human brain data.

The first three results were measured on the **v0.5.0 release candidate**.
The million-cell run used prepared RNA and 2,095 selected genes;
separate full-data preparation peaked at **71.49 GB**. These measurements do not
describe a complete million-cell spatial workflow.

Current release: `v0.5.0`. Python 3.10 to 3.13; Linux, macOS and Windows.
Core analysis runs without Java, dask, CUDA or Snakemake.

## Installation

```bash
pip install rustscenic
```

## Upgrading from 0.4.x

v0.5.0 reduces memory use and adds separate positively and negatively correlated
target sets. Gene-network fitting now follows arboreto's stopping rule; use
`early_stop_mode="legacy_inbag"` only to reproduce the previous behaviour.
See the [release notes](docs/releases/v0.5.0.md) for compatibility options and
the change to Apache-2.0. Earlier published releases remain under MIT.

## Benchmark Evidence

The SCENIC+ comparison starts from prepared matrices and measures gene-network
inference, enhancer links and activity scores. It excludes raw-data processing,
topic modelling and motif-database construction.

The tools use different methods for enhancer linking: RustScenic
uses correlation over the fixed search space, while SCENIC+ uses boosted trees plus
Pearson scoring for region-to-gene links.

Machine: Apple M5 laptop, 16 GB RAM, macOS arm64, 4 CPU threads. RustScenic
rows used Python 3.13.9; SCENIC+ reference rows used Python 3.11.8 for its
dependency stack.
Rows can be sampled subsets; the shape column is the actual benchmark input.

| Dataset | Shape | RustScenic | SCENIC+ | Speedup | Peak RSS (RustScenic / SCENIC+) |
| --- | ---: | ---: | ---: | ---: | ---: |
| PBMC3k dense | 2,000 cells, 4,000 genes, 8,000 peaks, 30 TFs | 4.98 s | 258.9 s | 52x | 1.21 / 1.26 GB |
| PBMC10k dense | 2,000 sampled cells, 4,000 genes, 8,000 peaks, 30 TFs | 21.5 s | 241.5 s | 11x | 2.37 / 2.63 GB |
| Mouse brain E18 | 1,500 cells, 3,000 genes, 6,000 peaks, 25 TFs | 2.82 s | 90.4 s | 32x | 1.65 / 2.10 GB |
| Human brain GEM-X | 2,000 cells, 4,000 genes, 8,000 peaks, 30 TFs | 7.41 s | 146.0 s | 19.7x | 2.18 / 2.19 GB |

Including data preparation, the human brain GEM-X row is `11.89 s` for
RustScenic versus `150.36 s` for SCENIC+.

Full commands, hardware, validation metrics and output signatures are in
[site_docs/benchmarks.md](site_docs/benchmarks.md).

## Stage Coverage

| Stage | RustScenic API | SCENIC ecosystem stage covered |
| --- | --- | --- |
| TF-to-gene GRN | `rustscenic.grn.infer` | GRNBoost2-style regulatory-network inference |
| AUCell | `rustscenic.aucell.score` | Per-cell regulon activity scoring |
| cisTarget | `rustscenic.cistarget.enrich` | Motif enrichment and support filtering |
| Topics | `rustscenic.topics.fit`, `fit_gibbs` | scATAC topic modelling |
| ATAC preprocessing | `rustscenic.preproc` | Fragment matrix building and QC |
| Enhancer links | `rustscenic.enhancer.link_peaks_to_genes` | Peak-to-gene linking |
| eRegulons | `rustscenic.eregulon.build_eregulons` | Enhancer-linked regulon assembly |
| Orchestration | `rustscenic.pipeline.run` | Staged workflow across RNA and multiome inputs |

## Quick Start

This example uses **v0.5.0**. It builds candidate
gene sets from network edges and scores their activity; these sets have not
been filtered for motif support or split by positive and negative correlation.

```python
import anndata as ad
import rustscenic.aucell
import rustscenic.data
import rustscenic.grn

adata = ad.read_h5ad("rna.h5ad")
tfs = rustscenic.data.tfs("hs")

grn = rustscenic.grn.infer(
    adata, tf_names=tfs, n_estimators=5000, top_targets_per_tf=50, seed=777
)
regulons = {
    tf: group["target"].tolist()
    for tf, group in grn.groupby("TF")
    if len(group) >= 10
}
auc = rustscenic.aucell.score(adata, regulons, top_frac=0.05)
```

Command line:

```bash
rustscenic pipeline --rna data.h5ad --tfs tfs.txt --output out/
rustscenic grn --expression rna.h5ad --tfs tfs.txt --output grn.parquet
rustscenic aucell --expression rna.h5ad --regulons grn.parquet --output aucell.parquet
rustscenic topics --expression atac.h5ad --output topics.parquet --n-topics 30
rustscenic cistarget --rankings rankings.feather --regulons regulons.tsv --output motifs.parquet
```

See [examples/pbmc3k_end_to_end.py](examples/pbmc3k_end_to_end.py) for a small
real-data RNA example.

v0.5.0 also provides `add_correlation`, `build_regulons`, and
`rustscenic add-cor` to separate target sets by correlation sign.
See the [API map](site_docs/api.md) for those features. A correlation sign is
not experimental proof of activation or repression.

## Validation

| Validation axis | Result |
| --- | --- |
| cisTarget kernel | Pearson `1.0000` against `ctxcore.recovery.aucs`; mean absolute difference about `2.4e-5`. |
| AUCell parity | Ziegler 2021 airway atlas mean per-cell Pearson `0.984`; `91.7%` of cells above `0.95`. |
| Human brain GEM-X benchmark | Region-to-gene Jaccard `1.000`; region AUCell mean Pearson `0.823`. |
| Collaborator analysis | A human brain RNA/chromatin workflow recovered `16/17` expected brain transcription factors. |
| Open parity targets | Gene AUCell Pearson `0.386` and eRegulon edge Jaccard `0.161` on the human brain GEM-X row. |

Validation artefacts live under [validation/](validation/). Public interpretation
lives in [site_docs/benchmarks.md](site_docs/benchmarks.md) and
[site_docs/validation.md](site_docs/validation.md).

## Current Boundaries

- The SCENIC+ speedups cover selected analysis stages, not every possible
  raw-data and motif-database workflow.
- GRN, gene AUCell and eRegulon edge agreement are not claimed to be
  bit-identical to SCENIC+; see [Benchmarks](site_docs/benchmarks.md) for the
  parity metrics.
- v0.5.0 changes `grn.infer` to arboreto-compatible early stopping.
  Use `early_stop_mode="legacy_inbag"` only to reproduce historical
  RustScenic stopping behaviour. The pipeline defaults to separate activator
  and repressor regulons; `grn_regulon_polarities="unsigned"` is the explicit
  compatibility path.
- Complete million-cell spatial workflows and an atlas-wide CELLxGENE resource
  remain outside the validated scope.

## Documentation

- [Installation](site_docs/installation.md)
- [Quickstart](site_docs/quickstart.md)
- [API map](site_docs/api.md)
- [HPC operation](site_docs/hpc.md)
- [Benchmarks](site_docs/benchmarks.md)
- [Validation](site_docs/validation.md)
- [Scope](site_docs/limitations.md)

## Citation

If you use RustScenic in a paper, report, benchmark, derivative package or lab
workflow, cite the exact release used. GitHub citation metadata is in
[CITATION.cff](CITATION.cff). Zenodo concept DOI:
[10.5281/zenodo.20246040](https://doi.org/10.5281/zenodo.20246040).

RustScenic is created and maintained by Ekin Kahraman. See [AUTHORS.md](AUTHORS.md)
for attribution.

