Metadata-Version: 2.4
Name: phd_ms
Version: 1.3
Summary: Identifying multiscale tissue domains for spatial transcriptomic data using persistent homology.
Author: Perry Beamer
License: MIT
Keywords: multiscale domains,spatial transcriptomics,topological data analysis
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Operating System :: POSIX :: Linux
Classifier: Operating System :: MacOS
Classifier: Topic :: Scientific/Engineering
Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE.md
Requires-Dist: gudhi>=3.8.0
Requires-Dist: leidenalg>=0.9.0
Requires-Dist: matplotlib>=3.5
Requires-Dist: mpl_point_clicker
Requires-Dist: numpy>=1.23.4
Requires-Dist: pandas>=2.2.0
Requires-Dist: pot>=0.8.0
Requires-Dist: scanpy>=1.9
Requires-Dist: scipy>=1.9.1
Requires-Dist: scikit-learn>=1.5.1
Requires-Dist: tqdm>=4.67
Dynamic: license-file


# PHD-MS: Persistent Homology for Domains at Multiple Scales

Multiscale domain identification for spatial transcriptomic data.


## Installation

Simply install with pip:

    pip install phd_ms

Or install from source by downloading, cd to this directory, using:

    pip install .

## Usage/Examples

Detailed jupyter notebook tutorials are available in the examples folder.
Or simply download the file titled `point_and_click.py` and run it to use our clickable graphical interface. When using `point_and_click.py`, set the path to your data by changing this line:
```python
DATA = '/path/to/your_data.h5ad'
```

Here we'll show a simple example with Visium DLPFC data, using default parameters.
First, import necessary components.
```python
import numpy as np
import phd_ms
```
Preprocessing steps here (note that we assume a spatially-aware embedding has already been computed for your data and stored in `adata.obsm`):
```python
INPUT_FILE = '/path/to/adata_151673.h5ad'
RESOLUTIONS = np.linspace(start=0.15, stop=.95, num=8)
RES_KEYS = [f'leiden_{r:.3f}' for r in RESOLUTIONS]
adata = phd_ms.tl.preprocess_leiden(INPUT_FILE, output_file=INPUT_FILE, emb='X_gst',
                                    resolution=RESOLUTIONS, res_keys=RES_KEYS)
```
Compute persistent homology and plot the 10 most prominent multiscale domains:
```python
cluster_complex, clusterings = phd_ms.tl.cluster_filtration(adata, res_keys=RES_KEYS)
domains = phd_ms.tl.map_multiscale(adata.obsm['spatial'], cluster_complex, clusterings, num_domains=10)
```
