Metadata-Version: 2.4
Name: phyloom
Version: 0.1.0a6
Summary: Generate phylogenetic datasets with minimal setup effort
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE.txt
Requires-Dist: joblib==1.5.2
Requires-Dist: pandas==2.3.3
Requires-Dist: pydantic==2.12.3
Requires-Dist: pyyaml==6.0.3
Requires-Dist: scipy>=1.15.3
Requires-Dist: tqdm==4.67.1
Dynamic: license-file

<p align="center">
  <img src="https://raw.githubusercontent.com/gabriele-marino/phyloom/main/logo.svg" alt="Phyloom" width="620" />
</p>

---

[![Python](https://img.shields.io/badge/python-3.10%2B-3776AB?style=flat-square)](https://www.python.org/)
[![Documentation](https://img.shields.io/badge/docs-website-6F42C1?style=flat-square)](https://gabriele-marino.github.io/phyloom/)

# Phyloom

[Phyloom](https://github.com/gabriele-marino/phyloom) is a Python framework for designing stochastic phylogenetic tree and network simulators, then using them to generate complete datasets with minimal setup. It provides reusable simulation machinery so you can focus on the scientific model and its events instead of rebuilding the simulation loop and dataset-generation boilerplate.

Build a model directly with Phyloom's Python API, extend it with your own event rules and hazards, or describe simulations in YAML configuration files. Phyloom currently supports **BDSH**, **epidemic**, and **coalescent** model families.

## Features

- **Build custom tree and network models.** Compose event rules that describe eligible state changes with hazards that determine when those changes occur. Add scheduled actions, state-dependent rate multipliers, and model-specific state. Phyloom handles event scheduling and simulation time.
- **Start with built-in model families.** BDSH includes birth, death, sampling, migration, and hybridization events. Epidemic models combine lineage states with population pools and events such as infection, recovery, migration, and flow. Coalescent models simulate backward-time retracing and coalescence.
- **Keep model code focused.** Extend the Python model and event-rule interfaces, then register the model for use in configurations. Phyloom takes care of the common simulation and configuration plumbing.
- **Generate datasets from YAML.** Define parameter distributions and derived values once; Phyloom materializes a context for each sample and uses it throughout the model, stopping and acceptance criteria, metadata logs, and output hooks.
- **Control the full run.** Generate named dataset splits, set random seeds, run samples in parallel, stop simulations by criteria or time limits, and reject simulations that do not meet acceptance criteria.
- **Keep useful metadata and outputs together.** Record sampled parameters and custom logs in `metadata.csv`. Output hooks can write model-specific files for each accepted sample, while timeout and rejection context logs help track unsuccessful trials.
- **Extend with plugins.** Register custom model families and event rules so they can be loaded by Phyloom's command-line runner.

## Installation

Phyloom requires Python 3.10 or newer. Install it from GitHub:

```bash
git clone https://github.com/gabriele-marino/phyloom.git
cd phyloom
pip install .
```

## Quick start

Run one of the example configurations from the repository:

```bash
phyloom examples/yaml/SEIR.yaml
```

The example defines an epidemic model, samples its parameters, runs the requested simulations, and writes accepted samples and metadata beneath the configured `output_dir`. To explore the other included configurations, see [`examples/yaml`](https://github.com/gabriele-marino/phyloom/tree/main/examples/yaml).

A configuration can define sample counts and splits, parameter distributions and expressions, a model and its events, stopping and acceptance criteria, metadata logs, and output hooks. For example, the included [BD configuration](https://github.com/gabriele-marino/phyloom/blob/main/examples/yaml/BD.yaml) samples model parameters for each dataset sample and writes Newick histories.

## Documentation and examples

Visit the [Phyloom documentation](https://gabriele-marino.github.io/phyloom/) for the Python API, configuration reference, built-in models, and guidance on extending the simulator. Browse the [example configurations](https://github.com/gabriele-marino/phyloom/tree/main/examples/yaml) to see complete dataset-generation workflows.

## License

Phyloom is released under the [MIT License](https://github.com/gabriele-marino/phyloom/blob/main/LICENSE.txt).

## Contact

For questions, bug reports, or feature requests, [open an issue on GitHub](https://github.com/gabriele-marino/phyloom/issues).
