Metadata-Version: 2.4
Name: salmopredict
Version: 0.3.0
Summary: AutoGluon-based Incidence predictor for Salmonella virulence-factor gene-frequency features
Author-email: Dongyan Shao <563608176@qq.com>
License: PolyForm-Noncommercial-1.0.0
Project-URL: Homepage, https://github.com/shaodongyan/SalmoPredict
Project-URL: Repository, https://github.com/shaodongyan/SalmoPredict
Project-URL: Issues, https://github.com/shaodongyan/SalmoPredict/issues
Keywords: salmonella,incidence,autogluon,prediction,bioinformatics
Classifier: License :: Other/Proprietary License
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
Requires-Python: <3.11,>=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: autogluon.tabular[fastai]==1.1.1
Requires-Dist: setuptools<81
Requires-Dist: pandas>=2.0
Requires-Dist: openpyxl>=3.0
Requires-Dist: rich-argparse>=1.4
Requires-Dist: streamlit>=1.30
Provides-Extra: gui
Requires-Dist: streamlit>=1.30; extra == "gui"
Dynamic: license-file

<p align="center">
  <img src="https://raw.githubusercontent.com/shaodongyan/SalmoPredict/main/salmopredict/gui/assets/salmopredict_icon.png" alt="salmopredict" width="200">
</p>

# salmopredict

AutoGluon-based **Incidence** predictor for *Salmonella* virulence-factor
gene-frequency features, with a command-line interface and a Streamlit GUI.

Given a feature table (rows = samples, columns = virulence-factor genes),
salmopredict aligns the columns to the features a pre-trained AutoGluon
`TabularPredictor` expects, runs the `WeightedEnsemble_L2` model, and writes a
single prediction file. It reproduces the alignment used by the original
`predict_autogluon.py`: column names are normalised R-`make.names`-style
(`/` and `-` become `.`), genes the model expects but the input lacks are filled
with `0` (a missing gene means frequency 0), and extra input columns are ignored.

**Feature CSV input/output** — the output columns depend on whether the input
has a `Sample` column:

| Input | Output columns |
|-------|----------------|
| No `Sample` column (features only) | `Incidence(%)` |
| Has a `Sample` column | `Sample`, `Incidence(%)` |
| Has a `Sample` column **and** `--attach meta.csv` | `Sample`, `Incidence(%)`, + the metadata's other columns |

Metadata is joined on the `Sample` key (the metadata CSV must also have a
`Sample` column), so attaching metadata requires a `Sample` column in the input.

## Install

salmopredict runs on **Python 3.10** and loads its model with **AutoGluon
1.1.1** — both are hard requirements, because the model is pickled with that
exact stack.

**New installation from this source directory.** Run these commands from the
directory containing `pyproject.toml`:

```bash
conda create -n salmopredict python=3.10
conda activate salmopredict
python -m pip install .
salmopredict check
```

This installs AutoGluon 1.1.1 with its **Torch and FastAI backends**, compatible
`setuptools<81`, the Streamlit GUI, and the bundled prediction model. The
backends are required by the bundled ensemble; base `autogluon.tabular` alone
does not install them. AutoGluon 1.1.1 also needs `pkg_resources`, which newer
setuptools releases no longer provide.

Both interfaces are then available:

```bash
salmopredict run -i features.csv -o results/   # command line
salmopredict gui                                # browser GUI
```

**Conda environment (alternative, from this source directory).** Pins Python 3.10 and installs
AutoGluon via pip inside the env (conda-installed AutoGluon does not resolve
cleanly for this project):

```bash
conda env create -f environment.yml
conda activate salmopredict
salmopredict check
```

**Editable / development install (from a clone).**

```bash
python -m pip install -e .          # installs the CLI and the Streamlit GUI
```

**Install the updated local wheel.** In a Python 3.10 environment:

```bash
python -m pip install --upgrade dist/salmopredict-0.3.0-py3-none-any.whl
salmopredict check
```

**Update an existing source installation.** Stop a running GUI with `Ctrl+C`,
then run from this source directory:

```bash
conda activate salmopredict
python -m pip install --upgrade -e .
salmopredict check
salmopredict gui
```

**PyPI installation / upgrade.** In a Python 3.10 environment:

```bash
python -m pip install --upgrade "salmopredict>=0.3.0"
salmopredict check
```

Version 0.3.0 includes two-table prediction and installs the required Torch,
FastAI and compatible setuptools dependencies automatically.

If the page opens but prediction reports `No module named 'pkg_resources'`,
`torch`, or `fastai`, use the update/repair command above and restart the GUI.
To verify actual prediction from a source checkout (use a new output folder):

```bash
salmopredict run -i examples/example_features.csv -o results_install_check/
```

## The model

The prediction model is **already bundled** with salmopredict — both in this
repository and inside the PyPI wheel — at `salmopredict/models/model_default`, a
30 MB deployment. salmopredict uses it automatically, so the tool works out of
the box with no extra download or build step.

Model resolution order is `--model`, then `$SALMOPREDICT_MODEL`, then the single
directory under the package `models/` folder; with nothing specified it uses the
bundled `model_default`. Pass `--model /path/to/other` to run a different
AutoGluon model.

## Usage

Ready-to-run inputs live in [`examples/`](examples/) (see its README):
`example_features.csv` (Type 1, no `Sample`), `example_with_sample.csv`
(Type 2, with `Sample`), and `example_meta.csv` (metadata to attach). Features
are `gene_frequency × log10(CFU dose)`, matching how the model was trained. Try
one immediately:

```bash
salmopredict run -i examples/example_features.csv -o results/
```

```bash
# Features only  -> output has just Incidence(%)
salmopredict run -i features.csv -o results/ --model /path/to/model

# With a Sample column  -> output has Sample, Incidence(%)
salmopredict run -i examples/example_with_sample.csv -o results/

# Attach metadata joined on the Sample key -> Sample, Incidence(%), + meta columns
salmopredict run -i examples/example_with_sample.csv -o results/ \
  --attach examples/example_meta.csv

# Launch the GUI, or check the environment/model
salmopredict gui
salmopredict check --model /path/to/model
```

Each feature-input run writes one `pred_<input-stem>.csv` to the output directory; the
prediction column is `Incidence(%)`. Features filled with `0` (genes the model
expects but the input lacks) are always reported, and a prominent warning
appears when more than `--missing-warn-frac` (default 0.3) of the model's
features are missing.

## Predict from samples and gene frequencies

Supply two CSV files instead of calculating features yourself:

* **Samples**: `Sample,dose_cfu,serotype`. `dose_cfu` contains raw CFU, e.g.
  `1000`, not `3`. Each Sample must be nonblank and unique.
* **Gene frequencies**: `Serotype` plus one column per gene, one row per
  serotype, with numeric frequencies from 0 to 1. This accepts the layout of
  `02_gene_frequencies.csv` directly.

```bash
salmopredict run \
  --samples examples/example_samples.csv \
  --gene-frequencies examples/example_gene_frequencies.csv \
  -o results_two_tables/
```

For each sample, the program looks up its serotype and calculates
`gene_frequency × log10(dose_cfu)`, then predicts incidence. It writes:

* `features_example_samples.csv`: `Sample` and the calculated gene features.
* `pred_example_samples.csv`: `Sample,dose_cfu,serotype,Incidence(%)`.

Sample order and identifier strings (including leading zeros) are preserved.
Required header names are case-insensitive. Serotype values match exactly after
trimming outer spaces; synonyms and spelling differences are not guessed.
Missing serotypes stop the run and list affected samples. Duplicate sample IDs
or serotypes, blank required values, nonfinite/nonpositive doses, and frequencies
outside [0, 1] also stop the run. Additional sample columns are ignored. Missing
model genes use the existing fill-and-warning behavior.

`--samples` and `--gene-frequencies` must be used together and cannot be combined
with `-i` or `--attach`. Use `--force` to replace existing output files.

In the **GUI**, choose **Samples + gene frequencies** under **Input mode**,
upload both CSVs, inspect their previews, choose an output folder and click
**Run prediction**. The result table and both CSV download buttons appear after
success. **Feature CSV** selects the existing single-table workflow.

The two example inputs were reconstructed from the matching sample/dose metadata
and dose-weighted features. See [examples/README.md](examples/README.md) for their
provenance and rounding tolerance.

## License

Licensed under the [PolyForm Noncommercial License 1.0.0](LICENSE): free to use,
modify, and share for any **noncommercial** purpose — including research,
teaching, and personal use, and by academic, government, public-health, and
other nonprofit organizations — but **commercial use is not permitted**.
Developed at the State Key Laboratory of Veterinary Public Health and Safety,
China Agricultural University, in collaboration with the China National Center
for Food Safety Risk Assessment (CFSA).

<p align="center">
  <img src="https://raw.githubusercontent.com/shaodongyan/SalmoPredict/main/salmopredict/gui/assets/vphs_logo.png" alt="State Key Laboratory of Veterinary Public Health and Safety" height="80">
  &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;
  <img src="https://raw.githubusercontent.com/shaodongyan/SalmoPredict/main/salmopredict/gui/assets/cfsa_logo.png" alt="China National Center for Food Safety Risk Assessment" height="80">
</p>
