Metadata-Version: 2.4
Name: apb2
Version: 0.1.0
Summary: Convert proteomics software output to AnnData (rules-driven parser, second generation)
Keywords: proteomics,mass spectrometry,anndata,mudata,quantification
Author: Witold Wolski
Author-email: Witold Wolski <wew@fgcz.ethz.ch>
License-Expression: MIT
License-File: LICENSE
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
Classifier: Typing :: Typed
Requires-Dist: anndata>=0.11
Requires-Dist: cyclopts>=3
Requires-Dist: duckdb>=1.4,<2
Requires-Dist: loguru>=0.7
Requires-Dist: mudata>=0.4,<1
Requires-Dist: numpy>=2
Requires-Dist: packaging>=24
Requires-Dist: pandas>=2.2
Requires-Dist: polars>=1.43,<2
Requires-Dist: fastexcel>=0.21,<1
Requires-Dist: python-calamine>=0.5,<1
Requires-Dist: pydantic>=2.10
Requires-Dist: pyarrow>=15
Requires-Dist: pyyaml>=6
Requires-Dist: scipy>=1.15,<2
Requires-Python: >=3.13
Project-URL: Documentation, https://anndata-omics-bridge.github.io/apb2/
Project-URL: Repository, https://github.com/anndata-omics-bridge/apb2
Description-Content-Type: text/markdown

# apb2

APB2 is a [rules-driven framework](https://anndata-omics-bridge.github.io/apb2/rule-based/) for converting outputs from proteomics
software into AnnData or MuData. It supports ion, peptidoform, peptide, protein, and fragment
quantification levels and can also store the parsed data in Parquet or DuckDB.

“Rules-driven” means that declarative rule documents describe each vendor table: which columns
contain identifiers, measurements, and metadata, how those columns should be reshaped, and which
constraints the result must satisfy. One shared parser applies those rules, so a new or revised
input format can usually be supported by adding or updating a rule instead of writing a dedicated
reader.

Read the rendered [APB2 documentation](https://anndata-omics-bridge.github.io/apb2/) or its
[source index](https://github.com/anndata-omics-bridge/apb2/blob/main/docs/index.md). The [supported-software matrix](https://anndata-omics-bridge.github.io/apb2/supported_software/) lists
every packaged software version, quantification level, vendor input type, table shape, and
parameter parser.

Choose the documentation for your interface:

- [CLI reference](https://anndata-omics-bridge.github.io/apb2/cli/) and [command-line guides](https://anndata-omics-bridge.github.io/apb2/conversion/)
- [Python API reference](https://anndata-omics-bridge.github.io/apb2/api/)

## Installation

APB2 requires Python 3.13 or later.

```bash
pip install apb2
```

To install only the `apb2` command, use `uv tool install apb2`.

## Motivation and origin

APB2 is a refactoring and performance improvement of the now-discontinued [AnnData Proteomics Bridge (APB v1)](https://github.com/anndata-omics-bridge/anndata-proteomics-bridge).

APB2 is based on the work of [ProteoBench](https://github.com/proteobench/proteobench): it ports ProteoBench's parsing infrastructure — search-parameter parsing, vendor file-format parsing, modification parsing — into one rules-driven converter. See the ProteoBench preprint: [ProteoBench: the community-curated platform for comparing proteomics data analysis workflows](https://www.biorxiv.org/content/10.64898/2025.12.09.692895v2) (bioRxiv, 2025, doi:10.64898/2025.12.09.692895).

Packaged conversion rules include AlphaDIA, AlphaPept, DIA-NN, FragPipe, i2MassChroQ, MaxQuant, MSAngel, PEAKS, ProteoBench Custom, ProlineStudio, quantms, Sage, Spectronaut, and WOMBAT; the complete version and input-format matrix is in [supported software](https://anndata-omics-bridge.github.io/apb2/supported_software/).

The work that became APB2 was discussed and started during the Copenhagen ProteoBench Hackathon,
13–17 April 2026, as one of the efforts to improve the backend of the
[ProteoBench platform](https://proteobench.cubimed.rub.de/). The hackathon included the public
[EuBIC-MS Seminar 2026 on 15 April](https://eubic-ms.org/events/latest-developments-and-tools-for-data-analysis/).

APB2 was also motivated by the vendor-specific readers maintained behind
[`prolfquapp::preprocess_software()`](https://github.com/prolfqua/prolfquapp/blob/master/R/preprocess_software.R#L137)
and in
[`prolfquappPTMreaders`](https://github.com/prolfqua/prolfquappPTMreaders). We plan to move their
remaining input variants and PTM/site-level formats into APB2 so one rules-driven parser can serve
both prolfquapp and ProteoBench, and hopefully other tools analysing quantification data.

## Command-line interface

### Convert

Use a packaged rule selected from the vendor parameter file and source header. `DATA` may be one vendor table or a vendor-result directory:

`--software` selects the parameter-file grammar and restricts result recognition to that vendor and its declared quantification software. For FragPipe parameters with DIA-NN output, pass `--software fragpipe`. Omit the hint for recognition across parameter-bearing rules; mismatches and ambiguity are errors, not a fallback to unrelated rules.

ProteoBench Custom uploads have no parameter file: `apb2 convert custom.txt ion --software pb_custom --output results/custom`.

```bash
apb2 convert DATA LEVEL --params PARAMETER_FILE [--software VENDOR] [--output BASENAME]
```

Omit `LEVEL` to convert every compatible level into one APB2 result:

```bash
apb2 convert DATA --params PARAMETER_FILE [--software VENDOR] [--format FORMAT] [--output BASENAME]
```

Use an explicit schema-0.8 rule document, with optional search-parameter evidence:

```bash
apb2 convert DATA LEVEL --rule-config RULES_JSON [--params PARAMETER_FILE] \
  [--software VENDOR] [--output BASENAME]
```

`LEVEL` is one of `ion`, `peptidoform`, `peptide`, `protein`, or `fragment`. MaxQuant accepts any nonempty subset of evidence, modification-specific peptide, peptide and protein-group exports. Evidence stays separate from the higher-level join; an omitted level converts every available level. `--format` selects `hdf5`, `parquet`, or `duckdb`. HDF5 uses `.h5ad` with an explicit level and `.h5mu` otherwise. Complete one-to-one observation aliases are aligned; fractionated or unmapped resolutions produce separate outputs such as `result.raw_file.h5mu` and `result.experiment.h5mu`. See [output naming](https://anndata-omics-bridge.github.io/apb2/conversion/#output-naming). `--strict` promotes layer-contract warnings to errors. `--timings-output PATH` optionally writes a separate versioned JSON file containing internal compile, read, parse and write durations plus per-level read/parse durations; it does not enter the APB result or its scientific representation. The command performs conversion only; FASTA annotation and protein inference are outside Parser V2.

### Reformat a parsed result

Change only the persisted format; no vendor parsing or annotation runs:

```bash
apb2 reformat SOURCE TARGET
```

The suffix selects h5ad, h5mu, an APB2 Parquet directory dataset, or DuckDB.

### Annotate samples

Attach a generic prolfquapp-style CSV/TSV table to any APB2 result format:

```bash
apb2 annotate INPUT ANNOTATION OUTPUT
```

The default prolfquapp behavior retains unmatched quantitative observations and writes null
annotation fields. `--unmatched error` requires complete coverage; `--unmatched drop` explicitly
subsets every observation-aligned value. ProteoBench-specific module annotation and scoring live
in the separate `apb-proteobench` package. See the
[sample-annotation guide](https://anndata-omics-bridge.github.io/apb2/sample_annotation/).

## Python API

The file-to-file facades mirror the CLI operations. The compiler/parser APIs expose
storage-neutral values for custom pipelines. Result formats also have explicit adapters:

```python
from pathlib import Path

from apb2.api import read_parsed_levels, write_parsed_levels

parsed = read_parsed_levels(Path("result.parquet"))
write_parsed_levels(parsed, Path("result.duckdb"))
```

Parquet and DuckDB preserve Polars result values exactly; h5ad and h5mu apply the stored
numeric/factor matrix projection. Every public result write also publishes an adjacent compact `.apb.json` scientific representation for inspection without loading the full result. See the [Python API reference](https://anndata-omics-bridge.github.io/apb2/api/) for vendor
conversion, annotation, result values, and errors.

## Architecture

The CLI delegates conversion to Parser V2 and annotation to the independent annotation facade.
The controlling designs and dependency boundaries are documented in
[`docs/architecture_converter.md`](https://anndata-omics-bridge.github.io/apb2/architecture_converter/) and
[`docs/architecture_annotation.md`](https://anndata-omics-bridge.github.io/apb2/architecture_annotation/).

## Development

```bash
uv sync --group dev
make check
make docs
.venv/bin/pre-commit install --hook-type pre-commit --hook-type pre-push
```

All Python commands run from the synchronized project `.venv`.
`make docs-serve` serves the user documentation locally. GitHub Actions publishes the strict
Zensical build to GitHub Pages from `main`.

The rule JSON Schema is a packaged artifact. Developers regenerate it from the Parser V2 rule
package rather than through a user-facing CLI command:

```bash
uv run python -c 'from apb2.parserV2.vendor_parse_rules.schema_artifact import write_artifact; write_artifact()'
```
