Metadata-Version: 2.4
Name: pad-causal
Version: 0.1.1
Summary: PaD: joint interpretation of single- and double-perturbation evidence
License-Expression: LicenseRef-Proprietary
Keywords: causal-inference,perturbation,gene-interactions,kernel-learning,PaD
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Operating System :: OS Independent
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE.txt
License-File: NOTICE.md
Requires-Dist: numpy<3,>=1.23
Requires-Dist: scipy<2,>=1.10
Requires-Dist: scikit-learn<2,>=1.3
Provides-Extra: test
Requires-Dist: pytest>=7; extra == "test"
Provides-Extra: release
Requires-Dist: build>=1.2; extra == "release"
Requires-Dist: twine>=6; extra == "release"
Dynamic: license-file

# PaD: joint interpretation of perturbation evidence

**PaD learns when double-perturbation context changes the meaning of a directional signal.** It combines a frozen Proposer score, directional double-perturbation features, and a product kernel in one learning objective.

The `pad-causal` distribution provides the algorithm as a reusable Python API and command-line tool. It supports **causal-direction scoring** and **ordered-reference ranking**, with explicit pair-exchange semantics. The kernel width, 35-candidate pool and paired one-standard-error selector match the frozen research implementation.

This is an anonymous review distribution. See `docs/ANONYMITY.md` and `docs/CHANGES.md` in the source package.

## Install

After this distribution has been published on PyPI:

```bash
python -m pip install pad-causal
```

Before publication, install the wheel supplied with the release:

```bash
python -m pip install pad_causal-0.1.1-py3-none-any.whl
```

Python 3.10 or newer is required. Runtime dependencies are NumPy, SciPy and scikit-learn. No benchmark archive, neural-network runtime or external baseline library is required.

## Run an example

```bash
pad-causal demo --output pad_demo
```

This trains the original joint kernel on a small synthetic example, predicts an independent query set, and writes a numeric model and a score table. It is a usage example, not a reproduction of a paper benchmark.

The same workflow in Python:

```python
import numpy as np
from pad_causal import PaD, endpoint_disjoint_splits
from pad_causal.demo import make_demo

train, y, pairs = make_demo(n=64, random_state=17)
query, _, _ = make_demo(n=32, random_state=29)
folds = endpoint_disjoint_splits(pairs, np.arange(len(train)) % 5)

model = PaD(task="direction").fit(train, y, pairs=pairs, cv=folds)
scores = model.score_pairs(query)
print(scores.direction)

model.save("pad_model.npz")
restored = PaD.load("pad_model.npz")
print(restored.predict(query))
```

## Use your evidence

`Evidence` is a label-free container. Supply measurements transformed consistently on the training and query sides:

```python
from pad_causal import Evidence, PaD

evidence = Evidence(
    p_score=p_score,       # n frozen Proposer scores
    p_summary=p_summary,   # n x 8 Proposer summary matrix
    d_odd=d_odd,           # n x q raw directional D features
    d_summary=d_summary,   # n x k transformed D summary, k >= q
)
model = PaD().fit(evidence, y, pairs=pairs, cv=folds)
```

In `p_summary`, the first coordinate changes sign when A and B are exchanged; the remaining coordinates are invariant. In `d_summary`, the first `q` coordinates change sign and the rest are invariant. `d_odd` and `p_score` also change sign. Reverse arrays are computed from these rules; explicitly supplied reverse arrays are validated.

The package accepts the prepared NPZ files from the companion reproducibility repository:

```python
from pad_causal import load_evidence
inputs = load_evidence("Costanzo.npz")
```

The original keys `p`, `xp`, `odd`, `xd`, `xpr`, and `xdr` are supported. Stored `y` is not loaded as a feature. Supply labels separately to `fit`.

NaN inputs are rejected at this interface. Missing measurements should first be encoded with the applicable finite-value/mask convention. The package does not infer assay-specific normalizations from test data.

## Direction and ranking

| Mode | Learning target | Returned task score |
|---|---|---|
| `task="direction"` | Orientation of a candidate relation | Antisymmetric direction margin |
| `task="ranking"` | Support of an ordered reference record | Full ordered score, including symmetric support |

`score_pairs` always returns four arrays: `score`, `reverse_score`, `direction=(score-reverse_score)/2`, and `support=(score+reverse_score)/2`. `decision_function` selects the mode's task score. `predict` applies a zero threshold, with ties assigned to label 1. Scores are **not calibrated probabilities**.

## Selection and fixed fits

The default `selection="cv"` evaluates the original seven kernel weights and five regularizers. Pass either pair endpoints or explicit training-side CV folds. When pair endpoints are supplied, train/validation endpoint overlap and reverse records split across validation groups are rejected. All folds must validate every training record exactly once and leave both labels in each fit.

`endpoint_disjoint_splits(pairs)` holds out one unordered pair at a time. Supply `fold_ids` to hold out larger groups. Reverse records must share a group. Final test records must remain outside `fit`.

For a previously selected configuration:

```python
model = PaD(
    task="ranking",
    selection="fixed",
    weights=(0.0, 0.5, 0.5),
    regularization=0.003,
).fit(train, y)
```

Fixed mode performs no model search. `selected_weights_`, `regularization_`, `cv_results_`, and `n_splits_` expose the fitted choice and selection trace.

## Proposer and local readout primitives

`FrozenProposer.transform` computes the inherited five-feature score from log2-expression responses and p-values, masking the candidate endpoints' intervention columns. It also returns the eight-coordinate summary used by the mechanism encoder. Its coefficients are frozen; no Proposer retraining occurs.

`double_features(single_a, single_b, double, null)` returns the original local single-effect contrasts, five double-perturbation contrasts and four quality coordinates. It is a readout primitive, not a universal complete-D adapter. The biological panels have different measurement/summary dimensions; the caller supplies the corresponding context and transformation conventions.

## CLI

```bash
pad-causal fit --input train.npz --pairs pairs.csv --fold-column fold --task direction --output model.npz
pad-causal predict --model model.npz --input query.npz --output scores.csv
```

The training NPZ needs `y`; the pair CSV needs `a`, `b`, and the requested fold column. A query NPZ does not need labels. `python -m pad_causal` exposes the same commands on systems where the console script is not on PATH.

## Reproducibility and scope

Version 0.1.1 is the anonymous, corrected toolkit accompanying GitHub reproducibility release v4 and frozen paper v08. The learning algorithm is inherited from release v3. Models are saved as numeric NPZ arrays plus JSON metadata, without pickle execution. Predictions use training-frozen scales and rectangular query-by-training kernels.

The companion research repository retains the full benchmark data, baselines, certificates and analysis. This wheel supplies the learning tool. It does not certify arbitrary new datasets, infer unobserved interventions, identify direct edges from total effects without additional assumptions, or establish that one set of inherited Proposer coefficients transfers to every assay.

Release checks and replay commands are recorded in `docs/VALIDATION.md` in the source distribution. The public API reproduced Costanzo's 611/657 and Jonikas's 0.881114 AUROC under the frozen protocols.

Original study: *PaD: When Perturbation Context Changes the Meaning of Directional Evidence*. Anonymous manuscript, paper v08.

## License

This anonymous review distribution preserves the original copyright status (`LicenseRef-Proprietary`); no open-source grant is added by the packaging work. See the included `LICENSE.txt` and `NOTICE.md`. The author may select another project license for a later release.
