Metadata-Version: 2.4
Name: wlearn
Version: 0.3.2
Summary: Portable ML computation primitives
Author: Anton Zemlyansky
License-Expression: Apache-2.0
Project-URL: Homepage, https://github.com/wlearn-org/wlearn
Project-URL: Repository, https://github.com/wlearn-org/wlearn
Keywords: machine-learning,ml,wasm,xgboost,random-forest,portable
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
License-File: NOTICE
Requires-Dist: numpy>=1.22
Provides-Extra: xgboost
Requires-Dist: xgboost>=2.0; extra == "xgboost"
Provides-Extra: sklearn
Requires-Dist: scikit-learn>=1.3; extra == "sklearn"
Provides-Extra: liblinear
Requires-Dist: liblinear-official>=2.50; extra == "liblinear"
Provides-Extra: libsvm
Requires-Dist: libsvm-official>=3.30; extra == "libsvm"
Provides-Extra: nanoflann
Requires-Dist: pynanoflann>=0.10; extra == "nanoflann"
Provides-Extra: ebm-fit
Requires-Dist: interpret>=0.6; extra == "ebm-fit"
Provides-Extra: tsetlin-fit
Requires-Dist: tmu<0.9,>=0.8.3; extra == "tsetlin-fit"
Provides-Extra: lightgbm
Requires-Dist: lightgbm>=4.0; extra == "lightgbm"
Provides-Extra: stochtree
Requires-Dist: stochtree>=0.1; extra == "stochtree"
Provides-Extra: nn
Requires-Dist: polygrad>=0.7.0; extra == "nn"
Provides-Extra: dimred
Requires-Dist: wlearn-dimred<0.3,>=0.2.0; extra == "dimred"
Provides-Extra: bo
Requires-Dist: wlearn-bo>=0.1.0; extra == "bo"
Provides-Extra: rf
Requires-Dist: wlearn-rf>=0.4.1; extra == "rf"
Provides-Extra: gam
Requires-Dist: wlearn-gam>=0.2.0; extra == "gam"
Provides-Extra: cluster
Requires-Dist: wlearn-cluster>=0.1.0; extra == "cluster"
Provides-Extra: basis
Requires-Dist: wlearn-basis>=0.1.0; extra == "basis"
Provides-Extra: sym
Requires-Dist: wlearn-sym[polygrad]>=0.1.0; extra == "sym"
Provides-Extra: uncertainty
Requires-Dist: wlearn-uncertainty>=0.1.0; extra == "uncertainty"
Provides-Extra: preprocess
Requires-Dist: tranfi<0.3,>=0.2.1; extra == "preprocess"
Provides-Extra: all
Requires-Dist: xgboost>=2.0; extra == "all"
Requires-Dist: scikit-learn>=1.3; extra == "all"
Requires-Dist: liblinear-official>=2.50; extra == "all"
Requires-Dist: libsvm-official>=3.30; extra == "all"
Requires-Dist: pynanoflann>=0.10; extra == "all"
Requires-Dist: interpret>=0.6; extra == "all"
Requires-Dist: tmu<0.9,>=0.8.3; extra == "all"
Requires-Dist: lightgbm>=4.0; extra == "all"
Requires-Dist: stochtree>=0.1; extra == "all"
Requires-Dist: polygrad>=0.7.0; extra == "all"
Requires-Dist: wlearn-bo>=0.1.0; extra == "all"
Requires-Dist: wlearn-rf>=0.4.1; extra == "all"
Requires-Dist: wlearn-gam>=0.2.0; extra == "all"
Requires-Dist: wlearn-cluster>=0.1.0; extra == "all"
Requires-Dist: wlearn-basis>=0.1.0; extra == "all"
Requires-Dist: wlearn-sym[polygrad]>=0.1.0; extra == "all"
Requires-Dist: wlearn-uncertainty>=0.1.0; extra == "all"
Requires-Dist: tranfi<0.3,>=0.2.1; extra == "all"
Provides-Extra: test
Requires-Dist: pytest>=7.0; extra == "test"
Requires-Dist: hypothesis>=6.0; extra == "test"
Provides-Extra: solver
Requires-Dist: z3-solver<4.15.4,>=4.12; extra == "solver"
Dynamic: license-file

# wlearn (Python)

Portable ML computation primitives for Python. Train models with native backends, save to cross-language `.wlrn` bundles, load bundles produced by JS `@wlearn/*` packages. The base package depends on NumPy because core task, prediction, measure, resampling, AutoML, and ensemble primitives operate on numeric arrays.

Part of [wlearn](https://wlearn.org) ([GitHub](https://github.com/wlearn-org), [all packages](https://github.com/wlearn-org/wlearn#repository-structure)).

## Install

```bash
pip install 'wlearn[xgboost]'
```

That command installs the backend used by the runnable quick start. Use
`pip install wlearn` for bundle/core utilities without a training backend.

Install with model backends as needed:

```bash
pip install wlearn[xgboost]      # XGBoost
pip install wlearn[liblinear]    # Logistic regression, linear SVM
pip install wlearn[libsvm]       # Kernel SVM
pip install wlearn[nanoflann]    # KNN
pip install wlearn[ebm-fit]      # EBM training (interpret)
pip install wlearn[lightgbm]     # LightGBM
pip install wlearn[stochtree]    # BART training (stochtree)
pip install wlearn[tsetlin-fit]  # Tsetlin machine training (tmu)
pip install wlearn[nn]           # Neural tabular models (polygrad)
pip install wlearn[preprocess]   # Tranfi-backed fitted preprocessing
pip install 'wlearn[sym]'         # Symbolic models and optional Polygrad scoring
pip install 'wlearn[uncertainty]' # Calibration and conformal prediction
pip install wlearn[bo]           # Bayesian AutoML strategy (wlearn-bo)
pip install 'wlearn[dimred]'     # PCA, t-SNE, UMAP, TriMap (wlearn-dimred)
pip install wlearn[all]          # All supported backends, including original C11 packages
```

Random forests are now distributed separately: install `wlearn-rf` and import
`RFModel` from `wlearn_rf`. The former `wlearn.rf` module is no longer included
in the core `wlearn` distribution.

The `rf`, `gam`, `cluster`, `basis` and `bo` extras install the original C11
Python packages and are included in `all`; dimensionality reduction is a separate
opt-in extra (`pip install 'wlearn[dimred]'`). These packages are tested on Linux;
Windows and macOS are not tested, and the `bo` native package does not support Windows.

## Quick start

```python
import numpy as np
import wlearn
from wlearn.xgboost import XGBModel

# Train
X = np.array([[-4.0], [-3.0], [-2.0], [-1.0], [1.0], [2.0], [3.0], [4.0]])
y = np.array([0, 0, 0, 0, 1, 1, 1, 1], dtype=np.int32)
X_test = np.array([[-2.5], [2.5]])

model = XGBModel.create({
    'objective': 'binary:logistic',
    'max_depth': 2,
    'eta': 1.0,
    'numRound': 8,
    'seed': 0,
    'nthread': 1,
})
model.fit(X, y)
print(model.predict(X_test).tolist())  # [0, 1]

# Save to .wlrn file (loadable from JS @wlearn/xgboost too)
model.save('model.wlrn')

# Load from bundle
restored = wlearn.load('model.wlrn')
print(restored.predict(X_test).tolist())  # [0, 1]
```

Importing `wlearn.xgboost` registers its WLRN loaders. In a fresh process, import
the relevant model module before calling generic `wlearn.load()`.

## If you know scikit-learn

The common lifecycle is familiar, but wlearn is artifact-first rather than a
drop-in sklearn clone:

- construct with `Model.create(params)`, then call `fit`, `predict`, and `score`;
- pass NumPy-compatible dense arrays rather than relying on pandas metadata;
- use `save()`/`load()` with portable WLRN bundles instead of pickle or joblib;
- pass model parameters as a mapping; wrappers retain some backend-native names;
- inspect `capabilities` before using optional methods such as `predict_proba`;
- reshape probability output to `(n_rows, len(model.classes))` because the
  cross-language representation is a flat row-major array.

For fitted tabular transformations, `wlearn.preprocess.Preprocessor` provides
sklearn-style `fit`, `transform`, and `fit_transform` over Tranfi prepared plans.

## API

### Model wrappers

Training-capable wrappers use native Python backends and produce `.wlrn` bundles
compatible with the corresponding JS `@wlearn/*` package. XLearn is currently a
load/inference adapter for bundles trained by `@wlearn/xlearn` in JavaScript.

| Module | Class | Backend | Tasks |
|--------|-------|---------|-------|
| `wlearn.xgboost` | `XGBModel` | xgboost | classification, regression |
| `wlearn.liblinear` | `LinearModel` | liblinear-official | classification, regression |
| `wlearn.libsvm` | `SVMModel` | libsvm-official | classification, regression |
| `wlearn.nanoflann` | `KNNModel` | pynanoflann | classification, regression |
| `wlearn.lightgbm` | `LGBModel` | lightgbm | classification, regression |
| `wlearn.ebm` | `EBMModel` | numpy (inference), interpret (fit) | classification, regression |
| `wlearn.xlearn` | `XLearnModel` | numpy load/inference for JS-trained xLearn bundles; no Python `fit()` | classification, regression |
| `wlearn.stochtree` | `BARTModel` | stochtree | classification, regression |
| `wlearn.tsetlin` | `TsetlinModel` | tmu | classification, regression |
| `wlearn.nn` | `MLPClassifier`, `MLPRegressor`, `TabMClassifier`, `TabMRegressor`, `NAMClassifier`, `NAMRegressor` | polygrad (ctypes) | classification, regression |

Training-capable model wrappers share the API below. Optional methods depend on
the fitted model's `capabilities`:

```text
model = Model.create(params)     # create unfitted
model.fit(X, y)                  # train
model.predict(X)                 # predict labels
model.predict_proba(X)           # optional; flat rows * n_classes probabilities
model.score(X, y)                # accuracy (clf) or R^2 (reg)
bundle = model.save()            # return .wlrn bytes
model.save('model.wlrn')         # write .wlrn file and return bytes
```

Models that wrap native handles also expose `dispose()` for deterministic cleanup in long-running processes; ordinary scripts can usually rely on Python object cleanup.

`XLearnModel` is the exception: import `wlearn.xlearn` to register its loaders,
then use `wlearn.load(...)` on a JavaScript-trained bundle before calling
`predict()`, `predict_proba()`, `score()`, or `save()`. Its `create()` method only
creates an unfitted placeholder and it does not implement `fit()`.

### Bundle format

The `.wlrn` bundle is a binary container (header + JSON manifest + TOC + blobs) designed for cross-language compatibility. Bundles produced in Python can be loaded in JS and vice versa.

```python
from wlearn import encode_bundle, decode_bundle, validate_bundle

# Store custom bytes without claiming they are a trained model.
data = encode_bundle(
    manifest={'typeId': 'myorg.payload@1', 'params': {}},
    artifacts=[{'id': 'payload', 'mediaType': 'application/octet-stream',
                'data': bytes([1, 2, 3])}],
)
manifest, toc, blobs = decode_bundle(data)
validate_bundle(data)

# For estimator persistence and registry loading, use model.save() and load(),
# as in the quick start. A model's native bytes and metadata must agree.
```

Canonical writers always include `bundleVersion`, `requires`, `params`, and an
`artifacts` declaration exactly matching the TOC. Fixed TOC and artifact
declaration records have no extension fields, and nested WLRN artifacts must also
be canonical. Call
`validate_bundle(data, allow_legacy_manifest=False)` to prove conformance.

The default decoder is intentionally a compatibility-ingestion path for older v1
artifacts. It permits missing `requires`, `params`, `artifacts`, and TOC
`mediaType`, non-canonical TOC ordering, and extensions on fixed records. Bounds,
portable JSON, exact blob coverage, hashes, recursive validation, and all present
fields remain enforced. An old bundle without `requires` cannot provide complete
nested-loader preflight. Writers never produce legacy bundles. Compatibility reads
are supported through this major release; removal requires a major release,
migration tooling, and advance notice.

### Registry

Model wrappers register their loaders automatically on import. The `load()` function reads the bundle's `typeId` and dispatches to the registered loader.

```text
from wlearn import register, load

# Custom loader
register('myorg.custom@1', lambda manifest, toc, blobs: MyModel(blobs[0]))
model = load(bundle_bytes)
```

### Pipeline

```python
from wlearn import Pipeline, load
from wlearn.xgboost import XGBModel
from wlearn.scalers import StandardScaler

scaler = StandardScaler()
model = XGBModel.create({'objective': 'binary:logistic'})
pipe = Pipeline([('scaler', scaler), ('clf', model)])

pipe.fit(X_train, y_train)
pipe.predict(X_test)
pipe.score(X_test, y_test)

# Save/load preserves the full pipeline
pipe.save('pipeline.wlrn')
restored = load('pipeline.wlrn')
```

### Task, prediction, measures, resampling, archive

Python exposes the same structured primitives as JS so agents and apps can use one mental model:

```python
from wlearn import (
    create_task, create_prediction, create_resampling_plan,
    evaluate_metric_set, Archive,
)

task = create_task(id='toy', X=X_train, y=y_train)
plan = create_resampling_plan(strategy='stratified_kfold', n=len(y_train), y=y_train, k=5)
pred = create_prediction(truth=y_test, response=response, proba=proba, classes=classes)
scores = evaluate_metric_set(['accuracy', 'log_loss', 'roc_auc_ovr'], pred)

archive = Archive(task_id=task.id, measures=['accuracy'], primary_measure='accuracy')
archive.add({'trial_id': 'xgb-0', 'candidate_id': 'xgb-depth6', 'scores': scores, 'status': 'ok'})
archive.leaderboard()
```

Measures support sample weights, multiclass AUC (`roc_auc_ovr`, `roc_auc_ovo`), and explicit undefined-metric policies. Resampling plans cover holdout/k-fold/group/time-series plus sliding row, index, and period windows for rolling validation.

`wlearn.automl.auto_fit()` returns `archive` as well as `leaderboard`, so failed candidates, params, fold scores, and timings are queryable after a run.

`wlearn.cv` owns splits, scalar metrics and CV execution; `wlearn.rng` owns the
deterministic generator, and `wlearn.measure` defines each metric's response and
direction. AutoML and ensembles use the same utilities. Explicit resampling plans
are accepted for evaluation; OOF predictions and stacking require every row to be
in a test fold exactly once. Temporal, index and period split generators are
experimental.

### AutoML preprocessing

Install the optional prepared-transform backend with `pip install wlearn[preprocess]`.
`auto_fit()` then fits a fresh Tranfi-backed preprocessor inside every CV and OOF
training fold, and returns the refitted winner as a `Pipeline`:

```python
from wlearn.automl import auto_fit
from wlearn.liblinear import LinearModel

result = auto_fit(
    [{'name': 'linear', 'classId': 'wlearn.liblinear.classifier@1',
      'cls': LinearModel}],
    X, y,
    preprocess={'impute': 'auto', 'encode': 'onehot', 'scale': 'standard'},
)

result['bestCandidate']       # structured model + resolved preprocessing identity
result['model'].save()        # preserves preprocessing inside pipelines/ensemble
```

`preprocess` also accepts `True` for defaults or a list of fixed templates with
`templateId`, `typeId='wlearn.preprocess.tabular@1'`, and `params`. All search
strategies cross candidates with every template. Candidate IDs are opaque
`wlc1_<sha256>` values; use the structured candidate instead of parsing the ID.
Raw-feature stacking passthrough is rejected when preprocessing is active because
fold-fitted feature spaces may differ.

For the simple path, ignore these primitives: call `Model.create()`, `fit()`, `predict()`, `score()`, and `save()`, or use `auto_fit()` and read `result['model']`, `result['leaderboard']`, and `result['bestScore']`. The archive is there when you need provenance or agent-readable run history.

## License

Apache-2.0
