Metadata-Version: 2.4
Name: windowsml
Version: 2.7.2021a0
Summary: Python bindings for Windows AI Machine Learning catalogs and Runtime
License: Copyright (C) Microsoft Corporation. All rights reserved.
Classifier: Development Status :: 3 - Alpha
Classifier: Operating System :: Microsoft :: Windows
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.9
Description-Content-Type: text/markdown
Requires-Dist: comtypes>=1.4.0; platform_system == "Windows"
Provides-Extra: with-ort
Requires-Dist: onnxruntime-windowsml==1.30.0.202609102321; extra == "with-ort"
Provides-Extra: llama
Requires-Dist: windowsml-llama==2.7.2021a0; extra == "llama"
Dynamic: provides-extra
Dynamic: requires-dist

# windowsml

Python bindings for the Windows AI Machine Learning API.

This package provides a ctypes-based wrapper around the WinML flat C API,
enabling Python applications to discover and manage execution providers and
model catalogs for ONNX Runtime on Windows.

## Installation

```bash
pip install windowsml
```

### ONNX Runtime dependency

To automatically install a matching `onnxruntime-windowsml` package, use the
`with-ort` extra:

```bash
pip install windowsml[with-ort]
```

If you choose **not** to use `[with-ort]`, install a compatible version of
`onnxruntime-windowsml` separately for Runtime preview ONNX workloads. Catalog
APIs do not require this optional dependency.

The Runtime preview projection does not import ONNX Runtime or load its native
sidecar during module import. At the first model load, `[with-ort]` loads the
exact native sidecar when Python ONNX Runtime is not already present. To share a
manually registered execution provider, import `onnxruntime` and complete
registration before calling any Runtime model-load entry point. The projection
applies this guard to every artifact format because the native load call detects
the format. Runtime then uses the same process-loaded ONNX Runtime instance and
registration state for ORT-backed stages.

```python
import onnxruntime as ort
from windowsml.runtime import Runtime

provider_path = r"C:\path\to\provider.dll"
ort.register_execution_provider_library("MyExecutionProvider", provider_path)

with Runtime() as runtime, runtime.load_model("model.onnx") as model:
    ...
```

The loader choice is process-wide. RuntimeCore finalizes its ORT API selection
when an ORT-backed operation first needs it.

Raw artifacts with external resources can supply a generic mapping of resource
names to contiguous buffer-protocol values. Runtime retains each zero-copy
buffer export through the model and its dependent stages:

```python
import mmap

from windowsml.runtime import Runtime

with open("weights.bin", "rb") as weights_file:
    weights = mmap.mmap(weights_file.fileno(), 0, access=mmap.ACCESS_READ)
    with Runtime() as runtime, runtime.load_model(
        "model.onnx",
        resources={"model.weight": memoryview(weights)},
    ) as model:
        ...
    weights.close()
```

## Quick Start

### EP Catalog — discover execution providers

```python
from windowsml import EpCatalog

with EpCatalog() as catalog:
    for ep in catalog.find_all_providers():
        print(f"{ep.name} v{ep.version} - {ep.ready_state.name}")

        # Ensure the provider is ready (downloads/installs if needed)
        ep.ensure_ready()
```

### Model Catalog — discover and download models

```python
from windowsml import ModelCatalogSource, ModelCatalog

with ModelCatalogSource.create_from_uri("https://example.com/catalog.json") as source:
    with ModelCatalog([source]) as catalog:
        for model in catalog.find_all_models():
            print(f"{model.name} v{model.version} by {model.publisher}")
            print(f"  Size: {model.model_size_in_bytes} bytes")
            print(f"  EPs:  {model.execution_providers}")
```

### Runtime preview — target-bound tensors

Dev-profile wheels also expose `windowsml.runtime`. NumPy and Pillow are
optional ecosystem dependencies; install only what the application uses:

```bash
pip install numpy pillow
```

```python
import numpy as np
from windowsml.runtime import (
    ImageTensorOptions,
    Runtime,
    TensorAccessMode,
    TensorDataType,
)

with Runtime() as runtime, runtime.create_cpu_target() as cpu:
    dense = runtime.tensor_from_numpy(
        np.zeros((1, 3, 224, 224), dtype=np.float32),
        target=cpu,
    )

    storage = bytearray(np.arange(4, dtype=np.float32).tobytes())
    borrowed = runtime.tensor_from_buffer(
        storage,
        TensorDataType.FLOAT32,
        (2, 2),
        access=TensorAccessMode.READWRITE,
        target=cpu,
    )

    image = runtime.tensor_from_image(
        np.zeros((224, 224, 3), dtype=np.uint8),
        options=ImageTensorOptions(size=(224, 224)),
        target=cpu,
    )

    image.close()
    borrowed.close()
    dense.close()
```

| Ingress | Ownership and copy behavior |
|---|---|
| `tensor_from_numpy` | Runtime-owned tensor; source data is copied synchronously |
| `tensor_from_buffer(..., COPY)` | Runtime-owned packed copy |
| `tensor_from_buffer(..., READ)` | Borrowed contiguous view; tensor retains the Python buffer export |
| `tensor_from_buffer(..., READWRITE)` | Borrowed writable view; mutations are shared |
| image/audio/token helpers | Native conversion creates a target-bound Runtime tensor |

Close tensors before resizing or releasing borrowed storage. Context managers
and explicit `close()` calls release native references deterministically.

### Optional native Task component

Wheels that include the native Task component also expose `windowsml.tasks`.
Task IDLs for Text Generation, Chat Completion, and Automatic Speech
Recognition are included in the generated low-level projection. The idiomatic
facade covers all three families and executes entirely in the bundled native
`WinMLTasks.dll`.

```python
from windowsml.tasks import (
    Tasks,
    TextGenerationOptions,
    TextGenerationOutputKind,
)

with Tasks(runtime) as tasks:
    configuration = (
        tasks.create_text_generation_configuration()
        .configure_unified(
            pipeline=pipeline,
            token_input_stage=stage,
            token_input_index=0,
            output_stage=stage,
            output_index=0,
            output_kind=TextGenerationOutputKind.LOGITS,
            state_owner_stage=stage,
            token_target=cpu,
            tokenizer=tokenizer,
        )
    )
    with tasks.create_text_generation_task(
        tokenizer,
        default_configuration=configuration,
    ) as task:
        result = task.generate(
            "Write one short greeting.",
            TextGenerationOptions(max_new_tokens=32),
        )
        print(result.text)
```

Use `task.stream(...)` for token/text deltas, or disclose
`task.default_session` and call `generate_tokens` / `continue_tokens` for raw
token control. Sessions, streams, configurations, and the owning Runtime remain
available through exact native-backed wrapper properties.

Task metadata and explicit session/configuration creation have no Task-creation-thread
affinity. Supplied and retained collaborators retain their own contracts, and each
session belongs to its creation thread. Convenience APIs share one lazy
`default_session`; callers synchronize default-configuration changes, lazy cache
access, and shared convenience use. `close()` must not overlap any operation on
the same Task wrapper, including through reentrant callbacks. No concurrent
Close or mutation safety is provided.

With a live cached session, call Task `close()` on that session's creation thread;
without one, Task `close()` has no creation-thread restriction. Call session,
stream, and live-input `close()` on their owning creation thread. Final native
`Release` is free-threaded, and immutable Text and Chat results have no
creation-thread affinity. Stream `cancel()` may be called from another thread
and signals a separate native cancellation source; it does not marshal
`close()` to the creation thread.

ONNX stages expose `stage.symbolic_dimensions` before pipeline construction.
For example, a Whisper encoder with symbolic feature axes can be specialized
without changing the model:

```python
encoder_stage.symbolic_dimensions["feature_size"] = 80
encoder_stage.symbolic_dimensions["encoder_sequence_length"] = 3000
encoder_pipeline = encoder_builder.build()
```

ASR runtime configurations are created by their owning Task. Assign the
validated configuration back to that same Task before using the convenience
methods:

```python
configuration = task.create_configuration()
configuration.configure_whisper(
    encoder_pipeline=encoder_pipeline,
    encoder_stage=encoder_stage,
    decoder_pipeline=decoder_pipeline,
    decoder_stage=decoder_stage,
    tensor_target=cpu,
)
configuration.validate()
task.default_configuration = configuration
result = task.transcribe(waveform, metadata)
```

`configure_whisper` uses one atomic typed native binding call. A failed
replacement leaves the previous bindings unchanged; successful session creation
consumes the configuration. `session.configuration.whisper_bindings` returns an
immutable record of independently owned Pipeline, Stage and ExecutionTarget
wrappers over the actual native collaborators, including optional preprocessing
and `force_encoder_output_copy`.

Use `session.supports_whisper_log_mel` and `session.supports_live_input` for
non-mutating capability QI. Current Whisper supports log-mel input but not live
input. `session.transcribe_whisper_log_mel(features)` requires the exact
FLOAT32 `[1,80,3000]` Whisper preprocessing semantics documented in the ASR API
reference; shape alone does not establish those semantics. No feature GUID,
lengths tensor, implementation selector or runtime-object bag is required.

## API Reference

### Enums

- **`EpReadyState`** — `Ready`, `NotReady`, `NotPresent`
- **`EpCertification`** — `Unknown`, `Certified`, `Uncertified`
- **`CatalogModelStatus`** — `Ready`, `NotReady`
- **`CatalogModelInstanceStatus`** — `Available`, `InProgress`, `Unavailable`

### Classes

#### `EpCatalog()`

Context manager for the EP catalog.

| Method | Description |
|---|---|
| `find_all_providers()` | Returns `list[ExecutionProvider]` |
| `close()` | Release the catalog |

#### `ExecutionProvider`

Handle to a discovered EP (owned by the catalog).

| Property / Method | Type |
|---|---|
| `name` | `str` |
| `version` | `str` |
| `package_family_name` | `str` |
| `library_path` | `str` |
| `package_root_path` | `str` |
| `ready_state` | `EpReadyState` |
| `certification` | `EpCertification` |
| `ensure_ready()` | Blocks until ready |
| `ensure_ready_async(on_complete, on_progress)` | Returns `AsyncOperation` |

#### `ModelCatalogSource`

Represents a catalog source endpoint. Create via static factory methods.

| Method / Property | Description |
|---|---|
| `create_from_uri(uri)` | Create from a URI or local path (blocking) |
| `create_from_uri_async(uri, ...)` | Create from a URI (async) |
| `create_from_uri_with_headers(uri, headers)` | Create with custom HTTP headers (blocking) |
| `create_from_uri_with_headers_async(uri, headers, ...)` | Create with custom HTTP headers (async) |
| `id` | `str` — source identifier |
| `uri` | `str` — source URI |
| `close()` | Release the source |

#### `ModelCatalog(sources)`

Context manager for the model catalog. Created from a list of `ModelCatalogSource`.

| Method / Property | Description |
|---|---|
| `execution_providers` | `list[str]` — configured EPs |
| `set_execution_providers(providers)` | Set EP filter for queries |
| `get_available_model(id_or_name)` | Returns `CatalogModelInfo` |
| `get_available_models()` | Returns `list[CatalogModelInfo]` |
| `find_model(id_or_name)` | Find a model (blocking) |
| `find_model_async(id_or_name, ...)` | Find a model (async) |
| `find_all_models()` | Find all models (blocking) |
| `find_all_models_async(...)` | Find all models (async) |
| `close()` | Release the catalog |

#### `CatalogModelInfo`

Metadata about a model in the catalog.

| Property / Method | Type |
|---|---|
| `id` | `str` |
| `name` | `str` |
| `publisher` | `str` |
| `source_id` | `str` |
| `version` | `str` |
| `license` | `str` |
| `license_uri` | `str` |
| `license_text` | `str` |
| `uri` | `str` |
| `model_size_in_bytes` | `int` |
| `execution_providers` | `list[str]` |
| `status` | `CatalogModelStatus` |
| `get_instance_async(...)` | Returns `AsyncOperation` → `CatalogModelInstance` |
| `get_instance_with_headers_async(headers, ...)` | Returns `AsyncOperation` → `CatalogModelInstance` |
| `close()` | Release the model info |

#### `CatalogModelInstance`

A downloaded or local instance of a catalog model.

| Property / Method | Type |
|---|---|
| `model_paths` | `list[str]` — local file paths |
| `model_info` | `CatalogModelInfo` |
| `close()` | Release the instance |

#### `AsyncOperation`

Returned by async methods. Supports context-manager usage.

| Method | Description |
|---|---|
| `wait()` | Block until complete |
| `get_status(wait=False)` | Poll or wait for status |
| `get_result()` | Extract the typed result (if applicable) |
| `cancel()` | Request cancellation |
| `close()` | Release async resources |
