Metadata-Version: 2.4
Name: vapoursynth-mlrt-ort
Version: 17.1
Summary: ONNX Runtime-based CPU/GPU inference plugin for VapourSynth
Author: AmusementClub
Maintainer-Email: Vardë <ichunjo.le.terrible@gmail.com>
License-Expression: GPL-3.0-or-later
License-File: vs-mlrt/LICENSE
Project-URL: Source Code, https://github.com/Ichunjo/vs-mlrt
Project-URL: Bug Tracker, https://github.com/AmusementClub/vs-mlrt/issues
Project-URL: Repository, https://github.com/Jaded-Encoding-Thaumaturgy/vs-wheels
Requires-Python: >=3.12
Requires-Dist: vapoursynth>=75
Provides-Extra: cuda
Requires-Dist: vapoursynth-mlrt-ort-cuda==17.1; extra == "cuda"
Description-Content-Type: text/markdown

# VapourSynth-MLRT-ORT

This package contains the ONNX Runtime backend implementation of the [vs-mlrt](https://github.com/Ichunjo/vs-mlrt) plugin.

## Installation

To install the standard CPU/DirectML/CoreML package:

```bash
pip install vapoursynth-mlrt-ort
```

To install the CUDA-enabled package:

```bash
pip install "vapoursynth-mlrt-ort[cuda]"
```

## Building from source

### Requirements

- **C++ Compiler**: C++20 compatible (e.g. MSVC 2019+, GCC, Clang)
- **Build Tools**: [uv](https://docs.astral.sh/uv/), CMake, Ninja
- **Dependencies** (must be built/installed before the plugin):
  - `onnxruntime` (ONNX Runtime SDK — built from source in CI)
  - `ONNX` (built from source, version matched to ONNX Runtime's pinned submodule)
- **Backend-specific Dependencies**:
  - **DirectML** (Windows only): DirectML SDK. Set `DML_DIR` to the SDK directory.
  - **CUDA** (Windows/Linux): `CUDAToolkit` and `cuDNN`. Ensure `CUDA_PATH` and `CUDNN_PATH`/`CUDNN_HOME` are set.
  - **CoreML** (macOS only): Enabled automatically, no extra setup needed.

### Compilation

1. **Initialize the submodule:**

   ```bash
   git submodule update --init --recursive vsmlrt/ort/vs-mlrt
   ```

2. **Build ONNX Runtime from source** (or install a pre-built SDK), then **build ONNX from source** using the version pinned by ONNX Runtime (`onnxruntime/cmake/external/onnx/`). Both must be discoverable by CMake (via `CMAKE_PREFIX_PATH` or installed system-wide).

3. **Set environment variables** for backend-specific SDKs:

   ```powershell
   # Windows DirectML
   $env:DML_DIR = "C:\Path\To\DirectML"
   ```

4. **Build the wheel:**

   ```bash
   uv build --package vapoursynth-mlrt-ort
   ```

   To disable bundling CUDA provider libraries into the main wheel (CI does this, shipping them in a separate `vapoursynth-mlrt-ort-cuda` package instead):

   ```bash
   uv build --package vapoursynth-mlrt-ort -C "cmake.define.INSTALL_CUDA=OFF"
   ```
---

Detailed parameter information from the parent project follows.

---
# VapourSynth ONNX Runtime

The vs-onnxruntime plugin provides optimized CPU & CUDA runtime for some popular AI filters.

## Usage

Prototype: `core.ort.Model(clip[] clips, string network_path[, int[] overlap = None, int[] tilesize = None, string provider = "", int device_id = 0, int verbosity = 2, bint cudnn_benchmark = False, bint fp16 = False, bint path_is_serialization = False, bint use_cuda_graph = False])`

Arguments:

- `clip[] clips`: the input clips, only 32-bit floating point RGB or GRAY clips are supported. For model specific input requirements, please consult our [wiki](https://github.com/AmusementClub/vs-mlrt/wiki).
- `string network_path`: the path to the network in ONNX format.
- `int[] overlap`: some networks (e.g. [CNN](https://en.wikipedia.org/wiki/Convolutional_neural_network)) support arbitrary input shape where other networks might only support fixed input shape and the input clip must be processed in tiles. The `overlap` argument specifies the overlapping (horizontal and vertical, or both, in pixels) between adjacent tiles to minimize boundary issues. Please refer to network specific docs on the recommended overlapping size.
- `int[] tilesize`: Even for CNN where arbitrary input sizes could be supported, sometimes the network does not work well for the entire range of input dimensions, and you have to limit the size of each tile. This parameter specify the tile size (horizontal and vertical, or both, including the overlapping). Please refer to network specific docs on the recommended tile size.
- `string provider`: Specifies the device to run the inference on.
  - `"CPU"` or `""`: pure CPU backend
  - `"CUDA"`: CUDA GPU backend, requires Nvidia Maxwell+ GPUs.
  - `"DML"`: DirectML backend
  - `"COREML"`: CoreML backend
- `int device_id`: select the GPU device for the CUDA backend.'
- `int verbosity`: specify the verbosity of logging, the default is warning.
  - 0: fatal error only, `ORT_LOGGING_LEVEL_FATAL`
  - 1: also errors, `ORT_LOGGING_LEVEL_ERROR`
  - 2: also warnings, `ORT_LOGGING_LEVEL_WARNING`
  - 3: also info, `ORT_LOGGING_LEVEL_INFO`
  - 4: everything, `ORT_LOGGING_LEVEL_VERBOSE`
- `bint cudnn_benchmark`: whether to let cuDNN use benchmarking to search for the best convolution kernel to use. Default False. It might incur some startup latency.
- `bint fp16`: whether to quantize model to fp16 for faster and memory efficient computation.
- `bint path_is_serialization`: whether the `network_path` argument specifies an onnx serialization of type `bytes`.
- `bint use_cuda_graph`: whether to use CUDA Graphs to improve performance and reduce CPU overhead in CUDA backend. Not all models are supported.
- `int ml_program`: select CoreML provider.
  - 0: NeuralNetwork
  - 1: MLProgram

When `overlap` and `tilesize` are not specified, the filter will internally try to resize the network to fit the input clips. This might not always work (for example, the network might require the width to be divisible by 8), and the filter will error out in this case.

The general rule is to either:

1. left out `overlap`, `tilesize` at all and just process the input frame in one tile, or
2. set all three so that the frame is processed in `tilesize[0]` x `tilesize[1]` tiles, and adjacent tiles will have an overlap of `overlap[0]` x `overlap[1]` pixels on each direction. The overlapped region will be throw out so that only internal output pixels are used.
