Metadata-Version: 2.2
Name: gaia-ml
Version: 0.13.21
Summary: GAIA-ML ONNX-to-AIE project generator
Author: I. Xiotidis
Maintainer: NGT WP2.1 Group
License: MIT
Project-URL: Homepage, https://gitlab.cern.ch/atlas-nextgen-wp21/aie/gaia-ml
Project-URL: Documentation, https://gitlab.cern.ch/atlas-nextgen-wp21/aie/gaia-ml/-/blob/main/README.md
Project-URL: Repository, https://gitlab.cern.ch/atlas-nextgen-wp21/aie/gaia-ml
Project-URL: Issues, https://gitlab.cern.ch/atlas-nextgen-wp21/aie/gaia-ml/-/issues
Keywords: aie,amd,cmake,machine-learning,onnx,vitis
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: POSIX :: Linux
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Code Generators
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: jinja2>=3.1
Requires-Dist: numpy>=1.22
Requires-Dist: onnx>=1.17

# GAIA-ML

GAIA-ML is an experimental compiler that turns ONNX models into optimized C++
kernels and graphs for AMD AI Engine v1, targeting the Versal Premium VP2802-2M.

## Features

- Dense/MLP networks, supported CNN patterns and bounded tree ensembles.
- Float32, int32, int16 and int8 neural-network kernels.
- Custom model-local ONNX functions using supported float32 arithmetic and tensor operations.
- Operation fusion, weight packing and implementation selection across AIE tiles.
- Scheduling for first-output latency, steady-state output interval or event bursts.
- New opt-in surrogate selection for supported FP32 CNN row kernels and tensor
  reductions, plus vectorized reductions of up to 4096 elements.

GAIA generates an ADF graph, C++ kernels, weights, test vectors and build files.
**Generating a project does not require Vitis.** Building and simulating the
emitted AIE code requires the AMD toolchain separately.

## Install

From a checkout, with Python 3.9 or newer:

```bash
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -e .
gaia_ml version
```

## Compile a model

Inspect the model and target constraints, then generate a project:

```bash
gaia_ml inspect model.onnx --stage constraints

gaia_ml compile model.onnx \
  --output gaia_project \
  --schedule-objective latency \
  --data-file input.csv \
  --simulation-events 10
```

Each CSV row is one flattened input event; provide at least as many rows as
`--simulation-events`. For dynamic inputs, supply a static `--input-shape`,
for example `--input-shape 1,3,18,18`. Use a new output directory or `--force`
to regenerate an existing project.

Useful compilation options:

| Option | Purpose |
|---|---|
| `--aggressive-latency-search` | Explore more supported implementation and tile choices. |
| `--latency-metric first-output` | Prioritize the first completed inference. |
| `--latency-metric interval` | Prioritize spacing between successive outputs. |
| `--latency-metric balanced` | Balance first-output latency and output interval. |
| `--latency-metric burst --burst-events 100` | Prioritize completion of a 100-event burst. |
| `--quantize int8` | Generate integer kernels; also accepts `int16` and `int32`. |
| `--max-tiles 16` | Limit the available compute tiles. |
| `--cnn-row-kernel surrogate --cnn-row-surrogate selection.json` | Use a supplied trained selection bundle for supported FP32 7×7, stride-3 CNN row kernels. |

Surrogate selection requires a trained bundle for the supported compute family.
Use `gaia_ml compile --help` for all options, including tensor microkernel selection.
The generated `manifest.json` records resource use, the selected schedule and
estimated latency.

## Build and validate

After loading your AMD Vitis environment:

```bash
cd gaia_project
make aie
make aiesim
cd ..
gaia_ml verify gaia_project
```

Latency estimates guide optimization; validate numerical accuracy and measured
timing for your model with the external AMD toolchain.

Developed by the **NGT WP2.1 Group**. **Core developer: Ioannis Xiotidis.**
Distributed under the [MIT license](LICENSE).
