Metadata-Version: 2.2
Name: gaia-ml
Version: 0.11.0
Summary: GAIA-ML ONNX-to-AIE project generator
Author: I. Xiotidis
Maintainer: NGT WP2.1 Group
License: MIT
Project-URL: Homepage, https://gitlab.cern.ch/atlas-nextgen-wp21/aie/gaia-ml
Project-URL: Documentation, https://gitlab.cern.ch/atlas-nextgen-wp21/aie/gaia-ml/-/blob/main/README.md
Project-URL: Repository, https://gitlab.cern.ch/atlas-nextgen-wp21/aie/gaia-ml
Project-URL: Issues, https://gitlab.cern.ch/atlas-nextgen-wp21/aie/gaia-ml/-/issues
Keywords: aie,amd,cmake,machine-learning,onnx,vitis
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: POSIX :: Linux
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Code Generators
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: jinja2>=3.1
Requires-Dist: numpy>=1.22
Requires-Dist: onnx>=1.14

# GAIA-ML

GAIA-ML is an experimental ONNX compiler that generates AMD AI Engine v1
graphs, vector kernels, packed weights, and simulation projects for low-latency
inference. Version 0.11.0 focuses on compact multilayer perceptrons, small
convolutional networks, and bounded additive tree ensembles.

GAIA-ML is developed by the NGT WP2.1 Group and is distributed under the MIT
license.

## What the compiler does

The compiler processes a model through these representations:

```text
ONNX model
  -> imported operator graph
  -> canonical Math IR
  -> mathematical rewrites
  -> Loop IR
  -> Schedule IR and implementation search
  -> AIE IR and target checks
  -> AMD ADF graph, AIE C++ kernels, packed weights, and test vectors
```

The implemented pipeline can:

- inspect ONNX graphs and inferred tensor shapes;
- simplify local algebraic identities;
- canonicalize `MatMul + Add` and `Gemm` into Dense operations;
- canonicalize ONNX convolution and layout helper patterns into Conv2D;
- fold BatchNorm constants into convolution weights and bias;
- fuse supported ReLU and sigmoid activations;
- lower equal-shape two-input Add as a native vector residual merge and
  materialize required fan-out with AIE-v1 binary broadcast trees;
- recognize depthwise convolution with a depth multiplier and fuse supported
  one-pixel float32 depthwise-Conv-to-Dense classifier tails;
- propagate physical tensor layouts through convolution chains;
- rank scalar, vector-dot, streaming, native-MMUL, sliding, and line-buffer
  microkernels;
- pad streams and reorganize weights for legal AIE vector and MMUL shapes;
- split work across tiles by output, spatial, batch, reduction, or channel
  dimensions when the selected backend supports it;
- validate AIE v1 tile memory, stack guidance, stream-port limits, and model
  input transfer bounds;
- lower bounded ONNX-ML tree ensembles through a QuickScorer bit-vector
  backend;
- emit float32, int32, int16, or int8 neural-network kernels; and
- generate Vitis x86 simulation and AIE simulation projects.

## Target

GAIA-ML 0.11.0 uses the Versal Premium VP2802-2M as its concrete AIE v1 target
profile, with:

- 1.25 GHz AIE cores;
- 59 columns by eight rows (472 compute tiles);
- 32 KiB memory per compute tile;
- 16 KiB program memory per compute tile;
- no memory tiles;
- at most two input and two output streams per kernel; and
- a fail-closed 36-net directional inter-column routing ceiling, applied to
  GAIA's conservative deterministic route-pressure envelope; and
- 32-bit PLIO at either 312.5 MHz or 500 MHz.

The target specification records the heterogeneous, overlapping
128/256/512/1024-bit vector aliases and 384/768-bit accumulator aliases. The
architecture analysis allocates against the 2048-bit overlapping physical
vector-alias capacity and 3072-bit accumulator-alias capacity. Alias counts
are never treated as additive homogeneous registers.

The target also records eight 4-KiB local-memory banks, vector load/store issue
widths, and the 384-bit cascade path. These values drive a structural issue,
bank, transport, and reduction analysis. The analysis is an optimization and
spill-risk signal; Vitis assembly and measured execution remain authoritative.
The modeled register aliases, load/store paths, memory banks, and cascade width
follow AMD's [AIE register-file description](https://docs.amd.com/r/en-US/am009-versal-ai-engine/Register-Files),
[AIE interfaces](https://docs.amd.com/r/en-US/am009-versal-ai-engine/AI-Engine-Interfaces),
and [AIE memory-module description](https://docs.amd.com/r/en-US/am009-versal-ai-engine/AI-Engine-Memory-Module).

Inspect registered profiles with:

```bash
gaia_ml show targets
gaia_ml show targets --json
```

Select the target explicitly with `--target vp2802-2m`. PLIO frequency,
interface type, tile count, and memory requests are validated against that
profile before scheduling.

## Architecture-aware optimization

GAIA lowers the emitted program to structural SIMD and AIE-v1 execution
analyses. For every kernel it reports overlapping vector/accumulator alias
pressure, load/store/compute issue lower bounds, recommended input/weight/
output/scratch memory-bank roles, aligned vector-stream eligibility, and
cascade-reduction opportunities. Graph edges additionally report producer and
consumer service cycles, sustained-rate balance, required consumer replication,
and an analytical FIFO requirement. A deeper FIFO is not presented as a fix
for a sustained rate mismatch.

The fastest supported float spatial-convolution paths consume complete input
row chunks with `readincr_v` into aligned circular storage and retain scalar
reads only for an unavoidable row tail. Sliding/gather row strides cover both
the complete physical input row and the widest operand span, preventing wider
model shapes from overwriting an adjacent circular-buffer row.

These metrics participate as late architecture-aware tie breakers after the
established latency/resource evidence. They do not displace a measured
known-good kernel merely because a structural model predicts an improvement.

## Search-space evidence

Schedule and constraint inspection now includes a versioned `search_space`
record. For each operation it lists every unique candidate constructed by the
registered backend, its target-legality result, whether it entered graph
search, and counts for duplicate, legality, local-dominance, local-cap, and
beam pruning. It also records the declared and retained Cartesian-product
sizes and the number of plans checked with exact dry-lowered resources.

When that declared product fits within the planner width, GAIA evaluates every
supported target-legal plan and marks the search exhaustive. The scope is
important: this means exhaustive over implementations represented by GAIA's
current capability registry, not over every program that could be handwritten
for the AIE. Larger searches state their exact completeness boundary and may
not be used to claim infeasibility.

## Latency evidence and planning bounds

GAIA reports three separate latency concepts for steady-state II and first
output: a physical lower bound where one is established, the emitted-program point
prediction, and an evidence-dependent guardbanded planning estimate. The
planning factors are explicit and versioned: 1.10 for matching-geometry
measurements, 1.25 for profile calibration, 1.60 for geometry scaling, 2.00
for analytical models, and 2.50 when evidence is unknown.

These factors are currently a provisional engineering policy, not validated
statistical coverage. Consequently the planning upper estimate is not a WCET
proof, and an observed AIESim maximum is not treated as an all-input timing
guarantee. Full-event PLIO traffic supplies the current provable steady-state
II floor. Ideal arithmetic issue time is reported separately and is not reused
as a first-output proof for streaming kernels. A hard deadline below an
available physical lower bound is provably impossible; a deadline inside the
planning envelope is reported separately from exact-plan measurement
verification.

## Hard-deadline resource minimization

When `--required-ii-us` or `--required-first-output-us` is present, GAIA first
filters graph plans using the evidence-class planning envelope. Among plans
that satisfy that envelope and all target constraints, it minimizes the exact
dry-lowered resource vector in this order: compute tiles, emitted kernels,
maximum per-tile memory, and FIFO bytes. Planning-envelope latency, routing,
and evidence quality break remaining ties.

This policy differs deliberately from a soft preference. A plan whose point
estimate meets a hard deadline but whose planning envelope misses it is not
admitted. `--preferred-ii-us` and legacy `--target-latency-us` continue to rank
point predictions without converting them into hard failures.

The compiler reports whether the selected plan is resource-minimal in the
explored domain and, when Phase-3 search is exhaustive, in the complete
declared supported domain. That result is conditional on the provisional
planning envelope; it is not global optimality or formal timing verification.

## Adaptive hard-deadline search

GAIA first evaluates the normal registered candidate domain. If that domain
does not satisfy a hard planning envelope, the compiler expands bounded,
codegen-supported dimensions rather than immediately failing:

- Dense: output and batch tile extents across registered vector/weight layouts;
- Conv2D: output-channel and spatial-height splits across compatible
  microkernels, vector widths, and physical layouts; and
- BDT: forest worker counts, trees per worker, and registered QuickScorer or
  traversal implementations.

Expansion remains bounded by the VP2802-2M target and the user's tile, kernel,
memory, precision, interface, and explicit microkernel/tile choices. Every
round reports its trigger, dimensions, candidate counts, contract bounds, and
termination. If the initial domain is feasible, expansion is skipped. If no
expanded plan is feasible, GAIA reports supported-space saturation instead of
implying that arbitrary handwritten AIE code is impossible.

## Optional exact-plan verification

Vitis is not a dependency of ONNX import, scheduling, constraint solving, or
project generation. After an externally run AIESim has completed, verify its
existing artifacts with:

```bash
gaia_ml verify gaia_project
```

The verifier checks complete-event timestamps against any hard II and
first-output requirements and compares AIESim output values with the generated
golden CSVs. It writes `verification_report.json`, attaches the result to the
compilation certificate and manifest, and records a successful exact-plan
observation in `.gaia_ml/aie_v1_feedback.json`. Use `--no-record` to suppress
the database update, or `--database PATH` to select another local database.

A verified result applies only to the observed events and exact plan
fingerprint; it is not an all-input WCET or numerical proof. A timing or
numerical failure is recorded as a counterexample for that exact plan. Missing
timestamps or golden outputs produce `incomplete`, never a successful result.
Discover installed verifier adapters with `gaia_ml show verifiers`.

For controlled compiler-certification studies, an unresolved hard-deadline
candidate can be materialized explicitly with
`--emit-unadmitted-candidate`. The command still returns an error and the
project is not admitted. After AIESim and `gaia_ml verify --database DB`, replay
the identical compile command with `--calibration-db DB`; only an exact
emitted-plan fingerprint match can satisfy the hard deadline. Ordinary failed
compilations do not emit a project.

### Counterexample-guided refinement

A failed verification is stored separately from successful latency feedback:

```bash
gaia_ml verify gaia_project \
  --refinement-database .gaia_ml/aie_v1_refinements.json

gaia_ml compile model.onnx \
  --output refined_project \
  --required-ii-us 5 \
  --refinement-db .gaia_ml/aie_v1_refinements.json
```

The second compilation excludes matching measured counterexamples before
Pareto and resource-minimal selection. A rejection is scoped to the exact
GAIA version, model checksum, target specification, and complete Schedule-IR
plan identity. It is never transferred to another model, target, compiler
version, or merely similar microkernel. If alternatives exist, the planner
selects the highest-priority remaining supported plan; if all explored plans
are rejected it reports refinement exhaustion instead of re-emitting a known
counterexample.

GAIA does not invoke Vitis or rerun AIESim automatically. This keeps external
tool licensing and availability outside the compiler while supporting a
reproducible compile--simulate--verify--refine loop.

### Evidence-corpus evaluation

Phase 9 aggregates generated projects and their exact-plan verification
reports into a reproducible metrics and claim-gate report:

```bash
gaia_ml evaluate project_a project_b project_c \
  --deadline-us 10 \
  --minimum-verified 3 \
  --output gaia_evidence.json \
  --require-ready
```

The report contains compiler prediction error, planning-envelope coverage,
false admissions, conservative rejections, numerical comparison coverage,
and resource-minimality scope. Its SHA-256 identity excludes machine-local
paths. `--require-ready` returns a failure status unless every supplied record
has exact-plan verification, passes the study deadline and numerical check,
was compiled with an applicable hard II contract, uses one exact target
specification and compiler version, and has no false admissions.

The observed deadline gate is deliberately narrower than a hard-real-time
proof: it describes only the supplied AIESim events. Resource minimality is
reported separately for the explored domain and the complete declared
supported domain; neither is described as global optimality over arbitrary
handwritten AIE programs.

Generated projects default to part `xcvp2802-vsva5601-2MHP-e-S`. Override it at
build time when using another compatible AIE v1 part:

```bash
make aie AIE_PART=<part-name>
```

## Installation

Install the release from PyPI:

```bash
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install gaia-ml==0.10.18
gaia_ml version
```

For development from a checkout:

```bash
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e .
gaia_ml version
```

The Python package requires Python 3.9 or newer. Generating a project does not
require Vitis. Compiling or simulating generated AIE sources requires a local
AMD Vitis installation; the generated build flow has been exercised with Vitis
2024.2.

## Inspecting a model

Inspect the complete compiler pipeline:

```bash
gaia_ml inspect model.onnx --stage all
```

Inspect an individual representation or pass:

```bash
gaia_ml inspect model.onnx --stage raw
gaia_ml inspect model.onnx --stage algebra
gaia_ml inspect model.onnx --stage canonicalize
gaia_ml inspect model.onnx --stage math
gaia_ml inspect model.onnx --stage loop
gaia_ml inspect model.onnx --stage schedule
gaia_ml inspect model.onnx --stage constraints
```

Use JSON output for scripts and regression tests:

```bash
gaia_ml inspect model.onnx --stage schedule --json > schedule.json
```

For a dynamic ONNX input, provide one static shape:

```bash
gaia_ml inspect model.onnx --input-shape 1,18,18,2 --stage constraints
```

List implemented compiler stages and scheduling controls:

```bash
gaia_ml show optimizers
gaia_ml show optimizers --json
```

## Compiling a model

The recommended starting point lets GAIA-ML choose legal implementations:

```bash
gaia_ml compile model.onnx \
  --output gaia_project \
  --force \
  --schedule-objective latency \
  --data-file model_input.csv \
  --simulation-events 10
```

The output directory contains:

```text
gaia_project/
├── graph/          ADF graph declaration and simulation host
├── kernels/        generated AIE C++ kernels
├── weights/        compile-time packed coefficient headers
├── testVectors/    scalar and AXI/PLIO input vectors
├── tools/          generated simulation-report helper
├── Makefile
├── CMakeLists.txt
├── manifest.json   selected IR, schedule, constraints, and estimates
├── compilation_certificate.json evidence-labelled compilation decision
└── compile.log     compiler decisions and diagnostics
```

`compile.log` is also printed to the terminal. `manifest.json` is the preferred
machine-readable record of the chosen schedule.

Both outputs distinguish the MAC-issue compute floor, compulsory and
selected-microkernel local-memory traffic, serial emitted PLIO cost,
hypothetical fully packed 2/4/8-port PLIO costs, generated topology transport,
the calibrated compiler prediction, and the requested deadline. GAIA reports
the dominant physical resource and the minimum compute/memory replication and
independent-port counts implied by a target. Hypothetical port scenarios do not
claim that the emitted graph already has those ports.

The machine-readable decomposition is under
`latency_estimates.resource_cost_model`:

```text
compute                    dtype-aware ideal issue floor
local_memory               compulsory bytes and selected-microkernel traffic
emitted_instruction_issue  load/store/compute/scalar issue floor
topology                   generated stitch/sum/broadcast service floor
io.serial                  actual one-boundary-port transfer cost
io.parallel_port_scenarios fully packed 1/2/4/8-port what-if costs
physical_lower_bound       dominant emitted-plan resource floor
prediction_gap             calibrated non-ideal residual above that floor
target_gap_analysis        target-implied replication and port minima
```

Memory-bank recommendations and emitted A/B/C DM resource annotations are
listed per kernel. They are compiler resource constraints, not claims about
literal physical-bank placement; the Vitis linker map remains authoritative.

From v0.10.20 this decomposition is part of whole-graph selection. Each
retained complete plan is dry-lowered before final ranking. Its emitted
compute, compulsory-memory, PLIO, internal-topology, and instruction-issue
floor can reject a hard-deadline plan; selected-microkernel memory traffic and
the same physical components also participate in Pareto ranking. The graph
planning report records the local prediction, physical floor, dominant
resource, and unexplained non-ideal residual for every evaluated plan.

`schedule.analysis.graph_planning.global_optimization` describes comparison
scope explicitly. GAIA jointly searches microkernel, SIMD/tail, tiling,
parallelism, layout, interface, and buffering choices and evaluates the actual
fusion, routing, placement, FIFO, register, memory, instruction, and I/O costs
created by dry lowering. Alternative placements, arbitrary fusion partitions,
independent multi-PLIO partitioning, inferred consumer replication, and full
token-level SDF retiming are reported as non-enumerated capabilities. Thus a
resource-minimal result is conditional on the declared retained search domain;
only an exhaustive search report supports minimality over that domain.

For comparisons with Vitis AI or another compiler, use
`latency_estimates.comparison_record`. It provides one normalized row of target
clock and PLIO parameters, compile constraints, predicted first-output latency,
steady-state II and throughput, physical resource floors, emitted tile/kernel/
memory/FIFO counts, evidence class, and empty measured-result fields. Populate
the measured fields only from an equivalent hardware or simulator experiment;
the record also lists the equivalence conditions that must be held constant.

### Objectives and explicit implementations

An objective ranks legal candidates; it is not itself a microkernel:

```bash
--schedule-objective latency
--schedule-objective throughput
--schedule-objective compact
--schedule-objective auto
```

Keep `--schedule-implementation auto` for normal compilation. An explicit
implementation is primarily useful for controlled benchmarking or debugging:

```bash
gaia_ml compile model.onnx \
  --output gaia_project \
  --schedule-objective latency \
  --schedule-implementation conv2d_line_buffer
```

Not every explicit implementation is legal for every operation, shape, or
datatype. GAIA-ML reports a structured compiler error instead of silently
falling back when a forced implementation is illegal.

### Target output interval

GAIA-ML distinguishes hard constraints from scheduling preferences. A required
steady-state initiation interval is expressed independently from first-output
latency:

```bash
gaia_ml compile model.onnx \
  --output gaia_project \
  --schedule-objective latency \
  --required-ii-us 3 \
  --required-first-output-us 8 \
  --max-tiles 16 \
  --max-kernels 20
```

Use `--preferred-ii-us` when the number is a ranking goal rather than a reason
to reject compilation. `--exact-tiles` requests an exact emitted compute-tile
count, while `--max-tiles` and `--max-kernels` impose upper bounds. Omitting a
constraint lets the compiler optimize that dimension according to
`--schedule-objective`.

`--target-latency-us` remains available for compatibility and is normalized as
a soft preferred steady-state II. Do not combine it with `--required-ii-us` or
`--preferred-ii-us`.

The normalized, versioned compilation contract and the observed emitted-plan
metrics are available from `inspect --stage constraints --json`, terminal
output, `compile.log`, and `manifest.json`. A violated hard constraint produces
`GAIA-CONTRACT-INFEASIBLE` before project code generation. Estimates still
require validation with AIE simulation or hardware, especially for tensor
geometries outside the calibrated microkernel set.

From version 0.10.2, hard requirements participate in candidate enumeration
and whole-graph beam search. The JSON schedule report contains
`analysis.constraint_search`, including the number of evaluated and feasible
plans and the selected pre-lowering feasibility evidence. Tile and kernel
counts at this stage are conservative compute-kernel lower bounds; routing,
reduction, and fusion kernels are counted exactly after AIE lowering.

From version 0.10.3, feasible graph plans are filtered through a Pareto
frontier before deterministic selection. The reported objective vector contains
steady-state II, first-output latency, tile/kernel lower bounds, maximum local
memory, layout adapters, and evidence risk. `latency` prioritizes steady II and
then first output; `throughput` prioritizes steady II before resource use;
`compact` prioritizes tiles, kernels, and memory before latency. The complete
priority order and non-dominated plans appear under
`schedule.analysis.graph_planning`.

Version 0.10.4 uses two resource fidelities during that search. Early beam
pruning uses inexpensive compute-worker lower bounds. Complete surviving plans
are then dry-lowered through AIE IR, so final Pareto selection includes exact
routing, stitch, reduction, fusion, kernel-count, and tile-memory effects. The
selected plan reports these under `emitted_resources`; a
`resource_uncertainty` value of zero means exact dry-lowered counts were used.

Version 0.10.5 moves deterministic placement out of C++ code generation and
into the AIE target model. Exact complete plans now report kernel coordinates,
internal Manhattan routes, total and maximum route hops, first-row occupancy,
and maximum shared-link load. These physical-pressure metrics participate in
the final Pareto frontier, while the validated implementation replacement gate
continues to prevent an analytical placement preference from silently replacing
a measured microkernel. Constraint reports, manifests, and emitted
`adf::tile(...)` declarations all consume the same placement plan.

The route model is deliberately structural rather than cycle accurate. It is a
compiler selection signal and debugging aid; Vitis routing, AIESim, and hardware
remain authoritative for latency and congestion measurements.

Version 0.10.6 similarly promotes stream/FIFO policy from graph C++ generation
to target analysis. Every exact emitted edge records its physical token count,
datatype, stream or window interface, recommended FIFO depth, buffering bytes,
boundary role, and structural rate-risk reason. Complete graph plans expose
total FIFO storage, maximum depth, reconvergent/fanout risk count, and maximum
event-to-buffer ratio to Pareto analysis. The ADF graph consumes these reported
depths by connection index, preventing the inspection result from diverging
from generated `adf::fifo_depth(...)` declarations.

This rate model deliberately preserves the validated v0.10.5 FIFO depths. A
fanout or reconvergent edge is now visible as an optimization venue, but GAIA
does not claim that increasing a FIFO improves latency without Vitis FIFO
guidance or AIESim/hardware evidence.

Version 0.10.8 adds a compilation certificate that distinguishes verified
feasibility, bounded feasibility, predicted feasibility inside the complete
guardbanded envelope, scoped/proven infeasibility, and an unresolved result.
Predicted feasibility leaves `contract_satisfied` unset until exact-plan
verification; a latency point estimate is never labelled as proof. The
selected kernel programming model remains AIE API C++; lower-level intrinsics
will be evaluated as microkernel candidates rather than replacing working
kernels globally.

Version 0.10.7 added an explicit SIMD microkernel IR after emitted AIE lowering.
For each kernel it records instruction families, datatype, lanes, vector bits,
register span, repeat count, vectorized axis, padding, scalar tails, estimated
live vector/accumulator pressure, and utilization of the configured 512-bit AIE
v1 register span. Complete-plan Pareto analysis can therefore distinguish an
otherwise equal vector-safe plan from one likely to spill, serialize, or waste
most of its allocated vector span.

The SIMD IR is currently a structural verification and selection layer over the
same AIE kernel objects consumed by the established C++ templates. It does not
yet replace those templates as the final code emitter, and its instruction
counts must not be presented as disassembly measurements. Vitis assembly,
profile reports, AIESim, and hardware remain authoritative.

### Datatypes and quantization

Select storage and arithmetic scheduling with `--schedule-dtype` or its
`--quantize` alias:

```bash
gaia_ml compile model.onnx --output float_project --schedule-dtype float32
gaia_ml compile model.onnx --output int16_project --schedule-dtype int16
gaia_ml compile model.onnx --output int8_project --schedule-dtype int8
```

Supported code-generation types are `float32`, `int32`, `int16`, and `int8`.
Integer kernels use wider accumulators and datatype-specific AIE primitives
where legal. Reduced precision does not automatically guarantee lower latency:
the available AIE instruction shape, padding, conversion work, routing, and
model topology all affect the result.

`int4` is accepted for schedule analysis but code generation deliberately
rejects it because packed-nibble storage and arithmetic are not implemented.

Compile-time quantization converts weights and generated input vectors using
the compiler's fixed-point policy. It is not a replacement for quantization-
aware training or a model-accuracy study. Always compare generated output with
an independently produced golden reference.

### Multi-tile schedules

GAIA-ML exposes the following tiling controls:

```text
none
output_split
spatial_split
reduction_split
batch_split
channel_unroll
```

With `--multi-tile none`, the automatic latency scheduler may still choose a
multi-tile topology when a single tile is illegal or a calibrated parallel
implementation ranks better. Explicit multi-tile modes constrain the search
and can require additional broadcast, stitch, or reduction kernels.

### Shape-driven float32 Conv selection

Schedule IR records a model-independent `shape_analysis` for each Conv2D. It
classifies the layer as depthwise-multiplier, depthwise, pointwise,
small-channel spatial, or channel-reduction; identifies the vector axis and
estimated native-lane utilisation; records border and line-buffer properties;
and lists independently tileable output dimensions and the required weight
transform.

For float32 depthwise convolution, a depth multiplier of four through eight
can be packed as `[input channel][kernel tap][8 multiplier lanes]`. GAIA then
computes all multiplier filters for one input channel with one native AIE1
vector accumulator. Smaller multipliers remain on reduction/spatial
vectorisation because filling fewer than half of the eight float lanes was not
profitable in Vitis-generated code. Explicit multi-tile requests currently use
the general depthwise backend so that tile boundaries cannot split a packed
multiplier group.

### Shape-driven Dense selection

Schedule IR also records a `shape_analysis` for every Dense operation. It
classifies dynamic, tiny, quantized, convolution-fed, native-MMUL, and general
vector Dense shapes; measures reduction- and output-lane utilisation; and
describes the required weight transform. The analysis recommends scalar,
vector-dot, streaming outer-product, buffered MMUL, or direct MMUL codegen and
an appropriate single-tile, output-split, or batch-split topology.

These recommendations seed the implementation search rather than replacing
it. GAIA's calibrated cost model remains the final ranking step. In
particular, the direct MMUL accumulator is restricted to a single two-output
block: previous AIE simulations showed that retaining several output blocks in
one kernel can increase register pressure and perform worse than buffered
MMUL. Streaming outer-product now uses the native AIE broadcast operation and
packed output-major weights, without a scalar temporary broadcast array.

Use `inspect --stage schedule` to see the classification, vector utilisation,
recommended microkernel and topology, and any cost-model override.

### AIE v1 capability contracts

Version 0.9.8 gives every current Dense, Conv2D, depthwise Conv2D, activation,
Softmax, and BDT implementation a formal capability declaration. A contract
records supported dtypes, interfaces, parallel modes, layouts, vector widths,
static-shape requirements, calibration provenance, fusion support, and the
existing AIE lowering backend. Each scheduled candidate is evaluated through
the contract for matching, target legality, code-generation readiness, cost,
latency, and its concrete lowering layout.

The AIE v1 registry now owns the implementation universe and candidate
admission. Dense and BDT builders enumerate their alternative tilings and
microkernels through the registry; Conv2D and activation selection must also
resolve to a registered strategy before lowering. Unsupported variants remain
visible as rejected candidates but cannot displace a supported implementation.

Fusion is explicit as well: compact CNN, Dense-chain, depthwise-to-Dense, and
streaming CNN-tail transformations are registered whole-region strategies.
This preserves the measured fast paths while removing hidden lowering choices.
Use JSON inspection to consume the complete report:

```bash
gaia_ml inspect model.onnx --stage schedule --json
```

The selected contracts are under `schedule.analysis.capability_contracts`, with
an aggregate `capability_summary`. `schedule.analysis.strategy_enumeration`
records declared, constructed, admitted, and rejected candidates. Human-readable
inspection prints both views.

Conv2D scheduling additionally emits `cross_validation` for every operation.
The compiler independently builds the applicable standard or depthwise
strategies, compares their topology, memory, cycle, and steady-II estimates,
and records whether each alternative has enough evidence to replace the current
winner. MAC-only estimates remain visible as analytical shadow candidates. They
cannot replace a Vitis/AIESim-validated implementation unless the alternative
has independent same-geometry profile evidence and improves the
confidence-adjusted objective by at least 20%. This guard prevents optimistic
instruction-count estimates from silently regressing measured float32 latency.

### Whole-graph planning

Version 0.10.0 includes an evidence-aware graph planner after local candidate
enumeration. It performs a bounded search across the complete operator chain
and compares steady-state pipeline II, first-output fill, tile count, stream
barriers, and producer/consumer layouts. Layout mismatches receive an explicit
adapter cost and cannot be selected until codegen can materialize the adapter.
Flatten-to-Dense permutations are treated specially: when mathematically
legal, GAIA reorganizes Dense weights offline instead of inserting a runtime
transpose.

Candidate estimates are classified as analytical, geometry-scaled,
profile-calibrated, or same-geometry measured. Ordinary ranking remains
conservative. A hard deadline permits any legal, codegen-ready candidate to
compete, but progressively larger uncertainty margins are applied as evidence
gets weaker. Fusion opportunities are rechecked by AIE lowering for memory and
semantic legality.

The JSON report is available under `schedule.analysis.graph_planning`. It
contains the incumbent and challenger signatures, transition costs, evidence
floor, evaluated-plan count, replacement decision, and eligible fusion regions.

Version 0.10.24 additionally emits `schedule.analysis.region_schedule`.
Residual diamonds and linear Conv chains carry an explicit partition axis,
worker count, physical stream layout, halo requirement, local operations, and
materialization points. This separates a layer schedule from the cross-layer
invariants needed for persistent streaming. The emitted constraint report also
contains `constraints.dataflow.rate_analysis`: event service controls the
predicted steady-state II, while a distinct first-token critical path models
pipeline fill. FIFO deficits and required consumer replication are reported;
these are schedule predictions, not formal WCET proofs.

For hard-II stream graphs, residual fanout uses native ADF multicast and the
reconvergent Add retains full-event FIFO elasticity. Conservative schedules use
explicit two-port-safe binary broadcast trees. Add followed immediately by
ReLU is fused into one vector kernel.

Version 0.10.25 adds exact-scope graph calibration and residual partition
alignment. A measured topology is reused only when its complete Conv geometry,
implementation families, and tile splits match. Equivalent residual stitch
maps are pushed below Add, retaining local shard ownership until the region
exit. Exact dry lowering also owns the full declared tile budget; the planner
does not reserve speculative mover tiles when it can count the emitted graph.

Schedule-side Conv alternatives also use the emitted-backend calibration model.
It charges scalar address generation, patch packing, padded stream traffic,
fixed prologues, and backend-specific vector work before graph ranking. This
corrects the canonical 18x18/K7/Cout5 direct-Conv shadow estimate from 1,102
cycles to 107,136 cycles (`85.7088 us` at 1.25 GHz), while leaving the selected
spatial-vector estimate at `2.6576 us`. Compile output includes
`latency_estimates.calibration_consistency`, which compares Schedule IR and
emitted-program steady-state II and reports their delta and ratio.

### Empirical compiler feedback

Version 0.10.0 closes the validation loop without allowing unrelated profiles
to influence scheduling. After a completed simulation, record its event timing
with:

```bash
gaia_ml profile test_prj/gaia_compile \
  --database .gaia_ml/aie_v1_feedback.json \
  --source "Vitis 2024.2 AIESim"
```

### Latency-bound benchmark and Pareto studies

Version 0.10.33 can construct a reproducible compiler study without requiring
Vitis during model generation or Schedule-IR exploration:

```bash
gaia_ml benchmark generate local_gaia_study/corpus \
  --profile paper --events 10

gaia_ml benchmark matrix local_gaia_study/corpus \
  --profile paper --output local_gaia_study/matrix.json

gaia_ml benchmark run local_gaia_study/matrix.json \
  --output local_gaia_study/compiler_results

gaia_ml benchmark analyze local_gaia_study/compiler_results/results.json \
  --output local_gaia_study/pareto_analysis.json
```

The corpus includes CNN, residual-CNN, DNN, and bounded-BDT families. Model
splits are assigned deterministically before measurements: calibration models
may fit uncertainty residuals, while evaluation models only measure transfer
error. A constraint point fixes hard steady-state II, maximum tiles, datatype,
PLIO frequency, and tile memory. The matrix records cases whose deadline is
already below serialized physical I/O as infeasible before schedule search.

Dense plan records include an explicit evidence lifecycle. An analytical
candidate is distinct from a dry-lowered, codegen-ready, resource-legal plan.
Only the latter can be marked `admitted`; `certified` remains false until an
external Vitis compile, simulation, and numerical check supplies that evidence.
When a lowering oracle is available, failure to obtain exact emitted-program
evidence is fail-closed for hard-contract admission.

Dense float32 and int16 code generation also shares one executable numerical
contract between test-vector generation and verification. Float32 preserves
the declared Dense operation order (`matmul`, bias, then activation) and uses
absolute and relative tolerances of `1e-5`. Signed int16 uses symmetric Q12
tensors (scale 4096, zero point zero), Q24 bias, a signed 48-bit accumulator,
nearest-ties-to-even input/parameter quantization, arithmetic-right-shift
requantization, and activation before signed-int16 saturation. The complete
contract is emitted as `numerical_contract` in `manifest.json`; int8 and int32
remain outside this Dense numerical certification milestone.

Small declared schedule products can be enumerated by raising
`--exhaustive-search-limit`. “Exhaustive” means the complete target-legal
candidate domain declared by GAIA, not every possible AIE instruction program.
For larger graphs, Pareto and optimality gaps are explicitly labelled as
explored-space results.

After AIESim projects have been verified, preserve their corpus split and fit
the uncertainty model with:

```bash
gaia_ml benchmark ingest local_gaia_study/corpus PROJECT... \
  --output local_gaia_study/measurements.json

gaia_ml benchmark fit local_gaia_study/corpus PROJECT... \
  --output local_gaia_study/cost_model.json

gaia_ml compile model.onnx -o project --force \
  --required-ii-us 5 --max-tiles 32 \
  --cost-model-db local_gaia_study/cost_model.json
```

The fitted guardband participates in hard-deadline admission. Held-out records
are never used to fit it.

GAIA derives a SHA-256 fingerprint from the complete AIE v1 schedule, tensor
geometry, physical layouts, kernel topology, and PLIO policy. Repeated runs are
aggregated with medians and assigned a confidence level. The fingerprint also
includes the ONNX model, emitted weight/table, and optional input-data hashes.
Simulation records capture the output/report artifact hash and detected Vitis
version; importing the same artifacts again is deduplicated. A JSON copy of
the observation is written to `gaia_feedback.json` in the generated project.

Pass the database to a later identical compilation to expose the measurement
alongside the analytical and compiler estimates:

```bash
gaia_ml compile model.onnx \
  --output test_prj/gaia_compile \
  --calibration-db .gaia_ml/aie_v1_feedback.json \
  --force
```

Only an exact emitted-plan fingerprint matches. GAIA does not extrapolate a
measurement across shapes, data types, tilings, layouts, interfaces, or kernel
topologies, and empirical feedback does not silently replace the selected
schedule. This makes discrepancies actionable while preserving the validated
fast paths.

### Bounded BDT support

GAIA can compile a single-output ONNX-ML `TreeEnsembleRegressor`. The scheduler
ranks three representations: legacy QuickScorer, cumulative-prefix QuickScorer,
and normalized direct traversal. Prefix QuickScorer replaces data-dependent
split scans and scattered tree updates with a threshold-bucket lookup followed
by contiguous bitwise masks. Its threshold, prefix-mask and leaf tables are
placed in separate AIE data-memory banks. Direct traversal covers deeper trees
without requiring one bit per leaf.

Generate a deterministic test model and 100 golden events with:

```bash
python tests/generate_bdt_models.py \
  --output-dir model \
  --features 32 \
  --depth 5 \
  --trees 8 \
  --events 100 \
  --force

gaia_ml compile model/test_bdt.onnx \
  --output test_prj/bdt \
  --data-file model/test_bdt_input.csv \
  --simulation-events 10 \
  --schedule-objective throughput \
  --multi-tile forest_split \
  --tile-output 2 \
  --force
```

This development backend supports float32 features, `BRANCH_LEQ`, `SUM` or
`AVERAGE`, one regression target, at most 32 features, depth at most eight, and
at most 1024 trees. Trees through depth five can use 32-bit QuickScorer masks;
deeper trees use direct traversal. `--multi-tile forest_split` partitions the
forest into independently scheduled workers, broadcasts each feature vector
through two-port-safe binary fan-out, and combines partial scores with a binary
reduction. `--tile-output N` bounds the trees assigned to each worker.
Multi-target/class-label output, other split predicates and integer thresholds
remain future work. Forest splitting expands ensemble capacity and throughput,
but each individual tree and worker table must still fit one AIE tile.

The BDT cost model ranks emitted implementations rather than treating the split
count as the latency by itself. It reports the arithmetic lower bound, physical
PLIO floor, calibrated compiler estimate, and requested steady-state II. The
current Vitis 2024.2 calibration includes feature ingest, dependent traversal,
prefix-mask updates, two-port-safe fan-out, reduction fill and per-worker table
memory. For the deterministic 32-feature/depth-5/eight-tree fixture, the
measured single-tile prefix backend is 2.131 us/event; four direct-traversal
workers measure 0.830 us/event with a 1.971 us first output. These are
calibration points, not universal guarantees—tree shape and target routing can
change the result, so publish AIESim or hardware measurements alongside the
estimate.

#### BDT 0.11.0 certification evidence

The B1--B10 certification programme separates importer semantics, exact
resource accounting, microkernel behavior, calibration, forest Pareto search,
capacity/routing limits, adversarial numerical diagnostics, held-out admission,
package audit, and homogeneous final closure. The frozen B8 evaluation contains
20/20 verified 100-event programs, including 16/16 genuinely held-out plans
over seeds 257, 307, and 353. Held-out measured p95 II spans 1.2352--9.6122 us.
Across 252 CLI constraint replays, the 184 held-out exact-actor comparisons
record TP=82, TN=89, FP=0, and FN=13.

The homogeneous 0.11.0 B10 closure verifies 12/12 held-out programs over 100
events each. Its measured p95 II spans 1.2352--9.6122 us (median 2.7632 us),
first output spans 2.7264--8.5344 us, the worst final-score absolute error is
1.188e-6, and all 12 compiler resource bounds cover the Vitis reports.

The evidence supports bounded numerical correctness and empirical sub-10-us
steady-state synthesis for the measured plans. It does not establish physical
WCET, arbitrary tree-model support, universal sub-10-us execution, or global
optimality. Depth-5 prefix ensembles are measured through 256 trees; the
current 512/1024-tree broadcast/reduction topology fails closed at the VP2802
directional-routing constraint. Raw certification data and paper-specific
material are kept outside the release repository under the study workspace.

### Stream and window interfaces

Streaming is the default:

```bash
--interface stream
```

Window interfaces are available for supported kernels:

```bash
--interface window
```

A stream does not inherently mean lower latency. Backpressure, reconvergent
paths, padding, and kernel production rates can dominate. Use the simulation
report before drawing conclusions from the interface name alone.

### Input data

`--data-file` accepts CSV rows where each row is one flattened input event:

```text
x0,x1,x2,...,xN
```

The logical row width must match the compiled model input. GAIA-ML performs any
required physical padding and emits both scalar test-vector files and Vitis
AXI/PLIO CSV files with command, data, `TLAST`, and `TKEEP` fields.

`--simulation-events N` selects the first `N` rows; it does not duplicate one
row to create additional events.

For float32 compilation GAIA evaluates the exact ONNX file with ONNX's
reference evaluator and emits that model-specific golden result. A sibling
`<stem>_output.csv` is cross-checked and rejected when it belongs to a different
model; the model SHA-256 and decision are recorded under
`input_data.golden_reference` in `manifest.json`. Integer sibling goldens remain
externally supplied references and are labelled unverified because storage
conversion is not yet an accuracy-preserving quantized reference interpreter.

## Building the generated project

Load the Vitis environment, enter the generated directory, and use the
generated Makefile:

```bash
source /opt/modules/Vitis/2024.2/settings64.sh
cd gaia_project

make x86
make x86sim
make aie
make aiesim
```

Useful additional targets include:

```bash
make aiesim_profile
make aiesim_report
make aie_fifo
```

The report helper summarizes output timestamps, first-output latency,
steady-state event intervals, large bubbles, PLIO throughput, stream stalls,
and available FIFO guidance.

If hardware compilation reports stack overflow, use the stack recommendation
in `compile.log` or `manifest.json`:

```bash
make aie AIE_STACK_SIZE=<recommended-by-gaia>
```

Do not increase the stack beyond the reported safe tile-memory headroom.

### AIE-v1 resource admission

Version 0.10.33 separates resources that the earlier aggregate 32-KiB check
could conflate. Each emitted kernel reports logical tensor storage, packed
vector padding, weights, biases, scratch, assigned FIFO bytes, stack
reservation, runtime data margin, and a separate program-memory bound. A plan
is rejected before project emission when its data admission exceeds the
32-KiB tile store, its code-template bound exceeds the 16-KiB AIE-v1 program
store, its stream-port count exceeds two inputs or two outputs, or placement
exceeds the VP2802-2M tile capacity.

Tensor and padding sizes are exact for the dry-lowered AIE IR. Stack and
program-store values are conservative admission bounds, not cycle-accurate
Vitis predictions. The resource-report reader compares them with
`report_heap.txt`, `report_stack.txt`, and `report_pm.txt` and records residual
margins; Vitis remains authoritative for a certified artifact.

### Dense certification calibration

The D4 study in version 0.10.34 measured 60 independent single-tile Dense
projects with Vitis 2024.2 AIESim and retained a deterministic held-out subset.
Hard-deadline admission now applies versioned upper planning envelopes for the
measured scalar/vector-dot and float32/int16 classes. The manifest records the
calibration key, content hash, training and held-out counts, and the zero
held-out-false-admission result. This calibration is deliberately excluded
from composed graphs, Conv/BDT operations, unmeasured datatypes, and multi-tile
plans; ordinary no-deadline ranking continues to use the compiler point model.

Version 0.10.35 also makes terminal Dense output splitting an end-to-end model
implementation: worker outputs are restored to one ordered tensor through an
AIE-v1-legal binary stitch tree. Tile budgets therefore include both compute
workers and stitch kernels instead of counting workers alone.

The D5 certification executed 12 such projects with two and four workers.
All projects compiled, simulated, and passed their numerical oracle; int16 was
exact and the largest float32 absolute error was `5.76e-6`. Four-worker
float32 cases measured `0.569`--`0.670` microseconds II. The measured Pareto is
not assumed monotonic: some int16 split plans were slower than their one-tile
counterpart, so tile multiplication remains a searched option rather than a
hard-coded latency rule. Version 0.10.36 begins certification of sequential
Dense composition using the same end-to-end evidence requirements.

D6 certified two-, three-, and four-layer float32 and int16 Dense pipelines.
All six projects passed; float32 sustained about `2.284` microseconds II and
int16 about `3.62` microseconds II as depth increased, while pipeline-fill
latency grew per stage. The phase also corrected sub-word internal streaming:
packed int8/int16 PLIO ingress may use vector reads, but a downstream kernel
must consume scalar producer transactions sample-by-sample unless the producer
explicitly emits packed vectors. Large fully-unrolled integer stream-outer
products are no longer selected automatically after Vitis 2024.2 demonstrated
pathological scheduling time; they remain available by explicit request.
Version 0.10.37 opens exact Q6 int8 Dense certification; int32 remains
fail-closed until an acc80-preserving reduction backend exists.

D7 certified four Q6 int8 Dense geometries with exact AIESim agreement. Their
measured steady II ranged from `0.148` to `0.649` microseconds.

D8 then evaluated ten held-out, non-grid Dense geometries with 100 events per
case. All ten Vitis 2024.2 AIESim projects passed: float32 measured at most
`2.691` microseconds II with maximum absolute error `8.38e-7`, while int16
measured at most `4.610` microseconds II with exact agreement. Replaying the
explicit `--required-ii-us 10` contract admitted all ten cases with zero false
admissions and zero false rejections. Version 0.10.49 also ensures that the
final emitted-program contract consumes the implementation-specific Schedule
IR envelope; it can no longer silently replace the stricter Dense calibration
with a generic evidence-class guardband. `--target-latency-us` remains a
legacy soft preference; use `--required-ii-us` for a hard steady-state bound.

## Supported model patterns

Version 0.11.0 targets feed-forward inference graphs composed from:

- Dense layers represented by `Gemm` or canonicalizable `MatMul + Add`;
- standard 2D convolution;
- depthwise 2D convolution represented by supported grouped convolution;
- ReLU and sigmoid activations;
- equal-shape two-input Add for residual/skip connections;
- terminal Softmax, including LUT-based implementations where selected;
- BatchNorm that can be folded into constant convolution parameters; and
- Flatten, Reshape, Transpose, and layout helpers that can be eliminated or
  represented as compiler layout metadata; and
- bounded additive `TreeEnsembleRegressor` models described above.

Supported mixtures include multilayer perceptrons, convolution chains, and
convolution-to-Dense classifier tails. Static shapes and constant trained
weights are strongly recommended.

## Current limitations

GAIA-ML 0.11.0 is an experimental compiler, not a general ONNX backend.

- Only the operators and canonicalizable patterns listed above are compiled.
  Pooling, recurrent networks, attention, arbitrary elementwise graphs, and
  general ONNX control flow are not implemented.
- Add currently requires exactly two tensors with identical static shape and
  dtype. General ONNX broadcasting is not yet lowered.
- Batch size one has the most mature latency path. Batch tiling exists, but is
  not as extensively calibrated as single-event streaming.
- AIE v1 float32 arithmetic flushes non-zero subnormal operands to signed zero.
  BDT diagnostics expose both strict ONNX float32 and `aie_v1_ftz` branch
  semantics; models that rely on subnormal distinctions are not bit-exact on
  this target.
- Dense output and batch splitting are supported. A Dense whose single output
  reduction cannot fit one tile is diagnosed as requiring reduction splitting;
  general Dense reduction-tree code generation is not implemented yet.
- Dynamic dimensions must be resolved with `--input-shape`; general dynamic-
  shape code generation is not supported.
- Grouped convolution support is primarily intended for standard or depthwise
  cases. Arbitrary group configurations may select conservative code or be
  rejected.
- Automatic layouts, fusion, and tiling are heuristic. The compiler does not
  exhaustively search every legal graph placement or routing solution.
- Performance estimates combine analytical bounds with a limited set of Vitis
  2024.2 AIE simulation calibrations. Estimates are most reliable near tested
  shapes and can be inaccurate for substantially different networks.
- A legal schedule is not guaranteed to be globally latency-optimal. Vitis
  placement, routing, FIFO behavior, stack use, and stream backpressure can
  change observed performance.
- Integer conversion is a compiler storage transformation, not automatic
  accuracy-preserving quantization. Accuracy must be validated by the user.
- int4 code generation is not implemented.
- Generated C++ targets AIE v1 APIs and has not been validated as an AIE-ML or
  future-architecture backend.
- Hardware execution, platform integration, and host application generation
  are outside the current release; GAIA-ML emits the AIE graph project.

For a new architecture, begin with `inspect --stage constraints`, compile with
automatic scheduling, run at least ten simulation events, and compare the
measured output and timing against an independent reference before forcing a
specific microkernel.

## Python inspection API

The stable Python-facing inspection objects are available from `gaia_ml`:

```python
from gaia_ml import GraphAnalyzer, OnnxModelLoader

graph = OnnxModelLoader("model.onnx").load()
report = GraphAnalyzer(graph).analyze()

print(report.operation_count)
print(report.operation_counts)
```

The IR and low-level scheduling classes remain available for compiler
development, but the command-line interface and generated manifest are the
recommended user interfaces for version 0.11.0.

## Development and release checks

Run the unit suite:

```bash
python -m unittest discover -s tests -v
```

The GitLab pipeline installs GAIA-ML, generates deterministic release models,
compiles representative Dense, Conv2D, depthwise, Conv-to-Dense, and BDT
projects, and validates the wheel metadata. Measurement-study harnesses and
paper artifacts are maintained outside the release repository.

Build release artifacts:

```bash
python -m pip install build
python -m build
```

The package version is defined in `python/gaia_ml/_version.py`.
