Metadata-Version: 2.4
Name: ameva-runtime
Version: 2.7.9
Summary: Unified Next-Gen Hardware Orchestration & AI Acceleration Runtime for Mobile & Edge
Home-page: https://github.com/uno-km/ameva-runtime
Author: Eunho Kim
Author-email: Eunho Kim <contact@uno-km.com>
License: Apache-2.0
Project-URL: Homepage, https://uno-km.vercel.app/lib/vulkan/
Project-URL: Repository, https://github.com/uno-km/ameva-runtime
Project-URL: Documentation, https://uno-km.vercel.app/lib/vulkan/
Keywords: vulkan-compute,mobile-gpu,hardware-acceleration,hardware-abstraction-layer,adreno-gpu,arm-mali,snapdragon-8-elite,exynos,termux,on-device-ai,edge-ai,tensor-acceleration,spir-v,compute-shaders,zero-silent-fallback,llamacpp,whisper-cpp,sherpa-onnx,stable-diffusion,vision-language-models,bitnet,gguf,ncnn,bionic-loader,arm64,aarch64,cgroup-management,cpu-neon,thermal-throttling,power-efficiency,smart-router,hardware-orchestration,multi-modal-ai,subgroup-operations,gemm-acceleration,mobile-vlm,speech-to-text,text-to-speech,image-generation,edge-inference,unprivileged-userspace,termux-wake-lock,phantom-process-killer,android-ai-runtime,valhall-gpu,adreno-830,mali-g68,ameva-foundation,uno-km,open-source-ai
Classifier: Development Status :: 5 - Production/Stable
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Operating System :: POSIX :: Linux
Classifier: Operating System :: Android
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.8
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.8
Description-Content-Type: text/markdown
License-File: LICENSE
Provides-Extra: stt
Requires-Dist: termux-stt>=1.2.0; extra == "stt"
Provides-Extra: diffusion
Requires-Dist: termux-diffusion>=1.5.0; extra == "diffusion"
Provides-Extra: bitnet
Requires-Dist: termux-bitnet>=1.2.0; extra == "bitnet"
Provides-Extra: llamacpp
Requires-Dist: termux-llamacpp>=1.3.0; extra == "llamacpp"
Provides-Extra: tts
Requires-Dist: termux-tts>=1.4.0; extra == "tts"
Provides-Extra: vision
Requires-Dist: termux-vision>=1.2.0; extra == "vision"
Provides-Extra: all
Requires-Dist: termux-stt>=1.2.0; extra == "all"
Requires-Dist: termux-diffusion>=1.5.0; extra == "all"
Requires-Dist: termux-bitnet>=1.2.0; extra == "all"
Requires-Dist: termux-llamacpp>=1.3.0; extra == "all"
Requires-Dist: termux-tts>=1.4.0; extra == "all"
Requires-Dist: termux-vision>=1.2.0; extra == "all"
Dynamic: author
Dynamic: home-page
Dynamic: license-file
Dynamic: requires-python

# AMEVA-Runtime

[![PyPI](https://img.shields.io/pypi/v/ameva-runtime.svg?style=flat-square&color=0369a1)](https://pypi.org/project/ameva-runtime/)
[![Python](https://img.shields.io/pypi/pyversions/ameva-runtime.svg?style=flat-square)](https://pypi.org/project/ameva-runtime/)
[![npm](https://img.shields.io/npm/v/@ameva/runtime.svg?style=flat-square&color=b91c1c)](https://www.npmjs.com/package/@ameva/runtime)
[![GitHub Release](https://img.shields.io/github/v/release/uno-km/ameva-runtime-releases?style=flat-square&color=0969da)](https://github.com/uno-km/ameva-runtime-releases/releases/tag/v2.7.9)
[![License](https://img.shields.io/badge/License-Apache_2.0-004499.svg?style=flat-square)](https://github.com/uno-km/ameva-runtime)
<img src="https://img.shields.io/badge/BitNet%201.58b-Vulkan%20Compute%20Accelerated-purple.svg?logo=vulkan&logoColor=white" alt="BitNet Vulkan">

> Next-Gen Unified On-Device Hardware Orchestration & 6-Modality AI Acceleration Runtime (with BitNet 1.58-bit Vulkan Compute) for Mobile & Edge

---

## Architecture & Overview

AMEVA-Runtime is a hardware abstraction layer (HAL) and compute orchestration engine engineered specifically for mobile ARM64 devices (Android Termux, Linux Edge). It continuously inspects underlying silicon topology (/dev/kgsl-3d0, /dev/mali0) to route tensor execution across Qualcomm Adreno, ARM Mali, and ARM Cortex CPU-NEON backends.

### 6-Modality Acceleration Matrix

| Modality | Engine Integration | Status (v2.7.9) | Hardware Acceleration Mechanism |
| :--- | :--- | :---: | :--- |
| **1. LLM (Text)** | Llama.cpp & Termux-BitNet (1.58-bit i2_s) | **Production (v2.7.9)** | Vulkan 25/25 layer VRAM offload (`ngl=999`), BitNet full pipeline & strict `AmbiguousModelMatchError` |
| **2. STT (Speech)** | Whisper.cpp (Large-v3-Turbo) | **Production (v2.7.9 / STT v1.3.2)** | Vulkan GPU acceleration with greedy decoding default (`-bs 1`), 6-SoC fleet validated |
| **3. TTS (Audio)** | MeloTTS / Piper / Kokoro / Supertonic | **Production (v2.7.9 / TTS v1.5.0)** | Bionic Direct Vulkan (`/system/lib64/libvulkan.so`) & HiFi-GAN 32MB buffer temporal tiling |
| **4. Diffusion (Image)** | SDXS / SD1.5 / Z-Image Turbo (6.0B DiT) | **Production (v2.7.9 / Diffusion v1.8.0)** | Z-Image Turbo 6.0B DiT, Flash Attention (`--diffusion-fa`), VAE Tiling, Zero-Collision Bionic Vulkan backend |
| **5. Vision (VLM)** | CLIP / MobileVLM / LLaVA | **Production (v2.7.9)** | GGML Vulkan vision encoder tensor bindings |
| **6. BitNet (1.58b)** | Termux-BitNet (i2_s) | **Production (v2.7.9)** | ARM64 NEON DotProd SIMD & Vulkan compute |

---

## Empirical Physical Device Benchmarks

Tested on physical devices running Android 16 under Termux ARM64:

### 1. LLM Generation (Qwen2.5-0.5B-Instruct Q4_K_M)
Measured 2026-10-04 with the v2.7.9 bundle through `ameva-run exec`, 4 threads, single runs; the generated text is compared with the CPU route's. The Vulkan rows run with flash attention off (`-fa 0`): with it the bundle's Vulkan backend returns corrupt text. The figures this table carried up to 2.7.8 (35.80 t/s on Galaxy S25, 4.44 t/s on Galaxy A35) are not reproduced with this bundle and are withdrawn.

| Target Device | Hardware Architecture | Route | Generation Speed | Text |
| :--- | :--- | :--- | :---: | :---: |
| **Galaxy S25** | Snapdragon 8 Elite / Adreno 830 | Vulkan (default route, all layers on the GPU) | 24 t/s | same as the CPU's |
| **Galaxy S25** | Snapdragon 8 Elite / Oryon CPU | CPU (`-b cpu`) | 94 t/s | reference |
| **Galaxy A35** | Exynos 1380 / ARM Mali-G68 MP5 | Vulkan (default route, all layers on the GPU) | 6.5 t/s | same as the CPU's |
| **Galaxy A35** | Exynos 1380 / CPU | CPU (`-b cpu`) | 39 t/s | reference |

For this model the bundle's Vulkan backend is slower than the CPU on both devices. Larger models, the other devices and the known defects are in the [v2.7.9 release notes](https://github.com/uno-km/ameva-runtime-releases/releases/tag/v2.7.9).

### 2. Speech-to-Text (Whisper Large-v3-Turbo Q5_0, 548MB)
| Target Device | Hardware Architecture | Backend Mode | Latency (1-min audio) | GPU Load | CPU Load | Speedup |
| :--- | :--- | :---: | :---: | :---: | :---: | :---: |
| **Galaxy S25** | Snapdragon 8 Elite / Adreno 830 | Vulkan GPU | **4.40 s (0.07x RTF)** | Adreno Turbo | ~12% | **18.5x (vs CPU)** |
| **Galaxy S22** | Snapdragon 8 Gen 1 / Adreno 730 | Vulkan GPU | **5.13 s (Encoder)** | Adreno Native | ~14% | **3.73x (vs CPU)** |
| **Galaxy A35** | Exynos 1380 / Mali-G68 MP5 | Vulkan GPU | **360.60 s (6m 00s)** | **949 MHz (100%)** | 20~30% | **2.26x (56% time saved)** |
| **Galaxy A35** | Cortex-A78 x4 Cores | CPU-NEON | 816.48 s (13m 36s) | 0% | 291% | Baseline |

### 3. Text-to-Speech (AMEVA Bionic Native Vulkan Fleet Benchmarks)
| Target Device | Hardware Architecture | Neural Engine / Model | Backend Mode | Latency | RTF | Forensics (RMS/Peak) | Status |
| :--- | :--- | :--- | :---: | :---: | :---: | :---: | :---: |
| **Galaxy S22** | Snapdragon 8 Gen 1 / Adreno 730 | MeloTTS Universal Bilingual | **Bionic Vulkan GPU** | **2,750 ms** | **0.88x** | 0.0814 / 0.6974 | Real-time Synthesized |
| **Galaxy S21** | Exynos 2100 / ARM Mali-G78 | Piper VITS (On-chip Tiled) | **Bionic Vulkan GPU** | **560 ms** | **0.18x** | 0.0921 / 0.7412 | 5.5x Faster than RT |
| **Galaxy S20** | Exynos 990 / ARM Mali-G77 | Piper VITS (On-chip Tiled) | **Bionic Vulkan GPU** | **750 ms** | **0.24x** | 0.0890 / 0.7105 | 4.1x Faster than RT |
| **Galaxy S25** | Snapdragon 8 Elite / Adreno 830 | Supertonic 3 Flow / Piper | **Bionic Vulkan GPU** | **380 ms** | **0.12x** | 0.1042 / 0.8120 | Studio Ultra-Fast |
| **Galaxy A35** | Exynos 1380 / ARM Mali-G68 MP5 | Piper VITS (`lessac-medium`) | **Vulkan GPU** | **5,180 ms** | **1.146x** | 0.0782 / 0.6540 | Validated |

---

### 4. BitNet 1.58-bit LLM (Microsoft BitNet b1.58 2B-4T i2_s)
| Target Device | Hardware Architecture | Active Backend | Generation Speed | Prompt Latency | Speedup | Status |
| :--- | :--- | :---: | :---: | :---: | :---: | :---: |
| **Galaxy S25** | Snapdragon 8 Elite / Adreno 830 | **AMEVA Vulkan GPU** | **17.558 t/s** | **205.9 ms** | **12.58x** | **Verified (Ground Truth)** |
| Galaxy S25 | Snapdragon 8 Elite / Oryon CPU | Native CPU (4 Threads) | 1.396 t/s | 2,041.0 ms | 1.00x | Baseline |
| **Galaxy A35** | Exynos 1380 / Mali-G68 MP5 | **AMEVA Vulkan GPU** | **3.471 t/s** | **1,552.8 ms** | **5.94x** | **Verified (Ground Truth)** |
| Galaxy A35 | Exynos 1380 / Cortex-A78 CPU | Native CPU (4 Threads) | 0.584 t/s | 8,775.0 ms | 1.00x | Baseline |

## Root-Cause Driver Solutions

1. **ARM Mali-G68 Valhall Integer Truncation**: Enforced medium tile matmul kernel dispatch (`loadstride_b = 4 > 0`), permanently eliminating shader zero-stride infinite loops on subgroup-16 hardware.
2. **Qualcomm Adreno 830 JIT Register Bug**: Bounded vector column specialization (`mul_mat_vec_max_cols = 2`), preventing compiler crash `VK_ERROR_UNKNOWN (-13)`.
3. **Bionic Direct Vulkan Linking**: Permanently bypasses Termux `$PREFIX/lib/libvulkan.so` Mesa `llvmpipe` CPU software emulator trap by dynamically binding `/system/lib64/libvulkan.so` directly, routing compute shader execution to physical Adreno/Mali silicon.
4. **Mobile GPU 32MB Memory Ceiling & Temporal Tiling**: Resolved MeloTTS HiFi-GAN 52.4MB buffer overflow (`VK_ERROR_OUT_OF_DEVICE_MEMORY` / kernel TDR kill) via temporal chunk slicing ($T_{\text{chunk}} \le 819$) and Piper VITS on-chip SRAM tiling.
5. **Zero-Silent-Fallback & Strict Model Resolution**: Guaranteed fail-fast architecture without silent CPU degradation. Explicitly raises `AmbiguousModelMatchError` on multi-model ambiguity and rejects silent guessing.

---

## Two-Track Supply Chain Architecture

To simultaneously guarantee **100% intellectual property & source code protection** and **frictionless, tokenless binary deployment**, AMEVA enforces a strict Two-Track Supply Chain Architecture:

```
[Track 1: Local Ground Truth & Offline Staging]
├── dev/ameva/ameva-runtime/releases/ (Consolidated 6-Modality SOTA Binaries)
└── dev/termux/termux-*/releases/     (Domain-Specific Compiled Binaries)
                     │
                     ▼ Synchronized via Git & Automated Release Workflows
[Track 2: Remote Public GitHub Releases  Cloud & CDN]
├── uno-km/ameva-runtime-releases    (Public Unified Binary Hub: Releases v2.7.9)
└── uno-km/termux-* (diffusion, stt) (Domain-Specific Public GitHub Releases)
```

1. **Core Source Code (`uno-km/ameva-runtime`)**: Kept **100% Private**. Architectural proprietary logic, HAL implementations, and optimization kernels are permanently protected.
2. **Public Binary Distribution Hub (`uno-km/ameva-runtime-releases`)**: **Public Registry**. Hosts verified precompiled binaries, hardware acceleration compute shaders, and wheels for anonymous 1-Click installation without requiring personal GitHub tokens.

---

## Installation & Engine Auto-Provisioning

### 1. Python SDK Installation
```bash
# Direct install from public release wheel
pip install --upgrade "https://github.com/uno-km/ameva-runtime-releases/releases/download/v2.7.9/ameva_runtime-2.7.9-py3-none-any.whl"

# Or standard PyPI install
pip install ameva-runtime
```

### 2. 1-Click Native Engine Provisioning
The `NativeAssetManager` automatically retrieves cryptographically verified native engine binaries from the public registry and establishes system links in `$PREFIX/bin`:

```bash
# Provision all hardware engines at once
ameva install --all

# Or provision specific engine modalities individually
ameva install --modality diffusion   # Stable Diffusion Vulkan Engine (sd-cli)
ameva install --modality stt         # Whisper STT Vulkan Engine (whisper-cli)
ameva install --modality llm         # LLaMA C++ Vulkan Engine (llama-cli)
ameva install --modality tts         # Sherpa-NCNN Neural Voice Engine
ameva install --modality bitnet      # BitNet 1.58-bit Acceleration Engine

# Validate system readiness with 12-stage hardware diagnostic
ameva doctor
```

---

## Quickstart

### Python SDK
```python
import ameva_runtime as ameva
from ameva_runtime import vulkan

# 1. Inspect on-device silicon topology
profile = ameva.detect_hardware()
print(f"SoC: {profile.soc_name} | GPU: {profile.gpu_vendor}")

# 2. Run hardware self-test
doc = vulkan.Doctor()
report = doc.run_self_test()
print(f"GPU: {report.device_name} (Passed: {report.passed_stages}/{report.total_stages})")
```

### Node.js / TypeScript
```typescript
import { Doctor, createContext } from '@ameva/runtime';

const doc = new Doctor();
const report = await doc.runSelfTest();
console.log(`Vulkan GPU: ${report.deviceName}`);
```

---

## Official Documentation & Benchmarks
- [Official Architecture & API Reference](https://uno-km.vercel.app/lib/vulkan/)
- [Ecosystem Metrics & Registry Stats](https://uno-km.vercel.app/foundation/metrics)
- [AMEVA Open-Source Foundation Portal](https://uno-km.vercel.app/foundation/index.html)

---

## License
Licensed under the Apache-2.0 License. Copyright (c) 2026 Eunho Kim ([@uno-km](https://github.com/uno-km)).
