Metadata-Version: 2.4
Name: vayrm
Version: 0.1.8
Summary: Adaptive Compute Runtime
Requires-Python: >=3.10
Description-Content-Type: text/markdown
Requires-Dist: psutil>=5.9.0
Requires-Dist: cryptography>=41.0.0
Requires-Dist: PyYAML>=6.0.1
Requires-Dist: aiohttp>=3.9.0
Requires-Dist: qrcode>=7.4
Requires-Dist: pillow>=10.0.0
Requires-Dist: python-dotenv>=1.0.0
Requires-Dist: rich>=13.0.0
Requires-Dist: requests>=2.31.0
Provides-Extra: macos
Provides-Extra: windows
Provides-Extra: linux
Provides-Extra: dev
Requires-Dist: pytest>=7.0.0; extra == "dev"
Requires-Dist: pytest-cov>=4.0.0; extra == "dev"
Requires-Dist: mypy>=1.0.0; extra == "dev"
Requires-Dist: ruff>=0.0.260; extra == "dev"

# PALADIN Adaptive Compute Runtime (ACR)

> **PALADIN ACR** is a heterogeneous, adaptive distributed-compute
> runtime combining local NodeBrain intelligence, capability-aware
> scheduling, adaptive workload reallocation, direct node-to-node
> training communication, persistent lifecycle/storage management,
> observability, and Tailscale-based production networking.

## 1. System Architecture

PALADIN is intentionally separated into control, compute, training-data,
persistence, and network/security planes.

``` mermaid
flowchart TB
    TS["TAILSCALE<br/>Production Network / Security Layer"]

    ROOT["ROOT<br/>Control Plane"]
    RB["RootBrain<br/>Cluster State"]
    SCH["Adaptive Scheduler"]
    TC["Training Coordinator"]
    DASH["Dashboard / API"]
    EP["Dynamic Endpoint Registry"]

    A["NODE A<br/>NodeBrain"]
    B["NODE B<br/>NodeBrain"]
    C["NODE C<br/>NodeBrain"]

    P2P["NODE ↔ NODE<br/>Training Data Plane"]
    STORAGE["Storage API<br/>Persistence Plane"]
    OMV["Persistent Storage / OMV"]

    TS --> ROOT
    TS --> A
    TS --> B
    TS --> C
    TS --> STORAGE

    ROOT --> RB
    RB --> SCH
    SCH --> TC
    ROOT --> DASH
    ROOT --> EP

    ROOT --> A
    ROOT --> B
    ROOT --> C

    A <--> P2P
    B <--> P2P
    C <--> P2P

    ROOT --> STORAGE
    STORAGE --> OMV
```

### Architectural planes

| Plane            | Components                              | Responsibility                                                 |
|------------------|-----------------------------------------|----------------------------------------------------------------|
| Control          | Root, RootBrain, Scheduler, Coordinator | Membership, scheduling, allocation, coordination               |
| Compute          | NodeBrain, local workers                | Hardware/runtime observation and execution                     |
| Training data    | Node ↔ Node                             | Gradients, tensors, model synchronization                      |
| Persistence      | Storage API, OMV/backend                | Datasets, models, checkpoints, artifacts, experiments, results |
| Observability    | Dashboard, telemetry                    | Runtime state, health, allocations, endpoints                  |
| Network/security | Tailscale                               | Private production connectivity and network boundary           |

**Root is not the training-data relay.** High-frequency
gradients/tensors/model traffic stays on the direct Node ↔ Node path.

------------------------------------------------------------------------

## 2. Core Runtime Loop

``` mermaid
flowchart LR
    H["Hardware / Runtime"] --> N["NodeBrain"]
    N --> CV["CapabilityVector"]
    W["Workload"] --> WS["WorkloadSignature"]

    CV --> M["Capability Matching"]
    WS --> M
    M --> S["Adaptive Scheduler"]
    S --> A["Allocation"]
    A --> E["Execution"]

    E --> O["Observed Performance"]
    O --> D["Degradation / Straggler Detection"]
    D --> C["AdaptiveTrainingController"]
    C --> G["Candidate Allocations"]
    G --> U["Utility / Communication Cost"]
    U --> S
```

The runtime continuously follows:

**Observe → Represent → Match → Allocate → Execute → Measure → Detect →
Adapt → Reallocate**

This turns scheduling into a feedback-driven runtime process rather than
a one-time placement decision.

------------------------------------------------------------------------

# 3. Design Principles

### Root = control plane

The Root performs cluster-level scheduling, allocation, coordination,
endpoint publication, and dashboard/API exposure.

### NodeBrain = local intelligence

Each worker observes its own hardware, runtime, health, capabilities,
and execution behavior.

### Node ↔ Node = training data plane

Workers communicate directly for high-frequency gradients, tensors,
model synchronization, and training payloads.

### Storage = persistence plane

Datasets, models, checkpoints, artifacts, experiments, and results are
persisted through the Storage API. Storage is not the training hot path.

### Tailscale = production network/security layer

Production communication uses Tailscale. Machine-specific Tailscale IPs
are not hard-coded; nodes/services dynamically discover their current
Tailscale address.

### Dynamic endpoints

Root, Dashboard, and Storage have independently allocated runtime ports.

``` text
Current Tailscale IP
        +
Dynamic service port
        ↓
Advertised endpoint
```

### Level-1 hierarchy

The current operational hierarchy is:

``` text
Root
 ├── Node
 ├── Node
 └── Node
```

Sub-root/hierarchical scheduling is reserved for future scope.

------------------------------------------------------------------------

# 4. Phase 1–2 — Architecture & NodeBrain Foundation

### Implemented

- Repository and architecture audit.
- Clear boundaries between PALADIN communication, ACR runtime, and Vayrm
  integration.
- NodeBrain local intelligence.
- Cross-platform hardware/runtime discovery.
- CPU/GPU/memory/health telemetry.
- Normalized node capability representation.
- Lightweight continuous monitoring.
- Capability-based node classification.

``` mermaid
flowchart LR
    NB["NodeBrain"] --> D["Discovery"]
    NB --> CAP["Capabilities"]
    NB --> TEL["Telemetry"]
    NB --> HEALTH["Health"]
    NB --> SELF["Local Self Model"]
    CAP --> ASSESS["Local Assessment"]
    TEL --> ASSESS
    HEALTH --> ASSESS
    ASSESS --> STATE["Node State"]
```

Platform-specific implementations are separated so Linux, macOS, and
Windows can expose a common capability model.

------------------------------------------------------------------------

# 5. Phase 3 — Workload & Capability Modeling

Two complementary quantitative representations are used.

``` text
Node
 ├── CPU
 ├── GPU / accelerator
 ├── Memory
 ├── Storage
 ├── Network
 └── Runtime capability
        ↓
CapabilityVector
```

``` text
Workload
 ├── Compute requirements
 ├── Memory requirements
 ├── Accelerator requirements
 ├── I/O characteristics
 └── Communication sensitivity
        ↓
WorkloadSignature
```

Matching:

``` mermaid
flowchart LR
    C["CapabilityVector"] --> MATCH["Capability Matcher"]
    W["WorkloadSignature"] --> MATCH
    MATCH --> OK["Compatible Nodes"]
    MATCH --> NO["Rejected / Incompatible Nodes"]
    OK --> SCH["Scheduler"]
```

This prevents heterogeneous machines from being treated as
interchangeable resources.

------------------------------------------------------------------------

# 6. Phase 4 — Binary Communication Foundation

Implemented communication primitives include:

- binary telemetry representation
- compact framing
- typed message/frame categories
- telemetry/control/gradient payload types
- bounded payload handling
- malformed-frame protection
- extensible framing

Conceptually:

``` text
┌──────────────┬──────────────┬──────────────┬───────────────┐
│ Type / Magic │ Version      │ Length       │ Payload       │
└──────────────┴──────────────┴──────────────┴───────────────┘
```

Bounded framing and partial-read handling provide predictable behavior
over stream-oriented transports.

------------------------------------------------------------------------

# 7. Phase 5–6 — Runtime & Scheduler Integration

``` mermaid
flowchart LR
    NR["Node Reports"] --> RB["RootBrain"]
    RB --> CS["Cluster State"]
    CS --> SCH["Adaptive Scheduler"]
    SCH --> DEC["Scheduling Decision"]
    DEC --> RF["Resource / Allocation Layer"]
    RF --> EXEC["Runtime Execution"]
    EXEC --> OBS["Observed Result"]
    OBS --> RB
```

Implemented concepts include:

- AdaptiveScheduler
- capability-gated scheduling
- scheduler/node-event integration
- RootBrain cluster state
- runtime execution mapping
- resource-aware filtering
- scheduling decision propagation

The core boundary is:

``` text
Scheduler decides
      ↓
Runtime executes
      ↓
Telemetry reports
      ↓
Scheduler adapts
```

------------------------------------------------------------------------

# 8. Phase 7 — Distributed Training Foundation

PALADIN includes a real PyTorch distributed-training path.

``` mermaid
sequenceDiagram
    participant R as Root
    participant A as Node A
    participant B as Node B
    participant C as Node C

    R->>A: Allocation / shard assignment
    R->>B: Allocation / shard assignment
    R->>C: Allocation / shard assignment

    A->>A: Forward + Backward
    B->>B: Forward + Backward
    C->>C: Forward + Backward

    A->>R: Gradient
    B->>R: Gradient
    C->>R: Gradient

    R->>R: Validate + Aggregate
    R->>R: Optimizer Step

    R->>A: Updated Model
    R->>B: Updated Model
    R->>C: Updated Model
```

Implemented:

- multi-worker shared-model training
- real forward/backward execution
- separate worker data shards
- gradient extraction
- shape/dtype/device validation
- model version and global-step management
- duplicate/stale gradient rejection
- optimizer-step ownership
- gradient aggregation
- model synchronization
- multi-worker validation

State progression:

``` text
Model N
  ↓
Worker gradients
  ↓
Validation
  ↓
Aggregation
  ↓
Optimizer step
  ↓
Model N+1
  ↓
Synchronization
```

------------------------------------------------------------------------

# 9. Phase 8 — Direct Training Transport

``` mermaid
flowchart LR
    ENGINE["Training Engine"]
    IF["TrainingTransport"]

    LOCAL["Local Test Transport"]
    TCP["TCP Transport"]
    TAIL["Tailscale TCP"]

    ENGINE --> IF
    IF --> LOCAL
    IF --> TCP
    IF --> TAIL
    TAIL --> PEER["Peer Node"]
```

Implemented:

- transport abstraction
- local loopback validation
- TCP transport
- Tailscale TCP transport
- binary tensor/gradient exchange
- model synchronization
- framing and partial-read reconstruction
- bounded queues/backpressure
- reconnect/backoff
- concurrent multi-worker validation

Production path:

``` text
Node A → Tailscale → Node B
        gradients / tensors / model data
```

The Root remains the control plane rather than becoming a mandatory
relay for every training payload.

------------------------------------------------------------------------

# 10. Phase 9 — Degradation Detection

Runtime adaptation is driven by observed execution behavior.

Tracked information includes:

- iteration time
- throughput
- compute time
- communication time
- observed performance
- expected performance
- node health

``` mermaid
flowchart LR
    RAW["Runtime Metrics"] --> EWMA["EWMA Performance Estimate"]
    EWMA --> EXPECT["Expected vs Observed"]
    EXPECT --> STRAG["Straggler / Degradation Detection"]
    STRAG --> STATE["Degradation / Recovery State"]
    STATE --> CTRL["Adaptive Controller"]
```

Implemented:

- `NodePerformanceState`
- EWMA performance estimation
- expected-vs-observed throughput
- iteration-time tracking
- straggler detection
- degradation/recovery states
- event-driven detection
- anti-thrashing state transitions

The adaptation model is not reduced to a single hard-coded utilization
threshold. It considers sustained observed behavior and the relationship
between expected and actual performance.

------------------------------------------------------------------------

# 11. Phase 10 — Adaptive Workload Reallocation

``` mermaid
flowchart TD
    OBS["Observed Performance"] --> DET["Degradation Detection"]
    DET --> GEN["Candidate Allocation Generation"]

    GEN --> A["Candidate A"]
    GEN --> B["Candidate B"]
    GEN --> C["Candidate C"]

    A --> U["Utility Evaluation"]
    B --> U
    C --> U

    U --> COMM["Communication Cost"]
    U --> GOOD["Expected Cluster Goodput"]
    U --> CAP["Capability / Capacity"]

    COMM --> DEC["Allocation Decision"]
    GOOD --> DEC
    CAP --> DEC

    DEC --> VER["Allocation Epoch / Version"]
    VER --> APPLY["Apply Reallocation"]
    APPLY --> TC["Training Coordinator"]
    TC --> EXEC["Workers"]
    EXEC --> OBS
```

Implemented:

- `AdaptiveTrainingController`
- dynamic candidate allocation generation
- utility-based allocation selection
- communication-aware decisions
- allocation epochs/versioning
- hysteresis/cooldown
- dynamic workload redistribution
- node disappearance handling
- node recovery handling
- TrainingCoordinator integration
- heterogeneous multi-worker adaptive validation

The intended behavior is:

``` text
Detect persistent imbalance
        ↓
Generate alternatives
        ↓
Estimate utility/cost
        ↓
Version allocation
        ↓
Apply
        ↓
Cooldown
        ↓
Observe again
```

This reduces oscillation and prevents stale allocation decisions from
being applied indefinitely.

------------------------------------------------------------------------

# 12. Lifecycle & Persistent Storage

The lifecycle layer coordinates:

``` mermaid
flowchart TB
    LIFE["Training Lifecycle"]
    LIFE --> DATA["Dataset Manager"]
    LIFE --> MODEL["Model Manager"]
    LIFE --> HP["Hyperparameters"]
    LIFE --> EXP["Experiment Manager"]
    LIFE --> CK["Checkpoint Lifecycle"]
    LIFE --> RES["Result Manager"]

    DATA --> API["Storage API"]
    MODEL --> API
    EXP --> API
    CK --> API
    RES --> API

    API --> DS["Datasets"]
    API --> MO["Models"]
    API --> CH["Checkpoints"]
    API --> AR["Artifacts"]
    API --> EX["Experiments"]
    API --> RS["Results"]
```

Supported lifecycle responsibilities include:

- dataset management
- model management
- experiment tracking
- result management
- checkpoint creation
- checkpoint metadata
- training-state restoration
- lifecycle coordination

Checkpoint metadata can include:

``` text
model version
global step
epoch
dataset
allocation
optimizer state
hyperparameters
world/node information
metrics
```

Model-loading paths include formats such as:

- Safetensors
- `.pt`
- `.pth`
- `.h5`
- `.hdf5`
- `.bin`

------------------------------------------------------------------------

# 13. Storage API & Dynamic Storage Endpoint

Storage is a separate plane:

``` mermaid
flowchart LR
    ROOT["Root"]
    CLIENT["Dashboard / Client"]
    API["Storage API"]
    PERSIST["Persistent Storage / OMV"]

    ROOT --> API
    CLIENT --> API
    API --> PERSIST
```

Namespaces are organized around:

``` text
/datasets
/models
/checkpoints
/artifacts
/experiments
/results
```

Supported operation categories include:

``` text
GET / PUT / DELETE / HEAD / LIST
```

The Storage API receives its own dynamically allocated port.

### Endpoint model

``` text
Tailscale IP
   │
   ├── Root dynamic port
   ├── Dashboard dynamic port
   └── Storage dynamic port
```

No service should depend on a machine-specific hard-coded Tailscale IP.

------------------------------------------------------------------------

# 14. Dashboard & Observability

The Dashboard is an observability/control surface over authoritative
runtime state.

It is intended to expose:

- node membership
- node capabilities
- health
- cluster state
- current allocations
- scheduler state
- telemetry
- service endpoints
- storage status

The scheduler/cluster state remains authoritative; the dashboard should
not invent duplicate allocation state.

``` text
Root / Scheduler
       ↓
Authoritative runtime state
       ↓
Dashboard API
       ↓
Dashboard
```

------------------------------------------------------------------------

# 15. Dynamic Endpoint & Network Architecture

``` mermaid
flowchart TB
    IP["Current Tailscale IP"]

    ROOTP["Dynamic Root Port"]
    DASHP["Dynamic Dashboard Port"]
    STORP["Dynamic Storage Port"]

    ROOT["Root"]
    DASH["Dashboard"]
    STORAGE["Storage API"]

    IP --> ROOTP --> ROOT
    IP --> DASHP --> DASH
    IP --> STORP --> STORAGE
```

Important invariant:

``` text
ROOT_PORT != DASHBOARD_PORT != STORAGE_PORT
```

Endpoint discovery is runtime-based:

``` text
Discover local Tailscale IP
        ↓
Bind service to appropriate interface
        ↓
Obtain runtime port
        ↓
Health-check
        ↓
Advertise endpoint
```

Hard-coded development machine addresses are not part of the production
architecture.

------------------------------------------------------------------------

# 16. Node Failure, Degradation & Recovery

``` mermaid
stateDiagram-v2
    [*] --> HEALTHY
    HEALTHY --> DEGRADED: sustained performance deviation
    DEGRADED --> HEALTHY: performance recovery
    DEGRADED --> UNAVAILABLE: node failure
    UNAVAILABLE --> RECOVERING: node returns
    RECOVERING --> HEALTHY: reconciliation complete
    DEGRADED --> REALLOCATION: adaptive controller
    UNAVAILABLE --> REALLOCATION: node removed
    REALLOCATION --> HEALTHY: new allocation active
```

The runtime accounts for:

- temporary degradation
- sustained straggling
- communication problems
- node disappearance
- node recovery
- allocation version changes

Hysteresis/cooldown prevents immediate repeated reallocations.

------------------------------------------------------------------------

# 17. Resource Fabric

The resource-fabric layer represents and manages heterogeneous resource
capacity.

``` text
Node
 │
 ├── CPU
 ├── GPU / accelerator
 ├── Memory
 ├── Network
 ├── Storage
 ├── Health / thermal state
 └── Runtime capability
       ↓
Capability / Capacity Model
       ↓
Filtering / Scoring
       ↓
Offer / Claim / Lease
       ↓
Allocation
```

Relevant mechanisms include:

- capacity models
- normalization
- filtering
- scoring
- resource offers
- resource claims
- quotas
- leases
- health
- preemption
- reconciliation
- cooldown

------------------------------------------------------------------------

# 18. Communication & Security Boundaries

``` mermaid
flowchart TB
    TS["TAILSCALE"]

    ROOT["Root"]
    N1["Node A"]
    N2["Node B"]
    N3["Node C"]
    ST["Storage API"]

    TS --- ROOT
    TS --- N1
    TS --- N2
    TS --- N3
    TS --- ST

    ROOT -. control .-> N1
    ROOT -. control .-> N2
    ROOT -. control .-> N3

    N1 <-->|training data| N2
    N2 <-->|training data| N3
    N1 <-->|training data| N3

    ROOT --> ST
```

Security/architecture invariants:

- Tailscale is mandatory for production networking.
- No machine-specific hard-coded Tailscale IPs.
- No insecure localhost/LAN production fallback.
- Control and training data paths are separated.
- Storage is separate from the training hot path.
- Binary framing is bounded.
- Network queues apply backpressure.
- Storage namespace/path protections are enforced.
- Root does not become an authoritative training/model database.

------------------------------------------------------------------------

# 19. Platform Architecture

``` mermaid
flowchart TB
    ACR["Common ACR Runtime"]

    L["Linux Provider"]
    M["macOS Provider"]
    W["Windows Provider"]

    ACR --> L
    ACR --> M
    ACR --> W

    L --> LH["Linux CPU / GPU / Runtime"]
    M --> MH["Apple Silicon / CPU / MPS"]
    W --> WH["Windows CPU / GPU / Runtime"]
```

Platform-specific hardware/runtime behavior is isolated from common
scheduling logic.

The target is one common capability model across heterogeneous
environments.

------------------------------------------------------------------------

# 20. Repository Structure

``` text
adaptive-runtime/
├── acr/
│   ├── cluster/
│   ├── communication/
│   ├── core/
│   ├── execution/
│   ├── integration/
│   ├── lifecycle/
│   ├── network/
│   ├── node/
│   ├── peers/
│   ├── platforms/
│   ├── registry/
│   ├── resource_fabric/
│   ├── resources/
│   ├── runtime/
│   ├── schedulers/
│   ├── scheduling/
│   ├── simulation/
│   ├── state/
│   ├── storage/
│   ├── telemetry/
│   ├── training/
│   ├── transfer/
│   ├── ui/
│   └── workloads/
├── ai/
├── apps/
│   ├── cli/
│   ├── node/
│   └── root/
│       └── api/
├── cluster/
├── core/
├── protocol/
├── vayrm/
├── docs/
└── tests/
```

### Major runtime areas

| Directory                 | Role                                                 |
|---------------------------|------------------------------------------------------|
| `acr/node/`               | NodeBrain and node-local intelligence                |
| `acr/resources/`          | Resource/capability state                            |
| `acr/workloads/`          | Workload signatures and descriptions                 |
| `acr/scheduling/`         | Adaptive allocation                                  |
| `acr/schedulers/`         | Scheduler strategies                                 |
| `acr/resource_fabric/`    | Resource capacity/allocation                         |
| `acr/training/`           | Distributed training                                 |
| `acr/training/transport/` | Training transport                                   |
| `acr/lifecycle/`          | Dataset/model/checkpoint/experiment/result lifecycle |
| `acr/storage/`            | Storage backends/API                                 |
| `acr/telemetry/`          | Runtime telemetry                                    |
| `acr/communication/`      | Binary communication/framing                         |
| `acr/platforms/`          | Linux/macOS/Windows support                          |
| `acr/ui/`                 | Runtime dashboards                                   |
| `acr/registry/`           | Cluster/runtime registries                           |
| `vayrm/`                  | Runtime/network integration facade                   |

------------------------------------------------------------------------

# 21. Architecture Integrity

The current design intentionally preserves these boundaries:

- `vayrm/` is treated as an integration/runtime dependency and is not
  rewritten as part of PALADIN feature work unless explicitly required.
- PALADIN core is not replaced by ACR.
- Tailscale is not bypassed for production networking.
- No local production fallback exposes runtime services insecurely.
- No Root-side authoritative training/model database is introduced.
- Training payloads are separated from Root control traffic.
- Storage is not the high-frequency training transport.
- Sub-root architecture remains future scope.
- Dynamic endpoint discovery replaces machine-specific IP assumptions.

------------------------------------------------------------------------

# 22. Validation Status

| Phase                                        | Status                      |
|----------------------------------------------|-----------------------------|
| Phase 1 — Architecture / audit               | **COMPLETE**                |
| Phase 2 — NodeBrain foundation               | **COMPLETE**                |
| Phase 3 — Capability / workload modeling     | **COMPLETE**                |
| Phase 4 — Binary communication               | **COMPLETE**                |
| Phase 5 — Runtime integration                | **COMPLETE where verified** |
| Phase 6 — Scheduler integration              | **COMPLETE where verified** |
| Phase 7 — Distributed training               | **VALIDATED**               |
| Phase 8 — Node ↔ Node TCP transport          | **VALIDATED**               |
| Phase 9 — Degradation detection              | **VALIDATED**               |
| Phase 10 — Adaptive reallocation             | **VALIDATED**               |
| Phase 11 — Physical heterogeneous validation | **NEXT**                    |

Previously validated development areas include:

- multi-worker training
- gradient validation/aggregation
- model synchronization
- adaptive allocation scenarios
- TCP training transport
- Tailscale transport abstraction
- dynamic endpoint discovery
- dynamic Storage API port allocation
- Storage API development GET/PUT
- lifecycle tests
- Vayrm compatibility tests

**Physical heterogeneous hardware validation remains the next major
validation step.**

------------------------------------------------------------------------

# 23. Phase 11 — Real Hardware Validation & Benchmarking

The next phase should turn runtime validation into reproducible research
evidence.

## Hardware matrix

A representative test matrix should include combinations such as:

``` text
macOS / Apple Silicon
Linux / NVIDIA
Windows / CPU or NVIDIA
```

Actual hardware should be recorded as experimental configuration rather
than hard-coded into the runtime.

## Baselines

Useful baselines include:

``` text
1. Static allocation
2. Round-robin allocation
3. Capability-aware allocation
4. Adaptive PALADIN allocation
```

## Metrics

Measure:

- iteration time
- samples/sec
- effective throughput
- compute time
- communication time
- synchronization time
- node utilization
- straggler duration
- degradation detection latency
- reallocation latency
- recovery time
- checkpoint overhead
- total training time
- allocation changes

## Controlled degradation experiment

``` mermaid
sequenceDiagram
    participant R as Root
    participant F as Fast Node
    participant S as Degraded Node

    R->>F: Initial allocation
    R->>S: Initial allocation

    F->>R: Performance telemetry
    S->>R: Performance telemetry

    Note over S: Controlled slowdown

    S->>R: Reduced observed throughput
    R->>R: EWMA / degradation detection
    R->>R: Candidate allocation generation
    R->>R: Utility + communication evaluation
    R->>F: Updated allocation
    R->>S: Updated allocation

    F->>R: New performance
    S->>R: New performance
```

The experiment should report the complete timeline and quantitative
changes, not only the final result.

------------------------------------------------------------------------

# 24. Research Context

PALADIN intersects with established systems research areas:

- introspective scheduling
- goodput-aware scheduling
- heterogeneous distributed training
- straggler mitigation
- adaptive resource allocation
- distributed training
- feedback-controlled runtime systems

Representative references include:

- **Gandiva**, OSDI 2018 — introspective cluster scheduling for deep
  learning.
- **Pollux**, OSDI 2021 — co-adaptive scheduling for goodput-oriented
  deep learning.
- **SADDLE**, 2025 — feedback-controlled heterogeneous distributed
  training.
- **PyTorch DistributedDataParallel** — established distributed-training
  infrastructure.

PALADIN should distinguish established mechanisms from the proposed
system-level contribution.

### Established concepts

``` text
EWMA
Straggler detection
Adaptive allocation
Gradient aggregation
Distributed training
Capability-aware scheduling
```

### Proposed integrated system contribution

``` text
Heterogeneous adaptive runtime architecture
        +
NodeBrain local intelligence
        +
Capability/workload representation
        +
Adaptive runtime reallocation
        +
Direct Node ↔ Node training plane
        +
Dynamic secure endpoints
        +
Persistent lifecycle/storage plane
```

Individual mechanisms should not be presented as novel merely because
they are implemented here. Novelty should be supported through
literature review and reproducible experimental comparison.

------------------------------------------------------------------------

# 25. Runtime Lifecycle

``` mermaid
flowchart TD
    START["Startup"]
    NET["Discover Tailscale"]
    EP["Allocate Dynamic Endpoints"]
    REG["Register Services / Nodes"]
    CAP["Collect Capabilities"]
    WORK["Describe Workload"]
    MATCH["Capability Matching"]
    ALLOC["Initial Allocation"]
    EXEC["Execute"]
    MON["Monitor"]
    DET["Detect Degradation"]
    ADAPT["Adaptive Reallocation"]
    CK["Checkpoint"]
    RES["Persist Results"]
    REC["Recovery / Reconciliation"]
    STOP["Shutdown"]

    START --> NET
    NET --> EP
    EP --> REG
    REG --> CAP
    CAP --> WORK
    WORK --> MATCH
    MATCH --> ALLOC
    ALLOC --> EXEC
    EXEC --> MON
    MON --> DET
    DET --> ADAPT
    ADAPT --> EXEC
    EXEC --> CK
    CK --> RES
    MON --> REC
    REC --> ALLOC
    RES --> STOP
```

------------------------------------------------------------------------

# 26. Final Architecture at a Glance

``` mermaid
flowchart TB
    TS["TAILSCALE<br/>Private Production Network"]

    subgraph CP["CONTROL PLANE"]
        ROOT["ROOT"]
        RB["RootBrain"]
        SCH["Adaptive Scheduler"]
        TC["Training Coordinator"]
        DASH["Dashboard"]
    end

    subgraph NODES["HETEROGENEOUS COMPUTE"]
        A["Node A<br/>NodeBrain"]
        B["Node B<br/>NodeBrain"]
        C["Node C<br/>NodeBrain"]
    end

    P2P["DIRECT NODE ↔ NODE<br/>Gradients • Tensors • Model Sync"]

    subgraph STORE["PERSISTENCE"]
        API["Storage API"]
        DB["Datasets / Models / Checkpoints / Artifacts / Results"]
    end

    TS --- CP
    TS --- NODES
    TS --- P2P
    TS --- STORE

    ROOT --> RB
    RB --> SCH
    SCH --> TC
    ROOT --> DASH

    ROOT --> A
    ROOT --> B
    ROOT --> C

    A <--> P2P
    B <--> P2P
    C <--> P2P

    ROOT --> API
    API --> DB
```

## One-sentence architecture

> **The Root decides, NodeBrain observes local conditions, heterogeneous
> nodes execute work, nodes exchange training data directly, Storage
> persists lifecycle artifacts, and observed runtime behavior
> continuously feeds back into adaptive scheduling decisions.**

------------------------------------------------------------------------

## Current Status

**Phases 1–10:** Implemented/validated where explicitly verified.

**Current focus:** **Phase 11 — real heterogeneous hardware validation
and baseline benchmarking.**

**Current hierarchy:** **Root → Node.**

**Training hot path:** **Node ↔ Node.**

**Persistence path:** **Storage API → persistent storage.**

**Production network:** **Tailscale.**

**Future scope:** hierarchical sub-roots, broader physical validation,
reproducible benchmark suite, and research evaluation.


# 14. Root Dashboard Architecture

The Root Dashboard is the operational view over the Root control plane. It does not become a second scheduler; it reads authoritative cluster/runtime state and presents it through the API/UI.

```mermaid
flowchart TB
    USER["Operator / Researcher"]

    subgraph DASH["ROOT DASHBOARD"]
        UI["Dashboard UI"]
        API["Dashboard API"]
        ENDPOINTS["Endpoint View"]
        NODES["Node View"]
        CLUSTER["Cluster View"]
        ALLOC["Allocation View"]
        SCHED["Scheduler View"]
        TEL["Telemetry View"]
        STORAGE["Storage View"]
    end

    subgraph ROOT["ROOT CONTROL PLANE"]
        RB["RootBrain"]
        SCH["Adaptive Scheduler"]
        TC["Training Coordinator"]
        STATE["Authoritative Cluster / Allocation State"]
        REG["Endpoint Registry"]
    end

    subgraph N["COMPUTE NODES"]
        NA["Node A / NodeBrain"]
        NB["Node B / NodeBrain"]
        NC["Node C / NodeBrain"]
    end

    subgraph STORE["PERSISTENCE"]
        SA["Storage API"]
        PS["Persistent Storage"]
    end

    USER --> UI
    UI --> API

    API --> ENDPOINTS
    API --> NODES
    API --> CLUSTER
    API --> ALLOC
    API --> SCHED
    API --> TEL
    API --> STORAGE

    ENDPOINTS --> REG
    NODES --> STATE
    CLUSTER --> STATE
    ALLOC --> STATE
    SCHED --> SCH
    TEL --> RB

    RB --> STATE
    SCH --> STATE
    TC --> STATE

    RB --> NA
    RB --> NB
    RB --> NC

    STORAGE --> SA
    SA --> PS
```

## Dashboard information flow

```mermaid
flowchart LR
    NODE["NodeBrain"] --> REPORT["Node Reports / Telemetry"]
    REPORT --> ROOT["Root / RootBrain"]

    ROOT --> STATE["Authoritative Runtime State"]

    STATE --> API["Dashboard API"]

    API --> NODES["Nodes"]
    API --> CLUSTER["Cluster"]
    API --> ALLOC["Allocations"]
    API --> SCHED["Scheduler"]
    API --> TEL["Telemetry"]
    API --> EP["Endpoints"]
    API --> ST["Storage"]

    NODES --> UI["Dashboard UI"]
    CLUSTER --> UI
    ALLOC --> UI
    SCHED --> UI
    TEL --> UI
    EP --> UI
    ST --> UI
```

### Root Dashboard responsibilities

```text
                 ROOT DASHBOARD
                       │
        ┌──────────────┼──────────────┐
        ↓              ↓              ↓
      Nodes         Cluster        Scheduler
        │              │              │
   capabilities      health       decisions
   health            state        allocations
   telemetry         membership   adaptation
        │              │              │
        └──────────────┼──────────────┘
                       ↓
                  Operator View
```

The dashboard can expose:

- current Root endpoint
- Dashboard endpoint
- Storage endpoint
- registered nodes
- node capabilities
- node health
- current allocations
- allocation epochs/versions
- scheduler state
- training state
- telemetry
- storage availability
- cluster-level runtime state

### Dashboard authority boundary

```text
NodeBrain
    ↓
Node report
    ↓
RootBrain / Cluster State
    ↓
Scheduler / Allocation State
    ↓
Dashboard API
    ↓
Dashboard UI
```

The Dashboard **observes and presents authoritative state**. It does not create an independent allocation database.

---



## Root Dashboard — Mermaid recreation of the terminal dashboard

> This diagram is deliberately a **dashboard/wireframe reconstruction**, not just an architecture graph. It mirrors the major regions of the attached Root terminal UI: header → live-worker table → three-column lower dashboard → events/pipeline/resource/prediction panels.

```mermaid
flowchart TB
    %% =========================================================
    %% HEADER
    %% =========================================================
    H["<b>ACR PALADIN ROOT</b><br/>node_id = ROOT &nbsp; | &nbsp; platform = macOS &nbsp; | &nbsp; PALADIN: RUNNING &nbsp; | &nbsp; workers = 3"]
    HS["<b>ROOT SELF</b>"]

    H --> HS

    %% =========================================================
    %% LIVE WORKERS - dashboard table
    %% =========================================================
    subgraph LIVE["LIVE WORKERS"]
        direction TB

        WH["<b>PEER</b> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <b>PLATFORM</b> &nbsp;&nbsp;&nbsp; <b>STATUS</b> &nbsp;&nbsp; <b>FRESH</b> &nbsp;&nbsp; <b>AGE</b> &nbsp;&nbsp; <b>CPU</b> &nbsp;&nbsp; <b>MEM</b> &nbsp;&nbsp; <b>GPU</b> &nbsp;&nbsp; <b>PREDICTION</b> &nbsp;&nbsp; <b>RISK</b> &nbsp;&nbsp; <b>SCHEDULER</b>"]

        W1["<b>WORKER A</b><br/>macOS<br/>ONLINE • FRESH<br/>CPU / MEM / GPU telemetry<br/>STABLE • SAFE • ELIGIBLE"]

        W2["<b>WORKER B</b><br/>Windows<br/>ONLINE • FRESH<br/>CPU / MEM / GPU telemetry<br/>STABLE • SAFE • ELIGIBLE"]

        W3["<b>WORKER C</b><br/>Windows<br/>ONLINE • FRESH<br/>CPU / MEM / GPU telemetry<br/>STABLE • SAFE • ELIGIBLE"]

        WH --- W1
        WH --- W2
        WH --- W3
    end

    HS --> LIVE

    %% =========================================================
    %% LOWER DASHBOARD
    %% =========================================================
    subgraph DASH["ROOT DASHBOARD"]
        direction LR

        %% LEFT COLUMN
        subgraph LEFT["LEFT COLUMN"]
            direction TB

            SUB["<b>SUBROOT HIERARCHY</b><br/><br/>No hierarchical<br/>subroots discovered"]

            TOPO["<b>PEER / P2P TOPOLOGY</b><br/><br/><b>LOCAL NODE</b><br/>node_id : ROOT<br/>role : ROOT<br/>tailscale_ip : dynamic<br/>endpoint : dynamic<br/><br/><b>PEERS</b><br/>Worker A • Worker B • Worker C"]

            EVENTS["<b>EVENTS</b><br/><br/>TELEM • PeerUpdated<br/>health_ok = true<br/>COMM_ACK<br/>PeerHealthMessage<br/>RootAnnouncement<br/>PeerDirectoryMessage"]

            SUB --- TOPO
            TOPO --- EVENTS
        end

        %% CENTER COLUMN
        subgraph CENTER["CENTER COLUMN"]
            direction TB

            ID["<b>IDENTITY & CONSTRAINTS</b><br/><br/>NODE: selected worker<br/>OS / platform<br/><br/><b>HARD CONSTRAINTS</b><br/>none / reported constraints<br/><br/><b>SOFT CONSTRAINTS</b><br/>none / reported constraints"]

            FABRIC["<b>RESOURCE FABRIC</b><br/><br/>Worker A → CPU / MEM / confidence / freshness<br/>Worker B → CPU / MEM / confidence / freshness<br/>Worker C → CPU / MEM / confidence / freshness<br/><br/><b>CLUSTER RESOURCE STATE</b>"]

            PIPE["<b>PIPELINE VALIDATION</b><br/><br/>PASS  PALADIN connected<br/>PASS  Endpoint valid<br/>PASS  Telemetry fresh<br/>PASS  NodeKnowledge updated<br/>PASS  Resource Fabric updated"]

            ID --- FABRIC
            FABRIC --- PIPE
        end

        %% RIGHT COLUMN
        subgraph RIGHT["RIGHT COLUMN"]
            direction TB

            PRED["<b>PREDICTIVE PREEMPTION — CLUSTER</b><br/><br/><b>ALL NODES NORMAL</b><br/><br/>DETAIL SOURCE: NodeBrain<br/>STATUS: NORMAL<br/>RESOURCE: CPU / MEMORY / GPU<br/>CURRENT: utilization<br/><br/>TRAJECTORY: STABLE / VOLATILE<br/>CONFIDENCE: prediction confidence<br/>FRESHNESS: telemetry freshness<br/>TIME TO LIMIT: predicted horizon"]

            HARNESS["<b>EXPERIMENT HARNESS</b><br/><br/>Decision Engine<br/>Experiment state<br/>Allocation / scheduling decision<br/>No decision engine attached"]

            PRED --- HARNESS
        end
    end

    LIVE --> DASH

    %% =========================================================
    %% INFORMATION FLOW INTO PANELS
    %% =========================================================
    W1 -. "telemetry" .-> FABRIC
    W2 -. "telemetry" .-> FABRIC
    W3 -. "telemetry" .-> FABRIC

    W1 -. "prediction" .-> PRED
    W2 -. "prediction" .-> PRED
    W3 -. "prediction" .-> PRED

    W1 -. "peer health" .-> TOPO
    W2 -. "peer health" .-> TOPO
    W3 -. "peer health" .-> TOPO

    W1 -. "events" .-> EVENTS
    W2 -. "events" .-> EVENTS
    W3 -. "events" .-> EVENTS

    %% =========================================================
    %% STYLING — terminal-like visual grouping
    %% =========================================================
    classDef header fill:#171717,stroke:#7c3aed,stroke-width:2px,color:#ffffff;
    classDef worker fill:#102018,stroke:#22d3ee,stroke-width:1.5px,color:#ffffff;
    classDef panel fill:#171717,stroke:#777777,stroke-width:1.5px,color:#ffffff;
    classDef identity fill:#171717,stroke:#d946ef,stroke-width:2px,color:#ffffff;
    classDef resource fill:#171717,stroke:#7c3aed,stroke-width:2px,color:#ffffff;
    classDef prediction fill:#171717,stroke:#a3e635,stroke-width:2px,color:#ffffff;
    classDef validation fill:#171717,stroke:#eab308,stroke-width:2px,color:#ffffff;
    classDef event fill:#171717,stroke:#64748b,stroke-width:1.5px,color:#ffffff;

    class H,HS header;
    class WH,W1,W2,W3 worker;
    class SUB,TOPO panel;
    class EVENTS event;
    class ID identity;
    class FABRIC resource;
    class PRED prediction;
    class PIPE,HARNESS validation;
```

### Dashboard layout correspondence

```mermaid
flowchart TB
    A["ROOT HEADER<br/>ACR PALADIN ROOT • identity • platform • runtime"] 
    B["LIVE WORKERS<br/>worker table"]

    subgraph ROW["LOWER DASHBOARD ROW"]
        direction LR
        C["LEFT<br/>Subroot Hierarchy<br/>Peer / P2P Topology<br/>Events"]
        D["CENTER<br/>Identity & Constraints<br/>Resource Fabric<br/>Pipeline Validation"]
        E["RIGHT<br/>Predictive Preemption — Cluster<br/>Experiment Harness"]
    end

    A --> B --> ROW
```

The important distinction is that this Mermaid is **not replacing the dashboard with a generic architecture diagram**. It documents the dashboard as a visual operator surface: the top worker table and the lower panel layout are represented explicitly, while the dotted links show where NodeBrain telemetry, predictions, peer state and events populate those panels.

# 15. NodeBrain Architecture

NodeBrain is the local intelligence component running on every compute node.

Its responsibility is to continuously understand:

- what hardware exists
- what resources are currently available
- what the node is capable of
- how the node is performing
- whether the node is healthy
- what the node observes about peers/runtime state
- what local execution is doing

```mermaid
flowchart TB
    NB["NODEBRAIN"]

    subgraph DISC["DISCOVERY"]
        HW["Hardware Discovery"]
        RT["Runtime Discovery"]
        OS["Platform Provider"]
    end

    subgraph MODEL["LOCAL MODEL"]
        SELF["Self Model"]
        CAP["Capability Vector"]
        SNAP["Resource Snapshot"]
        STATE["Node State"]
    end

    subgraph OBS["OBSERVATION"]
        TEL["Telemetry"]
        PERF["Performance"]
        HEALTH["Health"]
        PEER["Peer View / Peer Health"]
    end

    subgraph DEC["LOCAL ASSESSMENT"]
        ASSESS["Assessment"]
        TEMP["Temporal State"]
        MACHINE["State Machine"]
    end

    subgraph PUB["RUNTIME OUTPUT"]
        PUBSTATE["State Publisher"]
        REPORT["Node Reports"]
        EXEC["Local Execution / Workload"]
    end

    NB --> HW
    NB --> RT
    NB --> OS

    HW --> CAP
    RT --> SELF
    OS --> SELF

    CAP --> SNAP
    SELF --> SNAP

    SNAP --> ASSESS
    TEL --> ASSESS
    PERF --> ASSESS
    HEALTH --> ASSESS
    PEER --> ASSESS

    ASSESS --> TEMP
    TEMP --> MACHINE
    MACHINE --> STATE

    STATE --> PUBSTATE
    PUBSTATE --> REPORT
    STATE --> EXEC

    EXEC --> PERF
    EXEC --> TEL
    EXEC --> HEALTH
```

## NodeBrain internal feedback loop

```mermaid
flowchart LR
    HARDWARE["Hardware"] --> DISC["Discovery"]
    DISC --> CAP["Capability Model"]

    RUNTIME["Runtime"] --> SNAP["Resource Snapshot"]
    CAP --> SNAP

    EXEC["Local Workload"] --> METRIC["Observed Metrics"]
    SNAP --> EXEC

    METRIC --> PERF["Performance State"]
    PERF --> ASSESS["Node Assessment"]

    CAP --> ASSESS
    SNAP --> ASSESS
    ASSESS --> STATE["Node State"]

    STATE --> REPORT["Node Report"]
    REPORT --> ROOT["Root"]

    ROOT --> ALLOC["Allocation / Work Assignment"]
    ALLOC --> EXEC
```

## NodeBrain responsibilities by layer

| Layer | NodeBrain responsibility |
|---|---|
| Discovery | Discover platform, hardware, runtime and accelerator capabilities |
| Capability | Build normalized capability representation |
| Resource state | Track current CPU/GPU/memory/network/storage state |
| Telemetry | Collect lightweight runtime observations |
| Performance | Track iteration/throughput/performance information |
| Health | Detect local and peer health conditions |
| Assessment | Combine capability + state + performance evidence |
| State machine | Represent node lifecycle/degradation/recovery state |
| Publishing | Send relevant state to Root |
| Execution awareness | Understand local workload execution and feedback |

---


## Node-side NodeBrain — runtime/telemetry reconstruction

The node-side view is represented from the NodeBrain/communication trace: each worker maintains local resource intelligence, publishes health/prediction knowledge, receives the Root announcement and peer directory, and exchanges health information directly with other workers.

```mermaid
flowchart TB
    NODE["COMPUTE NODE<br/>Windows / macOS / Linux"]

    subgraph NB["NODEBRAIN — LOCAL INTELLIGENCE"]
        direction TB

        DISC["Hardware & Runtime Discovery<br/>CPU • Memory • GPU/accelerator • OS • runtime"]
        CAP["Capability Model<br/>capacity • device type • execution capability"]
        SNAP["Resource Snapshot<br/>current CPU • memory • GPU"]
        PRED["Predictive Resource Assessment<br/>predicted utilization • trajectory • confidence<br/>time-to-threshold • predicted shortage • urgency"]
        HEALTH["Node Health / Readiness<br/>communication status • resource health • resource pressure"]
        KNOW["NodeKnowledge<br/>peer/runtime observations • endpoint • confidence • freshness • generation"]
        PUB["State Publisher<br/>Root reporting + peer health broadcast"]

        DISC --> CAP
        CAP --> SNAP
        SNAP --> PRED
        SNAP --> HEALTH
        PRED --> KNOW
        HEALTH --> KNOW
        KNOW --> PUB
    end

    subgraph ROOTLINK["ROOT CONTROL / COORDINATION"]
        RA["RootAnnouncement<br/>root_node_id • root_endpoint • generation"]
        PD["PeerDirectoryMessage<br/>worker endpoints • status • endpoint validity"]
        ROOT["Root / RootBrain"]
    end

    subgraph P2P["DIRECT NODE ↔ NODE P2P"]
        PA["Peer A"]
        PB["Peer B"]
        PC["Peer C"]
        PH["PeerHealthMessage<br/>resource health • resource pressure<br/>knowledge_dict • confidence • freshness"]
    end

    NODE --> NB

    ROOT --> RA
    RA --> KNOW
    ROOT --> PD
    PD --> KNOW

    PUB --> ROOT
    PUB --> PH
    PH --> PA
    PH --> PB
    PH --> PC

    PA --> PH
    PB --> PH
    PC --> PH
```

### NodeBrain resource-prediction loop

The trace shows NodeBrain repeatedly producing a local assessment containing **current utilization, predicted utilization, trajectory, confidence, freshness, time-to-threshold, predicted-shortage state and urgency**.

```mermaid
flowchart LR
    S["Resource Snapshot"] --> E["Prediction / Estimation"]
    E --> T["Trajectory"]
    E --> C["Confidence"]
    E --> F["Freshness"]
    E --> L["Time to Threshold"]
    E --> PS["Predicted Shortage"]
    E --> U["Urgency"]

    T --> K["NodeKnowledge"]
    C --> K
    F --> K
    L --> K
    PS --> K
    U --> K

    K --> ROOT["Root"]
    K --> PEERS["Peer Nodes"]
```

### Node-side communication flow observed in runtime

```mermaid
sequenceDiagram
    participant NB as NodeBrain
    participant R as Root
    participant P as Peer Nodes

    R->>NB: RootAnnouncement(root_node_id, root_endpoint, generation)
    R->>NB: PeerDirectoryMessage(peers)
    NB->>NB: Update local peer topology
    NB->>NB: Sample CPU / memory / GPU
    NB->>NB: Predict trajectory + confidence + freshness
    NB->>R: NodeKnowledge / telemetry update
    NB->>P: PeerHealthMessage
    P-->>NB: COMM_ACK
    NB-->>R: COMM_ACK / health acknowledgement
    R-->>NB: Updated cluster / peer information
```

### NodeBrain state carried toward Root and peers

```mermaid
flowchart TB
    K["NodeKnowledge"]

    K --> RES["Resource Health<br/>CPU • Memory • GPU"]
    K --> PRED["Prediction<br/>current • predicted • trajectory"]
    K --> CONF["Confidence"]
    K --> FRESH["Freshness"]
    K --> PRESS["Resource Pressure"]
    K --> ENDPOINT["Endpoint<br/>dynamic Tailscale address + service port"]
    K --> STATUS["Communication / Health Status"]
    K --> GEN["Generation / State Version"]

    RES --> REPORT["Health / Telemetry Report"]
    PRED --> REPORT
    CONF --> REPORT
    FRESH --> REPORT
    PRESS --> REPORT
    ENDPOINT --> REPORT
    STATUS --> REPORT
    GEN --> REPORT

    REPORT --> ROOT["Root"]
    REPORT --> PEERS["Peer Nodes"]
```

### Root ↔ NodeBrain ↔ Peer relationship

```mermaid
flowchart LR
    ROOT["ROOT<br/>Control + Scheduling Plane"]

    N1["NodeBrain A"]
    N2["NodeBrain B"]
    N3["NodeBrain C"]

    ROOT -->|RootAnnouncement| N1
    ROOT -->|PeerDirectory| N1
    ROOT -->|RootAnnouncement| N2
    ROOT -->|PeerDirectory| N2
    ROOT -->|RootAnnouncement| N3
    ROOT -->|PeerDirectory| N3

    N1 -->|NodeKnowledge / telemetry| ROOT
    N2 -->|NodeKnowledge / telemetry| ROOT
    N3 -->|NodeKnowledge / telemetry| ROOT

    N1 <-->|PeerHealthMessage| N2
    N2 <-->|PeerHealthMessage| N3
    N3 <-->|PeerHealthMessage| N1

    N1 --> TRAIN["Direct Node ↔ Node Training Data Plane"]
    N2 --> TRAIN
    N3 --> TRAIN
```

**Interpretation:** Root distributes coordination information and consumes node state; NodeBrain owns local observation/prediction; worker nodes exchange peer-health knowledge directly. High-frequency training traffic remains on the direct node-to-node path rather than being routed through the Root dashboard/control plane.

# 16. NodeBrain ↔ Root Interaction

NodeBrain does not replace the Root scheduler.

The relationship is:

```mermaid
sequenceDiagram
    participant H as Node Hardware
    participant N as NodeBrain
    participant R as RootBrain
    participant S as Adaptive Scheduler
    participant E as Local Executor

    H->>N: Hardware / runtime state
    N->>N: Build capability + resource model
    N->>R: Node registration / state report
    R->>S: Cluster state
    S->>S: Capability matching
    S->>S: Allocation decision
    S->>R: Versioned allocation
    R->>N: Work / allocation
    N->>E: Local execution
    E->>N: Runtime metrics
    N->>N: Performance / health assessment
    N->>R: Updated node state
    R->>S: Updated cluster evidence
```

This creates the distributed feedback relationship:

```text
             ROOT
              │
       allocation decision
              ↓
          NodeBrain
              │
        local execution
              ↓
       local observation
              │
              ↓
             ROOT
              │
       adaptive decision
              ↓
          NodeBrain
```

---

# 17. Root + NodeBrain Combined Architecture

The most important control relationship in PALADIN is:

```mermaid
flowchart TB
    subgraph ROOT["ROOT — GLOBAL CONTROL / SCHEDULING"]
        RB["RootBrain"]
        SCH["Adaptive Scheduler"]
        RF["Resource Fabric"]
        TC["Training Coordinator"]
        STATE["Cluster / Allocation State"]
        API["Dashboard API"]
    end

    subgraph A["NODE A — LOCAL INTELLIGENCE"]
        NA["NodeBrain"]
        CA["Capability"]
        TA["Telemetry"]
        EA["Assessment"]
        XA["Execution"]
    end

    subgraph B["NODE B — LOCAL INTELLIGENCE"]
        NB["NodeBrain"]
        CB["Capability"]
        TB["Telemetry"]
        EB["Assessment"]
        XB["Execution"]
    end

    subgraph C["NODE C — LOCAL INTELLIGENCE"]
        NC["NodeBrain"]
        CC["Capability"]
        TCAP["Telemetry"]
        EC["Assessment"]
        XC["Execution"]
    end

    RB --> STATE
    STATE --> SCH
    SCH --> RF
    RF --> TC
    STATE --> API

    SCH --> NA
    SCH --> NB
    SCH --> NC

    NA --> CA
    NA --> TA
    NA --> EA
    EA --> NA
    NA --> XA

    NB --> CB
    NB --> TB
    NB --> EB
    EB --> NB
    NB --> XB

    NC --> CC
    NC --> TCAP
    NC --> EC
    EC --> NC
    NC --> XC

    NA --> RB
    NB --> RB
    NC --> RB
```

The conceptual division is:

```text
ROOT
 ├── Global view
 ├── Cluster state
 ├── Scheduling
 ├── Allocation
 └── Coordination

NODEBRAIN
 ├── Local view
 ├── Capability
 ├── Telemetry
 ├── Performance
 ├── Health
 └── Execution awareness
```

Neither side replaces the other:

```text
NodeBrain = local evidence + local intelligence
Root      = global coordination + scheduling
```

---

# 18. Complete Root–Node–Training Relationship

```mermaid
flowchart TB
    ROOT["ROOT"]
    RB["RootBrain"]
    SCH["Adaptive Scheduler"]
    TC["Training Coordinator"]

    NA["Node A<br/>NodeBrain"]
    NB["Node B<br/>NodeBrain"]
    NC["Node C<br/>NodeBrain"]

    P2P["DIRECT NODE ↔ NODE<br/>Training Data Plane"]

    ROOT --> RB
    RB --> SCH
    SCH --> TC

    TC --> NA
    TC --> NB
    TC --> NC

    NA <--> P2P
    NB <--> P2P
    NC <--> P2P

    NA --> RB
    NB --> RB
    NC --> RB

    P2P -. gradients .-> P2P
    P2P -. tensors .-> P2P
    P2P -. model sync .-> P2P
```

This emphasizes the central PALADIN separation:

```text
                ROOT
                 │
       CONTROL / ALLOCATION
                 │
        ┌────────┼────────┐
        ↓        ↓        ↓
      Node A   Node B   Node C
     NodeBrain NodeBrain NodeBrain
        │        │        │
        └────────┼────────┘
                 ↓
          DIRECT TRAINING
           NODE ↔ NODE
```

---

# 19. Dashboard + NodeBrain Operational View

The operator-visible system can be understood as:

```mermaid
flowchart LR
    subgraph LOCAL["NODE"]
        NB["NodeBrain"]
        CAP["Capabilities"]
        TEL["Telemetry"]
        PERF["Performance"]
        HEALTH["Health"]
    end

    subgraph ROOT["ROOT"]
        RB["RootBrain"]
        SCH["Scheduler"]
        ALLOC["Allocation"]
    end

    subgraph UI["ROOT DASHBOARD"]
        NODEUI["Node Status"]
        CLUI["Cluster Status"]
        ALUI["Allocation View"]
        TEUI["Telemetry"]
        EPUI["Endpoints"]
    end

    NB --> CAP
    NB --> TEL
    NB --> PERF
    NB --> HEALTH

    CAP --> RB
    TEL --> RB
    PERF --> RB
    HEALTH --> RB

    RB --> SCH
    SCH --> ALLOC

    RB --> NODEUI
    RB --> CLUI
    ALLOC --> ALUI
    TEL --> TEUI
    RB --> EPUI
```

This provides a direct explanation of what the operator sees versus what the runtime actually does:

```text
NodeBrain
   ↓
collects evidence

RootBrain
   ↓
aggregates evidence

Adaptive Scheduler
   ↓
makes allocation decisions

Dashboard
   ↓
visualizes authoritative state
```
