ASTRAL Developer Architecture Guide v0.2.1
Switch to User Guide →

ASTRAL System Architecture & Technical Specification

An in-depth guide to Adaptive Spectral Resonance Decomposition (SRD), Invariant Risk Minimization (IRM) stability diagnostics, Gram-Schmidt orthogonalization, and conformal uncertainty calibration for interpretable tabular machine learning.

01 Core Philosophy & Design Rationale

Tabular machine learning has long suffered an acute dichotomy between opaque gradient-boosted decision trees (GBDTs) / deep neural networks and rigid linear/logistic models:

⚡ The Ensembling Problem

GBDTs (XGBoost, LightGBM, CatBoost) achieve high empirical accuracy by recursively partitioning feature space, but generate thousands of discontinuous step functions that cannot be mathematically audited, verified for causal invariant behavior, or proven free of spurious shortcut correlations.

🧠 The Neural Tabular Problem

Deep tabular models (TabNet, FT-Transformer) incur immense computational training overhead, require delicate learning rate schedules, struggle with heterogeneous unscaled data, and remain black-box function approximators.

💎 The ASTRAL Alternative

ASTRAL operates in a continuous Reproducing Kernel Hilbert Space (RKHS) by dynamically synthesising empirical quantile basis functions, isolating non-linear cross-feature resonance, penalizing unstable shortcut features via IRM, and solving a closed-form convex regularized objective.

02 Mathematical Foundations

Let D = {(x_i, y_i)}_{i=1}^N denote a dataset where x_i ∈ R^d and y_i ∈ Y. ASTRAL models the response function f(x) as a linear combination of discovered non-linear basis functionals φ_k: R^d → R:

f(x) = w_0 + ∑_{k=1}^M w_k φ_k(x) = w_0 + w^T Φ(x)

Unlike fixed polynomial or Fourier expansions whose basis size suffers from the curse of dimensionality O(d^p), ASTRAL discovers an adaptive, highly sparse dictionary B = {φ_1, ..., φ_M} with M ≪ d^2 via mutual information resonance screening and Minimal Description Length (MDL) pruning.

03 Basis Function Dictionary

For each marginal feature X_j, ASTRAL computes Q empirical quantiles q_1 < q_2 < ... < q_Q and evaluates six complementary non-linear basis families:

Basis Family Mathematical Formulation Structural Target
Linear φ(x) = x_j Global monotonic trends and baseline linear relationships
Quantile Hat (Spline) Λ(x; q_{k-1}, q_k, q_{k+1}) = max(0, 1 - |x - q_k| / Δ) Piecewise linear local regimes and localized modal densities
Gaussian RBF κ(x; c, σ) = exp(- (x - c)^2 / (2σ^2)) Smooth bell-shaped clustering and non-linear localized attraction
Sigmoidal Transition σ(x; c, w) = 1 / (1 + exp(-w(x - c))) Smooth binary phase shifts and threshold saturation phenomena
Root & Power Transforms φ(x) = sign(x)√|x|,  φ(x) = x^2 Sub-linear diminishing returns and super-linear quadratic acceleration
Heaviside Step θ(x; q_k) = 1 if x ≥ q_k else 0 Hard administrative boundaries, policy thresholds, and categorical splits

04 Pairwise Non-Linear Resonance Screening

Given top informative marginal bases φ_i and φ_j, ASTRAL constructs candidate bivariate interaction functionals:

ψ_{mult} = φ_i(x) · φ_j(x),   ψ_{ratio} = φ_i(x) / (|φ_j(x)| + ε),   ψ_{min} = min(φ_i(x), φ_j(x)),   ψ_{diff} = |φ_i(x) - φ_j(x)|

A candidate interaction functional ψ is admitted into the dictionary if and only if it exhibits Resonance Gain over both individual constituents:

Gain(ψ; y) = I(ψ; y) - max( I(φ_i; y), I(φ_j; y) ) > τ_{resonance}

This guarantees that the model only incorporates interactions that capture genuine synergistic non-linear information, preventing combinatorial explosion.

05 Invariant Risk Minimization (IRM) Stability Diagnostics

Standard empirical risk minimization (ERM) eagerly exploits spurious correlations that happen to align with targets in the training sample. ASTRAL mitigates this by generating E synthetic perturbation environments E = {e_1, e_2, ..., e_E} through bootstrap resampling and adversarial feature noise injection.

For each candidate basis functional φ_k, ASTRAL computes the correlation vector across all environments ρ_k = [corr^{(e)}(φ_k, y)]_{e ∈ E} and calculates an invariance stability score s_k ∈ [0, 1]:

s_k = | μ(ρ_k) | · max(0, 1 - σ(ρ_k) / (ε + |μ(ρ_k)|))
Invariance Ingestion Rule: If a basis functional displays erratic correlation flips across resampled splits, its stability score s_k → 0. If it reliably maintains consistent directionality and high signal-to-noise ratio across all splits, s_k → 1.

06 Minimal Description Length & Gram-Schmidt Selection

To prevent collinear redundancy and over-parameterization, candidate bases are selected sequentially using greedy forward orthogonal residual pursuit:

  1. Initialize residual vector r_0 = y - mean(y) and active basis set S_0 = ∅.
  2. At step m, project each candidate basis φ_j onto the orthogonal complement of the subspace spanned by S_{m-1} using modified Gram-Schmidt:
    φ̃_j^{(m)} = φ_j - ∑_{k ∈ S_{m-1}} [ ⟨φ_j, u_k⟩ / ||u_k||^2 ] u_k
  3. Score candidate φ_j by trading off residual correlation against stability and Description Length penalty:
    Score(φ_j) = [ |⟨r_{m-1}, φ̃_j^{(m)}⟩| / ||φ̃_j^{(m)}|| ] · (0.5 + 0.5 s_j) - γ · MDL(φ_j)
  4. Add the maximizer φ* = argmax_j Score(φ_j) to S_m and update residual r_m. Terminate when gain drops below threshold or max bases budget is reached.

07 Stability-Weighted Regularized Fitting

Once the active basis design matrix Φ ∈ R^{N × M} is assembled, the model solves the convex stability-weighted regularized problem:

min_w   L(y, Φ w) + ½ w^T R w

where R is a diagonal regularization matrix whose penalty dynamically scales inversely with causal stability:

R_{jj} = λ / (0.25 + 0.75 s_j)

For regression, L is squared error or Huber loss, yielding the direct closed-form normal equations:

w* = ( Φ^T Φ + R )^{-1} Φ^T y

For classification, L is binary or multinomial cross-entropy, solved via Newton-CG optimization with analytical gradient and Hessian vector products.

08 Conformal Uncertainty & Abstention Engine

ASTRAL incorporates non-asymptotic distribution-free conformal calibration. Let α ∈ (0, 1) be the user error rate budget:

🎯 Conformal Prediction Intervals (Regression)

Computes absolute residual nonconformity scores R_i = |y_i - ŷ_i|. Finds the empirical (1 - α)(1 + 1/N) quantile q̂. Guarantees finite-sample coverage:

P( Y ∈ [ f(x) - q̂, f(x) + q̂ ] ) ≥ 1 - α

🛡️ Conformal Prediction Sets (Classification)

Computes cumulative softmax nonconformity. Returns guaranteed prediction sets C(x) ⊆ Y. If |C(x)| > 1 or predicted class probability falls below the confidence margin, triggers automated prediction abstention.

09 Six-Phase Execution Pipeline

[Input Data Matrix X (N x d), Target y]
                     │
    Phase 1 ─────────▼───────── Input Validation & Imputation
                     │          • Median imputation for NaNs / infinities
                     │          • Empirical quantile mesh construction
    Phase 2 ─────────▼───────── Univariate Basis Synthesis
                     │          • Linear, Quantile Hats, RBFs, Sigmoids, Roots, Steps
                     │          • Mutual information marginal filtering
    Phase 3 ─────────▼───────── Non-Linear Resonance Interaction
                     │          • Multiplicative, Ratio, Min, Difference combinations
                     │          • Positive resonance synergy screening: I(a,b; y) > max(I_a, I_b)
    Phase 4 ─────────▼───────── Perturbation Environments & IRM
                     │          • E=10 bootstrapped environments
                     │          • Stability score computation: s_j in [0, 1]
    Phase 5 ─────────▼───────── Orthogonal Gram-Schmidt & MDL Pruning
                     │          • Residual projection and collinearity elimination
                     │          • Sparse dictionary selection (M << d^2)
    Phase 6 ─────────▼───────── Stability-Weighted Optimization
                                • R_jj = lambda / (0.25 + 0.75 s_j)
                                • Direct normal equations (regression) or Newton-CG (classification)
                                • Calibrated conformal intervals & prediction sets

10 Performance & Scalability

🚀 Subsampled Discovery

For large tabular datasets (N > 25,000), ASTRAL executes basis screening and stability diagnostics on a stratified sub-sample (N_{sub} = 4,000 default, configurable). This slashes discovery latency from minutes to under 2.5 seconds on datasets with 600,000+ rows.

⚡ Zero C++ / Pure NumPy

The entire pipeline is vectorized with contiguous memory allocations, SIMD-friendly linear algebra, and zero external binary compiled dependencies.

11 Developer Extension Points

Developers can easily inject custom basis transforms or regularizers by extending the ResonanceBasis interface:

from astral import ResonanceBasis
import numpy as np

class CustomPeriodicBasis(ResonanceBasis):
    """Custom Fourier / Periodic resonance basis."""
    def __init__(self, feature_idx, frequency, name=None):
        super().__init__(feature_indices=[feature_idx], transform_type="periodic", name=name)
        self.frequency = frequency

    def evaluate(self, X):
        col = X[:, self.feature_indices[0]]
        return np.sin(self.frequency * col)

    def description_length(self):
        return 2.5  # Complexity penalty for periodic functions

12 Verification & Test Protocol

The package includes an end-to-end automated testing harness in test_astral.py covering metadata, classification, regression, conformal coverage, data quality reporting, and scikit-learn compliance.

# Run complete test suite
python test_astral.py

# Run multi-dataset benchmark comparison against 10 baselines
python benchmark_models.py