Three-tier neural architecture designed for sub-millisecond CPU execution without GPU dependencies. Routes incoming representations dynamically across three hierarchical tiers based on classification confidence:
- Tier-0: Normalized linear validator for microsecond classification, continuous scoring, or early exit (0.35 us).
- Tier-1: Decoupled Gated Delta Attention (DGDA) linear recurrence cell ($O(1)$ memory footprint, zero KV-cache growth).
- Tier-2: Fixed-length canvas ($L = 32$) decoded in parallel via conditional flow matching with multi-head latent attention (MLA).
Benchmark Summary
Measured on Intel Xeon Platinum 8370C (AVX-512, batch = 1, d = 256, L = 32 tokens):
| Model | Paradigm | P50 Latency | Throughput | Memory Footprint | KV-Cache Growth |
|---|---|---|---|---|---|
| maba-instant v1 (Tier-0 Exit) | Linear projection | 0.47 us | 1,811,636 tok/s | 80 KB static | 0 bytes |
| maba-instant v1 (Cascade) | 3-Tier JEV+DGDA+MLA | 3.34 ms | 8,362 tok/s | 80 KB static | 0 bytes |
| Mamba-3 SSM (Step) | Selective state space | 132.92 us | 7,298 steps/s | 64 KB static | 0 bytes |
| RWKV-v6 (Step) | Linear recurrence | 205.20 us | 4,422 steps/s | 1 KB static | 0 bytes |
| LLaMA-3 AR (32 tokens) | Causal autoregressive | 25.41 ms | 1,162 tok/s | Linear with L | Dynamic |
| Discrete Diffusion (32 tokens) | 16-step canvas flow | 26.85 ms | 1,133 tok/s | 64 KB static | 0 bytes |
Full benchmark breakdowns, memory tracking via /proc/self/status, and Level-2 order book simulation metrics are documented in report.md.
Formal mathematical definitions, state transition equations, and complexity proofs are documented in architecture.md.
Requirements
- Linux x86_64 with AVX2 or AVX-512
- GCC >= 11 or Clang >= 14 (C++20)
- OpenMP runtime
- Python >= 3.10 (optional, for ctypes bindings and training scripts)
Build
Compile the native C++ library and benchmark binaries:
make all
Build targets:
lib/libmaba_instant.so: Native C shared library with SIMD vectorization.bin/maba_bench: Component latency microbenchmark.bin/compare_arch: Comparative benchmark against baseline architectures.bin/trading_bench: Level-2 order book simulation.bin/maba_endurance: Multithreaded numerical stability runner.
Usage
C++
#include "maba_instant/pipeline.hpp"
maba::Config cfg;
cfg.dim_model = 256;
cfg.canvas_len = 32;
maba::MabaPipeline pipeline(cfg);
std::vector<float> input(cfg.dim_model, 0.05f);
const auto& res = pipeline.process(input.data(), maba::SchemaKind::Choice);
Python
from maba_instant import Config, MabaPipeline, SchemaKind
cfg = Config(dim_model=256, canvas_len=32)
pipe = MabaPipeline(cfg)
out = pipe.process([0.1] * 256, SchemaKind.CHOICE)
Training
Inference runs entirely on CPU via AVX-512 and OpenMP with zero GPU runtime dependency. Training is supported across both GPU (PyTorch with CUDA/MPS) and lightweight CPU modes.
GPU Training (PyTorch / CUDA)
# Train with auto-detected accelerator (CUDA or MPS)
python3 train_gpu.py --device auto --batch_size 32 --steps 200 --save_checkpoint checkpoint.pt
# Train on CUDA with DGDA recurrence and custom dimensions
python3 train_gpu.py --device cuda --dim 512 --layers 6 --canvas_len 32 --steps 500
# Train on CUDA with Mamba SSM recurrence
python3 train_gpu.py --device cuda --use_mamba --steps 500
CPU Standalone Training (Zero Dependencies)
python3 train_cpu.py --steps 100 --lr 0.01
Running Tests and Benchmarks
# Run unit tests
python3 -m unittest discover tests
# Run Python demo
python3 demo.py
# Run component microbenchmark
./bin/maba_bench
# Run cross-architecture benchmark
./bin/compare_arch
# Run order book simulation
./bin/trading_bench
License
This architecture and implementation are licensed under the MABA OPEN ARCHITECTURE LICENSE (MOAL-1.0). See LICENSE for terms.
- Downloads last month
- -