maba-instant logo

maba-instant v1 Architecture

Cascade Linear Recurrence & Non-Autoregressive Flow Architecture

License Latency Decode O(1) Zero KV-Cache Benchmark


Three-tier neural architecture designed for sub-millisecond CPU execution without GPU dependencies. Routes incoming representations dynamically across three hierarchical tiers based on classification confidence:

  1. Tier-0: Normalized linear validator for microsecond classification, continuous scoring, or early exit (0.35 us).
  2. Tier-1: Decoupled Gated Delta Attention (DGDA) linear recurrence cell ($O(1)$ memory footprint, zero KV-cache growth).
  3. Tier-2: Fixed-length canvas ($L = 32$) decoded in parallel via conditional flow matching with multi-head latent attention (MLA).

Benchmark Summary

Measured on Intel Xeon Platinum 8370C (AVX-512, batch = 1, d = 256, L = 32 tokens):

Model Paradigm P50 Latency Throughput Memory Footprint KV-Cache Growth
maba-instant v1 (Tier-0 Exit) Linear projection 0.47 us 1,811,636 tok/s 80 KB static 0 bytes
maba-instant v1 (Cascade) 3-Tier JEV+DGDA+MLA 3.34 ms 8,362 tok/s 80 KB static 0 bytes
Mamba-3 SSM (Step) Selective state space 132.92 us 7,298 steps/s 64 KB static 0 bytes
RWKV-v6 (Step) Linear recurrence 205.20 us 4,422 steps/s 1 KB static 0 bytes
LLaMA-3 AR (32 tokens) Causal autoregressive 25.41 ms 1,162 tok/s Linear with L Dynamic
Discrete Diffusion (32 tokens) 16-step canvas flow 26.85 ms 1,133 tok/s 64 KB static 0 bytes

Full benchmark breakdowns, memory tracking via /proc/self/status, and Level-2 order book simulation metrics are documented in report.md.

Formal mathematical definitions, state transition equations, and complexity proofs are documented in architecture.md.

Requirements

  • Linux x86_64 with AVX2 or AVX-512
  • GCC >= 11 or Clang >= 14 (C++20)
  • OpenMP runtime
  • Python >= 3.10 (optional, for ctypes bindings and training scripts)

Build

Compile the native C++ library and benchmark binaries:

make all

Build targets:

  • lib/libmaba_instant.so: Native C shared library with SIMD vectorization.
  • bin/maba_bench: Component latency microbenchmark.
  • bin/compare_arch: Comparative benchmark against baseline architectures.
  • bin/trading_bench: Level-2 order book simulation.
  • bin/maba_endurance: Multithreaded numerical stability runner.

Usage

C++

#include "maba_instant/pipeline.hpp"

maba::Config cfg;
cfg.dim_model = 256;
cfg.canvas_len = 32;

maba::MabaPipeline pipeline(cfg);

std::vector<float> input(cfg.dim_model, 0.05f);
const auto& res = pipeline.process(input.data(), maba::SchemaKind::Choice);

Python

from maba_instant import Config, MabaPipeline, SchemaKind

cfg = Config(dim_model=256, canvas_len=32)
pipe = MabaPipeline(cfg)

out = pipe.process([0.1] * 256, SchemaKind.CHOICE)

Training

Inference runs entirely on CPU via AVX-512 and OpenMP with zero GPU runtime dependency. Training is supported across both GPU (PyTorch with CUDA/MPS) and lightweight CPU modes.

GPU Training (PyTorch / CUDA)

# Train with auto-detected accelerator (CUDA or MPS)
python3 train_gpu.py --device auto --batch_size 32 --steps 200 --save_checkpoint checkpoint.pt

# Train on CUDA with DGDA recurrence and custom dimensions
python3 train_gpu.py --device cuda --dim 512 --layers 6 --canvas_len 32 --steps 500

# Train on CUDA with Mamba SSM recurrence
python3 train_gpu.py --device cuda --use_mamba --steps 500

CPU Standalone Training (Zero Dependencies)

python3 train_cpu.py --steps 100 --lr 0.01

Running Tests and Benchmarks

# Run unit tests
python3 -m unittest discover tests

# Run Python demo
python3 demo.py

# Run component microbenchmark
./bin/maba_bench

# Run cross-architecture benchmark
./bin/compare_arch

# Run order book simulation
./bin/trading_bench

License

This architecture and implementation are licensed under the MABA OPEN ARCHITECTURE LICENSE (MOAL-1.0). See LICENSE for terms.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support