How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf Lordnyx/qwen3loop-0.6b-sft-deep-supervision-v1:F16
# Run inference directly in the terminal:
llama cli -hf Lordnyx/qwen3loop-0.6b-sft-deep-supervision-v1:F16
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf Lordnyx/qwen3loop-0.6b-sft-deep-supervision-v1:F16
# Run inference directly in the terminal:
llama cli -hf Lordnyx/qwen3loop-0.6b-sft-deep-supervision-v1:F16
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf Lordnyx/qwen3loop-0.6b-sft-deep-supervision-v1:F16
# Run inference directly in the terminal:
./llama-cli -hf Lordnyx/qwen3loop-0.6b-sft-deep-supervision-v1:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf Lordnyx/qwen3loop-0.6b-sft-deep-supervision-v1:F16
# Run inference directly in the terminal:
./build/bin/llama-cli -hf Lordnyx/qwen3loop-0.6b-sft-deep-supervision-v1:F16
Use Docker
docker model run hf.co/Lordnyx/qwen3loop-0.6b-sft-deep-supervision-v1:F16
Quick Links

Qwen3Loop-0.6B-SFT-Deep-Supervision (v1.0 Production: Eurus-2 + Open-R1 CoT & Adaptive Halting)

Qwen3Loop-0.6B is a state-of-the-art recursive reasoning language model utilizing recurrent looped layers to achieve the reasoning density and depth of a ~2B parameter model within a 0.6B physical parameter footprint.

๐ŸŒŸ Core Architecture & Innovation

  • Physical Parameters: 28 transformer blocks (hidden_size: 1024, intermediate_size: 3072, heads: 16).
  • LoopSplit Architecture:
    • Prefix: Layers 0..6 (computed once).
    • Middle Stack (Miolo): Layers 7..20 (computed 3 times recurrently = 42 layer passes).
    • Suffix: Layers 21..27 (computed once).
    • Total Logical Depth: 56 layers ($7 + 14 \times 3 + 7 = 56$).
  • Curated Multi-Phase Dataset: Fine-tuned on 6,324 strictly filtered and curated samples (25.3 MB) from PRIME-RL/Eurus-2-SFT-Data (agentic multi-phase) and open-r1/Mixture-of-Thoughts (coding CoT), annotated across 9 meta-supervision dimensions with convaiinnovations/laya.
  • Adaptive Halting Probe: Includes latent_halting_probe.pt, an ultralight MLP head trained directly on intermediate prompt prefill states to decide optimal exit depth with 91.5% classification accuracy.

๐Ÿ“Š Training & Validation Metrics

Metric Result Context / Meaning
Validation Loss 0.4012 Down from 0.6705 (smooth, stable convergence)
Perplexity (PPL) 1.49 Near-optimal linguistic and logical fluency
Hardware Stability 5.58 GB VRAM Flat VRAM curve throughout 375 steps on RTX 3060 12GB
Halting Probe Routing Acc 91.5% Calibrated on prompt prefill representations (MSE: 0.0015)

๐ŸฅŠ Direct Confrontation: Base 28-Layer vs Qwen3Loop 56-Layer Unrolled

Evaluated under identical sampling parameters (temperature=0.1, Q8_0 quantization on RTX 3060):

Challenge Task Base Unsloth (Qwen3-0.6B-Q8_0) [28L] Qwen3Loop SFT (unrolled_q8_0) [56L] Outcome
Modular Arithmetic
($2^{100} \pmod 7$)
Incomplete / Verbose divergence
(Overflowed token budget without answer)
Exact: $\boxed{2}$ in 5 logical steps
(Detected optimal order $2^3 \equiv 1 \pmod 7$)
๐ŸŸข Qwen3Loop
Deductive Logic
(Sally Siblings Riddle)
Trapped in self-doubt loop Deduced correct sister relationship ๐ŸŸข Qwen3Loop
Linguistic Trick
(17 sheep, all but 9 die)
Arithmetic subtraction error ($17-9=8$) Arithmetic subtraction error ($17-9=8$) โšช Tie
Power Cycles
(Units digit of $3^{2025}$)
Stopped at cycle definition Full sequence & pattern $(3,9,7,1)$ ๐ŸŸข Qwen3Loop
Algorithm Formulation
(Balanced Parentheses)
Drafted stack concept, cut off Counter-based $O(n)$ time / $O(1)$ space logic ๐ŸŸข Qwen3Loop
Generation Speed 255.6 tokens/s (28 layers) 156.1 tokens/s (56 layers real pass) Hardware verified

๐Ÿ“ฆ Model Files & Download Options

1. Universal Unrolled Models (Compatible with stock LM Studio, Ollama, llama.cpp)

Run natively out of the box without any custom forks or patches (architecture mapped to qwen3 with 56 logical layers):

2. Compact Native Looped Models (For custom engines supporting cyclic execution)

3. Standalone PyTorch Weights & Halting Probe


๐ŸŽฏ Recommended Sampling Parameters

{
  "temperature": 0.6,
  "top_p": 0.95,
  "top_k": 40,
  "repeat_penalty": 1.08,
  "context_length": 32768
}

(For deterministic math and coding tasks, set temperature: 0.0 or 0.1).

Downloads last month
3,273
Safetensors
Model size
0.6B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Space using Lordnyx/qwen3loop-0.6b-sft-deep-supervision-v1 1

Collection including Lordnyx/qwen3loop-0.6b-sft-deep-supervision-v1