Qwen3Loop-0.6B-Tool-Calling-SFT (v3)

Qwen3Loop-0.6B-Tool-Calling-SFT is a narrow tool-calling specialist: the 0.6B-parameter recurrent Qwen3Loop architecture, fully fine-tuned on a curated curriculum of 2,498 strict tool-calling samples (single-turn user โ†’ assistant(tool_calls), mid-trajectory corrections, abstention pairs, finance decoys, terminal/memory/API calls, Qwen Hermes <tool_call> format).

This is the third data generation (v3). v1 (1,983 samples) and v2 (2,104) remain in git history.

Core architecture

  • Physical parameters: 28 transformer blocks (hidden_size: 1024, intermediate_size: 3072, heads: 16).
  • LoopSplit: prefix layers 0..6 (once) + middle stack 7..20 (3 recurrent passes = 42 layer passes) + suffix 21..27 (once) = 56 logical layers ($7 + 14 \times 3 + 7$).
  • Training: full fine-tuning (not LoRA) from the base SFT checkpoint, 2ร—T4 DDP, 1 epoch. Lineage: Qwen3Loop-0.6B SFT family.

Measured capabilities (Q8_0, RTX 3060, temperature=0.0)

121-item English holdout battery (50 tools, 8 domains):

Slice v1 v2 v3 (this)
Single-tool precise (L1) 30/30 27/30 29/30 (96.7%)
Multi-tool discrimination (L2) 35/41 32/41 36/41 (87.8%)
Python reasoning (L3) 10/30 14/30 19/30 (63.3%)
Negative abstention (L4) 10/20 14/20 12/20 (60.0%)
Overall strict 85/121 (70.2%) 87/121 (71.9%) 96/121 (79.3%)
Format / selection 83.5% / 81.8% โ€” / โ€” 91.7% / 90.9%

25-item probe (in-distribution, all English): 9/25 (Python 4/5, format 96%). 12-item multi-turn correction battery (identical prompts, GPU): v3 11/12 vs v2 10/12 vs v1 9/12 vs stock Qwen3-0.6B 5/12.

Note on L3: an earlier v2 training batch mirrored battery instances; v3 was retrained on fresh instances only and the gains above are clean (probe Python confirms out-of-battery).


Explicit limitations โ€” read before deploying as an agent

  1. English only. ~99% of training is English; non-English prompts measurably degrade tool selection and triggering.
  2. Abstention is unstable (12/20). It improved from v1 (10/20) but oscillates across data mixes (v2 hit 14/20, v3 fell back to 12/20): the act/abstain threshold has not converged. Gate it externally if calls have cost or side effects.
  3. Code reasoning, not triggering. It now reliably calls execute_python (zero no-calls on 30 items, up from 10 misses in v1), but ~1/3 of generated snippets are wrong (algorithm slips, float repr). Trigger solved; code correctness is the next ceiling.
  4. Unseen tools confuse it. Finance decoys improved, but tools outside the training catalog still cause wrong-tool or no-call failures. In-distribution invocation is 29/30; transfer is not.
  5. Validated single-turn only (plus 12 multiturn items). Long agentic loops are untested.
  6. 0.6B capacity ceiling. Fine-grained decoys oscillate across runs; obscure flags and invented entities reflect capacity limits, not just data.
  7. Stylistic biases. Verbose <think> preambles mirror training priors; harmless but noisy in logs.

Dynamic depth: entropy-gated loop skipping

latent_halting_probe.pt is an ultralight MLP head trained on intermediate prompt prefill states. At inference it estimates per-request entropy/complexity and skips the recurrent middle-stack passes when they are not needed, falling back to a shallower pass โ€” higher throughput on easy requests, full 56-layer depth when it matters. Wire it through the forward hooks in engine/qwen3loop_python/. (Note: routing accuracy was measured on the base-SFT lineage probe, not re-measured on this checkpoint; treat the gate as a speed heuristic and validate on your workload.)


Model files

1. Universal unrolled models (stock llama.cpp / Ollama / LM Studio, arch qwen3, 56 blocks)

2. Compact native looped models (28 physical blocks; need a loop-aware engine or the patch below)

3. PyTorch weights, probe & engine


Usage

PyTorch (custom modeling ships in-repo, no trust_remote_code needed):

import sys
sys.path.insert(0, "engine")
from transformers import AutoTokenizer
from qwen3loop_python.modeling_qwen3loop import Qwen3LoopForCausalLM
import torch

tok = AutoTokenizer.from_pretrained("Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1")
model = Qwen3LoopForCausalLM.from_pretrained(
    "Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1", torch_dtype=torch.bfloat16).cuda()
prompt = tok.apply_chat_template(
    [{"role": "user", "content": "Convert 42 km to miles."}],
    tools=[{"type": "function", "function": {
        "name": "convert_measurement_units",
        "description": "Converts physical quantities.",
        "parameters": {"type": "object", "properties": {
            "value": {"type": "number"}, "from_unit": {"type": "string"},
            "to_unit": {"type": "string"}}, "required": ["value", "from_unit", "to_unit"]}}}],
    tokenize=False, add_generation_prompt=True)

GGUF: load any unrolled_* file in Ollama / LM Studio / llama.cpp (-ngl 99 for full GPU offload).

{ "temperature": 0.0, "top_p": 1.0, "repeat_penalty": 1.0, "context_length": 8192 }

(Use temperature: 0.0 for deterministic tool calls; the model was evaluated greedy.)

Downloads last month
2,208
Safetensors
Model size
0.6B params
Tensor type
F16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Space using Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1 1

Collection including Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1