Instructions to use Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1:F16 # Run inference directly in the terminal: llama cli -hf Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1:F16 # Run inference directly in the terminal: llama cli -hf Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1:F16 # Run inference directly in the terminal: ./llama-cli -hf Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1:F16
Use Docker
docker model run hf.co/Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1:F16
- LM Studio
- Jan
- vLLM
How to use Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1:F16
- Ollama
How to use Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1 with Ollama:
ollama run hf.co/Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1:F16
- Unsloth Desktop
- Pi
How to use Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1:F16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1:F16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1 with Docker Model Runner:
docker model run hf.co/Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1:F16
- Lemonade
How to use Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1:F16
Run and chat with the model
lemonade run user.qwen3loop-0.6b-tool-calling-sft-v1-F16
List all available models
lemonade list
- Hermes Agent
How to use Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1:F16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1:F16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1:F16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1:F16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen3Loop-0.6B-Tool-Calling-SFT (v3)
Qwen3Loop-0.6B-Tool-Calling-SFT is a narrow tool-calling specialist: the 0.6B-parameter
recurrent Qwen3Loop architecture, fully fine-tuned on a curated curriculum of 2,498 strict
tool-calling samples (single-turn user โ assistant(tool_calls), mid-trajectory corrections,
abstention pairs, finance decoys, terminal/memory/API calls, Qwen Hermes <tool_call> format).
This is the third data generation (v3). v1 (1,983 samples) and v2 (2,104) remain in git history.
Core architecture
- Physical parameters: 28 transformer blocks (
hidden_size: 1024,intermediate_size: 3072,heads: 16). - LoopSplit: prefix layers 0..6 (once) + middle stack 7..20 (3 recurrent passes = 42 layer passes) + suffix 21..27 (once) = 56 logical layers ($7 + 14 \times 3 + 7$).
- Training: full fine-tuning (not LoRA) from the base SFT checkpoint, 2รT4 DDP, 1 epoch. Lineage: Qwen3Loop-0.6B SFT family.
Measured capabilities (Q8_0, RTX 3060, temperature=0.0)
121-item English holdout battery (50 tools, 8 domains):
| Slice | v1 | v2 | v3 (this) |
|---|---|---|---|
| Single-tool precise (L1) | 30/30 | 27/30 | 29/30 (96.7%) |
| Multi-tool discrimination (L2) | 35/41 | 32/41 | 36/41 (87.8%) |
| Python reasoning (L3) | 10/30 | 14/30 | 19/30 (63.3%) |
| Negative abstention (L4) | 10/20 | 14/20 | 12/20 (60.0%) |
| Overall strict | 85/121 (70.2%) | 87/121 (71.9%) | 96/121 (79.3%) |
| Format / selection | 83.5% / 81.8% | โ / โ | 91.7% / 90.9% |
25-item probe (in-distribution, all English): 9/25 (Python 4/5, format 96%). 12-item multi-turn correction battery (identical prompts, GPU): v3 11/12 vs v2 10/12 vs v1 9/12 vs stock Qwen3-0.6B 5/12.
Note on L3: an earlier v2 training batch mirrored battery instances; v3 was retrained on fresh instances only and the gains above are clean (probe Python confirms out-of-battery).
Explicit limitations โ read before deploying as an agent
- English only. ~99% of training is English; non-English prompts measurably degrade tool selection and triggering.
- Abstention is unstable (12/20). It improved from v1 (10/20) but oscillates across data mixes (v2 hit 14/20, v3 fell back to 12/20): the act/abstain threshold has not converged. Gate it externally if calls have cost or side effects.
- Code reasoning, not triggering. It now reliably calls
execute_python(zero no-calls on 30 items, up from 10 misses in v1), but ~1/3 of generated snippets are wrong (algorithm slips, float repr). Trigger solved; code correctness is the next ceiling. - Unseen tools confuse it. Finance decoys improved, but tools outside the training catalog still cause wrong-tool or no-call failures. In-distribution invocation is 29/30; transfer is not.
- Validated single-turn only (plus 12 multiturn items). Long agentic loops are untested.
- 0.6B capacity ceiling. Fine-grained decoys oscillate across runs; obscure flags and invented entities reflect capacity limits, not just data.
- Stylistic biases. Verbose
<think>preambles mirror training priors; harmless but noisy in logs.
Dynamic depth: entropy-gated loop skipping
latent_halting_probe.pt is an ultralight MLP head trained on intermediate
prompt prefill states. At inference it estimates per-request entropy/complexity and skips the
recurrent middle-stack passes when they are not needed, falling back to a shallower pass โ
higher throughput on easy requests, full 56-layer depth when it matters. Wire it through the
forward hooks in engine/qwen3loop_python/.
(Note: routing accuracy was measured on the base-SFT lineage probe, not re-measured on this
checkpoint; treat the gate as a speed heuristic and validate on your workload.)
Model files
1. Universal unrolled models (stock llama.cpp / Ollama / LM Studio, arch qwen3, 56 blocks)
unrolled_modelo_qwen3loop_tool_sft_q8_0.gguf(1.03 GB): Q8_0. Recommended for local engines.unrolled_modelo_qwen3loop_tool_sft_f16.gguf(1.94 GB): full-precision FP16.
2. Compact native looped models (28 physical blocks; need a loop-aware engine or the patch below)
modelo_qwen3loop_tool_sft_q8_0.gguf(0.64 GB)modelo_qwen3loop_tool_sft_f16.gguf(1.12 GB)
3. PyTorch weights, probe & engine
model.safetensors(1.11 GB): full fine-tuned weights (v3 data).latent_halting_probe.pt: entropy halting MLP (see above).config.json,tokenizer.json,tokenizer_config.json,generation_config.json,chat_template.jinja.engine/qwen3loop_python/: modeling + config + minimal forward.engine/llamacpp_patch/: llama.cpp loader patch reference for the looped architecture.
Usage
PyTorch (custom modeling ships in-repo, no trust_remote_code needed):
import sys
sys.path.insert(0, "engine")
from transformers import AutoTokenizer
from qwen3loop_python.modeling_qwen3loop import Qwen3LoopForCausalLM
import torch
tok = AutoTokenizer.from_pretrained("Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1")
model = Qwen3LoopForCausalLM.from_pretrained(
"Lordnyx/qwen3loop-0.6b-tool-calling-sft-v1", torch_dtype=torch.bfloat16).cuda()
prompt = tok.apply_chat_template(
[{"role": "user", "content": "Convert 42 km to miles."}],
tools=[{"type": "function", "function": {
"name": "convert_measurement_units",
"description": "Converts physical quantities.",
"parameters": {"type": "object", "properties": {
"value": {"type": "number"}, "from_unit": {"type": "string"},
"to_unit": {"type": "string"}}, "required": ["value", "from_unit", "to_unit"]}}}],
tokenize=False, add_generation_prompt=True)
GGUF: load any unrolled_* file in Ollama / LM Studio / llama.cpp (-ngl 99 for full GPU offload).
{ "temperature": 0.0, "top_p": 1.0, "repeat_penalty": 1.0, "context_length": 8192 }
(Use temperature: 0.0 for deterministic tool calls; the model was evaluated greedy.)
- Downloads last month
- 2,208