MiMo-V2.6-Distill-Qwen-9B-FP8

MiMo-V2.6-Distill-Qwen-9B-FP8 is a FP8 dynamic quantized version of XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B, a 9B agentic model developed by Xiaomi MiMo through supervised fine-tuning of Qwen/Qwen3.5-9B on MiMo-generated data. The base checkpoint covers coding, general-purpose agent tasks, visual coding, and cybersecurity, and is released as a starting point for open research in agentic reinforcement learning. This quantization reduces model size and memory footprint while preserving the model's agentic tool use, coding, long-form reasoning, and instruction-following capabilities, making deployment more accessible on smaller GPUs. The checkpoint includes its tokenizer and the MiMo v2.6 chat template, with thinking mode toggleable via enable_thinking.

Quantization to FP8 may introduce minor numerical differences relative to the bf16 source model. This is an SFT-only checkpoint (not RL-trained); behavior may differ from downstream RL-tuned MiMo-V2.6 releases.

Quantization Details

Quantization was performed using llmcompressor with the following recipe:

default_stage:
  default_modifiers:
    QuantizationModifier:
      targets: [Linear]
      ignore: ['re:.*lm_head', 're:.*embed_tokens$', 're:.*visual.*',
        're:.*model.visual.*', 're:.*linear_attn.*']
      scheme: FP8_DYNAMIC
      bypass_divisibility_checks: false
      requires_calibration_data: false

Linear layers are quantized to FP8 with dynamic per-tensor activation scaling, so no calibration dataset is required (requires_calibration_data: false). The lm_head, embedding table, any vision-tower (visual) components, and linear_attn layers are excluded from quantization and remain at full precision to preserve output-head fidelity and numerical stability.

Base model XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
Original backbone Qwen/Qwen3.5-9B
Quantization scheme FP8_DYNAMIC (Linear layers only)
Format compressed-tensors
Calibration data required No (dynamic activation scaling)
Excluded from quantization lm_head, embed_tokens, visual (if present), linear_attn

Use with vLLM

MiMo-V2.6-Distill-Qwen-9B-FP8 is served through vLLM with native support for compressed-tensors FP8 checkpoints.

Requirements

  • torch >= 2.11.0
  • vllm >= 0.19.1
  • A GPU with FP8 support recommended (Hopper or Blackwell class) for best throughput; also runs on Ampere with FP8 dequantized on the fly.

Serve

vllm serve prithivMLmods/MiMo-V2.6-Distill-Qwen-9B-FP8 \
  --max-model-len 32768

Client request

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")

messages = [
    {
        "role": "user",
        "content": "What is 15% of 240?"
    }
]

response = client.chat.completions.create(
    model="prithivMLmods/MiMo-V2.6-Distill-Qwen-9B-FP8",
    messages=messages,
    temperature=0.0,
    max_tokens=2048,
    extra_body={"chat_template_kwargs": {"enable_thinking": True}},
)

message = response.choices[0].message
print("Thinking:", getattr(message, "reasoning_content", "") or "")
print("Answer:", message.content or "")

Quick Start with Transformers

pip install transformers accelerate
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model = AutoModelForCausalLM.from_pretrained(
    "prithivMLmods/MiMo-V2.6-Distill-Qwen-9B-FP8",
    torch_dtype="auto",
    device_map="auto"
)

tokenizer = AutoTokenizer.from_pretrained(
    "prithivMLmods/MiMo-V2.6-Distill-Qwen-9B-FP8"
)

messages = [
    {
        "role": "user",
        "content": "What is 15% of 240?"
    }
]

text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
    enable_thinking=True
)

inputs = tokenizer(text, return_tensors="pt").to(model.device)

outputs = model.generate(
    **inputs,
    max_new_tokens=2048
)

print(
    tokenizer.decode(
        outputs[0][inputs["input_ids"].shape[-1]:],
        skip_special_tokens=True
    )
)

Evaluation (Base Model)

Results for the released SFT checkpoint (bf16 source), as reported in the MiMo-V2.6 technical report. The FP8 quantized checkpoint is expected to track these results closely; minor numerical differences from quantization should be validated for production use cases.

Domain Benchmark Metric Qwen3.5-9B MiMo-V2.6-Distill-Qwen-9B (SFT)
Code SWE Verified avg@3 60.0 61.1
Code SWE Pro avg@3 32.0 44.6
Code MiMo Code (mini)† avg@3 19.5 51.6
Cyber MiMo Cyber (mini)† avg@3 5.7 31.3
General AutomationBench v1.0.6 avg@1 5.0 30.3
General Terminal Bench 2.1 avg@1 27.0 37.1
General Toolathlon-Verified avg@1 25.9 35.2
General OfficeQA avg@1 9.0 19.5
General JobBench avg@1 2.6 18.3
General MiMo General (mini)† avg@1 28.5 62.2
Visual MiMo Visual Coding (mini)† avg@1 61.7 64.0

Training Details (Base Model)

Setting Value
Base Model Qwen/Qwen3.5-9B
Developer Xiaomi MiMo Team
Training Method Supervised Fine-Tuning (SFT) on MiMo-generated data
Total SFT Tokens 77.4B (27.2B loss-bearing)
Domains Code, Cyber, General, Visual
Chat Template MiMo v2.6 (thinking toggle via enable_thinking)
Purpose Starting point for open agentic reinforcement learning research

SFT Data Mixture

Domain Total tokens (B) Token share (%) Loss-bearing tokens (B)
Code 23.2 29.9 7.3
Cyber 11.0 14.2 4.8
General 22.0 28.5 5.7
Visual 21.2 27.4 9.4
Total 77.4 100.0 27.2

Intended Use and Limitations

Intended use, known limitations, training data composition, and responsible use guidance are unchanged from the base model. See the MiMo-V2.6-Distill-Qwen-9B model card for full details.

  • Agentic RL Research: Open SFT starting point for agentic reinforcement learning experiments.
  • Coding Assistance / SWE Tasks: Repository-level coding and software engineering workflows.
  • Tool Use & General Agents: Multi-step tool calling, terminal, and automation tasks.
  • Visual Coding: Front-end and UI-oriented visual coding generation.
  • Cybersecurity Research: Structured security analysis and problem solving.
  • Efficient Local Deployment: Reduced memory footprint enables 9B agentic inference on smaller GPUs.

Limitations

  • Experimental Model: Behavior may differ from the base model in certain scenarios.
  • SFT-only Checkpoint: This is a pre-RL SFT release; downstream RL-tuned MiMo models may exhibit stronger agentic performance.
  • Reasoning Artifacts: Complex reasoning chains may occasionally produce incorrect intermediate steps or conclusions.
  • Internal Benchmarks: Some reported evaluation sets (†) are internal and not publicly reproducible.
  • Quantization Drift: FP8 dynamic quantization may introduce minor numerical differences relative to the bf16 source model; downstream accuracy should be validated for production use cases.

License

Released under the license terms consistent with the base model. Refer to XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B for the authoritative licensing information.

Acknowledgements

Citation

If you use this model, please cite the original MiMo-V2.6 release:

@misc{mimo2026v26,
  title={MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement},
  author={{Xiaomi MiMo Team}},
  year={2026},
  howpublished={\url{https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL}},
}
Downloads last month
190
Safetensors
Model size
9B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/MiMo-V2.6-Distill-Qwen-9B-FP8

Finetuned
Qwen/Qwen3.5-9B
Quantized
(56)
this model

Collections including prithivMLmods/MiMo-V2.6-Distill-Qwen-9B-FP8