blink-mimo-9b-GGUF

blink-mimo-9b is a personal research release by thegovind, a LoRA fine-tune (rank 16, 43.3M merged parameters) of MiMo-V2.6-Distill-Qwen-9B built for one-pass typed decisions rather than text generation: given a text/JSON state and up to 512 questions of type choice (≤255 options), noul (yes/no), or score (2–10 ordered levels), it returns FP32-softmax probabilities over only the offered option-letter logits from lm_head in a single forward pass, with the 27-block vision tower left byte-for-byte unchanged and all other language-model tensors bit-identical to the MiMo base. Trained on one pre-registered epoch (123,195 rows, 615 steps) mixing public-source data, program-generated reasoning, decision worlds, and Qwen3.8-27B-authored teacher questions, it lifted the MiMo base's Decision Index 0.1 score from 47.09 (DI-S) to a full-suite score of 56.53 — placing third among locally-compared models behind Jev 1.13.0 (59.51) and the author's own blink-27b, and ahead of larger open entries like Jevfire (27B) and Decider (35B-A3B) — though the authors flag meaningful public-benchmark training overlap (ContractNLI, iSarcasmEval, VAST) that would drop the score to 54.96 if those areas were excluded. On JevBench's public items it solved 48/48 easy, 70/72 standard, and 77/111 hard cases with a hard ECE of 0.136 (less calibrated than the sibling blink-4b's 0.067), and it ships as a Docker-deployable server exposing a Jev-compatible /v1/systemone endpoint with SHA256 weight verification at startup; the MiMo base is declared MIT and its Qwen3.5-9B ancestor is Apache-2.0, but the blink LoRA weights themselves are licensed for non-commercial research and evaluation only.

Model Files

File Name Quant Type File Size File Link Description
blink-mimo-9b.BF16.gguf BF16 17.9 GB Link Full BF16 weights. Highest quality, largest file size.
blink-mimo-9b.Q3_K_L.gguf Q3_K_L 4.93 GB Link Lower quality but usable, good for low RAM availability.
blink-mimo-9b.Q3_K_M.gguf Q3_K_M 4.62 GB Link Low quality.
blink-mimo-9b.Q4_K_M.gguf Q4_K_M 5.63 GB Link Good quality, default size for most use cases, recommended.
blink-mimo-9b.Q4_K_S.gguf Q4_K_S 5.35 GB Link Slightly lower quality with more space savings, recommended.
blink-mimo-9b.Q5_K_M.gguf Q5_K_M 6.47 GB Link High quality, recommended.
blink-mimo-9b.Q5_K_S.gguf Q5_K_S 6.31 GB Link High quality, recommended.
blink-mimo-9b.Q6_K.gguf Q6_K 7.36 GB Link Very high quality, near perfect, recommended.

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Downloads last month
796
GGUF
Model size
9B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

6-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/blink-mimo-9b-GGUF

Finetuned
Qwen/Qwen3.5-9B
Quantized
(1)
this model

Collection including prithivMLmods/blink-mimo-9b-GGUF