MTPLX: the fastest way to run MiMo-V2.6-Distill-Qwen-9B on a Mac. Native multi-token-prediction speculative decoding on Apple Silicon, exact at any temperature.

MiMo V2.6 Qwen 9B Optimized Speed

6-bit quantization. Xiaomi's agentic coding distill of Qwen 3.5 9B.

Xiaomi MiMo retrained Qwen 3.5 9B for coding, agent work, visual coding and security tasks, and the result is a big step up from the model it started from. This pack runs it on MTPLX with a multi-token-prediction head drafting ahead, on any Apple Silicon Mac with 16 GB or more.

Available from MTPLX 2.12.0. The Mac app lists it, and on M3, M4 and M5 Macs with 16 to 31 GB it is the second suggestion, right after Bonsai 2 27B.

Why this model

This is the MiMo-V2.6-Distill-Qwen-9B MLX pack for MTPLX, the fastest way to run MiMo V2.6 on a Mac. Xiaomi's results against the Qwen 3.5 9B it started from, from Xiaomi's model card:

Benchmark Qwen3.5-9B MiMo-V2.6-Distill-Qwen-9B
SWE Pro 32.0 44.6
Terminal Bench 2.1 27.0 37.1
Toolathlon-Verified 25.9 35.2
SWE Verified 60.0 61.1

Same size as Qwen 3.5 9B, a much stronger coding and agent model.

How it is built

  • 6-bit quantization with 64-weight groups across the model, the recipe of our Qwen 3.5 9B Optimized Speed. Against the full-precision model it keeps 97.3 percent top-1 agreement, with a KL divergence of 0.0054.
  • The Qwen 3.5 9B multi-token-prediction head drafts ahead, and the full model verifies its drafts in one pass.
Download 8.7 GB
Peak while serving 8.7 GiB at a 15K-token context
Recommended Macs 16 GB+ unified memory (18 GB+ for agent clients such as OpenCode)
Context window 20,480 tokens on 16 GB, 45,056 on 18 GB, 192,512 on 24 GB, up to 262,144
MTP depth 2
Sampling temperature 0.6, top-p 0.95, top-k 20 (Xiaomi's settings)

The serving contract ships inside mtplx_runtime.json. MTPLX reads it on load. Drafts are accepted with the probability-ratio rule plus residual resampling, so the output follows the model's own distribution at any temperature.

Use it

Mac app (MTPLX 2.12.0 or later): download at mtplx.com, pick "MiMo V2.6 Qwen 9B Optimized Speed".

Command line (MTPLX 2.12.0 or later):

pip install mtplx
mtplx serve --model Youssofal/MiMo-V2.6-Qwen-9B-MTPLX-Optimized-Speed

Siblings: Qwen 3.5 9B Optimized Speed and Bonsai 2 27B Optimized Speed (a 27B-class model in about the same memory).

Base model: XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B (MIT), a fine-tune of Qwen/Qwen3.5-9B (Apache 2.0), whose draft head this pack uses.

Downloads last month
698
Safetensors
Model size
9B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Youssofal/MiMo-V2.6-Qwen-9B-MTPLX-Optimized-Speed

Finetuned
Qwen/Qwen3.5-9B
Quantized
(53)
this model