Qwen3.8-Flash-Next HFQ — canonical hipfire artifact

File Size SHA-256 MD5 Status
qwen3.8-flash-next.mq4 125288540696 bytes 8aa01cf41bf2a90b319b9a1c70837baf51a2f92811af551f59b59d418518f650 001878abd9b68218876ac0ae732e481b canonical since 2026-09-28
qwen3.8-flash-next.mq6q8-pleq8 125467331096 bytes c0628b848077f02afed9ce5a4daa0598c379aa1d1e5ca0c0b24ddd04b49773a9 8cccd6e1d5c7472ed9a756a099356519 previous rung, kept

Container revision 2026-09-29. Both files were re-tagged in place: the three raw-I64 PLE metadata records (layer_multipliers, ngram_heads_vocab_sizes, ngram_heads_offsets) moved from HFQM quant type 52 to 54, because hipfire now assigns qt=52 to MQ4G256V2L. Only those three index bytes changed; every tensor payload is byte-identical to the 2026-09-28 upload, so the measurements below still apply. Builds before hipfire cb566dab9 refuse the re-tagged files, and builds after the qt=54 cleanup refuse the old ones.

Both are Qwen/Qwen3.8-Flash-Next converted to hipfire's HFQM container for the Qwen4 architecture (arch 16). They are not official Qwen checkpoints, and they must be read by a hipfire build that supports this layout.

qwen3.8-flash-next.mq4 is the same recipe as mq6q8-pleq8 except for one tensor: the language head ships MQ6G256V2 (6.25 bpw) instead of Q8F16 (8.5 bpw). It was quantized from the BF16 checkpoint, not converted from the previous file. The name follows the routed experts, which are 97% of the bytes.

What the recipe is

1325 tensors, four declared tiers:

Class Format Detail
Trunk attention/GDN, language head MQ6G256V2 (qt47) 240 rank-2 projections plus lm_head; 6-bit payload in 256-wide rotated groups, 200 B per group
Routed experts MQ4G256V2 (qt44) + MQ4G128V2 (qt53) gate_up_proj reduces over hidden (2560, aligned to the 256-wide group); down_proj reduces over moe_intermediate_size (640) and takes the row-local 128-wide group
Embedding, MTP attention, PLE n-gram rows Q8F16 (qt3) f16 scale + 32 int8 per block, 34 B per 32 weights; embed_tokens, the MTP layer's five attention projections, and all 128 PLE row shards
Hyper-connection, MoE router, shared expert, PLE projections, norms BF16 keep source bytes

The PLE n-gram table (54.4 GB) is leased from the mapped artifact in pages rather than held resident. The exact per-tensor census is in provenance.json.

Measured against the previous rung

gfx1151 (Strix Halo), same hipfire binary (commit c7c8c52f1), 1131-token committed prompt, greedy, 128 tokens, max_seq 2048, kv q8, graph off, MTP off, fresh process per artifact, 1 warmup + 5 runs:

mq6q8-pleq8 mq4
Decode 32.78 tok/s 33.42 tok/s (two 5-run sets: 33.46, 33.39)
Prefill 1253 tok/s 1255 tok/s
KLD vs the BF16 source, decode route (32 × 512 wikitext-2) 0.07513 0.07602
KLD vs the BF16 source, prefill route (32 × 512) 0.07347 0.07432
Language head read per token 0.68 GB 0.50 GB

The decode numbers include runtime work shipped in the same hipfire build: on gfx1151, single-token forwards read the hyper-connection read projections and the shared expert as Q8_0 copies made at load (+0.9 GB device memory), while prefill reads the BF16 source. Both files get that; the head is the difference between the two columns.

How each tier was decided

  • Trunk at MQ6G256V2. A four-bit trunk (MFP4-E8) hard-failed coherence, and a five-bit trunk (MQ5, requantized from MQ6) measured +0.018 decode KLD for +3% decode; MQ6 ships.
  • Language head at MQ6G256V2. +0.0005 decode KLD against the Q8F16 head for 0.18 GB less read per token. Four bits (MQ4) cost +0.014 KLD, so six is the floor.
  • Experts at MQ4G256V2 / MQ4G128V2. A three-bit simulation of either expert matrix cost +0.06 to +0.11 decode KLD.
  • Embedding, MTP attention, PLE n-gram rows at Q8F16. The embedding and the PLE rows are read one row per token, so their tier is a size decision.
  • Everything else stays BF16 in the file. Storing the hyper-connection read projections, the shared expert and the PLE key/value projections as Q8F16 was built and measured: decode already reads Q8_0 copies of the matrices it streams, while prefill then lost its BF16 source (prefill-route KLD +0.0035) and short prompts lost their fast exact multi-row GEMM (+46% prefill time at ~100 tokens). The router at Q8 cost +0.0025 KLD.

Requirements and caveats

  • Needs a hipfire build at or after commit cb566dab9 (qt=54 I64 metadata records; see the container revision note above). Earlier builds refuse these files at admission.
  • Tested only on 128 GB-class AMD Strix Halo (gfx1151) with ordinary HIP inference. The catalog's 128 GB gate is conservative, not a measured minimum.
  • Quality evidence: the KLD figures above against the BF16 source, and a five-prompt serve battery (code, reasoning, factual, prose, instruction) with 5/5 turns finished stop, 0 runaway, 0 empty and 0 attractor turns, answers read and correct.
  • The MTP figures below need hipfire feat/qwen38-flash-next at 86461c4a1 or later (PR #774, not yet on master). MTP is on by default from ec856487a.
  • Treat this as an experimental artifact and preserve the upstream license notice and attribution when redistributing it.

Native MTP (speculative decoding): 1.77x AR on decode, same tokens

The artifact embeds the trained MTP head.

  • Default: MTP attaches by default from hipfire ec856487a (speculation.mtp = auto). Earlier builds need speculation.mtp = on.
  • Greedy only: MTP runs on greedy requests. Sampled requests, including the recommended temperature 1.0, stay on the AR route, and attaching the head did not measurably change prefill.

Measured on 2026-09-28.

  • Setup: gfx1151, hipfire feat/qwen38-flash-next at 86461c4a1 (daemon md5 d0fc9051a4244f131871e601773ac536), greedy, max_seq 2048, kv q8, graph off.
  • Workload: 9 committed prompts x 200 tokens, a fresh daemon per session, and 3 rounds with the arm order rotated.
  • Statistic: geomean of each prompt's best-of-3 decode tok/s.
file AR MTP MTP / AR mean tau
qwen3.8-flash-next.mq4 33.69 59.47 1.77 2.10
qwen3.8-flash-next.mq6q8-pleq8 (Q8F16 head)¹ 33.14 59.18 1.79 2.09

¹ This row was measured on a copy of mq4 whose head was replaced by the Q8F16 head, byte-identical to this rung's head. The two published files differ only in that tensor.

Greedy MTP emits exactly AR's tokens. Token ids were equal to AR in every session. The few-row verify forward is bitwise equal to single-token decode.

Per prompt on mq4 (AR and MTP are best-of-3; tau is the median):

prompt AR MTP MTP / AR tau
merge_sort_thinking_off 34.0 72.1 2.12 2.90
code_edit_rewrite_copy 33.9 71.1 2.10 2.90
lru_cache_pep8_strict 33.3 66.0 1.98 2.65
trains-meet 33.5 64.5 1.92 2.26
mixed_code_then_prose 33.9 61.7 1.82 2.11
glimmer_prefill_1024 32.8 55.6 1.70 1.65
humaneval_3_below_zero 34.1 52.0 1.52 1.87
bare_factual 33.7 50.2 1.49 1.38
fiction_lighthouse 34.1 47.7 1.40 1.21

How the draft route works:

  • Draft ranking. An MQ2 copy of the language head ranks the vocabulary. Its top 8 are then re-scored exactly against the head's own rows, which here are MQ6G256V2.
    • The MQ6 re-score landed in 86461c4a1. Builds before it drafted from the MQ2 argmax on this file, at 55.4 tok/s (1.64x).
    • The Q8F16 head of the previous rung had the re-score all along.
  • Draft depth. Each window picks its depth, or the interleaved route, from per-depth draft agreement. It stops drafting once the drafts' exact logit margins make the prefix unlikely to be accepted.
  • Verify cost. Verification runs 2..8 target rows in one forward. On a rejection, the kept rows' GDN recurrence is re-run in one launch.

Caveats.

Known instrument limitation

hipfire bench cannot measure this artifact. The Qwen4 admission contract requires max_seq to be exactly 2048, and bench has no way to ask for that: it requests the configured memory.max_seq (32768 by default) raised by the request's token budget, so it fails closed at load, before any measurement, with

qwen4: max_seq must be exactly 2048 (got 32768)

and with memory.max_seq forced to 2048 it still asks for 5120 and fails the same way. This is a limitation of the bench instrument, not of MTP: it is permanent for as long as the architecture pins max_seq, and the fix, if one is wanted, is a --max-seq (or contract-aware) knob on bench, not a change to the MTP route. Until then, do not read any hipfire bench --spec mtp number for this model as an MTP number. The MTP figures in this card come from the serve path, because that is an instrument that can load the artifact.

Note on the recipe

The MTP layer's five attention projections are stored at Q8F16 (qt=3), and its routed experts keep the trunk's 4-bit tiers. That is a deliberate size decision and the encoding is correct — an independent decode of those five tensors against the training-precision source agrees to one quantization step, with exact row geometry. But eight bits is not a "rounding error" for acceptance the way it is for artifact size: this project's own DeepSeek V4 recipe already documents that Q8 quant noise on MTP attention projections costs acceptance, and that published 60-80% acceptance figures assume training precision. Precision alone was measured, though, and it does not account for the remaining gap:

  • Routed experts at Q8. Re-quantizing the MTP head's routed experts to Q8 from the BF16 source left pooled draft agreement unchanged (0.818 vs 0.819).
  • Draft-side tiers. Other draft-side tiers moved agreement only by losing it (hyper-connection matrices at Q8: 0.758).

How to pull and run

hipfire pull qwen3.8:flash-next   # lands at ~/.hipfire/models/qwen3.8-flash-next.mq4
hipfire pull qwen3.8:flash-next-mq6q8-pleq8   # the previous rung

Recommended sampling: temperature 1.0, top_p 0.95, top_k 20. Qwen4 requires max_seq 2048 and the legacy (contiguous) KV backend; decode numbers on other configurations will differ.

Upstream provenance and license

The source checkpoint is Qwen/Qwen3.8-Flash-Next at pinned revision de4b8e4d43b917e7706784d8bb445c9af86a3540. The model remains under the Qwen Community License 1.0; this repository does not relicense the upstream weights.

The quantized container and runtime integration are hipfire work, licensed separately under Apache-2.0 and MIT. No hipfire source code is included in this model repository.

Publication record

provenance.json carries the artifact identity (size, MD5, SHA-256), the per-class tensor census, the loader contract, the validation boundary, and the exact upload command. It is the authoritative record for this file; where this card and that file disagree, the file wins.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for hipfire-models/qwen3.8-flash-next

Finetuned
(67)
this model