Qwen3.8-Flash-Next HFQ — canonical hipfire artifact
| File | Size | SHA-256 | MD5 | Status |
|---|---|---|---|---|
qwen3.8-flash-next.mq4 |
125288540696 bytes | 8aa01cf41bf2a90b319b9a1c70837baf51a2f92811af551f59b59d418518f650 |
001878abd9b68218876ac0ae732e481b |
canonical since 2026-09-28 |
qwen3.8-flash-next.mq6q8-pleq8 |
125467331096 bytes | c0628b848077f02afed9ce5a4daa0598c379aa1d1e5ca0c0b24ddd04b49773a9 |
8cccd6e1d5c7472ed9a756a099356519 |
previous rung, kept |
Container revision 2026-09-29. Both files were re-tagged in place: the
three raw-I64 PLE metadata records (layer_multipliers,
ngram_heads_vocab_sizes, ngram_heads_offsets) moved from HFQM quant type
52 to 54, because hipfire now assigns qt=52 to MQ4G256V2L. Only those three
index bytes changed; every tensor payload is byte-identical to the 2026-09-28
upload, so the measurements below still apply. Builds before hipfire
cb566dab9 refuse the re-tagged files, and builds after the qt=54 cleanup
refuse the old ones.
Both are Qwen/Qwen3.8-Flash-Next converted to hipfire's HFQM container for the Qwen4 architecture (arch 16). They are not official Qwen checkpoints, and they must be read by a hipfire build that supports this layout.
qwen3.8-flash-next.mq4 is the same recipe as mq6q8-pleq8 except for one
tensor: the language head ships MQ6G256V2 (6.25 bpw) instead of Q8F16
(8.5 bpw). It was quantized from the BF16 checkpoint, not converted from the
previous file. The name follows the routed experts, which are 97% of the bytes.
What the recipe is
1325 tensors, four declared tiers:
| Class | Format | Detail |
|---|---|---|
| Trunk attention/GDN, language head | MQ6G256V2 (qt47) | 240 rank-2 projections plus lm_head; 6-bit payload in 256-wide rotated groups, 200 B per group |
| Routed experts | MQ4G256V2 (qt44) + MQ4G128V2 (qt53) | gate_up_proj reduces over hidden (2560, aligned to the 256-wide group); down_proj reduces over moe_intermediate_size (640) and takes the row-local 128-wide group |
| Embedding, MTP attention, PLE n-gram rows | Q8F16 (qt3) | f16 scale + 32 int8 per block, 34 B per 32 weights; embed_tokens, the MTP layer's five attention projections, and all 128 PLE row shards |
| Hyper-connection, MoE router, shared expert, PLE projections, norms | BF16 | keep source bytes |
The PLE n-gram table (54.4 GB) is leased from the mapped artifact in pages
rather than held resident. The exact per-tensor census is in provenance.json.
Measured against the previous rung
gfx1151 (Strix Halo), same hipfire binary (commit c7c8c52f1), 1131-token
committed prompt, greedy, 128 tokens, max_seq 2048, kv q8, graph off, MTP off,
fresh process per artifact, 1 warmup + 5 runs:
mq6q8-pleq8 |
mq4 |
|
|---|---|---|
| Decode | 32.78 tok/s | 33.42 tok/s (two 5-run sets: 33.46, 33.39) |
| Prefill | 1253 tok/s | 1255 tok/s |
| KLD vs the BF16 source, decode route (32 × 512 wikitext-2) | 0.07513 | 0.07602 |
| KLD vs the BF16 source, prefill route (32 × 512) | 0.07347 | 0.07432 |
| Language head read per token | 0.68 GB | 0.50 GB |
The decode numbers include runtime work shipped in the same hipfire build: on gfx1151, single-token forwards read the hyper-connection read projections and the shared expert as Q8_0 copies made at load (+0.9 GB device memory), while prefill reads the BF16 source. Both files get that; the head is the difference between the two columns.
How each tier was decided
- Trunk at MQ6G256V2. A four-bit trunk (MFP4-E8) hard-failed coherence, and a five-bit trunk (MQ5, requantized from MQ6) measured +0.018 decode KLD for +3% decode; MQ6 ships.
- Language head at MQ6G256V2. +0.0005 decode KLD against the Q8F16 head for 0.18 GB less read per token. Four bits (MQ4) cost +0.014 KLD, so six is the floor.
- Experts at MQ4G256V2 / MQ4G128V2. A three-bit simulation of either expert matrix cost +0.06 to +0.11 decode KLD.
- Embedding, MTP attention, PLE n-gram rows at Q8F16. The embedding and the PLE rows are read one row per token, so their tier is a size decision.
- Everything else stays BF16 in the file. Storing the hyper-connection read projections, the shared expert and the PLE key/value projections as Q8F16 was built and measured: decode already reads Q8_0 copies of the matrices it streams, while prefill then lost its BF16 source (prefill-route KLD +0.0035) and short prompts lost their fast exact multi-row GEMM (+46% prefill time at ~100 tokens). The router at Q8 cost +0.0025 KLD.
Requirements and caveats
- Needs a hipfire build at or after commit
cb566dab9(qt=54 I64 metadata records; see the container revision note above). Earlier builds refuse these files at admission. - Tested only on 128 GB-class AMD Strix Halo (
gfx1151) with ordinary HIP inference. The catalog's 128 GB gate is conservative, not a measured minimum. - Quality evidence: the KLD figures above against the BF16 source, and a
five-prompt serve battery (code, reasoning, factual, prose, instruction) with
5/5 turns finished
stop, 0 runaway, 0 empty and 0 attractor turns, answers read and correct. - The MTP figures below need hipfire
feat/qwen38-flash-nextat86461c4a1or later (PR #774, not yet onmaster). MTP is on by default fromec856487a. - Treat this as an experimental artifact and preserve the upstream license notice and attribution when redistributing it.
Native MTP (speculative decoding): 1.77x AR on decode, same tokens
The artifact embeds the trained MTP head.
- Default: MTP attaches by default from hipfire
ec856487a(speculation.mtp = auto). Earlier builds needspeculation.mtp = on. - Greedy only: MTP runs on greedy requests. Sampled requests, including the
recommended
temperature 1.0, stay on the AR route, and attaching the head did not measurably change prefill.
Measured on 2026-09-28.
- Setup: gfx1151, hipfire
feat/qwen38-flash-nextat86461c4a1(daemon md5d0fc9051a4244f131871e601773ac536), greedy,max_seq 2048,kv q8, graph off. - Workload: 9 committed prompts x 200 tokens, a fresh daemon per session, and 3 rounds with the arm order rotated.
- Statistic: geomean of each prompt's best-of-3 decode tok/s.
| file | AR | MTP | MTP / AR | mean tau |
|---|---|---|---|---|
qwen3.8-flash-next.mq4 |
33.69 | 59.47 | 1.77 | 2.10 |
qwen3.8-flash-next.mq6q8-pleq8 (Q8F16 head)¹ |
33.14 | 59.18 | 1.79 | 2.09 |
¹ This row was measured on a copy of mq4 whose head was replaced by the Q8F16
head, byte-identical to this rung's head. The two published files differ only
in that tensor.
Greedy MTP emits exactly AR's tokens. Token ids were equal to AR in every session. The few-row verify forward is bitwise equal to single-token decode.
Per prompt on mq4 (AR and MTP are best-of-3; tau is the median):
| prompt | AR | MTP | MTP / AR | tau |
|---|---|---|---|---|
merge_sort_thinking_off |
34.0 | 72.1 | 2.12 | 2.90 |
code_edit_rewrite_copy |
33.9 | 71.1 | 2.10 | 2.90 |
lru_cache_pep8_strict |
33.3 | 66.0 | 1.98 | 2.65 |
trains-meet |
33.5 | 64.5 | 1.92 | 2.26 |
mixed_code_then_prose |
33.9 | 61.7 | 1.82 | 2.11 |
glimmer_prefill_1024 |
32.8 | 55.6 | 1.70 | 1.65 |
humaneval_3_below_zero |
34.1 | 52.0 | 1.52 | 1.87 |
bare_factual |
33.7 | 50.2 | 1.49 | 1.38 |
fiction_lighthouse |
34.1 | 47.7 | 1.40 | 1.21 |
How the draft route works:
- Draft ranking. An MQ2 copy of the language head ranks the vocabulary.
Its top 8 are then re-scored exactly against the head's own rows, which
here are MQ6G256V2.
- The MQ6 re-score landed in
86461c4a1. Builds before it drafted from the MQ2 argmax on this file, at 55.4 tok/s (1.64x). - The Q8F16 head of the previous rung had the re-score all along.
- The MQ6 re-score landed in
- Draft depth. Each window picks its depth, or the interleaved route, from per-depth draft agreement. It stops drafting once the drafts' exact logit margins make the prefix unlikely to be accepted.
- Verify cost. Verification runs 2..8 target rows in one forward. On a rejection, the kept rows' GDN recurrence is re-run in one launch.
Caveats.
- These numbers are specific to this fixture and gfx1151.
- The GPU was shared with an external process; best-of-3 absorbs that.
- Method and raw per-run data:
docs/perf-checkpoints/2026-09-28-qwen4-mtp-mq6-rescore-gfx1151.mdin hipfire PR #774, which is not yet onmaster.
Known instrument limitation
hipfire bench cannot measure this artifact. The Qwen4 admission contract
requires max_seq to be exactly 2048, and bench has no way to ask for that: it
requests the configured memory.max_seq (32768 by default) raised by the
request's token budget, so it fails closed at load, before any measurement, with
qwen4: max_seq must be exactly 2048 (got 32768)
and with memory.max_seq forced to 2048 it still asks for 5120 and fails the
same way. This is a limitation of the bench instrument, not of MTP: it is
permanent for as long as the architecture pins max_seq, and the fix, if one is
wanted, is a --max-seq (or contract-aware) knob on bench, not a change to the
MTP route. Until then, do not read any hipfire bench --spec mtp number for this
model as an MTP number. The MTP figures in this card come from the serve path,
because that is an instrument that can load the artifact.
Note on the recipe
The MTP layer's five attention projections are stored at Q8F16 (qt=3), and its
routed experts keep the trunk's 4-bit tiers. That is a deliberate size decision
and the encoding is correct — an independent decode of those five tensors against
the training-precision source agrees to one quantization step, with exact row
geometry. But eight bits is not a "rounding error" for acceptance the way it is
for artifact size: this project's own DeepSeek V4 recipe already documents that
Q8 quant noise on MTP attention projections costs acceptance, and that published
60-80% acceptance figures assume training precision. Precision alone was
measured, though, and it does not account for the remaining gap:
- Routed experts at Q8. Re-quantizing the MTP head's routed experts to Q8 from the BF16 source left pooled draft agreement unchanged (0.818 vs 0.819).
- Draft-side tiers. Other draft-side tiers moved agreement only by losing it (hyper-connection matrices at Q8: 0.758).
How to pull and run
hipfire pull qwen3.8:flash-next # lands at ~/.hipfire/models/qwen3.8-flash-next.mq4
hipfire pull qwen3.8:flash-next-mq6q8-pleq8 # the previous rung
Recommended sampling: temperature 1.0, top_p 0.95, top_k 20. Qwen4
requires max_seq 2048 and the legacy (contiguous) KV backend; decode numbers on other
configurations will differ.
Upstream provenance and license
The source checkpoint is
Qwen/Qwen3.8-Flash-Next at
pinned revision
de4b8e4d43b917e7706784d8bb445c9af86a3540.
The model remains under the Qwen Community License 1.0; this
repository does not relicense the upstream weights.
The quantized container and runtime integration are hipfire work, licensed separately under Apache-2.0 and MIT. No hipfire source code is included in this model repository.
Publication record
provenance.json carries the artifact identity (size, MD5, SHA-256), the
per-class tensor census, the loader contract, the validation boundary, and the
exact upload command. It is the authoritative record for this file; where this
card and that file disagree, the file wins.
Model tree for hipfire-models/qwen3.8-flash-next
Base model
Qwen/Qwen3.8-Flash-Next