Benchmarks · measured 2026-10-07
hipfire 0.4.1 benchmarks
hipfire 0.4.1 tunes Qwen3.8-Flash (qwen3.8:flash-next) for agent work on
Strix Halo and Radeon AI PRO R9700: native speculative decoding that now covers sampled requests with
penalties, a session cache that keeps many prompts warm, exact tokenization, and PeaceMaker-built kernels.
These are the measured numbers.
Strix Halo (gfx1151) · agent tasks
Ciru's TC70–84 tool-use panel and the Hermes-20 agent suite, both run end to end with release defaults. Seconds are the sum of per-scenario (TC) or per-case (Hermes) wall time.
| suite | result | seconds |
|---|---|---|
| TC70–84 · sampled · seed 123 | 23/30 | 169.9 |
| TC70–84 · sampled · seed 124 | 24/30 | 158.2 |
| TC70–84 · sampled · seed 125 | 21/30 | 162.9 |
| TC70–84 · sampled · median of 3 seeds | 23/30 | 162.9 |
| TC70–84 · greedy · 1 seeda | 21/30 | 126.4 |
| Hermes-20 · pass 1 | 95 (18/20) | 729.8 |
| Hermes-20 · pass 2 | 96 (18/20) | 782.0 |
TC70–84 sampled: seeds 123 / 124 / 125, seconds 158.2–169.9, T0.7 / top-p 0.8 / top-k 20 / presence 1.5, thinking off, 32 turns. Hermes-20: thinking on (xhigh), seeds 160916 / 160917, one run per pass; Hermes scores vary by several points from run to run. TC-75 and TC-76 score 0 on every sampled seed. a Greedy panel (T0, top-k 1, 8 turns): a single seed, measured on an earlier 0.4.1 build; the release notes list the build for each cell.
Cold prefill and decode
| cell | median tok/s | min – max |
|---|---|---|
| synthetic pp8192, cold prefill | 2072.9 | 2072.7 – 2076.9 |
| Ciru 32K prompt, cold prefill | 1776.5 | 1776.3 – 1780.3 |
| decode, greedy AR | 33.78 | 33.77 – 33.88 |
| decode, greedy native MTP | 58.67 | 58.54 – 58.76 |
| decode, sampled native MTP + presence penalty | 56.98 | 45.55 – 57.21 |
Fresh-process medians of three repetitions, warm sample discarded. Sampled MTP uses T0.7 / top-p 0.8 / top-k 20 and a presence penalty of 1.5; one of its three repetitions ran slow (45.55 tok/s), so its range is wide. Penalties cost about nothing on AR and at most about 1.5 % on sampled MTP.
Measured 2026-10-07 on Strix Halo (gfx1151), Qwen3.8-Flash GPTQ3 MQ4 (qwen3.8:flash-next), with Ciru's
recommended TC70–84 protocol and the Hermes-20 suite. Ciru's
Strix Showdown board is the public reference for those protocols and for how other
engines score on them; different hosts make it a reference, not a controlled comparison.
Why it's fast
Native sampled MTP, with penalties
Flash-Next's own MTP head now verifies sampled requests by speculative rejection sampling and applies repeat, presence and frequency penalties to both the target and draft distributions. Agent harnesses sample, and Ciru's panel adds a presence penalty of 1.5; before this release every temperature > 0 request ran plain autoregressive decode. A GPU penalty prepass keeps the penalty cost small; AR stays the floor when speculation would lose.
A session cache that survives the next prompt
An engine-owned session cache keeps many prompts warm instead of one, so a second session or a sub-agent no longer evicts the first. It is built on @fivetide's session-cache work (PRs #825 / #826): chunk-boundary snapshots with delta state. Live continuation inside a conversation is layered on top. Thank you.
Tokenizer fidelity
Our byte-level BPE pre-tokenizer was missing Hugging Face's \s+(?!\S) whitespace rule, so
indented prompts tokenized differently from the reference. Exact semantics now match HF on 945 of 945
corpus cases (Hermes prompts, every TC request, bench prompts, synthetic whitespace); in the Hermes runs the
fix removed an indentation-error and request-loop failure mode.
PeaceMaker kernels
The gathered QSA attention, its score/select step on Strix Halo, and the MQ6 trunk now run from certified PeaceMaker-built code objects, byte-for-byte replacements for the hipcc kernels, each with an env opt-out. The dense GDN scan and symmetric-IU4 MoE are the exceptions: they trade KLD for prefill speed and are KLD-gated, not bit-exact.
Reproduce it
# install hipfire 0.4.1 and pull the exact artifact $ curl -L https://raw.githubusercontent.com/warpfront/hipfire/v0.4.1/scripts/install.sh | bash $ hipfire pull qwen3.8:flash-next # 116.7 GiB # serve by tag: a model id containing "gpt" makes Hermes inject GPT prompt blocks $ hipfire serve qwen3.8:flash-next # Ciru's TC70–84, recommended sampler, seed 123 (tool-eval-bench v2.0.7 from Ciru's repo) $ tool-eval-bench --model qwen3.8:flash-next --base-url http://127.0.0.1:11435/v1 --api-key local \ --backend llamacpp --hardmode-only --no-think --temperature .7 --top-p .8 --top-k 20 --min-p 0 \ --repeat-penalty 1 --seed 123 --max-turns 32 --trials 1 --parallel 1 --timeout 1800 \ --structured-response-format json_object \ --backend-kwargs '{"temperature":0.7,"top_p":0.8,"top_k":20,"min_p":0,"presence_penalty":1.5,"frequency_penalty":0,"repeat_penalty":1,"seed":123,"cache_prompt":true,"max_tokens":-1,"chat_template_kwargs":{"enable_thinking":false,"preserve_thinking":true}}' \ --scenarios TC-70 TC-71 TC-72 TC-73 TC-74 TC-75 TC-76 TC-77 TC-78 TC-79 TC-80 TC-81 TC-82 TC-83 TC-84 # Hermes-20: Ciru's verifier image and case set, thinking on (xhigh), seeds 160916 / 160917
| artifact | checksum |
|---|---|
qwen3.8-flash-next-gptq3.mq4 sha256 | 8b15b6fede7d7c5bfed0db4720a8295bedda51bc93e545fa242bd50d0f200972 |
qwen3.8-flash-next-gptq3.mq4 md5 · 125,288,544,792 B | be007fc3219e9f6cdb1d4dfa8380625f |
daemon md5 (Strix Halo measurement build) | 040e12ee04030c182277ba364c4dfc27 |
hipfire md5 (Strix Halo measurement build) | dacebd7d2a50eca134280c804ca9bfac |
Model: hipfire-models/qwen3.8-flash-next (the sha256 is the Hugging Face LFS object
id; the file size matches the measured copy). Harness:
ciru-ai/strix-showdown-20261001. The two binary md5s are the build the Strix Halo
numbers were measured with; a build of the release commit can differ byte for byte. The release page carries
the kernel-pack checksums (release page). Protocol:
BENCHMARKS.md,
CHANGELOG.
Radeon AI PRO R9700 (gfx1201)
Same artifact on a discrete card, with VMM context storage on by default. Medians of three fresh-process
repetitions, 2026-10-07, measured on an earlier 0.4.1 build; the R9700 Flash-Next lane was not repeated
on the final release build (the release notes list the build for each cell).
cell median tok/s min – max synthetic pp8192, cold prefill 3817.0 3801.6 – 3819.9 Ciru 8K prompt, cold prefill 2082.6 2078.0 – 2087.8 Ciru 32K prompt, cold prefill 2051.3 2047.6 – 2053.6 Ciru 128K prompt, cold prefill 1814.6 1811.9 – 1816.5 decode, greedy AR 34.92 34.88 – 35.01 decode, greedy native MTP 40.71 40.12 – 40.86 decode, sampled native MTP 41.09 40.99 – 41.62
Tool Gauntlet on the same card, one sampled seed (123): 24/30 in 256.6 s;
Hermes smoke 100 (3/3) in 94.6 s. A
larger R9700 prefill chunk is not promoted: its scratch costs expert residency and regresses decode.
27B · Qwen3.8 27B MQ4 XTS
qwen3.8:27b-mq4-xts (14.0 GiB, sha256 3e38ccbae3776470eb5a89344d300e9279d6b9ab6c31fd40ca1758c4f7c6f8ae), medians of
three fresh-process repetitions per GPU, 2026-10-07. Decode is eight prompts, greedy, 256 tokens. Native MTP uses
the model's MTP sidecar and is the default when it is installed; DFlash uses the paired draft and is opt-in.
GPU prompt tok/s (8K) decode tok/s + native MTP + DFlash (8 mixed prompts) Radeon RX 7900 XTX gfx1100 3021.5 51.63 87.52 τ 2.42 131.69 τ 7.23 Radeon AI PRO R9700 gfx1201 5166.1 41.06 67.95 τ 2.35 123.30 τ 7.15 Strix Halo gfx1151 1192.7 15.08 27.45 τ 2.40 43.21 τ 7.23 Radeon AI PRO R9700 gfx1201 · DFlash by prompt type
code prompts: 238.7 tok/s median
(5 code prompts × 3 runs, single stream, greedy);
prose 58.4
R9700 DFlash code prompts, per-prompt median tok/s:
code-edit copy 271.2 · merge sort 269.2 · glimmer coding 238.7 · HumanEval below_zero 229.2 · LRU cache 201.4; prose (fiction)
58.4. The table's DFlash column is the eight-mixed-prompt median.
τ is the mean number of tokens emitted per target forward. Speculative speed depends on the prompt: code and copy-heavy
text accept long drafts, open-ended prose accepts fewer. The first DFlash run on gfx1100 and gfx1201 compiles a kernel
that the 0.4.1 kernel pack does not yet include.
Known limits
- Tool score. 23/30 median on the sampled TC70–84 panel; TC-75 and TC-76 score
0 on every sampled seed, also with the cache off.
- Hermes pass 1. 95 with 18/20 cases passing, against 96 for pass
2. Runs are single and noisy; the pass-1 time is dominated by generation volume, not decode speed.
- Greedy chains. A five-turn greedy chain can end turn 4 inside its reasoning. It is
deterministic, identical with MTP on and off and with the cache off, so it is the model's trajectory and not
cache state (release notes, known issues).
Next
0.4.2 is about more prefill and decode speed on Strix Halo. No numbers until they are measured.
Earlier snapshots
Dated, before 0.4.1. These fixtures predate the 0.4.1 release matrix and were not re-measured on 0.4.1. Each figure is scoped to its fixture and date; /docs/benchmarks is the build-time live ledger.
Three independent fixtures — different models, quants, and hardware setups. Rows are heterogeneous fixtures. Compare cells only when model, quant/mode, prompt, method, and backend match.
Qwen3.6 35B-A3B MQ4R · single-GPU AR
Ordinary autoregressive decode, single GPU; no MTP, DFlash, reduced-output bench, or manual clock pinning. TG128 = three-run medians. Multi-turn columns from clean eight-turn serving runs.
| GPU | arch | TG128 AR | 8-turn avg | final turn |
|---|---|---|---|---|
| Radeon RX 7900 XTX | gfx1100 | 253.3 tok/s | 191.0 tok/s | 160.3 @ 18.2K |
| Radeon 8060S / Strix Halo | gfx1151 | 115.1 tok/s | 92.2 tok/s | 82.5 @ 21.3K |
| Radeon AI PRO R9700 | gfx1201 | 203.9 tok/s | 169.5 tok/s | 146.7 @ 22.2K |
Source: README · Qwen3.6 35B-A3B MQ4R · live ledger: /docs/benchmarks
Qwen3.8-27B MQ4V2 · 2026-08-20 · hiptrx · gfx1201 · Radeon AI PRO R9700
Product-ladder checkpoint on the hiptrx gfx1201 / Radeon AI PRO R9700 fixture. AR decode, prefill, DFlash decode, mean accepted draft length (τ), and bits-per-weight.
| tier | AR decode | prefill | DFlash | τ | bpw |
|---|---|---|---|---|---|
| XT | 35.3 tok/s | 490.7 | 251.6 tok/s | 11.70 | 4.456 |
| Base | 33.2 tok/s | 479.0 | 263.3 tok/s | 13.11 | 4.659 |
| Pro | 31.8 tok/s | 473.8 | 258.2 tok/s | 13.11 | 4.897 |
Dated 2026-08-20. Source:
2026-08-20 Qwen3.8 MQ-V2 product ladder
(docs/perf-checkpoints/2026-08-20-qwen38-mq-v2-product-ladder.md)
DeepSeek V4 Flash MQ2R · 4× Radeon AI PRO R9700 (gfx1201)
n=3 fresh-process medians; greedy; speculative decode off; KV f32; 2052-token prompt.
| topology | decode | prefill |
|---|---|---|
| TP3/EP | 53.1 tok/s | 481 tok/s |
| TP4/EP | 54.3 tok/s | 389 tok/s |
Source: README · RDNA4 gfx1201 R9700 · live ledger: /docs/benchmarks