hipfire

Strix Halo (gfx1151) · agent tasks

Ciru's TC70–84 tool-use panel and the Hermes-20 agent suite, both run end to end with release defaults. Seconds are the sum of per-scenario (TC) or per-case (Hermes) wall time.

suite result seconds
TC70–84 · sampled · seed 123 23/30 169.9
TC70–84 · sampled · seed 124 24/30 158.2
TC70–84 · sampled · seed 125 21/30 162.9
TC70–84 · sampled · median of 3 seeds 23/30 162.9
TC70–84 · greedy · 1 seeda 21/30 126.4
Hermes-20 · pass 1 95 (18/20) 729.8
Hermes-20 · pass 2 96 (18/20) 782.0

TC70–84 sampled: seeds 123 / 124 / 125, seconds 158.2–169.9, T0.7 / top-p 0.8 / top-k 20 / presence 1.5, thinking off, 32 turns. Hermes-20: thinking on (xhigh), seeds 160916 / 160917, one run per pass; Hermes scores vary by several points from run to run. TC-75 and TC-76 score 0 on every sampled seed. a Greedy panel (T0, top-k 1, 8 turns): a single seed, measured on an earlier 0.4.1 build; the release notes list the build for each cell.

Cold prefill and decode

cell median tok/s min – max
synthetic pp8192, cold prefill 2072.9 2072.7 – 2076.9
Ciru 32K prompt, cold prefill 1776.5 1776.3 – 1780.3
decode, greedy AR 33.78 33.77 – 33.88
decode, greedy native MTP 58.67 58.54 – 58.76
decode, sampled native MTP + presence penalty 56.98 45.55 – 57.21

Fresh-process medians of three repetitions, warm sample discarded. Sampled MTP uses T0.7 / top-p 0.8 / top-k 20 and a presence penalty of 1.5; one of its three repetitions ran slow (45.55 tok/s), so its range is wide. Penalties cost about nothing on AR and at most about 1.5 % on sampled MTP.

Measured 2026-10-07 on Strix Halo (gfx1151), Qwen3.8-Flash GPTQ3 MQ4 (qwen3.8:flash-next), with Ciru's recommended TC70–84 protocol and the Hermes-20 suite. Ciru's Strix Showdown board is the public reference for those protocols and for how other engines score on them; different hosts make it a reference, not a controlled comparison.

Why it's fast

Native sampled MTP, with penalties

Flash-Next's own MTP head now verifies sampled requests by speculative rejection sampling and applies repeat, presence and frequency penalties to both the target and draft distributions. Agent harnesses sample, and Ciru's panel adds a presence penalty of 1.5; before this release every temperature > 0 request ran plain autoregressive decode. A GPU penalty prepass keeps the penalty cost small; AR stays the floor when speculation would lose.

A session cache that survives the next prompt

An engine-owned session cache keeps many prompts warm instead of one, so a second session or a sub-agent no longer evicts the first. It is built on @fivetide's session-cache work (PRs #825 / #826): chunk-boundary snapshots with delta state. Live continuation inside a conversation is layered on top. Thank you.

Tokenizer fidelity

Our byte-level BPE pre-tokenizer was missing Hugging Face's \s+(?!\S) whitespace rule, so indented prompts tokenized differently from the reference. Exact semantics now match HF on 945 of 945 corpus cases (Hermes prompts, every TC request, bench prompts, synthetic whitespace); in the Hermes runs the fix removed an indentation-error and request-loop failure mode.

PeaceMaker kernels

The gathered QSA attention, its score/select step on Strix Halo, and the MQ6 trunk now run from certified PeaceMaker-built code objects, byte-for-byte replacements for the hipcc kernels, each with an env opt-out. The dense GDN scan and symmetric-IU4 MoE are the exceptions: they trade KLD for prefill speed and are KLD-gated, not bit-exact.

Reproduce it

# install hipfire 0.4.1 and pull the exact artifact
$ curl -L https://raw.githubusercontent.com/warpfront/hipfire/v0.4.1/scripts/install.sh | bash
$ hipfire pull qwen3.8:flash-next       # 116.7 GiB

# serve by tag: a model id containing "gpt" makes Hermes inject GPT prompt blocks
$ hipfire serve qwen3.8:flash-next

# Ciru's TC70–84, recommended sampler, seed 123 (tool-eval-bench v2.0.7 from Ciru's repo)
$ tool-eval-bench --model qwen3.8:flash-next --base-url http://127.0.0.1:11435/v1 --api-key local \
    --backend llamacpp --hardmode-only --no-think --temperature .7 --top-p .8 --top-k 20 --min-p 0 \
    --repeat-penalty 1 --seed 123 --max-turns 32 --trials 1 --parallel 1 --timeout 1800 \
    --structured-response-format json_object \
    --backend-kwargs '{"temperature":0.7,"top_p":0.8,"top_k":20,"min_p":0,"presence_penalty":1.5,"frequency_penalty":0,"repeat_penalty":1,"seed":123,"cache_prompt":true,"max_tokens":-1,"chat_template_kwargs":{"enable_thinking":false,"preserve_thinking":true}}' \
    --scenarios TC-70 TC-71 TC-72 TC-73 TC-74 TC-75 TC-76 TC-77 TC-78 TC-79 TC-80 TC-81 TC-82 TC-83 TC-84

# Hermes-20: Ciru's verifier image and case set, thinking on (xhigh), seeds 160916 / 160917
artifactchecksum
qwen3.8-flash-next-gptq3.mq4 sha2568b15b6fede7d7c5bfed0db4720a8295bedda51bc93e545fa242bd50d0f200972
qwen3.8-flash-next-gptq3.mq4 md5 · 125,288,544,792 Bbe007fc3219e9f6cdb1d4dfa8380625f
daemon md5 (Strix Halo measurement build)040e12ee04030c182277ba364c4dfc27
hipfire md5 (Strix Halo measurement build)dacebd7d2a50eca134280c804ca9bfac

Model: hipfire-models/qwen3.8-flash-next (the sha256 is the Hugging Face LFS object id; the file size matches the measured copy). Harness: ciru-ai/strix-showdown-20261001. The two binary md5s are the build the Strix Halo numbers were measured with; a build of the release commit can differ byte for byte. The release page carries the kernel-pack checksums (release page). Protocol: BENCHMARKS.md, CHANGELOG.

Radeon AI PRO R9700 (gfx1201)

Same artifact on a discrete card, with VMM context storage on by default. Medians of three fresh-process repetitions, 2026-10-07, measured on an earlier 0.4.1 build; the R9700 Flash-Next lane was not repeated on the final release build (the release notes list the build for each cell).

cellmedian tok/smin – max
synthetic pp8192, cold prefill 3817.0 3801.6 – 3819.9
Ciru 8K prompt, cold prefill 2082.6 2078.0 – 2087.8
Ciru 32K prompt, cold prefill 2051.3 2047.6 – 2053.6
Ciru 128K prompt, cold prefill 1814.6 1811.9 – 1816.5
decode, greedy AR 34.92 34.88 – 35.01
decode, greedy native MTP 40.71 40.12 – 40.86
decode, sampled native MTP 41.09 40.99 – 41.62

Tool Gauntlet on the same card, one sampled seed (123): 24/30 in 256.6 s; Hermes smoke 100 (3/3) in 94.6 s. A larger R9700 prefill chunk is not promoted: its scratch costs expert residency and regresses decode.

27B · Qwen3.8 27B MQ4 XTS

qwen3.8:27b-mq4-xts (14.0 GiB, sha256 3e38ccbae3776470eb5a89344d300e9279d6b9ab6c31fd40ca1758c4f7c6f8ae), medians of three fresh-process repetitions per GPU, 2026-10-07. Decode is eight prompts, greedy, 256 tokens. Native MTP uses the model's MTP sidecar and is the default when it is installed; DFlash uses the paired draft and is opt-in.

GPU prompt tok/s (8K) decode tok/s + native MTP + DFlash (8 mixed prompts)
Radeon RX 7900 XTX gfx1100 3021.5 51.63 87.52 τ 2.42 131.69 τ 7.23
Radeon AI PRO R9700 gfx1201 5166.1 41.06 67.95 τ 2.35 123.30 τ 7.15
Strix Halo gfx1151 1192.7 15.08 27.45 τ 2.40 43.21 τ 7.23
Radeon AI PRO R9700 gfx1201 · DFlash by prompt type code prompts: 238.7 tok/s median (5 code prompts × 3 runs, single stream, greedy); prose 58.4

R9700 DFlash code prompts, per-prompt median tok/s: code-edit copy 271.2 · merge sort 269.2 · glimmer coding 238.7 · HumanEval below_zero 229.2 · LRU cache 201.4; prose (fiction) 58.4. The table's DFlash column is the eight-mixed-prompt median.

τ is the mean number of tokens emitted per target forward. Speculative speed depends on the prompt: code and copy-heavy text accept long drafts, open-ended prose accepts fewer. The first DFlash run on gfx1100 and gfx1201 compiles a kernel that the 0.4.1 kernel pack does not yet include.

Known limits

  • Tool score. 23/30 median on the sampled TC70–84 panel; TC-75 and TC-76 score 0 on every sampled seed, also with the cache off.
  • Hermes pass 1. 95 with 18/20 cases passing, against 96 for pass 2. Runs are single and noisy; the pass-1 time is dominated by generation volume, not decode speed.
  • Greedy chains. A five-turn greedy chain can end turn 4 inside its reasoning. It is deterministic, identical with MTP on and off and with the cache off, so it is the model's trajectory and not cache state (release notes, known issues).

Next

0.4.2 is about more prefill and decode speed on Strix Halo. No numbers until they are measured.

Earlier snapshots

Dated, before 0.4.1. These fixtures predate the 0.4.1 release matrix and were not re-measured on 0.4.1. Each figure is scoped to its fixture and date; /docs/benchmarks is the build-time live ledger.

Three independent fixtures — different models, quants, and hardware setups. Rows are heterogeneous fixtures. Compare cells only when model, quant/mode, prompt, method, and backend match.

Qwen3.6 35B-A3B MQ4R · single-GPU AR

Ordinary autoregressive decode, single GPU; no MTP, DFlash, reduced-output bench, or manual clock pinning. TG128 = three-run medians. Multi-turn columns from clean eight-turn serving runs.

GPU arch TG128 AR 8-turn avg final turn
Radeon RX 7900 XTX gfx1100 253.3 tok/s 191.0 tok/s 160.3 @ 18.2K
Radeon 8060S / Strix Halo gfx1151 115.1 tok/s 92.2 tok/s 82.5 @ 21.3K
Radeon AI PRO R9700 gfx1201 203.9 tok/s 169.5 tok/s 146.7 @ 22.2K

Source: README · Qwen3.6 35B-A3B MQ4R · live ledger: /docs/benchmarks

Qwen3.8-27B MQ4V2 · 2026-08-20 · hiptrx · gfx1201 · Radeon AI PRO R9700

Product-ladder checkpoint on the hiptrx gfx1201 / Radeon AI PRO R9700 fixture. AR decode, prefill, DFlash decode, mean accepted draft length (τ), and bits-per-weight.

tier AR decode prefill DFlash τ bpw
XT 35.3 tok/s 490.7 251.6 tok/s 11.70 4.456
Base 33.2 tok/s 479.0 263.3 tok/s 13.11 4.659
Pro 31.8 tok/s 473.8 258.2 tok/s 13.11 4.897

Dated 2026-08-20. Source: 2026-08-20 Qwen3.8 MQ-V2 product ladder (docs/perf-checkpoints/2026-08-20-qwen38-mq-v2-product-ladder.md)

DeepSeek V4 Flash MQ2R · 4× Radeon AI PRO R9700 (gfx1201)

n=3 fresh-process medians; greedy; speculative decode off; KV f32; 2052-token prompt.

topology decode prefill
TP3/EP 53.1 tok/s 481 tok/s
TP4/EP 54.3 tok/s 389 tok/s

Source: README · RDNA4 gfx1201 R9700 · live ledger: /docs/benchmarks