hipfire
/docs/benchmarks

benchmarks

Build-time live ledger of cross-arch tok/s on hipfire, plus curated fixture-scoped snapshots from warpfront/hipfire beta. 146 approved runs across 6 architectures, fetched from localmaxxing.com. Every live cell links to its /runs/<id> for reproducibility.

Freshness / lifecycle. Optimization on the hipfire beta branch is continuous. Every figure below is fixture-scoped, and a date is shown only when its source carries a measurement date. These source-published snapshots are not immutable or lasting claims. Numbers update frequently. /docs/benchmarks is the build-time live ledger.

Rows are heterogeneous fixtures. Compare cells only when model, quant/mode, prompt, method, and backend match. Engine sources and methodology docs point at warpfront/hipfire@beta.

Last build: 2026-09-09T04:53:12.638Z · live source: GET /api/leaderboard?engineName=hipfire · engine: github.com/warpfront/hipfire branch beta

live localmaxxing ledger

Architecture tables below are approved community runs pulled at this site build. They are a separate surface from the featured beta snapshots above — different harness, different fixtures. Do not treat a live cell and a dated snapshot as the same measurement.

gfx1100 — RDNA 3

hardware model q mode decode tok/s prefill AR base × AR τ run
RX 7900 XTX LFM2.5-350M MQ4 AR 1470.8 — — — — cmsnp1w730…
RX 7900 XTX LFM2.5-350M MQ4 AR 885.3 — — — — cmsnp2nkg0…
RX 7900 XTX LFM2.5-350M MQ4 AR 858.4 — — — — cmsnp2opb0…
RX 7900 XTX Qwen3.5-9B MQ4 DFlash 576.9 755 122.9 4.70× 13.18 cmpdzs3fq0…
RX 7900 XTX Qwen3.5-9B MQ4 DFlash 575.3 — — — — cmofzk74w0…
RX 7900 XTX Qwen3.5-9B MQ4 DFlash 575.2 — 123.1 4.67× 13.18 cmp8fqdz20…
RX 7900 XTX Qwen3.5-0.8B MQ4 AR 550.2 — — — — cmsnp33j90…
RX 7900 XTX Qwen3.6-35B-A3B MQ4R MTP 494.6 1468 — — — cmsnp3cmd0…
RX 7900 XTX LFM2.5-350M MQ4 AR 458.7 — — — — cmsnp3h5g0…
RX 7900 XTX LFM2.5-350M MQ4 AR 398.4 — — — — cmsnp3mx40…
RX 7900 XTX Qwen3.5-9B MQ4 DFlash 337.0 — — — — cmofdp1ve0…
RX 7900 XTX Qwen3.5-9B MQ4 DFlash 322.6 — — — — cmofkl1gm0…
RX 7900 XTX Qwen3.6-35B-A3B MQ4R AR 272.6 — — — — cmsnp41sa0…
RX 7900 XTX Qwen3.6-35B-A3B MQ4R AR 270.8 — — — — cmsnp42xp0…
RX 7900 XTX Qwen3.5-27B MQ4 DFlash 254.4 — 44.6 5.70× 13.18 cmp8fozao0…
RX 7900 XTX Qwen3.6-27B MQ4 DFlash 252.1 367 — — — cms8csy5d0…
RX 7900 XTX Qwen3.5-27B MQ4 DFlash 250.3 — — — — cmofyatrj0…
RX 7900 XTX Qwen3.6-27B MQ4-AWQ DFlash 232.1 335 — — — cmppe779e0…
RX 7900 XTX Qwen3.6-27B MQ4 DFlash 216.7 — 44.8 4.84× 10.93 cmp8fw36n0…
RX 7900 XTX Qwen3.6-27B MQ4 DFlash 216.6 — 44.6 4.85× 10.93 cmp8fnkw00…
RX 7900 XTX Qwen3.6-35B-A3B MQ4R AR 214.7 913 — — — cmsnp49si0…
RX 7900 XTX Qwen3.5-27B MQ4 DFlash 201.1 — — — — cmofkrpp00…
RX 7900 XTX Qwen3.5-27B MQ4 DFlash 182.0 — — — — cmofe2efg0…
RX 7900 XTX Qwen3.5-35B-A3B MQ4 DFlash 154.6 — — — — cmofkye0g0…
RX 7900 XTX Qwen3.8-27B MQ4 DFlash 154.2 51 — — — cmt1ff27w0…
RX 7900 XTX Qwen3.6-35B-A3B MQ4R AR 150.0 — — — — cmsnp4ebk0…
RX 7900 XTX Qwen3.5-35B-A3B MQ4 DFlash 140.6 — — — — cmofefr0j0…
RX 7900 XTX Qwen3.5-35B-A3B MQ4 AR 135.9 — — — — cmofe92tt0…
RX 7900 XTX Qwen3.6-35B-A3B MQ4 AR 135.3 — — — — cmofezrun0…
RX 7900 XTX Ornstein3.6-35B-A3B MQ4 AR 134.0 — — — — cmofgd9iz0…
RX 7900 XTX Qwen3.5-9B MQ4 AR 123.1 — — — — cmp8ft6rt0…
RX 7900 XTX Qwen3.5-9B MQ4 AR 122.3 — — — — cmofdidib0…
RX 7900 XTX Qwen3.6-27B MQ4 DFlash 118.2 — — — — cmofet3j80…
RX 7900 XTX Qwen3.6-27B MQ4 DFlash 118.1 — — — — cmofl528t0…
RX 7900 XTX Qwen3.6-35B-A3B MQ4 DFlash 68.6 — — — — cmoff6g560…
RX 7900 XTX Qwen3.8-27B-GGUF MQ4 AR 46.2 1012 — — — cmsv00obw0…
RX 7900 XTX Qwen3.6-27B MQ4 AR 44.8 — — — — cmp8frsd70…
RX 7900 XTX Qwen3.6-27B MQ4 AR 44.6 — — — — cmp8ful710…
RX 7900 XTX Qwen3.5-27B MQ4 AR 44.6 — — — — cmp8fl9rq0…
RX 7900 XTX Qwen3.5-27B MQ4 AR 43.7 — — — — cmofdvq7c0…
RX 7900 XTX Qwen3.6-27B MQ4 AR 43.6 — — — — cmofemf840…

Engine: hipfire@0.3.0+b0bcc3f91506 · Submitter: @schuttdev

gfx1201 — RDNA 4

hardware model q mode decode tok/s prefill AR base × AR τ run
Radeon AI Pro R9700 Qwen3.6-35B-A3B MQ4R AR 6001.4 — — — — cmsnp1rko0…
Radeon AI Pro R9700 Qwen3.6-35B-A3B MQ4R AR 5993.3 — — — — cmsnp232a0…
RX 9070 XT LFM2.5-230M MQ4 AR 4816.1 — — — — cmsnp1spu0…
Radeon AI Pro R9700 LFM2.5-230M MQ4 AR 4347.3 — — — — cmsnp1tv30…
Radeon AI Pro R9700 Qwen3.6-35B-A3B MQ4R AR 4207.7 — — — — cmsnp24750…
RX 9070 XT LFM2.5-350M MQ4 AR 3669.8 — — — — cmsnp1v0f0…
Radeon AI Pro R9700 LFM2.5-350M MQ4 AR 3283.3 — — — — cmsnp25c40…
RX 9070 XT LFM2.5-350M MQ4 AR 1876.9 — — — — cmsnp26gr0…
RX 9070 XT LFM2.5-230M MQ4 AR 1704.6 — — — — cmsnp27le0…
Radeon AI Pro R9700 LFM2.5-230M MQ4 AR 1597.7 — — — — cmsnp28q50…
Radeon AI Pro R9700 LFM2.5-230M MQ4 AR 1585.6 — — — — cmsnp29vb0…
RX 9070 XT LFM2.5-230M MQ4 AR 1523.1 — — — — cmsnp2b0g0…
RX 9070 XT LFM2.5-230M MQ4 AR 1457.0 — — — — cmsnp2c5d0…
RX 9070 XT LFM2.5-230M MQ4 AR 1376.7 — — — — cmsnp2d9x0…
Radeon AI Pro R9700 LFM2.5-350M MQ4 AR 1291.9 — — — — cmsnp2eef0…
Radeon AI Pro R9700 LFM2.5-350M MQ4 AR 1278.9 — — — — cmsnp2fk00…
RX 9070 XT LFM2.5-350M MQ4 AR 1258.5 — — — — cmsnp2gp30…
RX 9070 XT LFM2.5-230M MQ4 AR 1246.8 — — — — cmsnp2hup0…
RX 9070 XT LFM2.5-350M MQ4 AR 1237.7 — — — — cmsnp2j030…
Radeon AI Pro R9700 Qwen3.6-35B-A3B MQ4R AR 1115.7 — — — — cmsnsxbrr0…
RX 9070 XT LFM2.5-230M MQ4 AR 1044.8 — — — — cmsnp2k4r0…
RX 9070 XT LFM2.5-230M MQ4 AR 1017.0 — — — — cmsnp2l9o0…
Radeon AI Pro R9700 Qwen3.6-35B-A3B MQ4R AR 969.4 — — — — cmsnsxd8w0…
Radeon AI Pro R9700 LFM2.5-230M MQ4 AR 597.8 — — — — cmsnp2s4l0…
Radeon AI Pro R9700 LFM2.5-230M MQ4 AR 597.2 — — — — cmsnp2t9n0…
Radeon AI Pro R9700 LFM2.5-230M MQ4 AR 596.6 — — — — cmsnp2ue80…
Radeon AI Pro R9700 LFM2.5-230M MQ4 AR 594.1 — — — — cmsnp2viy0…
RX 9070 XT LFM2.5-230M MQ4 AR 574.3 — — — — cmsnp2wnr0…
RX 9070 XT LFM2.5-230M MQ4 AR 567.5 — — — — cmsnp2xsx0…
Radeon AI Pro R9700 Qwen3.5-0.8B MQ4 AR 563.5 — — — — cmsnp20sc0…
Radeon AI Pro R9700 Qwen3.5-0.8B MQ4 AR 562.4 — — — — cmsnp2yyv0…
Radeon AI Pro R9700 Qwen3.5-0.8B MQ4 AR 562.0 — — — — cmsnp30410…
Radeon AI Pro R9700 Qwen3.5-0.8B MQ4 AR 554.5 — — — — cmsnp32eb0…
RX 9070 XT Qwen3.5-0.8B MQ4 AR 545.3 — — — — cmsnp34or0…
Radeon AI Pro R9700 LFM2.5-350M MQ4 AR 542.5 — — — — cmsnp35tm0…
Radeon AI Pro R9700 LFM2.5-350M MQ4 AR 542.1 — — — — cmsnp36y20…
Radeon AI Pro R9700 LFM2.5-350M MQ4 AR 539.6 — — — — cmsnp382m0…
Radeon AI Pro R9700 LFM2.5-350M MQ4 AR 538.3 — — — — cmsnp397g0…
RX 9070 XT LFM2.5-350M MQ4 AR 520.8 — — — — cmsnp3ac90…
Radeon AI Pro R9700 LFM2.5-230M MQ4 AR 520.0 — — — — cmsnp3bh60…
RX 9070 XT LFM2.5-230M MQ4 AR 481.8 — — — — cmsnp3dqp0…
Radeon AI Pro R9700 LFM2.5-350M MQ4 AR 466.0 — — — — cmsnp3g0m0…
RX 9070 XT LFM2.5-350M MQ4 AR 437.7 — — — — cmsnp3jic0…
Radeon AI Pro R9700 Qwen3.5-9B MQ4 DFlash 371.8 1160 99.4 3.74× 13.18 cmpdzu8ig0…
Radeon AI Pro R9700 Qwen3.5-27B MQ4-AWQ DFlash 286.6 436 — — — cmppftgt80…
Radeon AI Pro R9700 Qwen3.8-27B INT4 DFlash 286.3 435 — — — cmt62jgz20…
Radeon AI Pro R9700 Qwen3.8-27B INT4 DFlash 286.2 434 — — — cmt61ypmz0…
Radeon AI Pro R9700 Qwen3.8-27B INT4 DFlash 285.5 432 — — — cmt62jh500…
Radeon AI Pro R9700 Qwen3.8-27B INT4 DFlash 284.5 431 — — — cmt62jgsc0…
Radeon AI Pro R9700 Qwen3.8-27B MQ4V2 DFlash 282.4 451 — — — cmt0pjfxq0…
Radeon AI Pro R9700 Qwen3.6-35B-A3B MQ4R MTP 280.4 1469 — — — cmsnp3tqo0…
Radeon AI Pro R9700 Qwen3.6-35B-A3B MQ4R MTP 279.4 1469 — — — cmsnp3uvq0…
Radeon AI Pro R9700 Qwen3.6-27B MQ4 DFlash 278.0 467 — — — cms8cr7d70…
Radeon AI Pro R9700 Qwen3.6-35B-A3B MQ4R MTP 277.4 1469 — — — cmsnp3x5g0…
Radeon AI Pro R9700 Qwen3.6-35B-A3B MQ4R MTP 274.9 1469 — — — cmsnp3zfg0…
Radeon AI Pro R9700 Qwen3.6-35B-A3B MQ4-AWQ MTP 259.5 1094 259.5 — — cmppmb47l0…
Radeon AI Pro R9700 Qwen3.6-27B MQ4-AWQ DFlash 254.8 435 — — — cmppfddi30…
Radeon AI Pro R9700 LFM2.5-8B-A1B MQ4 AR 241.1 242 — — — cmpruarq60…
Radeon AI Pro R9700 Qwen3.6-35B-A3B MQ4R AR 233.6 — — — — cmsnp46b50…
Radeon AI Pro R9700 Qwen3.6-35B-A3B MQ4R AR 230.9 — — — — cmsnp47g50…
Radeon AI Pro R9700 Qwen3.6-27B MQ4 SPEC 213.7 750 — — — cmrgp8xgt0…
Radeon AI Pro R9700 Qwen3.5-27B MQ4 DFlash 196.2 438 35.4 5.54× 13.18 cmpdzvqpx0…
Radeon AI Pro R9700 Qwen3.6-35B-A3B MQ4R AR 127.3 — — — — cmsnp4fgp0…
Radeon AI Pro R9700 DeepSeek-V4-Flash MQ2R AR 54.3 389 — — — cmsnp21x40…
Radeon AI Pro R9700 DeepSeek-V4-Flash MQ2R AR 53.1 481 — — — cmsnp4iw40…

Engine: hipfire@0.3.0+b0bcc3f91506 · Submitter: @schuttdev

gfx1151 — RDNA 3.5

hardware model q mode decode tok/s prefill AR base × AR τ run
Strix Halo (Radeon 8060S Graphics) Qwen3.5-9B MQ4 DFlash 255.8 406 45.6 5.60× 13.18 cmpdzx93g0…
Strix Halo (Radeon 8060S Graphics) Qwen3.5-27B MQ4 DFlash 104.5 135 — — 13.18 cmpe035k30…
Strix Halo (Radeon 8060S Graphics) Qwen3.6-27B MQ4 DFlash 88.3 134 14.8 5.96× 10.93 cmpe04nu50…
Radeon 8060S DeepSeek-V4-Flash MQ2-Lloyd + Q8 KV + MTP sidecar SPEC 19.0 146 — — — cmpqeql2m0…
Radeon 8060S Qwen3.5-27B MQ4 AR 14.8 — — — — cmoknenfb0…

Engine: hipfire@0.2.0+1a378379 · Submitter: @schuttdev

gfx1030 — RDNA 2

hardware model q mode decode tok/s prefill AR base × AR τ run
RX 6900 XT LFM2.5-350M MQ4 AR 584.2 — — — — cmsnp1znf0…
RX 6900 XT LFM2.5-350M MQ4 AR 557.2 — — — — cmsnp319s0…
RX 6900 XT Qwen3.5-0.8B MQ4 AR 346.1 — — — — cmsnp3qbh0…
RX 6900 XT LFM2.5-350M MQ4 AR 275.8 — — — — cmsnp3yab0…
RX 6950 XT Qwen3.5-9B MQ4 DFlash 222.0 479 75.1 2.96× 13.18 cmpe01n920…

Engine: hipfire@0.3.0+b0bcc3f91506 · Submitter: @schuttdev

gfx1010 — RDNA 1

hardware model q mode decode tok/s prefill AR base × AR τ run
RX 5700 XT LFM2.5-230M MQ4 AR 648.3 — — — — cmsnp1yij0…
RX 5700 XT LFM2.5-230M MQ4 AR 432.3 — — — — cmsnp3kn50…
RX 5700 XT LFM2.5-230M MQ4 AR 428.0 — — — — cmsnp3ls00…
RX 5700 XT LFM2.5-350M MQ4 AR 394.0 — — — — cmsnp3o210…
RX 5700 XT LFM2.5-350M MQ4 AR 390.8 — — — — cmsnp3p6o0…
RX 5700 XT LFM2.5-230M MQ4 AR 279.3 — — — — cmsnp3w0e0…
RX 5700 XT Qwen3.5-0.8B MQ4 AR 273.6 — — — — cmsnp40li0…
RX 5700 XT LFM2.5-350M MQ4 AR 257.3 — — — — cmsnp44230…
RX 5700 XT LFM2.5-230M MQ4 AR 226.3 — — — — cmsnp48nj0…
RX 5700 XT LFM2.5-350M MQ4 AR 207.7 — — — — cmsnp4ax10…
RX 5700 XT Qwen3.5-9B MQ4 AR 61.6 210 — — — cmpdzyrb70…

Engine: hipfire@0.3.0+b0bcc3f91506 · Submitter: @schuttdev

? — ?

hardware model q mode decode tok/s prefill AR base × AR τ run
Ryzen AI Max (Ryzen AI Max 395+) LFM2.5-350M MQ4 AR 1181.6 — — — — cmsnp1xcx0…
Ryzen AI Max (Ryzen AI Max 395+) LFM2.5-230M MQ4 AR 958.0 — — — — cmsnp2mfn0…
Ryzen AI Max (Ryzen AI Max 395+) LFM2.5-350M MQ4 AR 777.0 — — — — cmsnp2pua0…
Ryzen AI Max (Ryzen AI Max 395+) LFM2.5-350M MQ4 AR 766.3 — — — — cmsnp2qzh0…
Ryzen AI Max (Ryzen AI Max 395+) LFM2.5-230M MQ4 AR 470.3 — — — — cmsnp3ew10…
Ryzen AI Max (Ryzen AI Max 395+) LFM2.5-350M MQ4 AR 439.1 — — — — cmsnp3ian0…
Ryzen AI Max (Ryzen AI Max 395+) LFM2.5-350M MQ4 AR 343.1 — — — — cmsnp3rgh0…
Ryzen AI Max (Ryzen AI Max 395+) Qwen3.5-0.8B MQ4 AR 291.4 — — — — cmsnp3slf0…
Ryzen AI Max (Ryzen AI Max 395+) Qwen3.6-35B-A3B MQ4R MTP 248.2 1468 — — — cmsnp456p0…
Ryzen AI Max (Ryzen AI Max 395+) Qwen3.6-35B-A3B MQ4R AR 161.8 — — — — cmsnp4c290…
Ryzen AI Max (Ryzen AI Max 395+) Qwen3.6-35B-A3B MQ4R AR 161.4 — — — — cmsnp4d790…
Ryzen AI Max (Ryzen AI Max 395+) Qwen3.6-27B MQ4 DFlash 112.0 159 — — — cms8cugd00…
Ryzen AI Max (Ryzen AI Max 395+) Qwen3.8-27B MQ4V2 DFlash 103.9 144 — — — cmt6wvmy40…
Ryzen AI Max (Ryzen AI Max 395+) Qwen3.6-35B-A3B MQ4R AR 101.7 536 — — — cmsnp4glp0…
Ryzen AI Max (Ryzen AI Max 395+) Qwen3.6-35B-A3B MQ4R AR 91.4 — — — — cmsnp4hqe0…
Ryzen AI Max (Ryzen AI Max 395+) DeepSeek-V4-Flash MQ2R SPEC 39.0 — — — — cmsnu5ajq0…
Ryzen AI Max (Ryzen AI Max 395+) DeepSeek-V4-Flash MQ2R AR 30.8 202 — — — cmsnp4k0u0…
Strix Halo (Ryzen AI Max+ 395) Qwen3.6-27B MQ4 AR 17.0 — — — — cmokadzkk0…
Ryzen AI Max (Ryzen AI Max 395+) DeepSeek-V4-Flash MQ2-Lloyd MTP 15.2 145 — — — cmrh7bdda0…

Engine: hipfire@0.3.0+b0bcc3f91506 · Submitter: @schuttdev

methodology

Live localmaxxing cells follow a shared protocol so cross-arch comparisons inside that harness are meaningful. Featured beta snapshots use the method notes printed beside each table — those fixtures are not interchangeable with the live ledger.

  • Prompt: benchmarks/prompts/merge_sort_thinking_off.txt, md5 253c7ac50857fe6d0e10fb0d2c5e35c0, 27 input tokens. One newline-shape variant locks in a stable τ across runs. Source tree: warpfront/hipfire@beta.
  • Config: --max 256 --temp 0.0 --no-chatml --kv-mode q8 --ctx 4096, prompt_normalize=on (default since 2026-04-26).
  • Runs: one --max 16 warmup per cell (discarded), then 3 timed runs at --max 256, median reported. AR baseline is a paired --ar-baseline run from the same binary on the same hardware.
  • Coherence: every submitted run produced readable merge_sort code — no attractors, token loops, or special-token leaks.
  • GPU isolation: HIP_VISIBLE_DEVICES per cell on hosts with multiple GPUs; no two cells share a GPU concurrently.

Why this matters: a single newline change in the prompt can swing τ by 17% on 27B DFlash — same model, same flags, different token sequence. The prompt md5 is part of the claim.

Engine docs and performance checkpoints live on warpfront/hipfire branch beta — see docs/ and docs/perf-checkpoints/.

reproducibility

Every live row carries the exact launch command in its localmaxxing notes. To repro any cell:

$ hipfire pull <model>

$ hipfire bench <model> --spec dflash --runs 5 --warmups 3 \
    --max-tokens 256 --backend noslots --workload stateless --json

Prefer hipfire bench over examples/dflash_spec_demo: the demo runs raw open-ended continuation and under-reports by ~40% because it never hits the chatml + structured-thinking serve path (see AGENTS.md on beta).

For DFlash, the draft model is hosted at z-lab/Qwen3.5-9B-DFlash and z-lab/Qwen3.5-27B-DFlash. The native-head MQ4 conversion is shipped with hipfire's draft pulls (hipfire pull qwen3.5:9b-dflash).

submit your own

Anyone with a hipfire build can submit. Install lmx-bench for a curated harness, or POST directly to localmaxxing.com/api/benchmarks with the schema from the API docs. See an existing row's engineFlags.commandSnippet for the exact invocation to mirror. Build from warpfront/hipfire beta.