LLM inference for AMD RDNA GPUs.
A Rust engine over hand-written HIP kernels — two
dlopen calls into ROCm at runtime, nothing else.
No PyTorch. No Python in the hot path. Single binary.
performance snapshots
Beta moves fast. Optimization on the hipfire beta branch is continuous. Every figure below is fixture-scoped, and a date is shown only when its source carries a measurement date. These source-published snapshots are not immutable or lasting claims. Numbers update frequently. /docs/benchmarks is the build-time live ledger.
Three independent fixtures — different models, quants, and hardware setups. Rows are heterogeneous fixtures. Compare cells only when model, quant/mode, prompt, method, and backend match.
Qwen3.6 35B-A3B MQ4R · single-GPU AR
Ordinary autoregressive decode, single GPU; no MTP, DFlash, reduced-output bench, or manual clock pinning. TG128 = three-run medians. Multi-turn columns from clean eight-turn serving runs.
| GPU | arch | TG128 AR | 8-turn avg | final turn |
|---|---|---|---|---|
| Radeon RX 7900 XTX | gfx1100 | 253.3 tok/s | 191.0 tok/s | 160.3 @ 18.2K |
| Radeon 8060S / Strix Halo | gfx1151 | 115.1 tok/s | 92.2 tok/s | 82.5 @ 21.3K |
| Radeon AI PRO R9700 | gfx1201 | 203.9 tok/s | 169.5 tok/s | 146.7 @ 22.2K |
Source: beta README · Qwen3.6 35B-A3B MQ4R · live ledger: /docs/benchmarks
Qwen3.8-27B MQ4V2 · 2026-08-20 · hiptrx · gfx1201 · Radeon AI PRO R9700
Product-ladder checkpoint on the hiptrx gfx1201 / Radeon AI PRO R9700 fixture. AR decode, prefill, DFlash decode, mean accepted draft length (τ), and bits-per-weight.
| tier | AR decode | prefill | DFlash | τ | bpw |
|---|---|---|---|---|---|
| XT | 35.3 tok/s | 490.7 | 251.6 tok/s | 11.70 | 4.456 |
| Base | 33.2 tok/s | 479.0 | 263.3 tok/s | 13.11 | 4.659 |
| Pro | 31.8 tok/s | 473.8 | 258.2 tok/s | 13.11 | 4.897 |
Dated 2026-08-20. Source:
beta · 2026-08-20 Qwen3.8 MQ-V2 product ladder
(docs/perf-checkpoints/2026-08-20-qwen38-mq-v2-product-ladder.md)
DeepSeek V4 Flash MQ2R · 4× Radeon AI PRO R9700 (gfx1201)
n=3 fresh-process medians; greedy; speculative decode off; KV f32; 2052-token prompt.
| topology | decode | prefill |
|---|---|---|
| TP3/EP | 53.1 tok/s | 481 tok/s |
| TP4/EP | 54.3 tok/s | 389 tok/s |
Source: beta README · RDNA4 gfx1201 R9700 · live ledger: /docs/benchmarks
quickstart
Linux with ROCm 6+ and an RDNA GPU.
$ curl -L https://raw.githubusercontent.com/warpfront/hipfire/beta/scripts/install.sh | bash $ hipfire pull qwen3.5:9b $ hipfire run qwen3.5:9b "What is the capital of France?" $ hipfire serve -d # background daemon, OpenAI-compatible API on :11435 $ hipfire pull flux.schnell:1 $ hipfire img flux.schnell:1 "a tiny lighthouse on a rock at sunset" --out lighthouse.png
Windows, source builds, and NixOS: see Getting started and NixOS.
hardware support
| arch | example | status |
|---|---|---|
| gfx1100 (RDNA 3) | Radeon 7900 XTX / XT | primary target — tuned |
| gfx1201 (RDNA 4) | Radeon 9070 XT, R9700 | supported |
| gfx1151 (RDNA 3.5) | Strix Halo APU | supported |
| gfx1030 (RDNA 2) | Radeon 6950 XT | supported |
| gfx1010 (RDNA 1) | Radeon 5700 XT | experimental |
| gfx906 (Vega 20) | MI50 / MI60 | community-driven |
| gfx942 (CDNA 3) | MI300X | rocBLAS path, partial |
Per-arch kernel variants live in kernels/src/*.gfxNNNN.hip; dispatch is
feature-gated (has_wmma_f16, has_dot2_f32_f16),
not chip-string matching.
HIP, not ROCm-the-stack
llama.cpp + ROCm works on RDNA, but it leans on a userspace
stack — rocBLAS, MIOpen, hipBLASLt, Tensile — that AMD
officially supports on a handful of datacenter cards. Consumer RDNA is
a second-class citizen there.
hipfire takes a different path: a Rust orchestrator that dlopens
libamdhip64.so directly, plus an in-tree set of HIP C++
kernels that we tune per chip. librocblas.so is loaded
lazily and only used on MI300X-class hardware; absence is recoverable.
We don't call MIOpen, RCCL, hipBLASLt, or Composable Kernel.
The pattern is borrowed from ncdrone/rustane's approach to Apple's Neural Engine: safe Rust over a thin FFI to whatever the driver runtime actually exposes.
what's in the box
Hand-written HIP kernels
Architecture-gated paths from RDNA 1 through RDNA 4, plus gfx906 and CDNA support, with ordinary HIP fallback when a validated retained route cannot run.
Magnum V2 quantization
MQ3–MQ6 V2 for dense Qwen 3.8, plus quality and speed tiers for MoE models.
DFlash speculative decode
Registry-paired draft sidecars, explicit opt-in controls, and prompt-cache-safe multi-turn serving. Gains remain prompt- and genre-dependent.
FlashAttention & asym KV
Asym2 / Asym3 / Asym4 KV-cache compression keeps long contexts cheap on consumer VRAM budgets.
Multi-GPU pipeline parallel
Including mixed-arch — e.g. gfx1010 + gfx1030 + gfx1151 + gfx1201 in one rig.
OpenAI-compatible API
hipfire serve -d exposes /v1/chat/completions on :11435.
MoE & hybrid models
Qwen 3.6 35B-A3B and Ornith 1.5, with DeltaNet recurrent-state caching across turns.
Image generation
FLUX.1 schnell + FLUX.2 Klein on RDNA3/3.5 — hipfire img and OpenAI-compatible /v1/images/generations + /edits.
Vision sidecars
Qwen 3.8 27B vision is an opt-in shared sidecar, so text-only serving pays no vision-tower VRAM cost.
BYO models
Quantize from HuggingFace safetensors or llama.cpp GGUF via hipfire quantize — CPU-side, no GPU required.
NixOS first-class
Flake, dev shell, and a services.hipfire NixOS module.
learn
Short, dense decks on the silicon underneath. The HIP/ROCm side of the
conversation that brrr-style CUDA explainers don't cover.
RDNA & WMMA
wave32, 16×16×16 WMMA, gfx11 half16 vs gfx12 half8 packing, occupancy and bandwidth models, plus beta fixture examples across gfx1010 → gfx1201.
CDNA & MFMA
The Instinct lineage. wave64, HBM3, v_mfma_* tiles,
and the rocprofv3 counters that diagnose where cycles go.
See the full index at /learn.
documentation
- Getting started — install, first run
- CLI reference — every subcommand
- Models — curated tags & BYO
- Quantize — HF / safetensors / GGUF → hipfire formats
- Serve — OpenAI-compatible HTTP API
- Architecture — load sources, carriers, graph, and dispatch
- Quantization design — MQ4 / HFQ4 / asym KV / FWHT math
- Multi-GPU — pipeline-parallel, memory budget
- Benchmarks — live rows and fixture-scoped snapshots
Looking for the “how GPUs actually work” pictures? See /learn — RDNA / WMMA and CDNA / MFMA decks.