hipfire
hipfire

LLM inference for AMD RDNA GPUs.

263.3
tok/s DFlash decode Qwen3.8-27B MQ4V2 Base · gfx1201 R9700 · 2026-08-20
54.3
tok/s TP4/EP decode DeepSeek V4 Flash MQ2R · 4× R9700 · TP4/EP
253.3
tok/s AR on 7900 XTX Qwen3.6 35B-A3B MQ4R · AR · Q8 KV · gfx1100

A Rust engine over hand-written HIP kernels — two dlopen calls into ROCm at runtime, nothing else. No PyTorch. No Python in the hot path. Single binary.

no PyTorch no Python in the hot path single binary Ollama-style UX

performance snapshots

Beta moves fast. Optimization on the hipfire beta branch is continuous. Every figure below is fixture-scoped, and a date is shown only when its source carries a measurement date. These source-published snapshots are not immutable or lasting claims. Numbers update frequently. /docs/benchmarks is the build-time live ledger.

Three independent fixtures — different models, quants, and hardware setups. Rows are heterogeneous fixtures. Compare cells only when model, quant/mode, prompt, method, and backend match.

Qwen3.6 35B-A3B MQ4R · single-GPU AR

Ordinary autoregressive decode, single GPU; no MTP, DFlash, reduced-output bench, or manual clock pinning. TG128 = three-run medians. Multi-turn columns from clean eight-turn serving runs.

GPU arch TG128 AR 8-turn avg final turn
Radeon RX 7900 XTX gfx1100 253.3 tok/s 191.0 tok/s 160.3 @ 18.2K
Radeon 8060S / Strix Halo gfx1151 115.1 tok/s 92.2 tok/s 82.5 @ 21.3K
Radeon AI PRO R9700 gfx1201 203.9 tok/s 169.5 tok/s 146.7 @ 22.2K

Source: beta README · Qwen3.6 35B-A3B MQ4R · live ledger: /docs/benchmarks

Qwen3.8-27B MQ4V2 · 2026-08-20 · hiptrx · gfx1201 · Radeon AI PRO R9700

Product-ladder checkpoint on the hiptrx gfx1201 / Radeon AI PRO R9700 fixture. AR decode, prefill, DFlash decode, mean accepted draft length (τ), and bits-per-weight.

tier AR decode prefill DFlash τ bpw
XT 35.3 tok/s 490.7 251.6 tok/s 11.70 4.456
Base 33.2 tok/s 479.0 263.3 tok/s 13.11 4.659
Pro 31.8 tok/s 473.8 258.2 tok/s 13.11 4.897

Dated 2026-08-20. Source: beta · 2026-08-20 Qwen3.8 MQ-V2 product ladder (docs/perf-checkpoints/2026-08-20-qwen38-mq-v2-product-ladder.md)

DeepSeek V4 Flash MQ2R · 4× Radeon AI PRO R9700 (gfx1201)

n=3 fresh-process medians; greedy; speculative decode off; KV f32; 2052-token prompt.

topology decode prefill
TP3/EP 53.1 tok/s 481 tok/s
TP4/EP 54.3 tok/s 389 tok/s

Source: beta README · RDNA4 gfx1201 R9700 · live ledger: /docs/benchmarks

quickstart

Linux with ROCm 6+ and an RDNA GPU.

$ curl -L https://raw.githubusercontent.com/warpfront/hipfire/beta/scripts/install.sh | bash

$ hipfire pull qwen3.5:9b
$ hipfire run  qwen3.5:9b "What is the capital of France?"
$ hipfire serve -d   # background daemon, OpenAI-compatible API on :11435

$ hipfire pull flux.schnell:1
$ hipfire img  flux.schnell:1 "a tiny lighthouse on a rock at sunset" --out lighthouse.png

Windows, source builds, and NixOS: see Getting started and NixOS.

hardware support

archexamplestatus
gfx1100 (RDNA 3)Radeon 7900 XTX / XTprimary target — tuned
gfx1201 (RDNA 4)Radeon 9070 XT, R9700supported
gfx1151 (RDNA 3.5)Strix Halo APUsupported
gfx1030 (RDNA 2)Radeon 6950 XTsupported
gfx1010 (RDNA 1)Radeon 5700 XTexperimental
gfx906 (Vega 20)MI50 / MI60community-driven
gfx942 (CDNA 3)MI300XrocBLAS path, partial

Per-arch kernel variants live in kernels/src/*.gfxNNNN.hip; dispatch is feature-gated (has_wmma_f16, has_dot2_f32_f16), not chip-string matching.

HIP, not ROCm-the-stack

llama.cpp + ROCm works on RDNA, but it leans on a userspace stack — rocBLAS, MIOpen, hipBLASLt, Tensile — that AMD officially supports on a handful of datacenter cards. Consumer RDNA is a second-class citizen there.

hipfire takes a different path: a Rust orchestrator that dlopens libamdhip64.so directly, plus an in-tree set of HIP C++ kernels that we tune per chip. librocblas.so is loaded lazily and only used on MI300X-class hardware; absence is recoverable. We don't call MIOpen, RCCL, hipBLASLt, or Composable Kernel.

The pattern is borrowed from ncdrone/rustane's approach to Apple's Neural Engine: safe Rust over a thin FFI to whatever the driver runtime actually exposes.

what's in the box

Hand-written HIP kernels

Architecture-gated paths from RDNA 1 through RDNA 4, plus gfx906 and CDNA support, with ordinary HIP fallback when a validated retained route cannot run.

Magnum V2 quantization

MQ3–MQ6 V2 for dense Qwen 3.8, plus quality and speed tiers for MoE models.

DFlash speculative decode

Registry-paired draft sidecars, explicit opt-in controls, and prompt-cache-safe multi-turn serving. Gains remain prompt- and genre-dependent.

FlashAttention & asym KV

Asym2 / Asym3 / Asym4 KV-cache compression keeps long contexts cheap on consumer VRAM budgets.

Multi-GPU pipeline parallel

Including mixed-arch — e.g. gfx1010 + gfx1030 + gfx1151 + gfx1201 in one rig.

OpenAI-compatible API

hipfire serve -d exposes /v1/chat/completions on :11435.

MoE & hybrid models

Qwen 3.6 35B-A3B and Ornith 1.5, with DeltaNet recurrent-state caching across turns.

Image generation

FLUX.1 schnell + FLUX.2 Klein on RDNA3/3.5 — hipfire img and OpenAI-compatible /v1/images/generations + /edits.

Vision sidecars

Qwen 3.8 27B vision is an opt-in shared sidecar, so text-only serving pays no vision-tower VRAM cost.

BYO models

Quantize from HuggingFace safetensors or llama.cpp GGUF via hipfire quantize — CPU-side, no GPU required.

NixOS first-class

Flake, dev shell, and a services.hipfire NixOS module.

learn

Short, dense decks on the silicon underneath. The HIP/ROCm side of the conversation that brrr-style CUDA explainers don't cover.

See the full index at /learn.

documentation

Looking for the “how GPUs actually work” pictures? See /learn — RDNA / WMMA and CDNA / MFMA decks.