hipfire
/docs · beta branch · view source · edit on GitHub

HFP4

HFP4 floating-point format specification.

Status (member metadata — INDEX truth labels):

VariantWire qtState
HFP4G3221shipped / ref-pinned — encoder, GEMV, WMMA prefill on RDNA3/4
MFP4G3224shipped / ref-pinned — HFP4G32 + offline FWHT (MQ4-class drop-in)
MFP4G32Lloyd / MFP4G32P / MFP4G32E8 (+ SoA) / MFP3G32E8 / MFP2G32E832–37shipped / ref-pinned encoders + selective kernels; experimental opt-in qualifier — not CLI defaults
HFP4G16 / G64 / MX / NV aliases, online rotation (MFP4G32R), HFP8see ID tableplanned (unbuilt) or reassigned wire slots — do not claim shipped

Broader production index: docs/QUANTIZATION.md. CLI: docs/QUANTIZE.md · quantizer crates/hipfire-quantize · kernels kernels/src/gemv_hfp4g32*.hip, gemm_*_hfp4g32_wmma*.hip.


Mission

HFP4 is hipfire’s RDNA-oriented answer to OCP MXFP4 and NVIDIA NVFP4. Those reference formats target non-AMD silicon and do not exploit RDNA-specific ISA levers (v_ldexp_f32 for UE8M0 dequant, native FP8 WMMA on gfx12, V_PERMLANE16 scale broadcast, VOPD dual-issue on RDNA3+).

HFP4 keeps the OCP E2M1 nibble lattice (same sixteen codes). There is no implemented MXFP4 or NVFP4 checkpoint importer in the quantizer today — foreign checkpoints require full re-quantization into HFP4/MFP4. NVFP4’s G16 / E4M3 scaling cannot generally become HFP4 G32 / UE8M0 by transforming scales while preserving codes. Documented so other AMD-side projects can adopt the layout.

Format taxonomy

HFP4G32   — E2M1 + UE8M0 g32 + FP16 row scale       canonical (qt 21)
MFP4G32   — HFP4G32 + offline FWHT-256              drop-in MQ4-class (qt 24)
MFP4G32P  — MFP4 with E4M3 (non-PoT) block scale    experimental (qt 33)
MFP4G32E8 — E8 lattice codewords in MFP4+P frame    experimental (qt 34+)
MFP4G32Lloyd — per-tensor 16-entry Lloyd codebook   experimental (qt 32)

# Reserved / not product defaults
HFP4G16   — was qt 22 reservation; **ID 22 reassigned to TidI32** (DeepSeek tables)
HFP4G64   — **reserved** qt 23 (ablation; not built as product)
HFP4G32MX / HFP4G16NV / HFP8* — former 25–27 reservations; still **reserved**, unbuilt
MFP4G32R (online-R) — former qt 29; **reassigned to PARO4G128T** — online-R has no live reserved ID

Element format (OCP E2M1)

4-bit signed FP. Sixteen codes; eight magnitudes:

nibblevaluenibblevalue
0000+0.01000−0.0
0001+0.51001−0.5
0010+1.01010−1.0
0011+1.51011−1.5
0100+2.01100−2.0
0101+3.01101−3.0
0110+4.01110−4.0
0111+6.01111−6.0

Locked to spec: changing magnitudes breaks RDNA4 hardware paths (v_cvt_pk_fp8_e2m1 where used), LUT-decode strategies, and MX/NV interop.

Block scale — UE8M0

Every block of g elements (canonical g = 32) carries one UE8M0 byte: unsigned exponent-only. Encoded e ∈ [0, 254] means 2^(e − 127). 0xFF is block-NaN.

On RDNA, v_ldexp_f32(acc, e − 127) is one VALU op (no multiply). UE8M0 alone is coarse (OCP MX notes ~1–2% PPL loss vs FP16 in literature), so HFP4 adds a FP16 row scale.

Per-row second-level scale (FP16)

Each weight row has a 16-byte aligned header with row_scale_a, a reserved row_scale_b slot, block count, and format_flags. Effective dequant (v1):

value = row_scale_a * 2^(block_e − 127) * E2M1_LUT[nibble]

row_scale_a hoists outside the K loop (GEMV finalize or WMMA output stage). row_scale_b is reserved: the encoder always writes zero, and current kernels read only row_scale_a — it does not control or skip a second scale today. Format flag bit 1 remains defined for a future dual-output use.

Finer blocks (g=32 vs HFQ/MQ G256) catch within-row outliers a single row scale would clip; the FP16 row scale restores dynamic range without paying a multiply in the inner dequant loop beyond ldexp.

Byte layout — HFP4G32 (canonical)

For a row of K elements (format is K%32-aligned; v1 GEMV/WMMA path requires K%256==0 — quantizer and loaders enforce the kernel constraint):

Per row (16 B, aligned):
  +0  : f16  row_scale_a
  +2  : f16  row_scale_b      // reserved; encoder writes 0; kernels ignore
  +4  : u8   block_count_lo   // K/32 low
  +5  : u8   block_count_hi
  +6  : u8   format_flags
              bit 0: rotation present
              bit 1: row_scale_b used
              bits 2-3: rotation_kind
                00 = off
                01 = offline FWHT (MFP4)
                10 = online block-diag-128 (planned)
                11 = online HadaCore-16 (planned)
              bits 4-7: reserved
  +7  : u8   reserved
  +8  : u32  reserved
  +12 : u32  reserved

Per block × (K/32):
  +0  : u8   block_e          // UE8M0
  +1  : u8[16] nibbles        // 32 E2M1 codes; low nibble = even index

Total per row: 16 + 17 × (K / 32) bytes. Effective bpw: 4.25 from per-block payload + 128/K from the row header.

MFP4 stamps format_flags = 0x05 (rotation present + offline FWHT) after the same row pack; runtime uses the shared MQ FWHT signs (seeds 42 / 1042).

Nibble packing

Byte b encodes elements 2b (low nibble) and 2b+1 (high nibble). Matches HFQ4 extract patterns and MX/NVFP4 wire ordering.

Quantization recipe (HFP4G32)

For each row W[K]:

  1. row_scale_a = max_abs(W) / 6.0 (E2M1 max ±6).
  2. Normalize: W_n = W / row_scale_a.
  3. Per 32-element block: pick UE8M0 e so 2^e covers block_max/6; divide; round-to-nearest on the E2M1 lattice; pack nibbles.

Round-trip per-block max-abs error is bounded by the local half-gap of the E2M1 lattice under nearest rounding. Gaps are not uniform: adjacent magnitudes include steps of 0.5, 1.0, and 2.0 (between 4 and 6), so the largest half-gap is 1.0 scaled unit. A valid global bound is therefore row_scale_a · 2^(block_e − 127) · 1.0 (not · 0.5). Example: a normalized value of 5 can round to 4 or 6 with unit error 1.0 before scales.

MFP4 applies offline FWHT-256 on the row (same signs as MQ4) before the HFP pack.

Rotation modes

ModeStorageflags bits 2–3Status
offunrotated00HFP4 default
offline-fwhtpre-rotated weights; mq_rotate_x on x01MFP4 shipped
online-bd128unrotated + fused online Hadamard10planned
hadacore-16WMMA-fragment Hadamard11planned / research

Online modes block fused QKV/gate_up until fused-rotation siblings exist. Offline mode reuses MQ infrastructure (mq_rotate_x, fused rmsnorm/silu+rotate, signs1/signs2) unchanged.

Quant-type IDs

IDVariantStatus
21HFP4G32shipped / ref-pinned
22TidI32 (DeepSeek tid2eid)reassigned from former HFP4G16 reservation — do not emit HFP4G16 as 22
23HFP4G64reserved ablation (not product)
24MFP4G32shipped / ref-pinned
25–27MX/NV/HFP8 reservationsreserved / unbuilt (planned product intent only)
28PARO4G128 (unrelated family)shipped / ref-pinned load path — listed so IDs are not squatted
29PARO4G128T (unrelated)reassigned from former MFP4G32R online-R slot
30MQ4G256Lloyd (unrelated)shipped / ref-pinned research MQ opt-in — renumbered off 21
32MFP4G32Lloydshipped / ref-pinned impl; experimental opt-in
33MFP4G32Pshipped / ref-pinned impl; experimental opt-in
34MFP4G32E8shipped / ref-pinned impl; experimental opt-in
35MFP4G32E8SOAshipped / ref-pinned impl; experimental opt-in
36MFP3G32E8shipped / ref-pinned impl; experimental cold-tier opt-in
37MFP2G32E8shipped / ref-pinned impl; experimental cold-tier opt-in

Current reserved wire IDs: 23 and 25–27. IDs 22, 29, and 30 were reassigned. Planned online-R / HFP8 variants have no reserved ID beyond this table. Future PRs must not reuse reserved HFP slots for unrelated formats except the documented reassignments above.

Configurability

Runtime env (live)

VariableEffect
HIPFIRE_FP8_WMMA=1Live. Default off. On gfx12 with WMMA w32, enables the FP8-WMMA HFP path when batch_size >= FP8_WMMA_MIN_BATCH (1024); otherwise FP16-WMMA is used. Sources: feature_flags.rs, dispatch.rs, gemm.rs.

There are no current source definitions for HFP4_BLOCK_SIZE, HFP4_SCALE_FORMAT, HFP4_SECOND_LEVEL, HFP4_ROTATION_KIND, HFP4_R, HIPFIRE_HFP_USE_FP8_WMMA, HIPFIRE_HFP_ROTATION, or HIPFIRE_HFP_VARIANT. Treat any of those names as planned only — do not document them as live controls.

Quantize CLI / binary

FlagEffect
--format hfp4 (hfp4g32, hf4p, fp4)HFP4G32
--format mfp4 (mfp4g32, mf4p)MFP4G32 + offline FWHT
--format mfp4p / mfp4e8 / mfp4l / …experimental siblings (binary)

Thin hipfire quantize help emphasizes mq/hf defaults; HFP/MFP are available on the full hipfire-quantize format alias set.

Runtime support boundaries

ConcernBehavior
Batched prefillHFP4G32 / MFP4G32 are runtime batch-eligible under is_batchable_la on WMMA arches (gfx11/115x/12) only — runtime batch eligibility, not registry admission (admissions.yml)
Decode GEMVgemv_hfp4g32 family (scalar / arch variants); MFP uses prerotated path
Fused QKV / gate_up prefillWMMA-only HFP keys on RDNA3/4 — no scalar fused fallback
K alignmentLoaders refuse K%256≠0 for v1 kernels
CDNA / pre-WMMAFall back or reject batch path; do not claim parity

v1 correctness-anchor inner loop

Mirrors HFQ4G256 multi-accumulator structure with E2M1 LUT + ldexpf instead of INT4×scale+zp:

// preamble: shared E2M1 lut[16]
// per block: sc = row_scale_a * ldexpf(1.0f, block_e - 127);
// per nibble: value = sc * (float)lut[nibble];

No zero-point term (E2M1 is signed).

Validation (how to claim quality/perf)

Do not treat the historical checklist below as a universal gate (see VALIDATION.md). For HFP work, prefer:

  1. CPU round-trip max-abs ≤ local half-gap (global ≤ row_scale · 2^(e−127) · 1.0).
  2. CPU vs kernel element error on fixed (M,K) tensors.
  3. NRMSE / KLD vs FP16 or MQ4 baselines on named models — report as measured with fixture identity.
  4. Prefill path proof: rocprof / internal profile shows HFP WMMA symbols when claiming WMMA prefill.
  5. No coherence-gate battery as product admission.

Aspirational speed numbers (e.g. TFLOPS vs third-party MXFP4 demos) are not route-certified here.

Comparison to neighbors

PropertyHFQ4G256MQ4G256MXFP4 (strict)NVFP4HFP4G32
ElementINT4INT4 + FWHTE2M1E2M1E2M1
Block size256256321632
Block scalef32 affine scale + f32 minf32 affine scale + f32 minUE8M0E4M3UE8M0
Secondary scalenonenonenoneFP32 tensorFP16 row (row_scale_a; row_scale_b reserved)
~bpw4.254.254.25~4.5~4.25
Rotationnoneoffline FWHT (R=D2·H·D1/16)nonenoneoffline FWHT on MFP4; online planned
MX/NV importre-quantre-quantno importer (re-quant)no importer (re-quant; G16/E4M3 ≠ G32/UE8M0 scale-only)re-quant only

Provenance / references

HFP4 is an original hipfire layout specification; E2M1 magnitudes follow OCP MX. MQ/FWHT attribution: MagnumQuant rotation (seeds 42/1042) shared with the MQ weight family in QUANTIZATION.md.