hipfire
/docs · beta branch · view source · edit on GitHub

Multi-GPU

Pipeline and expert parallelism with memory budgeting.

Operator reference for hipfire multi-device modes. Mutable env inventory lives in env-vars.md. Validation route selection lives only in VALIDATION.md — this page does not invent minimum routes. Bring-up narrative is historical: multi-gpu-bringup-lessons.md.

FieldValue
Inventory date2026-07-19
Audited source ref692a726dde53508cb53de1a74c720e75a7c9f33e
Page stateshipped / ref-pinned for source-wired multi-device behavior (see INDEX.md); performance tables in the appendix are historical only
Orchestrationcrates/hipfire-runtime/src/multi_gpu.rs
PP load pathcrates/hipfire-loader/src/carriers.rs (load_qwen35_pp)
Daemon load / refusalscrates/hipfire-daemon/src/main.rs
Supporting multi-GPU scriptscripts/pp-gate.sh (not a VALIDATION selector minimum route)
Admissionsadmissions.yml — schema v2, exactly one single-GPU retained-PM4 record; no multi-GPU records (fail closed; none inferred)

Modes

Two multi-device modes exist. They are mutually exclusive at load (tp > 1 && pp > 1 → error).

ModeLoad knobWhat it doesSource-wired runtime surface (audited ref)
Pipeline parallel (PP)daemon load params.pp (N, default 1)Contiguous layer bands across N devices; residual stream crosses bands via boundary_copyQwen3.5 / 3.6 HFQ only (arch_id 5 dense, 6 MoE/A3B) via load_qwen35_pp
Expert parallel (EP / tp)params.tp, or CLI hipfire serve --tp N → HIPFIRE_TPWithin-layer expert sharding + all-reduce; every rank runs every layer (Gpus::init_tp)MiniMax-M2 (arch_id 10) and DeepSeek V4 Flash (arch_id 9) via load_model_ep

pp = 1 and tp = 1 are single-GPU. Behavior matches the pre-multi-GPU paths.

None of the listed PP/EP routes is an admission or product default. admissions.yml has no multi-GPU records at schema v2 (the sole earned row is single-GPU pp=tp=1). Source-wired means the load path exists in runtime at the audited ref — not that it is promoted.

PP is not tensor-parallel serving. It does not give multi-user throughput. It is a capacity tool: fit larger context / larger HFQ weights by splitting layers. No speedup is promised for models that already fit on one card; current historical measurements on one 2× gfx1100 station were slower under sequential PP=2 (see Appendix A).

CLI note: the shipped CLI forwards tp (HIPFIRE_TP / --tp) on serve load messages. params.pp is not a first-class CLI flag today — set it on a raw daemon JSONL load (or any client that builds that message). Examples and scripts/pp-gate.sh do this directly.

Topology (PP)

Source of truth: Gpus in multi_gpu.rs.

  1. Device pick — hardware.devices = "0,1,..." is the physical visibility list. Startup installs it as ROCR_VISIBLE_DEVICES and gives HIP the matching post-filter logical list 0..N-1; this avoids compounded nonzero filters while keeping both backends on the same physical GPUs.
  2. Layer map
    • Default: Gpus::init_uniform(pp, n_layers) — contiguous bands, base = n_layers / N, remainder distributed so max−min ≤ 1 layer.
    • Escape hatch: HIPFIRE_PP_LAYERS=a,b,… → Gpus::init_layers (length must equal pp, sum must equal n_layers). Skips the uniform free-VRAM delta check; still enforces arch match unless overridden.
  3. Placement convention (Variant 2) — output_device = last device holds output_norm + lm_head. Device 0 holds the embedding side of the split.
  4. Boundary traffic — at each band edge, boundary_copy moves the residual (hipMemcpyPeerAsync when peer access is up; otherwise HIP host-staging). Caller waits with wait_boundary.
  5. Peer access — enable_peer_all must run after weights/KV/scratch that need peer maps are allocated. Incomplete peer matrix → host-staging fallback (slower, still correct). Partial pair failure does not abort capable pairs.
  6. Preflight
    • Default: exact Gpu.arch string match across devices (d.arch != arch0 in preflight_vram / multi_gpu.rs). This is not a loose “family” compare.
    • Free-VRAM delta ≤ HIPFIRE_UNIFORM_VRAM_TOLERANCE_GB (default 2.0) for init_uniform / init_tp only.
    • HIPFIRE_ALLOW_MIXED_ARCH=1 opts into mixed-arch pairs (JIT per arch; peer may host-stage).
  7. Threading — HIP work is single-threaded for the daemon lifetime. bind_thread before peer enable and before device-bound work. No rayon/tokio HIP callers in v1.

EP topology is different: init_tp sets every device’s layer map as “all layers on rank 0” for PP helpers, while the EP forward ignores bands and shards experts. RCCL all-reduce is used unless HIPFIRE_TP_USE_RCCL=0 (host fallback not implemented — that opt-out errors). Non-standard ROCm layouts (e.g. nixpkgs splitting librccl out of the ROCm root) set HIPFIRE_RCCL_LIB to the full librccl.so path; the loader tries that before the ROCm root candidates.

Peer / fabric checks (host)

rocm-smi --showtopo
rocm-smi --showtoponuma
rocm-smi --showtopoaccess

A full True peer-access matrix is ideal. Missing peer access does not block load; copies fall back to host staging.

Launch and config

PP load (daemon JSONL)

{"type":"load","model":"/path/to/qwen3.5-9b.mq4","params":{"max_seq":16384,"pp":2}}

Example process:

# Persist one physical device list for both HIP and ROCr.
hipfire config set hardware.devices 0,1

# Optional: asymmetric bands (must sum to n_layers, length == pp)
# export HIPFIRE_PP_LAYERS=16,16

# Bit-stable k-split reduction for pp=1 vs pp=2 parity work
export HIPFIRE_DETERMINISTIC=1

cargo run --release --features deltanet -p hipfire-runtime --example daemon
# then send the load JSON above on stdin

EP load (CLI)

HIP_VISIBLE_DEVICES=0,1 hipfire serve <minimax-or-deepseek4-tag> --tp 2
# equivalent: HIPFIRE_TP=2 …

EP does not honor a non-default --kv-mode the same way single-GPU load does today — the CLI warns when both are set. Reload without a DFlash draft for EP.

Environment (operator-facing)

Canonical table: env-vars.md (MULTI-GPU group). Short map:

VariableRole
hardware.devicesPersistent physical device list; lowers to ROCr physical selectors and matching HIP logical selectors before initialization
HIP_VISIBLE_DEVICES / ROCR_VISIBLE_DEVICESLegacy one-shot filters; compatible pairs are normalized and ambiguous pairs fail closed
HIPFIRE_DEVICESLegacy compatibility alias for hardware.devices
HIPFIRE_PP_LAYERSExplicit per-device layer counts for PP
HIPFIRE_UNIFORM_VRAM_TOLERANCE_GBinit_uniform / init_tp free-VRAM delta
HIPFIRE_ALLOW_MIXED_ARCHOpt into mixed-arch device sets
HIPFIRE_DETERMINISTICDeterministic WMMA reduction path (parity / bisect)
HIPFIRE_TPEP degree (CLI --tp sets this)
HIPFIRE_TP_USE_RCCL0 opts out of RCCL (errors; no host AR yet)
HIPFIRE_RCCL_LIBExplicit librccl.so path tried before the ROCm root (nixpkgs / split RCCL installs)
HIPFIRE_PP_PFLASH=1Experimental — accept PFlash compose with pp>1 (not a product default; not route-certified)
HIPFIRE_PP_DFLASH=1Experimental — accept DFlash draft field with pp>1 (cross-card spec generate is not fully implemented; see daemon refusal text)

Test-only: HIPFIRE_HAVE_2_GPU, HIPFIRE_PP_PARITY_MODEL (pp_parity harness).

Refusals, silent paths, and experimental exceptions (pp > 1)

Many non-support cases are enforced at load (daemon and/or carrier) and fail closed — do not expect silent degrade to single-GPU for those. Exceptions and carrier-specific wording matter; do not treat the table as one universal string or a blanket “always refuse at load.”

ConditionBehavior
tp > 1 and pp > 1Error: mutually exclusive
DFlash draft set and HIPFIRE_PP_DFLASH unsetError: DFlash requires pp=1
DFlash draft set and HIPFIRE_PP_DFLASH=1Experimental exception: load may accept; cross-card speculative generate is not fully implemented (daemon message states PR2–4 of the hetero PFlash/DFlash plan are incomplete). Not an admission.
CASK / TriAttention sidecar setError: requires pp=1
PFlash drafter / mode on and HIPFIRE_PP_PFLASH unsetError: PFlash requires pp=1
PFlash on and HIPFIRE_PP_PFLASH=1Experimental exception: opt-in only; not a product default and not route-certified performance.
Non-PP carriers at pp>1Carrier-specific error strings (not one universal quote). Examples at the audited ref: qwen2: pipeline-parallel (pp>1) unsupported; llama: / dots_ocr: / deepseek4: / minimax: / lfm2moe: pipeline-parallel (pp>1) unsupported on HFQ (and distinct safetensors + pp>1 unsupported on dirs); cohere2moe: pp>1 unsupported via registry (different wording).
Qwen3.5 safetensors directory + pp>1Error: qwen35: safetensors + pp>1 unsupported
Qwen3.5 / 3.6 VL HFQ + pp>1Not a hard refuse. load_qwen35_pp is the text HFQ loader only — vision weights are not loaded on that path, so VL artifacts silently become text-only under PP. Serve real VL at pp=1.
bench_prefill / multi-GPU EPDaemon refuses bench_prefill when pp>1 or EP is active

EP-only: tp>1 with a DFlash draft → refused; non-EP arch → load_model_ep error.

Architectural limits (current)

  • Homogeneous exact arch string by default (ALLOW_MIXED_ARCH is opt-in).
  • No automatic VRAM-weighted split.
  • PP decode is sequential across bands (no async multi-band pipeline / per-band graph capture as a documented product path).
  • Experimental HIPFIRE_PP_* flags are not admissions and are not route-certified performance features.

Memory budget

Weights, KV, and scratch land on the devices that own each band. Last device also carries output_norm + lm_head. Per-card headroom depends on quant, KV mode, max_seq, and split.

Reproduce on your hardware (2+ visible GPUs, deltanet feature):

HIP_VISIBLE_DEVICES=0,1 cargo run --release --features deltanet \
  -p hipfire-runtime --example pp2_vram_probe -- \
  ~/.hipfire/models/qwen3.5-9b.mq4 4096

Historical VRAM and throughput tables from the original multi-GPU PP doc are preserved verbatim in Appendix A. They are historical only: they do not carry a complete measured-identity manifest (measurement date + binary md5 + model identity on the same report), so they must not be labeled measured, must not be treated as floors, and must not be edited in place. Re-run the probe / perf protocol before claiming fit or speed on other SKUs.

PP KV policy for Qwen3.5 multi-GPU defaults through QWEN35_PP_POLICY (see kv_mode.rs) — do not assume single-GPU auto KV defaults apply unchanged.

Validation routes

There is no universal GPU gate (VALIDATION.md). Multi-GPU claims are manual and path-specific. scripts/pp-gate.sh is not named in the VALIDATION claim→route selector and therefore is not a canonical minimum route (VALIDATION.md § retired / unnamed gate scripts). Treat it only as supporting manual evidence.

Map claims through the selector’s classes:

Claim class (selector)What applies for multi-GPUNotes
Forward / fusion / KV numerical or state parityPath-specific parity/state oracle when one exists for the surfaceSupporting tool today: pp_parity_chatml example (also invoked by pp-gate.sh). If no oracle exists for a surface → blocked. Not serve_harness.py.
Forward / serve user-facing semanticsscripts/serve_harness.py with the exact model (after parity if numbers/state can break)Semantics only
Perf improvement under PP or EPmethodology/perf-benchmarking.md + stationary matched runs; speed-gate.sh / gates.sh perf arm when applicableBench numbers without protocol/identity are not promotion evidence. Historical appendix rows are not floors.
Arch port (new device family behavior)methodology/arch-port-validation.mdChannel + speed; no retired coherence battery as acceptance
Model/route admissionRow in admissions.ymlSchema v2 exact-row only — multi-GPU PP/EP are not admitted
Docs-only editsNo-GPU CI / scripts/no-gpu-ci.shNever substitutes for GPU parity
Unknown multi-device surfaceBlocked until VALIDATION grows a rowFail closed

Supporting manual commands (not selector minimums):

# Supporting multi-GPU battery (parity + daemon e2e + refusals). Skips cleanly
# with <2 usable devices. Not automatic merge proof; not a VALIDATION minimum.
./scripts/pp-gate.sh

# Faster: parity example only
./scripts/pp-gate.sh --skip-end-to-end

# Topology filter dry-run (no GPU work)
./scripts/pp-gate.sh --dry-run

# Direct parity example
HIP_VISIBLE_DEVICES=0,1 cargo run --release --features deltanet \
  -p hipfire-runtime --example pp_parity_chatml -- \
  ~/.hipfire/models/qwen3.5-0.8b.mq4

pp-gate knobs (see script header): PP_GATE_DEVICES, HIPFIRE_PP_GATE_INCLUDE_IGPU, HIPFIRE_PP_GATE_HETEROGENEOUS, HIPFIRE_PP_GATE_MODEL, HIPFIRE_PP_GATE_REQUIRE_SYSFS. Filters drop known APU iGPUs and skip heterogeneous ISA families unless overridden.

Pre-commit (when hooks installed) runs pp-gate if staged paths match the multi-GPU hotspot regex in .githooks/pre-commit. That is a path-gated local guard, not full product admission.

Retired scripts/coherence-gate-*.sh batteries are not acceptance for PP (VALIDATION.md § retired).


Appendix A — Historical PP evidence (immutable)

Lifecycle — historical only. The block below is the prior multi-GPU PP document body retained for provenance. It is not current procedure, not a product floor, and not an admission. Truth state: historical (not measured — the retained tables lack a complete same-report measurement date + binary identity + model identity manifest required by INDEX.md). Do not edit the evidence body; amend only by adding new dated sections outside this appendix. Warnings in the active sections above supersede any stronger claim language inside the retained body (including “measured”, “Status: v1 feature-complete”, absolute speedup wording, and gate-as-acceptance framing).

Multi-GPU Pipeline-Parallel

Status: v1 feature-complete on feat/multi-gpu-pp branch — tracking issue #58. Stages 0–9 of the v2 plan are merged; refusal contracts (DFlash / VL / CASK + pp>1) are wired and validated. This doc is the source of truth for memory budget, deployment recipes, throughput, and known limitations.

Why PP

hipfire on a single 24 GB card hits VRAM walls on:

  • 27B at --max-ctx ≥ 16K with kv_mode=asym3 (AGENTS.md:356)
  • 35B-A3B at --max-ctx ≥ 4K with FP32 KV
  • hypothetical 80B-A3B at any context

Pipeline-Parallel (PP) shards layers across N devices. Each device owns a contiguous “band” of consecutive layers. The residual stream s.x flows through the bands sequentially: dev_0 runs layers 0..k1, copies s.x to dev_1, dev_1 runs layers k1..k2, and so on. Final output_norm + lm_head run on the last device (dev_last) — its s.logits is read by the sampler in place.

What PP gives you on 2× 24 GB:

  • Run 27B / 35B-A3B that don’t fit on one card with extended context
  • Unlock max_ctx on 27B beyond single-GPU OOM limits
  • ~50-70% of single-GPU throughput on already-fitting models (sequential PP=2 is slower per token)

What PP does NOT give you:

  • Faster multi-user serving — that’s TP (tensor parallel), separate roadmap
  • Speedup on models that already fit on one card

Memory budget (per-card, PP=2)

Numbers below are measured on 2× Radeon RX 7900 XTX (gfx1100, 25.8 GiB VRAM each) via crates/hipfire-runtime/examples/pp2_vram_probe.rs — hipMemGetInfo deltas captured at each allocation stage (load_weights_multi, Qwen35ScratchSet, KvCache::new_gpu_asym3_capped_multi, DeltaNetState). Per-card columns report the worst-of-two (the device that holds more — typically dev_last, which carries output_norm + lm_head). total is the sum across both cards.

Modelquantn_layersdimKV modectxweightsKV/cardscratch+DN/cardtotalper-card maxfits 24 GiB?
qwen3.5:0.8bmq4241024asym340961.3 GB50 MB8 MB1.3 GB0.7 GByes
qwen3.5:4bmq4322560asym340964.0 GB134 MB15 MB4.0 GB2.0 GByes
qwen3.5:9bmq4324096asym340965.6 GB134 MB19 MB5.6 GB2.8 GByes
qwen3.5:9bmq4324096asym316K6.2 GB436 MB46 MB6.2 GB3.1 GByes
qwen3.5:9bmq3324096asym340964.4 GB134 MB19 MB4.4 GB2.2 GByes
qwen3.5:27b (via 3.6 proxy)mq4645120asym3409615.5 GB268 MB42 MB15.5 GB7.8 GByes
qwen3.5:27b (via 3.6 proxy)mq4645120asym316K16.8 GB872 MB80 MB16.8 GB8.4 GByes
qwen3.5:27bmq3645120asym3409612.6 GB268 MB40 MB12.6 GB6.3 GByes
qwen3.5:27bmq3645120asym316K13.8 GB872 MB78 MB13.8 GB6.9 GByes
qwen3.6:35b-a3bmq4402048asym3409623.5 GB103 MB21 MB23.5 GB11.8 GByes
(hypothetical) 80B-a3bmq4808192asym34096~42 GB~250 MB~2 GB~44 GB~22 GBestimate, tight

Notes:

  • qwen3.5:27b mq4 rows use qwen3.6:27b mq4 as a measurement proxy (same n_layers=64, dim=5120, head_dim=256, n_kv_heads=4 — VRAM-equivalent). When qwen3.5:27b mq4 lands as a downloadable artifact the rows can be re-measured directly.
  • qwen3.6:35b-a3b mq4 is the current-generation A3B; qwen3.5:35b-a3b mq4 ships local-only with the same MoE shape — measurement carries over.
  • The 80B row stays an estimate — no public artifact exists.
  • “scratch+DN/card” combines Qwen35ScratchSet::per_device[i] (residual stream, attention/FFN scratch, flash partials, logits) and the DeltaNetState slice owned by that device’s LA-layer band.

Asymmetry under Variant 2 (lm_head on dev_last):

  • dev_0 carries token_embd
  • dev_last carries output_norm + lm_head
  • Layer count shifts by ±1 for n_layers % 2 != 0 (uniform split formula base + (i < rem ? 1 : 0))

Probed on this hardware: per-card max < 24 GiB at every measured shape. The A3B model at 11.8 GB/card has the largest headroom consumer; everything else stays under 8.5 GB/card with 4 K context, under 9 GB/card at 16 K.

To reproduce on your hardware:

HIP_VISIBLE_DEVICES=0,1 cargo run --release --features deltanet \
    -p hipfire-runtime --example pp2_vram_probe -- \
    ~/.hipfire/models/qwen3.5-9b.mq4 4096

Deployment recipes

The daemon takes a pp field in the load message (default 1 = single-GPU, identical behavior to pre-PP code paths):

# Filter to the two 7900 XTX (drop the iGPU)
HIP_VISIBLE_DEVICES=0,1 hipfire run qwen3.5:27b --max-ctx 16384 "Hi"

# Bypass the inherited `gemm_..._wmma_ksplit` non-determinism (k-split
# atomicAdd reduction varies by warp scheduling — see commit f54ca71
# and kernels/src/gemm_hfq4g256_residual_wmma_ksplit.hip:22). Required
# for byte-equivalent output across processes / pp configurations.
HIPFIRE_DETERMINISTIC=1 HIP_VISIBLE_DEVICES=0,1 hipfire run qwen3.5:9b "Hi"

# Override uniform-VRAM tolerance (default 2 GiB; arches must match)
HIPFIRE_UNIFORM_VRAM_TOLERANCE_GB=4 hipfire run qwen3.5:27b "Hi"

Direct daemon JSON (driving without the CLI):

{"type":"load","model":".../qwen3.5-27b.mq4","params":{"max_seq":16384,"pp":2}}

Environment variables

VariableEffect
hardware.devices = "3,1"ROCr physical filter 3,1; HIP and the engine receive matching logical devices 0,1
HIPFIRE_DETERMINISTIC=1Force k2 WMMA reduction (no atomicAdd) — bit-identical across processes/pp configs at ~33% perf cost on small-batch decode
HIPFIRE_UNIFORM_VRAM_TOLERANCE_GB=NPre-flight VRAM-asymmetry tolerance for Gpus::init_uniform (default 2.0)
HIPFIRE_PREFILL_BATCHED=0Disable batched WMMA prefill (per-token fallback). Diagnostic for ksplit non-det isolation
HIPFIRE_PREFILL_MAX_BATCH=NOverride per-chunk prefill batch. When unset/invalid: arch defaults are 512 on exact gfx1100, 384 on exact gfx1201, else 256 (PREFILL_MAX_BATCH). Under TP, prefill_max_batch_tp uses default×tp (cap 2048). An explicit HIPFIRE_PREFILL_MAX_BATCH wins over both the arch default and the TP scale. Chunks > N split with peer-copy at the boundary
HIPFIRE_WO_WMMA_VARIANT={k2,ksplit,k4,…}Manual override of the wo-residual GEMM variant — see dispatch.rs auto-dispatch

Refusal matrix at load (pp > 1)

FeatureBehaviorWhy
arch_id ∈ {5, 6} (Qwen3.5 dense + MoE/A3B)AcceptedValidated end-to-end
arch_id = others (LLaMA / Qwen3)RefusedSingle-GPU only in v1
VL models (vision_config + vision tensors)Refusedv1.1
DFlash draft (draft field set)Refusedv1.1 — see feedback_cask_mfold_dflash_broken.md for the v1 ship-blocker
CASK / TriAttention sidecarRefusedEviction context is single-device — v1.1

Throughput baseline (gfx1100 × 2)

Measured on 2× Radeon RX 7900 XTX, ROCm 6.4.3, with HIPFIRE_DETERMINISTIC=1 (bit-equivalent pp=1 ↔ pp=2 output).

ModelPromptpp=1 prefillpp=2 prefillpp=1 decodepp=2 decodepp=2/pp=1 decode
0.8B mq422 tok838 tok/s588 tok/s332 tok/s227 tok/s68%
0.8B mq4322 tok (chunked)6493 tok/s5490 tok/s315 tok/s212 tok/s67%
35B-A3B mq4 (MoE)15 tok331 tok/s258 tok/s142 tok/s97 tok/s68%

The pp=2 decode penalty is inherent to v1: per-token forward_scratch_multi pays one HIP launch per kernel per layer with no graph capture (vs pp=1 which captures + replays the AR-step graph after warmup). Pipelined decode + per-band graph capture lift this in v1.1.

Limitations (v1)

Refused at load time:

  • pp > 1 + DFlash speculative decode
  • pp > 1 + CASK/TriAttention sidecar (eviction is single-device)
  • pp > 1 + VL models (vision encoder is single-device)
  • pp > 1 + arch_id ∉ {5, 6} (LLaMA / Qwen3 dense are pp=1 only)

Architectural limits in v1:

  • Homogeneous arch only (init_uniform hard-fails on arch mismatch)
  • Uniform layer split — init_layers(per_device) is the manual escape hatch
  • Per-token decode (no async stream pipeline / per-band graph capture) — v1.1
  • Pipelined prefill (chunk N+1 on dev_0 while chunk N processes on dev_1) — v1.1

Validation (Stage 9)

# Multi-GPU gate. Skips silently when fewer than 2 GPU visible.
./scripts/pp-gate.sh

# Just the parity smoke (no daemon end-to-end), faster
./scripts/pp-gate.sh --skip-end-to-end

# Underlying byte-equivalence example
HIP_VISIBLE_DEVICES=0,1 cargo run --release --features deltanet \
    -p hipfire-runtime --example pp_parity_chatml -- \
    ~/.hipfire/models/qwen3.5-0.8b.mq4

The pp-gate.sh battery checks:

  1. Per-token forward_scratch_multi ≡ forward_scratch bit-exact (pp_parity_chatml)
  2. Daemon pp=1 ≡ pp=2 byte-identical with HIPFIRE_DETERMINISTIC=1 (greedy ChatML)
  3. DFlash + pp=2 refusal at load
  4. CASK + pp=2 refusal at load

A pre-commit hook calls pp-gate.sh automatically when staged files match the multi_gpu|pp_|peer_access|pipeline|stages hotspot regex.

Verifying peer access on your hardware

rocm-smi --showtopo            # weights, hops, link types
rocm-smi --showtoponuma        # NUMA placement
rocm-smi --showtopoaccess      # peer accessibility matrix

A True in every cell of --showtopoaccess for the cards you plan to use means peer-access should work. If not, hipfire falls back to host-staging via pinned buffers (slower but correct).

Open questions (will be filled in as Stages land)

  • DFlash + PP integration scope — pending maintainer guidance on issue #58
  • Whether mixed-arch should be soft-warn or hard-fail — currently hard-fail
  • hardware.devices is the physical visibility source of truth; startup lowers ROCr physical selectors to matching HIP logical 0..N-1