hipfire
/docs · beta branch · view source · edit on GitHub

Models

Curated model tags and bring-your-own weights.

Owner: registry-backed model surface (docs/INDEX.md). Machine sources: curated registry/models.json; generated and bundled registry/v1.json (loaded by hipfire-registry). Last checked: 2026-08-24 against the published Qwen3.8 MQ V2 ladder.

This page projects registry availability: tags, default artifact filenames, declared download size, and declared VRAM floor. It is not a product admission table and not a guarantee that every GPU/route runs every tag.

ConceptMeaning
Registry tagPull/list name resolved through the bundled v1 registry (+ aliases).
Default artifactfile field — what hipfire pull <tag> fetches into ~/.hipfire/models/.
Runtime supportWhether the daemon/loader/arch crate can load and run the artifact shape (arch_id, kernels, Cargo features). Source-of-truth: runtime crates + architecture-ids.md.
AdmissionExplicit product decision in admissions.yml. Schema v2 holds exactly one evidence-bound record; no inferred admissions beyond that row.

hipfire list -r prints the live registry plus local availability. Prefer that command when sizes change; this page is a checked narrative, not a second registry.


Pull and run

hipfire pull qwen3.5:9b
hipfire run qwen3.5:9b "hello"
hipfire list -r

Default serve pre-warm tag is qwen3.5:9b (CONFIG.md → default_model). Per-tag sampling defaults come from registry recommended_settings only (applied by the native CLI request resolver). Cards may set intentional reasoning_effort (semantic prompt strength only — never a budget). Effort-native families (Qwen3.8, Ornith 1.5, DeepSeek V4, Muse Glimmer) omit thinking_budget; absence means no implicit cap. Qwen3.8 and Ornith 1.5 accept hipfire’s explicit integer Qwen continuation cap. DeepSeek V4 and Muse Glimmer cap fields are dropped+warned rather than translated into hand-written tokenizer closure. Registry sampling blocks are legacy metadata and are not promoted by the native request resolver. Full contract: CONFIG.md; HTTP examples: SERVE.md.


Registry tags (from registry/models.json)

Fields: Tag, File (file), Size GB (size_gb), Min VRAM GB (min_vram_gb), Default KV (default_kv_mode when set; else empty — global kv_cache=auto resolves to q8), Notes (desc, truncated).

Qwen 3.5 dense / hybrid

TagFileSize GBMin VRAMDefault KVNotes
qwen3.5:0.8bqwen3.5-0.8b.mq40.552.0q8MQ4 default small
qwen3.5:0.8b-mq6qwen3.5-0.8b.mq60.672.2q8MQ6
qwen3.5:2bqwen3.5-2b.mq41.292.8MQ4 (registry desc still mentions legacy HF4 naming)
qwen3.5:2b-hf6qwen3.5-2b.hf61.63HF6
qwen3.5:2b-mq3qwen3.5-2b.mq31.162.7MQ3
qwen3.5:2b-mq6qwen3.5-2b.mq61.633.1MQ6
qwen3.5:4bqwen3.5-4b.mq42.594.1q8MQ4
qwen3.5:4b-mq3qwen3.5-4b.mq32.253.8q8MQ3
qwen3.5:4b-mq6qwen3.5-4b.mq63.485.0q8MQ6
qwen3.5:9bqwen3.5-9b.mq45.316.8q8MQ4; common default
qwen3.5:9b-mq3qwen3.5-9b.mq34.576.1q8MQ3 alpha (gfx11/gfx12 noted in desc)
qwen3.5:9b-mq6qwen3.5-9b.mq67.38.8q8MQ6
qwen3.5:27bqwen3.5-27b.mq415.016q8MQ4
qwen3.5:27b-mq3qwen3.5-27b.mq310.712MQ3 alpha
qwen3.5:27b-mq6qwen3.5-27b.mq621.424MQ6

Qwen 3.5 / 3.6 MoE (A3B)

Sizes below are registry declarations, not a substitute for runtime MoE layout checks.

TagFileSize GBMin VRAMDefault KVNotes
qwen3.5:35b-a3bqwen3.5-35b-a3b.mq419.722q835B / 3B-active
qwen3.6:35b-a3bqwen3.6-35b-a3b.mq4p19.822q8Default graded mq4p SKU
qwen3.6:35b-a3b-mq2qwen3.6-35b-a3b.mq211.614Floor SKU
qwen3.6:35b-a3b-mq3pqwen3.6-35b-a3b.mq3p17.220MQ3+P graded
qwen3.6:35b-a3b-mq4pqwen3.6-35b-a3b.mq4p19.822MQ4+P graded
qwen3.6:35b-a3b-mfp4qwen3.6-35b-a3b.mfp420.222MFP4-E8
qwen3.6:35b-a3b-mq4rqwen3.6-35b-a3b.mq4r18.722Uniform MQ4G256V1/qt13 Redline speed SKU; zero qt44/graded experts. Dated tok/s in registry text is not a live baseline.
qwen3.6:35b-a3b-mq5qwen3.6-35b-a3b.mq523.726Quality SKU
qwen3.6:35b-a3b-mq6qwen3.6-35b-a3b.mq627.730Max quality

Several A3B entries carry an mtp.file sidecar name (qwen3.6-35b-a3b.mtp). MTP enablement is config/runtime gated (mtp_mode, env); registry presence alone is not admission.

Qwen 3.6 dense

TagFileSize GBMin VRAMDefault KVNotes
qwen3.6:27bqwen3.6-27b.mq415.016q8Ships triattn.file in registry; template not effort-native — reasoning_effort dropped+warned, never converted to a cap
qwen3.6:27b-mq3qwen3.6-27b.mq310.712MQ3 alpha

Qwen 3.8 dense

TagFileSize GBMin VRAMDefault KVNotes
qwen3.8:27b-mq3-xtqwen3.8-27b.mq3-xt11.7813q8MQ3V2 XT
qwen3.8:27b-mq3qwen3.8-27b.mq312.6214q8MQ3V2 base
qwen3.8:27b-mq3-proqwen3.8-27b.mq3-pro13.1815q8MQ3V2 Pro
qwen3.8:27b-mq4-xtqwen3.8-27b.mq4-xt14.9816q8MQ4V2 XT (speed; supersedes legacy .mq4r)
qwen3.8:27bqwen3.8-27b.mq415.6617q8MQ4V2 base; default; effort-native (low/medium/xhigh, default xhigh); uncapped think span unless explicit integer cap
qwen3.8:27b-mq4-proqwen3.8-27b.mq4-pro16.4618q8MQ4V2 Pro
qwen3.8:27b-mq5-xtqwen3.8-27b.mq5-xt18.1820q8MQ5V2 XT
qwen3.8:27b-mq5qwen3.8-27b.mq518.7120q8MQ5V2 base
qwen3.8:27b-mq5-proqwen3.8-27b.mq5-pro19.3221q8MQ5V2 Pro
qwen3.8:27b-mq6-xtqwen3.8-27b.mq6-xt21.3923q8MQ6V2 XT
qwen3.8:27b-mq6qwen3.8-27b.mq621.7523q8MQ6V2 base
qwen3.8:27b-mq6-proqwen3.8-27b.mq6-pro22.1724q8MQ6V2 Pro

MQ2V2 is not registered. Explicit qwen3.8:27b-mq4 aliases to qwen3.8:27b. Legacy qwen3.8:27b-fast / qwen3.8:fast alias to qwen3.8:27b-mq4-xt. Reasoning contract: CONFIG.md (effort is semantic only; named thinking_budget dropped on this family).

DFlash draft artifacts (registry)

TagFileSize GBMin VRAMPairs with (by name)
qwen3.5:9b-draftqwen35-9b-dflash-mq4.hfq0.556qwen3.5:9b
qwen3.5:27b-draftqwen35-27b-dflash-mq4.hfq0.9216qwen3.5:27b
qwen3.5:27b-draft-mq3qwen35-27b-dflash-mq3.hfq0.6712qwen3.5:27b (mq3 draft)
qwen3.6:27b-draftqwen36-27b-dflash-mq4.hfq0.9216qwen3.6:27b
qwen3.6:27b-draft-mq3qwen36-27b-dflash-mq3.hfq0.6712qwen3.6:27b
qwen3.8:27b-draft-mq3qwen38-27b-dflash-mq3.hfq0.9816qwen3.8:27b* (same-bit alt)
qwen3.8:27b-draft-mq4qwen38-27b-dflash-mq4.hfq1.2116qwen3.8:27b* (recommended controller)
qwen3.8:27b-draft-mq5qwen38-27b-dflash-mq5.hfq1.4316qwen3.8:27b* (same-bit alt)
qwen3.8:27b-draft-mq6qwen38-27b-dflash-mq6.hfq1.6616qwen3.8:27b* (same-bit alt)
muse-glimmer:draftmuse-glimmer-30b-dflash.mq41.3626muse-glimmer / muse-glimmer:fast

Draft loading is registry-driven: hipfire pull <tag> fetches the draft sidecar alongside the target, and dflash_mode / speculation (CONFIG.md, env-vars.md) decides what happens next — default dflash_mode is off (pull ≠ enable); auto uses the sidecar when present (AR otherwise); on requires it and fails the load without it. Override the sidecar with developer.dflash_draft / HIPFIRE_DFLASH_DRAFT or run --model-draft. Watch for DFlash draft loaded: in load output. Several tags may share one dflash.file; hipfire rm keeps that file while another installed declarer still needs it (CLI.md).

Vision-tower loading is registry-driven the same way: every qwen3.8:27b* tier declares the shared vision.file (qwen3.8-27b-vision.hfq, llama.cpp mmproj-style), so hipfire pull <tag> fetches the tower once and every text quant tier serves images without requantizing the trunk. The loader opens the sidecar as a separate pack and applies it with the trunk’s vision config. vision_mode (CONFIG.md, env-vars.md) decides what happens next — default off never loads the tower (even an explicit --vision is skipped, the daemon’s dflash_mode=off-style hard override, so text loads pay no tower VRAM); auto uses the registry/sibling sidecar when present and stays silently text-only when absent; on requires the declared sidecar and fails the load closed without it. A trunk with an embedded tower declares no sidecar and is unaffected by this key. Override per load with run --vision / serve --vision or HIPFIRE_VISION_SIDECAR (empty opts out); a <trunk-stem>-vision.hfq file beside the trunk is also discovered. hipfire rm keeps the shared file while another installed declarer still needs it (CLI.md).

Qwen3 (non-3.5) dense HF4

TagFileSize GBMin VRAMNotes
qwen3:0.6bqwen3-0.6b.hf40.41standard attention
qwen3:8bqwen3-8b.hf44.16standard attention

Fine-tunes on Qwen 3.5 / 3.6 families

TagFileSize GBMin VRAMNotes
carnice:9bcarnice-9b.mq45.06Hermes tool-use; default_tool_format=hermes
carnice:9b-mq6carnice-9b.mq67.38Hermes MQ6
carnice:27bcarnice-27b.mq415.016Hermes 27B
carnice:27b-mq6carnice-27b.mq621.424Hermes 27B MQ6
qwopus:4bqwopus-4b.mq42.64Qwopus3.5 v3
qwopus:4b-mq6qwopus-4b.mq63.85
qwopus:9bqwopus-9b.mq45.36
qwopus:9b-mq6qwopus-9b.mq67.38
qwopus:27bqwopus-27b.mq415.016
qwopus:27b-mq6qwopus-27b.mq621.424
qwopus3.6:27b-coderqwopus3.6-27b-coder.mq415.016q8 default KV; agentic coder finetune
nex-n2:mininex-n2-mini.mq4p19.8222q8 default KV; Qwen3.5-35B-A3B agentic MoE finetune
ornith-1.5:35b-a3bornith-1.5-35b-a3b.mq419.0222q8 default KV; MQ4G256V2 quality trunk with selective MQ6/Q8 protection; semantic low/medium/xhigh effort (default xhigh), uncapped unless an explicit integer cap is set
ornith-1.5:35b-a3b-mq4rornith-1.5-35b-a3b.mq4r18.7022q8 default KV; uniform MQ4G256V2 Redline SKU, 20,871 qt44 and zero qt13/qt15; same effort contract as the quality trunk. Speed SKU aliases: ornith-1.5:fast / ornith-1.5:35b-a3b-fast → this tag (Muse/Qwen3.8 :fast pattern)

Other families (registry)

TagFileSize GBMin VRAMNotes
deepseek-v4-flashdeepseek-v4-flash-0731.mq2lloyd86.296current 0731 release; DSpark sidecar; default low reasoning_effort, uncapped think span (no named budget)
deepseek-v4-flash:mq2rdeepseek-v4-flash-0731.mq2r8296current 0731 golden MQ2R; MQ2-Lloyd routed experts + MFP4-E8 dense route; matching .mq2r DSpark sidecar
deepseek-v4-flash-previewdeepseek-v4-flash.mq2lloyd8296prior nwoolmer preview package, retained under an explicit preview identity
minimax-m2.7MiniMax-M2.7.mq279.296arch_id=10 Mixtral-style MoE
north-mini-codenorth-mini-code.mq4.hfq1624Cohere2-MoE arch_id=12; registry sampling block is inert metadata today
vibethinker:3bvibethinker-3b.mq4.hfq1.823.5Qwen2 MQ4
vibethinker:3b-mq6vibethinker-3b.mq6.hfq2.515.0Qwen2 MQ6
muse-glimmermuse-glimmer-30b.mq418.612630B dense + perception encoder; MQ4 quality trunk; always-on Onyx reasoning with strength dial; default uncapped think span
muse-glimmer:fastmuse-glimmer-30b.mq4r16.2624MQ4R speed SKU (MQ4 body and attention, Q8 lm_head); same Onyx reasoning contract as muse-glimmer

LFM2.5 (registry)

TagFileSize GBMin VRAMNotes (registry desc)
lfm2.5:350mlfm2.5-350m.q80.381.9350M dense; default artifact is Q8 file
lfm2.5:1.2blfm2.5-1.2b.mq40.72.21.2B Instruct dense
lfm2.5:1.2b-thinkinglfm2.5-1.2b-thinking.mq40.72.21.2B Thinking dense
lfm2.5:8b-a1blfm2.5-8b-a1b.mq44.666.28B-A1B MoE

Registry recommended_settings for LFM tags is low temperature (0.05–0.2) with repeat_penalty 1.05 — applied by the CLI resolver. Do not treat a registry sampling field as active defaults.

Gemma4 (registry — no published artifact yet)

No hipfire-quantized Gemma4 .hfq/.mq4 artifact has been published. The rows below are the intended tag layout for when the quantize + upload lane publishes; the current registry/models.json intentionally lists zero gemma4: tags so that no registry row can point at a 404. See the report at the bottom of this section for the exact repos/files to publish.

Architecture ground truth (crates/hipfire-arch-gemma4/src/config.rs, crates/hipfire-arch-gemma4/src/lowered.rs, and crates/hipfire-arch-gemma4/src/gemma4.rs):

  • Hybrid 5:1 sliding:global — every 6th layer is global (full) attention. The per-layer layer_types array is authoritative; on the 12B dense text config (hidden_size=3840, num_hidden_layers=48, vocab_size=262144) that is 40 sliding + 8 full. Do not assume the period in code.
  • Sliding-window layers: window = 1024 (sliding_window), head_dim = 256 (sliding_head_dim / head_dim), RoPE θ = 10_000 (sliding_rope_theta), RopeType::Default on full head_dim, attention scale = 1.0 (kernels bake 1/√d so decode pre-scales Q by √head_dim), Q/K per-head RMSNorm.
  • Global (full) layers: head_dim = 512 (global_head_dim / full_head_dim), window = 0 (full causal), RoPE θ = 1_000_000 (full_rope_theta), RopeType::Proportional with partial_rotary_factor = 0.25 (only first 64 of 512 dims rotate; remainder NoPE), attention_k_eq_v = true — V shares the pre-k_norm output of k_proj (no v_proj on those layers; v_norm is weight-less, implemented with a ones-filled scratch buffer).
  • KV head counts (variant-dependent): 12B text — num_attention_heads = 16, num_key_value_heads = 8 (sliding) and num_global_key_value_heads defaults to sliding when absent; 31B target (per lowered.rs comments) — n_heads = 32, sliding_n_kv_heads = 16, full_n_kv_heads = 4, sliding_head_dim = 256, full_head_dim = 512, hidden_dim = 21504. The 12B hidden_dim is intermediate_size = 15360, dim = 3840.
  • FFN: SwiGLU with gelu_pytorch_tanh activation, intermediate_size (above) per-layer.
  • Norms: Sandwich RMSNorm — input_layernorm, post_attention_layernorm, pre_feedforward_layernorm, post_feedforward_layernorm per layer plus a learned scalar layer_scalar [1] at layer end. Gemma4 uses plain x * w (weights init 1.0); the HIPFIRE_GEMMA4_NORM_PLUS_ONE=1 toggle bakes the Gemma-2/3 x * (1+w) form at load time.
  • Embeddings: embed_scale = sqrt(hidden_size) multiplied on every lookup; lm_head is TIED to embed_tokens (single GPU allocation, aliased WeightTensor; free skips the head to avoid double-free).
  • Output: final_logit_softcapping = 30.0 — tanh(logits/30)*30 before sampling. norm_eps = 1e-6, max_position_embeddings = 262144 (12B).
  • Stop / EOS: HF eos_token_id is the list [1, 106] (<eos> = 1 and <end_of_turn> / <turn|> = 106); parsing it as a scalar drops 106 and lets decode loop on <turn|> forever. The loader resolves eos_tok against ["<end_of_turn>", …] and masks both.
  • Sampling defaults (Gemma card): temperature 1.0, top_p 0.95, top_k 64 — the registry will carry these in recommended_settings / sampling_profiles when tags land (today no Gemma rows exist to carry them).
  • Thinking: boolean only (official default off). No native reasoning_effort; unsupported effort/named-budget values are dropped with a warning. See CONFIG.md / SERVE.md.
  • Crate: hipfire-arch-gemma4 (Gemma4Config, Gemma4Weights, Gemma4State), Gemma4Bundle { config, weights, state, eos_tok } in hipfire-loader. Gemma4Carrier claims arch ids 13 (gemma4_text) and 22 (gemma4_unified_assistant EAGLE drafter) — see architecture-ids.md.
Intended tagIntended fileRepo (to publish)Notes
gemma4:12bgemma4-12b.mq4schuttdev/hipfire-gemma4-12b or hipfire-models/hipfire-gemma4-12b12B dense text, MQ4 default; hipfire quantize google/gemma-4-12B-it --arch-id 13 --format mq4
gemma4:12b-mq4gemma4-12b.mq4same as abovealias-style explicit quant suffix
gemma4:12b-hfqgemma4-12b.hfqsame repoHFQ4 variant if published
gemma4:12b-draftgemma4-12b-draft.mq4.hfqsame repo or …-drafter side-repoEAGLE draft (arch_id 22) paired with gemma4:12b

No Gemma4 rows are registered until the files above (or the 27B/31B MoE variants) are actually on the Hub and pass scripts/registry_gen.py LFS probing — that script is the admission gate and will fail-closed on a missing file or mismatched size_gb. Publishing steps: hipfire quantize → upload to the Hub → add a models.json entry mirroring a dense neighbour like qwen3.6:27b (fields: repo, file, size_gb, min_vram_gb, desc, recommended_settings with temp 1.0/top_p 0.95/top_k 64) → run scripts/registry_gen.py to stamp sha256/size_bytes/arch_id/quant.


Aliases

String redirects in registry/models.json → aliases (not separate downloads). Partial table — for the complete surface read that file or run hipfire list -r.

AliasResolves to
qwen3.5qwen3.5:4b
qwen3.5:latestqwen3.5:9b
qwen3.5:smallqwen3.5:0.8b
qwen3.5:largeqwen3.5:27b
qwen3.6 / qwen3.6:a3bqwen3.6:35b-a3b
ornith / ornith-1.5 / ornith1.5 / ornith1.5:35b-a3bornith-1.5:35b-a3b
ornith-1.5:fast / ornith-1.5:35b-a3b-fastornith-1.5:35b-a3b-mq4r
qwen3.8 / qwen3.8:latestqwen3.8:27b
qwen3.8:fast / qwen3.8:27b-fastqwen3.8:27b-mq4-xt
qwen3.8:27b-mq4qwen3.8:27b
qwen3.8:draft / qwen3.8:27b-draftqwen3.8:27b-draft-mq4
muse-glimmer:latest / muse-glimmer:quality / muse-glimmer:30bmuse-glimmer
qwen3qwen3:8b
carnicecarnice:9b
qwopusqwopus:9b
qwopus:{4b,9b,27b}-{mq4,hf4}matching primary qwopus:{4b,9b,27b} tag
deepseek4 / deepseek-v4deepseek-v4-flash
deepseek4:mq2r / deepseek-v4:mq2rdeepseek-v4-flash:mq2r
deepseek4:preview / deepseek-v4:previewdeepseek-v4-flash-preview
vibethinkervibethinker:3b
qwen3.5:*-mq4 / *-hf4 / several *-hf6same-size primary or mq6 tag (see registry)
qwen3.5:9b:draft etc.matching *-draft tags

Runtime family map (source, not registry)

Runtime dispatch uses HFQ arch_id (architecture-ids.md). Summary for operators:

Familyarch_idCrateRegistry examples
LLaMA / Mistral / plain Qwen3 path0 / 1hipfire-arch-llamaqwen3:8b, many GGUF/HF4 dense
Qwen3.5 dense hybrid5hipfire-arch-qwen35qwen3.5:*, qwen3.6:27b, carnice/qwopus dense
Qwen3.5 / 3.6 MoE A3B6hipfire-arch-qwen35*:35b-a3b*, ornith-1.5:35b-a3b, nex-n2:mini
Qwen27hipfire-arch-qwen2vibethinker:3b, vibethinker:3b-mq6 (support, not admission)
DeepSeek V4 Flash9hipfire-arch-deepseek4deepseek-v4-flash
MiniMax-M210hipfire-arch-minimaxminimax-m2.7
LFM2.5 dense and MoE11hipfire-arch-lfm2moeall lfm2.5:*
Cohere2-MoE12hipfire-arch-cohere2moenorth-mini-code
Gemma4 text (dense)13hipfire-arch-gemma4(none yet — awaiting publish; intended gemma4:12b)
Gemma4 unified-assistant (EAGLE drafter)22hipfire-arch-gemma4 drafter(none yet; intended gemma4:12b-draft sidecar for 13)

Dense LFM2.5 is supported on arch_id 11. The LFM config parser treats num_experts == 0 as dense SwiGLU on every layer (crates/hipfire-arch-lfm2moe/src/config.rs). Do not claim dense LFM is unsupported.

hipfire-arch-lfm2moe is a non-optional dependency of hipfire-loader / daemon load paths on this tree (see crate Cargo.toml graphs). Feature flags on hipfire-runtime default set do not list a separate arch-lfm2moe toggle the way some other arches do — loader always links the crate.

Capability features (DFlash, CASK, PP, MTP, batched prefill, n-gram) are per-path and often narrower than “model loads”. Spec inventory history: speculation-support-inventory.md (historical). Product claims need source + admissions.yml.

LFM optimized prefill — branch-only scope

Branch-only; not shipped on origin/beta@202282de8759dfa6963ea5184ad2bf2b9259cef6.

Audited branch wording allowed for optimized LFM prefill (and nothing broader):

  • Exact cohort: 350M dense MQ4 fixture path used by the branch runtime fixture validation/guard (lfm2.5-350m.mq4 shape checks in hipfire-arch-lfm2moe forward), not a generic “all LFM” claim.
  • GPU: gfx1201 only for the batched opt-in path.
  • Flag: explicit opt-in HIPFIRE_LFM2_PREFILL_BATCH=1 (default off). Optional chunk override HIPFIRE_LFM2_PREFILL_MAX_BATCH (default 256, hard cap 512 in source).
  • Pin when citing branch implementation: lfm-redline@692a726dde53508cb53de1a74c720e75a7c9f33e (or later branch commits only if re-grounded).

Planned (not implemented claims here): Q8-first generic completion of the optimized path, wider LFM cohorts (1.2B / 8B-A1B), multi-GPU, and Phase-4 default-on. Admitted (exact one row): admissions.yml schema v2 admits only the sealed gfx1201 LFM2.5-350M MQ4 retained-PM4 plain-AR product route; nothing else. Not a current baseline: any exploratory tok/s tables in designs/plans.

Eager per-token prefill / decode remains the portable LFM path when the opt-in flag is off or the GPU is not gfx1201. On gfx1201 with HIPFIRE_LFM2_PREFILL_BATCH=1, the daemon selects the batched path from GPU+flag alone and has no post-selection fallback: requests outside the exact 350M dense MQ4 fixture fail closed at the runtime fixture guard. Source symbol validate_350m_mq4_admission names that fixture check only — it does not create a product admission; admissions.yml remains the sole authority (schema v2, exactly one earned retained-PM4 product row for this sealed fixture).


Bring your own

HuggingFace → quantize → register

hipfire quantize Jackrong/Qwopus3.5-4B-v3 \
  --format mq4 --install --register qwopus:4b

See QUANTIZE.md and QUANTIZATION.md.

Local safetensors directory

Requires config.json + .safetensors. Architectures the engine loads are those with arch crates / loaders above; the quantizer may accept more shapes than inference can run.

GGUF

hipfire quantize ./model.Q4_K_M.gguf --install --register my:tag

Dequant path support is format-specific (common Q4_0 / Q8_0 / Q4_K / Q6_K / F16 / BF16 / F32). Unsupported GGUF quants fail closed in the quantizer.


On-disk layout

~/.hipfire/models/
  <registry file names>
  optional sibling drafts / .triattn*.bin sidecars

Extension hints (loader recognizes several): .mq4, .mq6, .mq4p, .mq4r, .mq2, .mq2lloyd, .mq2r, .mfp4, .hf4, .hf6, .hfq, .q8, and related graded names as produced by quant tooling. Exact dtype routing is loader/kernel source, not this table.


Thinking / chat framing

Reasoning is a three-axis contract (mode / semantic effort / hard cap) owned by config and the OpenAI request layer — not by inventing per-tag field meanings:

  • Axes, defaults, and family table — CONFIG.md
  • HTTP fields, warn+drop metadata, curl examples — SERVE.md
  • Chat template overrides — chat_template, default_chatml / env in env-vars.md

Registry cards may publish reasoning_effort defaults. Effort-native tags should not pin a hipfire named thinking_budget; absence means uncapped. Ornith 1.5 uses Qwen3.8-compatible low/medium/xhigh semantic prompt steering (default xhigh) — never a budget reinterpretation. Qwen <think> framing is not universal — Gemma uses boolean thinking (default off), Glimmer uses Onyx channel strength, DeepSeek uses thinking.type + its own effort ladder.


TopicOwner
Config keys / defaultsCONFIG.md
Env varsenv-vars.md
CLI pull/run/listCLI.md
Arch IDsarchitecture-ids.md
Admissionsadmissions.yml (schema v2; exactly one earned record)
Validation routesVALIDATION.md
Dated benchesBENCHMARKS.md (tables remain historical regardless of admission; admission and measurement classification are independent)