hipfire is alpha. Real-world testing on cards we don't have, kernel work on archs we don't ship for, bug reports with full reproduction, and new model architecture support — all welcome.
Both paths below use only the installer-provided binaries (the
hipfire wrapper, daemon, and quantizer dropped into ~/.hipfire/bin/
by scripts/install.sh). No cargo, no ROCm SDK, no source build.
If you have an RDNA card the maintainer doesn't, the highest-leverage thing is running the standard tester workflow and posting numbers.
hipfire diag # ROCm + arch detection sanity check
hipfire pull qwen3.5:0.8b # ~0.5 GB; fits any RDNA card
hipfire pull qwen3.5:9b # ~5.3 GB; needs 6 GB+ VRAM
hipfire bench qwen3.5:0.8b --runs 5 # decode + prefill tok/s over 5 runs
hipfire bench qwen3.5:9b --runs 5For 16 GB+ cards, also pull and bench qwen3.5:27b. For mixed-arch
or non-Linux setups, see hipfire diag --help for environment-
specific guidance.
Open an issue titled Benchmarks: <your GPU> and paste the diag
output + each bench block. Results land in
docs/BENCHMARKS.md. The hipfire-tester skill
in .skills/hipfire-tester/ walks an AI agent through this end-to-
end if you want help.
hipfire diag # capture everything first
Open an issue with: GPU + ROCm version, exact command, full error
output (not just the last line), and the diag output. The
hipfire-autoheal skill (in .skills/hipfire-autoheal/) is a
fix-catalog walkthrough that an agent can apply on your behalf for
common runtime issues; if it doesn't resolve cleanly, that's exactly
the case we want filed.
git clone https://github.com/warpfront/hipfire
cd hipfire
cargo build --release --features deltanet --example daemon -p hipfire-runtime
cargo build --release --features deltanet --example test_kernels -p hipfire-runtime
cargo build --release -p hipfire-quantize
./scripts/install-hooks.shRequires Rust 1.75+ and ROCm 6+ (the dev workflow needs hipcc for
kernel JIT). Pre-compiled kernel blobs ship for gfx1010 / gfx1030 /
gfx1100 / gfx1200; other arches JIT-compile on first load.
scripts/install-hooks.sh is idempotent; it sets
core.hooksPath=.githooks and makes the local pre-commit hook
executable.
The default CI path intentionally avoids AMD GPU access and model downloads:
./scripts/no-gpu-ci.shIt runs cargo check --workspace --examples, no-GPU Rust unit tests,
CPU Python tests, and the env/docs drift check. Hardware-relevant changes
(kernel, dispatch, quant, forward-pass, load, serve, spec-decode) must pass
the required hw-gate CI check — see
docs/VALIDATION.md § hw-gate.
./target/release/examples/test_kernels # ~30s, no model neededValidates every dispatched kernel against a CPU reference on the detected arch. This is the load-bearing correctness gate for any arch port; if it fails on your hardware we want to hear about it (see issue template / autoheal skill).
Any change to kernels, dispatch, fusion, rotation, rmsnorm, sampling,
the spec-decode path, loader/daemon, or the forward pass MUST validate the
actual path under test. CI acceptance is hw-gate
(.github/workflows/hw-gate.yml,
scripts/hw-gate/); a maintainer applies the hw-run
label to authorize the hardware run. The fixed coherence-gate*.sh batteries
are retired and must not be used as acceptance evidence. Optional local
python3 -m tools.change_gate is not CI evidence.
python3 scripts/redline_daemon_harness.py --model /path/to/model --pm4
python3 scripts/serve_harness.py --model /path/to/model --mode battery --sampling greedy
./scripts/speed-gate.sh --fast # 4B prefill+decode regression vs tests/speed-baselines/<arch>.txtDon't bypass with --no-verify. A regression the gate catches is
information. Authorized exceptions need explicit written sign-off from
the maintainer for that specific change. Read
docs/methodology/perf-benchmarking.md
before claiming any perf win — within-session A/B noise is ±10–15% on
gfx1100, and the bench harness has a stale-binary trap that's bitten
us before.
# Kernel source: kernels/src/<name>.hip
# Per-arch overrides: kernels/src/<name>.gfx12.hip (family tag)
# kernels/src/<name>.gfx1100.hip (chip tag)
# Register in: crates/rdna-compute/src/kernels.rs
# Wire dispatch in: crates/rdna-compute/src/dispatch.rs
# After editing any .hip file, regenerate hashes for the pre-compiled
# blob loader (otherwise the runtime falls back to JIT):
./scripts/write-kernel-hashes.sh
# Compile-check across the supported arch matrix:
./scripts/compile-kernels.sh gfx1010 gfx1030 gfx1100 gfx1200 gfx1201Architecture deep-dive: docs/ARCHITECTURE.md.
Quantization design (MQ4 / HF4 / asym KV math):
docs/QUANTIZATION.md. Tuning an existing
kernel for perf (multi-row, K-tile depth, wave64 port, prefetch,
ISA flags) — see .skills/hipfire-kernel-tuning/,
which catalogs the empirical methodology + every lever this repo's
git log has actually used.
The .skills/hipfire-arch-port/ directory is the canonical entry
point — playbook, WMMA matrix, validation procedure, contributor
onboarding. Don't write code before reading it; six-week silent-
corruption bugs from getting WMMA C-mappings wrong are how every
prior arch port has gone sideways.
Recent reference: PR #56 (RobinVanCauter, gfx1201 / 9070 XT) walked the skill end-to-end and shipped a full validated 5-kernel port + 6 channel tests in one round. That's the bar.
Start with docs/ARCHITECTURE.md's "Two model
paths" section — crates/hipfire-runtime/src/llama.rs is the template
for dense Llama-style models, crates/hipfire-arch-qwen35/src/qwen35.rs
is the Qwen 3.5 hybrid path with DeltaNet linear attention. Add the architecture string to
from_gguf / from_hfq and patch the tensor-shape divergences.
For a new GGUF dequant type (Q5_K, IQ-quants, etc.), port from
llama.cpp's ggml-quants.c into
crates/hipfire-quantize/src/gguf_input.rs. ~150 lines per format.
| Type | Pattern |
|---|---|
| Feature | feature/<short-name> |
| Bug fix | fix/<short-name> |
| Arch port | port/<arch>-<kernel> |
| Benchmark contribution | bench/<gpu-name> |
Concise description, before/after numbers if perf-sensitive, mention
which gates passed. One logical change per PR. Run cargo fmt and
cargo clippy before submit; CI enforces both.
For perf claims: include the binary md5
(md5sum target/release/examples/bench_qwen35_mq4) and the prompt
md5 if the bench is prompt-dependent. Without these, the result is
unreproducible — see
docs/methodology/perf-benchmarking.md.
cargo fmt— required, CI-enforced.cargo clippy— no new warnings.- No Python in the inference hot path. Python is fine for tooling, benchmarks, and offline analysis; never in the engine.
- Comment HIP kernel parameters: VGPR budget, wave occupancy, LDS
usage, K-tile depth — anything a reader needs to understand the
perf shape without inspecting
--save-tempsoutput.
The 0.1.20 modularization split crates/engine/ into a runtime crate
and per-arch crates. The post-modular workspace:
crates/
hip-bridge/ HIP/ROCm FFI
rdna-compute/ kernel dispatch + per-RDNA-arch routing
(gfx1100/01/02/1150/1151/1152/1200/1201)
hipfire-runtime/ LM runtime: KV cache, sampler, loop_guard,
prompt_frame, eos_filter, spec decode,
eviction, paging
hipfire-arch-qwen35/ Qwen3.5 family (DeltaNet hybrid, MoE)
hipfire-arch-qwen35-vl/ Qwen3.5-VL (vision)
hipfire-arch-llama/ Llama-family (currently a facade — see
PR 14 for physical split)
hipfire-arch-toy/ minimal stub arch (reference for porters)
hipfire-quantize/ safetensors → .mq4 / .hfq quantizer CLI
- "I want to add a new model architecture" → new
crates/hipfire-arch-<name>/crate, implementArchitecturetrait. Copycrates/hipfire-arch-toy/as a template. - "I want to fix a kernel bug or add a kernel" →
kernels/src/*.hipfor the kernel +crates/rdna-compute/src/dispatch.rsfor the dispatch wiring. Stays in rdna-compute regardless of which arch uses it. - "I want to tune sampler / repeat_penalty / blocked tokens
behavior" →
crates/hipfire-runtime/src/sampler.rs. - "I want to add an end-of-turn marker for an arch" → arch crate's
eos_filter_overrides()returningEosFilterOverrides { stop_at: ..., holdback_prefixes: ... }. - "I want to add a CLI feature / daemon API endpoint" →
crates/hipfire-runtime/examples/daemon.rs. (Or, if it's CLI-side, the cli crate / TUI.) - "I want to optimize for a specific RDNA generation" →
crates/rdna-compute/src/dispatch.rs. NEVER inside an arch crate (that fragments per-arch knowledge across the workspace). - "I want to add a new quant format" →
kernels/src/for the kernel +crates/rdna-computefor dispatch routing +crates/hipfire-quantizefor the quantizer CLI. Arches consume via the runtime API automatically.
Every arch crate impl Architecture for Foo and may override four
behavior structs. Defaults assume Qwen3.5 family conventions (ChatML
prompt frame, <think> strip, default sampler/loop-guard config).
| Override | When to use | Example |
|---|---|---|
LoopGuardOverrides |
Base/instruct model legitimately repeats short phrases (structured output, code boilerplate) | LoopGuardOverrides { ngram_threshold: Some(8), .. } |
SamplerOverrides |
Add arch-specific blocked tokens (e.g. <tool_call> openers) or per-arch repeat_penalty |
SamplerOverrides { blocked_tokens: vec![99999], .. } |
PromptFrameOverrides |
Non-ChatML completion model (no `< | im_start |
EosFilterOverrides |
Arch-specific end-of-turn markers (Gemma's <end_of_turn>, etc.) |
EosFilterOverrides { stop_at: vec![b"<end_of_turn>".to_vec()], .. } |
Field-level docs live on
crates/hipfire-runtime/src/arch.rs; worked examples are in
crates/hipfire-arch-toy/src/arch.rs (one of each, default-bodied).
| Skill | When to use |
|---|---|
hipfire-tester |
First-time bringup + bench submission on a new GPU. |
hipfire-diag |
"Hipfire isn't working — what's wrong?" Captures GPU/HIP/kernel state. |
hipfire-autoheal |
Runtime issue triage: daemon hangs, JIT failures, port conflicts, OOM. |
hipfire-arch-port |
Porting hipfire to a new GPU arch (gfx12, gfx1152, gfx94x, …). |
hipfire-kernel-tuning |
Optimize an existing kernel — pick a lever (multi-row, K-tile depth, prefetch, wave64 port, WMMA/MFMA, fused projections, ISA flags) and validate the win across the supported arch matrix. |
Each skill has a SKILL.md (or skill.json + sibling .md files)
that any agent framework can load. Designed for Claude Code / Cursor /
Codex but framework-agnostic.
These three are real open questions where contributor input would land cleanly:
- Issue #57 — gfx12 (RDNA4) WMMA dispatch wiring + perf measurement vs the dot2 fallback. Needs R9700 / 9070 XT hardware. PR #56 landed the kernels; #57 measures and flips the dispatch.
- Issue #58 — multi-GPU support roadmap. Pipeline-parallel first cut design open for discussion. Mostly weighing in vs writing code, unless you have a multi-GPU rig and want to prototype the device-aware tensor allocator.
- Issue #50 — gfx1152 (Strix Halo APU) crash — awaiting a reproducer + dmesg from the original reporter. If you have one, comment there.
For everything else, hipfire list -r + a benchmark issue on any
arch we don't have local numbers for is welcome.
hipfire is licensed under the Apache License 2.0 (see LICENSE-APACHE) as of v0.3.0.
Individual files whose substantive authors have not elected Apache-2.0 remain under the MIT License (see LICENSE-MIT) and are identified by their per-file SPDX header. Those grants are preserved unchanged — nothing is relicensed in absentia. See LICENSE and NOTICE for attribution details and per-file SPDX semantics. The decision record (the 2026-05-19 course correction from a unilateral relicense to dual licensing, and the v0.3.0 move to outbound Apache-2.0) is at docs/governance/relicense-2026-05.md.
- New contributions default to Apache-2.0. By submitting a
contribution and signing off via
git commit -s(Developer Certificate of Origin — https://developercertificate.org/), you certify that you have the right to license your contribution under Apache-2.0 and that you intend to do so. - All commits MUST be signed off via
git commit -s. PRs without a DCO sign-off line on every commit will be asked to amend. - Contributors may explicitly elect MIT-only for their
contribution. State this in the PR description (a short note like
"license: MIT only" is enough) and the merger will tag the relevant
files accordingly. The SPDX header will read
SPDX-License-Identifier: MIT. The project ships under Apache-2.0 overall; your specific files stay MIT-only and keep that grant. - Add an SPDX header to every new source file you create. Templates
live in docs/governance/relicense-2026-05.md.
For sole-author files the default (Apache-2.0) template is:
// SPDX-License-Identifier: Apache-2.0 // Copyright (c) 2026 <Your Name> // hipfire — see LICENSE and NOTICE in the project root. - For substantial modifications (>30% of lines rewritten in an existing file), add your own copyright line BELOW existing ones. Do NOT remove existing copyright lines.
Your prior contributions remain licensed exactly as you originally submitted them — under the MIT license that hipfire used at the time. Nothing in the dual-licensing transition revokes or modifies that grant.
If you would like your prior contributions to be available under
Apache-2.0 as well (so that hipfire's NOTICE-backed attribution
machinery applies to them and downstream forks pulling under
Apache-2.0 get your code too), you can opt in by commenting on the
relicense tracking issue (link to be added here once the issue is
opened). After opt-in, the maintainer will re-run
scripts/governance/apply_spdx_headers.py --rewrite-spdx to refresh
the SPDX tags on the files where you are a substantive author.
Opt-in is entirely voluntary. Files where you are the sole
substantive author currently carry SPDX-License-Identifier: MIT;
that stays unless you elect otherwise. Files of mixed authorship
where you are one of multiple substantive authors carry
SPDX-License-Identifier: MIT OR Apache-2.0 until everyone
involved on that file has either opted in or declined.
hipfire as a whole is offered under Apache-2.0, so redistribution must satisfy Apache License 2.0 § 4 for the work — including shipping a readable copy of NOTICE. Individual MIT-tagged files carry an additional, narrower obligation:
- For files whose SPDX header reads MIT, the LICENSE-MIT text applies: preserve the copyright notice and permission text.
- If you redistribute under Apache-2.0, the LICENSE-APACHE text
applies, including § 4 obligations:
- (a) Include a copy of the Apache-2.0 license to recipients.
- (b) Mark modified files prominently as modified.
- (c) Preserve per-file copyright, patent, trademark, and attribution notices in the Source form. Do not strip SPDX headers, copyright lines, or the CREDITS.md inventory.
- (d) Include a readable copy of NOTICE in derivative-work distributions.
For files tagged SPDX-License-Identifier: MIT OR Apache-2.0 you
may pick either license for your use of that file. Files tagged
SPDX-License-Identifier: MIT are MIT-only — you do not get an
Apache-2.0 patent grant from them — and files tagged
SPDX-License-Identifier: Apache-2.0 are Apache-only.
Stripping attribution when redistributing hipfire code is a license
violation, treated as copyright infringement per Jacobsen v. Katzer, 535 F.3d 1373 (Fed. Cir. 2008). The point of the
dual-license transition is accreditation protection, not IP
control: forks remain welcome under either license at the
recipient's option, but attribution MUST travel with the code.
See AGENTS.md for the project-level notice addressed to AI agents helping users derive from this codebase, and PRIOR-ART.md for the inventory of original architectural innovations originating in hipfire (with dates and canonical commit hashes) that derivative works should attribute even when no code is copied verbatim.