Skip to content

Commit 7f52de6

Browse files
committed
feat: v0.9.12 — Q6_K codegen + AOT validation (closes the Q6_K arc)
Builds on v0.9.11's foundation: emits a Q6_K dot kernel into AOT binaries, wires is_q6k / lm_is_q6k dispatch through every projection / lm_head emission site, threads --force-q6k through forge compile and forge export-weights, and validates the codegen end-to-end via --score-corpus on Llama-3.2-1B. Result on the standard 4.9 KB factual corpus: - AOT-Q6_K --score-corpus : ppl 4.2647, BPB 0.4000, 13 tok/s - Interpreter Q8_0 bench-ppl : ppl 4.2294, BPB 0.3976, 13 tok/s Q6_K vs Q8_0 ppl: +0.8% — exactly the expected high-bit quantization gap. Better than AOT-Q4_K (4.57) by 7% as expected (Q6_K is 6.5 bits/elem vs Q4_K's 4.5). Storage: Q6_K row is 23% smaller than Q8_0 (1680 vs 2176 bytes for k=2048). Throughput is currently scalar-only (no NEON dot4 yet); ~3x slower than the AOT-Q4_K NEON sdot path on the same model. v0.9.13 will add the NEON kernel + per-tensor dtype mixing for real Q4_K_M GGUFs. 339/339 workspace tests still pass.
1 parent 6d3ecaa commit 7f52de6

6 files changed

Lines changed: 407 additions & 45 deletions

File tree

‎CHANGELOG.md‎

Lines changed: 66 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -2,6 +2,72 @@
22

33
All notable changes to ForgeLLM are documented here.
44

5+
## [0.9.12] — 2026-05-01 — Q6_K codegen + AOT validation (closes the Q6_K arc)
6+
7+
Builds on v0.9.11's foundation: emits a Q6_K dot kernel into AOT
8+
binaries, wires `is_q6k` / `lm_is_q6k` dispatch through every
9+
projection / lm_head emission site, threads `--force-q6k` through
10+
`forge compile` and `forge export-weights`, and validates the codegen
11+
end-to-end via `--score-corpus`.
12+
13+
### Result — AOT Q6_K codegen mathematically validated
14+
15+
Llama-3.2-1B-Q8_0 → `--force-q6k` → AOT-Q6_K, on 4.9 KB factual
16+
corpus:
17+
18+
| path | ppl | BPB | tok/s |
19+
| ------------------------------------- | ------ | ------ | ----- |
20+
| AOT-Q6_K `--score-corpus` | 4.2647 | 0.4000 | 13 |
21+
| Interpreter Q8_0 `bench-perplexity` | 4.2294 | 0.3976 | 13 |
22+
23+
Q6_K vs Q8_0 ppl: **+0.8%** — exactly the expected high-bit quant
24+
quality gap. The +0.6% over AOT-Q4_K (4.5733) is also as expected
25+
(Q6_K is 6.5 bits/elem vs Q4_K's 4.5). Storage: Q6_K row is 23%
26+
smaller than Q8_0 (1680 vs 2176 bytes for k=2048).
27+
28+
Throughput is currently scalar-only (no NEON dot4 yet); ~3× slower
29+
than the AOT-Q4_K NEON sdot path on the same model. v0.9.13 will
30+
add the NEON Q6_K kernel.
31+
32+
### Added
33+
34+
- **`emit_q6_k_kernel`** in codegen-cpu — emits scalar
35+
`dot_q6_k_q8_0` into the AOT crate. Reuses `f16_bits_to_f32` and
36+
`quantize_to_q8_0_blocks_into` shared with the Q8_0/Q4_0/Q4_K kernel
37+
families when any of them are also active.
38+
- **`emit_specialized_q6_k_matmul_functions`** — generates
39+
`matmul_vec_q6_k_KxN` per (in, out) projection size. Parallelizes
40+
over output rows when total weight bytes > 1 MB.
41+
- **`is_q6k` / `lm_is_q6k` dispatch** in `emit_forward_function` and
42+
`emit_prefill_function` (both QKV / o_proj / gate-up / down_proj /
43+
lm_head sites + the `proj_type` / `lm_head_type` selection).
44+
- **Q6_K paths in `forgellm-codegen-cpu/src/project.rs`** — row-byte
45+
helpers, weight-loading branch, weight-slicing for both regular and
46+
`--embed-weights` builds.
47+
- **`forge compile --force-q6k`** and **`forge export-weights
48+
--force-q6k`** — mirrors the `--force-q4k` flag. Mutually exclusive
49+
with `--force-q4k`.
50+
51+
### Validated
52+
53+
- All 339 workspace tests still pass; codegen tests still build
54+
syntactically valid Rust for the Q6_K-containing emission.
55+
- Greedy continuation on Llama-3.2-1B-Q6_K produces coherent text
56+
("The capital of France is Paris. The capital of Germany is Berlin.
57+
The capital of Italy is Rome.").
58+
- Per-chunk perplexity matches expected Q6_K quality: +0.8% over Q8_0,
59+
and notably better than Q4_K (4.26 vs 4.57).
60+
61+
### Deferred to v0.9.13
62+
63+
- **NEON Q6_K dot kernel** — current scalar path runs at ~13 tok/s on
64+
Llama-1B; AOT-Q4_K with NEON sdot runs at ~40 tok/s. Adding
65+
`dot4_q6_k_q8_0` with sdot ILP across 4 output rows should close
66+
most of that gap.
67+
- **Per-tensor dtype mixing** — required to load real Q4_K_M GGUFs
68+
(which mix Q4_K + Q6_K + F32 per tensor) without `--force-*`.
69+
Currently codegen requires uniform projection dtype.
70+
571
## [0.9.11] — 2026-05-01 — Native Q6_K quantizer + dot kernel (foundation)
672

773
Q4_K_M GGUFs ship `attn_v` and `ffn_down` as Q6_K, which historically

‎Cargo.lock‎

Lines changed: 9 additions & 9 deletions
Some generated files are not rendered by default. Learn more about customizing how changed files appear on GitHub.

‎Cargo.toml‎

Lines changed: 8 additions & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -13,7 +13,7 @@ members = [
1313
resolver = "2"
1414

1515
[workspace.package]
16-
version = "0.9.11"
16+
version = "0.9.12"
1717
edition = "2021"
1818
license = "MIT"
1919
repository = "https://github.com/sauravpanda/forge-llm"
@@ -24,13 +24,13 @@ readme = "README.md"
2424

2525
[workspace.dependencies]
2626
# Internal crates
27-
forgellm-frontend = { version = "0.9.11", path = "crates/forgellm-frontend" }
28-
forgellm-optimizer = { version = "0.9.11", path = "crates/forgellm-optimizer" }
29-
forgellm-codegen-cpu = { version = "0.9.11", path = "crates/forgellm-codegen-cpu" }
30-
forgellm-codegen-wasm = { version = "0.9.11", path = "crates/forgellm-codegen-wasm" }
31-
forgellm-codegen-gpu = { version = "0.9.11", path = "crates/forgellm-codegen-gpu" }
32-
forgellm-codegen-metal = { version = "0.9.11", path = "crates/forgellm-codegen-metal" }
33-
forgellm-runtime = { version = "0.9.11", path = "crates/forgellm-runtime" }
27+
forgellm-frontend = { version = "0.9.12", path = "crates/forgellm-frontend" }
28+
forgellm-optimizer = { version = "0.9.12", path = "crates/forgellm-optimizer" }
29+
forgellm-codegen-cpu = { version = "0.9.12", path = "crates/forgellm-codegen-cpu" }
30+
forgellm-codegen-wasm = { version = "0.9.12", path = "crates/forgellm-codegen-wasm" }
31+
forgellm-codegen-gpu = { version = "0.9.12", path = "crates/forgellm-codegen-gpu" }
32+
forgellm-codegen-metal = { version = "0.9.12", path = "crates/forgellm-codegen-metal" }
33+
forgellm-runtime = { version = "0.9.12", path = "crates/forgellm-runtime" }
3434

3535
# Serialization
3636
serde = { version = "1", features = ["derive"] }

‎crates/forgellm-cli/src/main.rs‎

Lines changed: 62 additions & 18 deletions
Original file line numberDiff line numberDiff line change
@@ -78,6 +78,12 @@ enum Commands {
7878
/// source GGUFs without needing a Q4_K_M file on disk.
7979
#[arg(long)]
8080
force_q4k: bool,
81+
82+
/// Force projection weights to be re-quantized to Q6_K (v0.9.12+).
83+
/// Same purpose as --force-q4k but routes through the native Q6_K
84+
/// kernel. Mutually exclusive with --force-q4k.
85+
#[arg(long)]
86+
force_q6k: bool,
8187
},
8288

8389
/// Export model weights as a flat binary file for AOT binaries
@@ -94,6 +100,10 @@ enum Commands {
94100
/// `forge compile --force-q4k`).
95101
#[arg(long)]
96102
force_q4k: bool,
103+
104+
/// Force projection weights to be re-quantized to Q6_K.
105+
#[arg(long)]
106+
force_q6k: bool,
97107
},
98108

99109
/// Export a model to ONNX format
@@ -313,21 +323,31 @@ fn main() -> Result<()> {
313323
cross_target,
314324
lora,
315325
force_q4k,
316-
} => cmd_compile(CompileArgs {
317-
model_path: &model,
318-
target: &target,
319-
output_path: &output,
320-
run,
321-
prompt: prompt.as_deref(),
322-
tokenizer_opt: &tokenizer,
323-
embed_weights,
324-
cross_target: cross_target.as_deref(),
325-
lora_path: lora.as_deref(),
326-
force_q4k,
327-
})?,
326+
force_q6k,
327+
} => {
328+
if force_q4k && force_q6k {
329+
bail!("--force-q4k and --force-q6k are mutually exclusive");
330+
}
331+
cmd_compile(CompileArgs {
332+
model_path: &model,
333+
target: &target,
334+
output_path: &output,
335+
run,
336+
prompt: prompt.as_deref(),
337+
tokenizer_opt: &tokenizer,
338+
embed_weights,
339+
cross_target: cross_target.as_deref(),
340+
lora_path: lora.as_deref(),
341+
force_q4k,
342+
force_q6k,
343+
})?
344+
}
328345

329-
Commands::ExportWeights { model, output, force_q4k } => {
330-
cmd_export_weights_impl(&model, &output, None, force_q4k)?
346+
Commands::ExportWeights { model, output, force_q4k, force_q6k } => {
347+
if force_q4k && force_q6k {
348+
bail!("--force-q4k and --force-q6k are mutually exclusive");
349+
}
350+
cmd_export_weights_impl(&model, &output, None, force_q4k, force_q6k)?
331351
}
332352

333353
Commands::ExportOnnx { model, output } => {
@@ -771,6 +791,9 @@ struct CompileArgs<'a> {
771791
/// If true, override the source GGUF dtype to Q4_K and requant projection
772792
/// weights via `quantize_f32_to_q4_k`.
773793
force_q4k: bool,
794+
/// If true, override the source GGUF dtype to Q6_K and requant projection
795+
/// weights via `quantize_f32_to_q6_k`.
796+
force_q6k: bool,
774797
}
775798

776799
fn cmd_compile(args: CompileArgs<'_>) -> Result<()> {
@@ -785,6 +808,7 @@ fn cmd_compile(args: CompileArgs<'_>) -> Result<()> {
785808
cross_target,
786809
lora_path,
787810
force_q4k,
811+
force_q6k,
788812
} = args;
789813
println!("Loading model config from {model_path}...");
790814
let mut config = load_model_config(model_path)?;
@@ -797,6 +821,15 @@ fn cmd_compile(args: CompileArgs<'_>) -> Result<()> {
797821
config.dtype = DType::Q4_K;
798822
config.lm_head_dtype = Some(DType::Q4_K);
799823
}
824+
if force_q6k {
825+
use forgellm_frontend::ir::DType;
826+
println!(
827+
"--force-q6k: overriding projection dtype {:?} → Q6_K (re-quantizing via in-tree quantizer)",
828+
config.dtype
829+
);
830+
config.dtype = DType::Q6_K;
831+
config.lm_head_dtype = Some(DType::Q6_K);
832+
}
800833
config
801834
.validate()
802835
.map_err(|e| anyhow::anyhow!("invalid model config: {e}"))?;
@@ -818,7 +851,7 @@ fn cmd_compile(args: CompileArgs<'_>) -> Result<()> {
818851
if embed_weights {
819852
println!("Exporting weights for embedding...");
820853
let weights_path = output_dir.join("weights.bin");
821-
cmd_export_weights_impl(model_path, &weights_path.to_string_lossy(), lora_path, force_q4k)?;
854+
cmd_export_weights_impl(model_path, &weights_path.to_string_lossy(), lora_path, force_q4k, force_q6k)?;
822855

823856
println!("Copying tokenizer for embedding...");
824857
let tokenizer_src = resolve_tokenizer(tokenizer_opt, model_path)?;
@@ -934,7 +967,7 @@ fn cmd_compile(args: CompileArgs<'_>) -> Result<()> {
934967
let weights_path = output_dir.join("weights.bin");
935968
let weights_path_str = weights_path.to_string_lossy().to_string();
936969
println!("[1/4] Exporting weights...");
937-
cmd_export_weights_impl(model_path, &weights_path_str, lora_path, force_q4k)?;
970+
cmd_export_weights_impl(model_path, &weights_path_str, lora_path, force_q4k, force_q6k)?;
938971

939972
println!("[2/4] Resolving tokenizer...");
940973
let tokenizer_src = resolve_tokenizer(tokenizer_opt, model_path)?;
@@ -2343,14 +2376,15 @@ fn cmd_models(dir: &str) -> Result<()> {
23432376
}
23442377

23452378
fn cmd_export_weights(model_path: &str, output_path: &str) -> Result<()> {
2346-
cmd_export_weights_impl(model_path, output_path, None, false)
2379+
cmd_export_weights_impl(model_path, output_path, None, false, false)
23472380
}
23482381

23492382
fn cmd_export_weights_impl(
23502383
model_path: &str,
23512384
output_path: &str,
23522385
lora_path: Option<&str>,
23532386
force_q4k: bool,
2387+
force_q6k: bool,
23542388
) -> Result<()> {
23552389
use forgellm_frontend::ir::DType;
23562390
use forgellm_frontend::weight_loader::{load_from_file_mixed_with_target, WeightData};
@@ -2365,16 +2399,26 @@ fn cmd_export_weights_impl(
23652399
config.dtype = DType::Q4_K;
23662400
config.lm_head_dtype = Some(DType::Q4_K);
23672401
}
2402+
if force_q6k {
2403+
eprintln!(
2404+
"--force-q6k: overriding projection dtype {:?} → Q6_K",
2405+
config.dtype
2406+
);
2407+
config.dtype = DType::Q6_K;
2408+
config.lm_head_dtype = Some(DType::Q6_K);
2409+
}
23682410

23692411
let is_q8 = config.dtype == DType::Q8_0;
23702412
let is_q4 = config.dtype == DType::Q4_0;
23712413
let is_q4k = config.dtype == DType::Q4_K;
2414+
let is_q6k = config.dtype == DType::Q6_K;
23722415

2373-
if is_q8 || is_q4 || is_q4k {
2416+
if is_q8 || is_q4 || is_q4k || is_q6k {
23742417
let quant_label = match config.dtype {
23752418
DType::Q8_0 => "Q8_0",
23762419
DType::Q4_0 => "Q4_0",
23772420
DType::Q4_K => "Q4_K",
2421+
DType::Q6_K => "Q6_K",
23782422
_ => "?",
23792423
};
23802424
if lora_path.is_some() {

0 commit comments

Comments
 (0)