Inferno is a lightweight Rust inference engine built specifically for running GLM-5.2 Q2 on Apple Silicon through native Metal kernels. Its primary target is a MacBook Pro with 64 GB of unified memory.
The name reflects the engineering challenge: the model is much larger than the available memory, so useful local inference requires careful coordination of Metal, unified memory, and SSD streaming. Inferno stays deliberately narrow. It supports one model layout and optimizes that path instead of becoming a general inference framework.
Warning
Inferno is experimental and under active development. It is not production ready, and correctness, stability, long-context behavior, and performance are still being validated.
- Apple Silicon and Metal only.
GLM-5.2-UD-Q2_K_RoutedQ2K.ggufonly.- Exact top-8 routed-expert execution.
- Native Metal execution with no CPU compute fallback.
- Interactive chat and one-shot generation.
- SSD-streamed routed experts and compressed KV history.
- Greedy decoding.
- Experimental MTP speculative decoding, disabled by default because it is currently slower than ordinary decode on the 64 GB target.
Inferno is not a general model loader and does not support arbitrary GGUF architectures or quantization formats.
| Metric | Result |
|---|---|
| Decode throughput | 1.711 tokens/s |
Measured with exact top-8 routing on a 64 GB Apple Silicon MacBook Pro using a release build. Prompt prefill is excluded. Results depend on prompt length, expert-cache state, SSD activity, and memory pressure.
- A Mac with Apple Silicon and Metal support.
- 64 GB of unified memory is the current development target.
- Rust and Cargo.
- The Hugging Face CLI for the download commands below.
- More than 503 GB of free SSD space for the source GGUF and ExpertPack, excluding temporary KV data and build outputs.
Install the Hugging Face CLI with Homebrew if needed:
brew install hfInferno uses the Q2 GGUF published by Antirez:
GLM-5.2-UD-Q2_K_RoutedQ2K.gguf
Source: antirez/glm-5.2-gguf
Inferno also supports a lossless derived ExpertPack. It preserves the original Q2 expert bytes but stores each routed expert's gate, up, and down matrices next to each other. This layout reduces scattered SSD reads. The ExpertPack is optional, but the documented performance assumes it is present.
ExpertPack: allemanfredi/inferno-glm-5.2-q2-expertpack
Create the model directory:
mkdir -p models/glm-5.2Download the source model files:
hf download antirez/glm-5.2-gguf \
config.json \
generation_config.json \
tokenizer.json \
GLM-5.2-UD-Q2_K_RoutedQ2K.gguf \
--local-dir models/glm-5.2Download the prebuilt ExpertPack:
hf download allemanfredi/inferno-glm-5.2-q2-expertpack \
GLM-5.2-UD-Q2_K_RoutedQ2K-Inferno-ExpertPack-v1.q2pack \
--local-dir models/glm-5.2The final directory must contain:
models/glm-5.2/
config.json
generation_config.json
tokenizer.json
GLM-5.2-UD-Q2_K_RoutedQ2K.gguf
GLM-5.2-UD-Q2_K_RoutedQ2K-Inferno-ExpertPack-v1.q2pack
The two weight artifacts occupy approximately 503 GB in total. Inferno validates the ExpertPack against the source GGUF before using it.
To recreate the ExpertPack locally instead of downloading it:
cargo run --release -p inferno --example pack_experts -- \
--model models/glm-5.2Creation is resumable at complete expert-record boundaries.
cargo build --releaseThe production binary is:
target/release/inferno
Start an interactive chat using the default models/glm-5.2 directory:
target/release/infernoUse /clear to reset the conversation and /exit to close the session. The
model and expert cache stay alive across turns, but each turn currently
re-prefills the accumulated conversation into a new request KV cache.
Use an explicit model directory when needed:
target/release/inferno chat --model /path/to/modeltarget/release/inferno generate \
--model models/glm-5.2 \
--prompt "Tell me the capital of Italy."Without --max-new-tokens, generation continues until an EOS token or the
context limit.
Inferno can act as the local model provider for Codex. Start the persistent model process in one terminal:
target/release/inferno serve --model models/glm-5.2Install the included Codex profile:
mkdir -p ~/.codex
cp examples/inferno.config.toml ~/.codex/inferno.config.toml
cp examples/inferno.models.json ~/.codex/inferno.models.jsonThen start Codex in a repository:
codex --profile infernoCodex sends each turn through its streaming Responses API. Inferno translates the conversation and direct function tools into GLM-5.2's native prompt, generates either text or a tool call, and returns that action to Codex. Codex executes the tool locally and sends the result back for the next model turn. The model, Metal backend, and expert cache remain alive between requests.
Requests are processed one at a time because they share one Metal runtime.
The profile exposes only exec_command and write_stdin to GLM. Codex still
enforces its sandbox and approval policy, but omitting unrelated tool schemas
keeps the expensive GLM prefill small. Plugin namespaces and hosted web search
are not part of this first integration.
Measure decode throughput:
target/release/inferno generate \
--model models/glm-5.2 \
--prompt "Tell me the capital of Italy." \
--max-new-tokens 8 \
--measure-tokens-per-secondRecord memory telemetry without mixing it into generated text:
target/release/inferno generate \
--model models/glm-5.2 \
--prompt "Tell me the capital of Italy." \
--telemetry-file /tmp/inferno-memory.logRecord adaptive cache decisions:
target/release/inferno generate \
--model models/glm-5.2 \
--prompt "Tell me the capital of Italy." \
--enable-unified-memory-controller \
--memory-controller-log /tmp/inferno-memory-controller.tsv| Option | Purpose |
|---|---|
--max-new-tokens <N> |
Limits the number of generated tokens. |
--measure-tokens-per-second |
Reports time to first token and decode throughput. |
--expert-cache-gb <GB> |
Pins the routed-expert RAM budget. |
--hot-kv-cache-gb <GB> |
Pins the hot Metal KV budget. |
--enable-unified-memory-controller |
Enables adaptive expert and hot-KV memory rebalancing. |
--enable-telemetry |
Prints runtime memory telemetry. |
--telemetry-file <PATH> |
Writes memory telemetry to a file. |
--memory-controller-log <PATH> |
Writes adaptive memory decisions to TSV. |
--speculative-mtp |
Enables the experimental MTP path. |
Use the executable help as the authoritative CLI reference:
target/release/inferno --help
target/release/inferno generate --help
target/release/inferno chat --help
target/release/inferno serve --helpInferno uses the SSD as a large, slower storage tier and unified memory as a smaller, faster tier. It maintains two separate caches.
Expert cache. For every token, the router selects eight experts. If an expert is already resident in a Metal cache slot, Inferno uses it immediately. Otherwise, Inferno reads that expert's Q2 weights from the ExpertPack on SSD, places them in a reusable slot, and executes it. When the cache is full, a less useful expert is replaced. Reusing a resident expert avoids another SSD read.
selected expert -> cache hit -> execute from Metal
-> cache miss -> read from SSD -> cache -> execute
KV cache. The KV cache is the model's memory of previous tokens. Inferno stores the complete compressed history in an append-only Q8 block store on SSD. A bounded Metal window contains only the rows currently needed by attention. If a row is not in that window, Inferno reads it from SSD and loads it into the window. DSA reduces this work by selecting only the most relevant history rows for long-context attention.
required KV row -> hot hit -> read from Metal
-> cold miss -> read from SSD -> Metal window -> use
GLM-5.2 MLA makes this practical by storing 512 latent values and 64 RoPE values per token and layer instead of expanded K/V for every attention head. The source GGUF remains memory-mapped so macOS can manage always-used weights through its page cache.
The expert cache and hot KV window compete for the same physical unified
memory. Inferno's fixed default starts from 30 expert slots per routed layer
and a 512 MiB hot KV budget. --enable-unified-memory-controller enables a
native controller that samples Mach and Metal counters without starting
subprocesses, tracks separate prefill and decode Metal high-water marks, and
releases prefill-only buffers before decode.
During decode, the controller changes one cache at a time and measures the next eight-token window. It keeps a change only when decode throughput improves; otherwise it restores the previous budget. Expert slots and hot KV are tuned independently. The controller targets 4 GiB of effective headroom, treats 3 GiB as the hard floor, and shrinks immediately when swap or compression grows. Explicit cache-size options pin that cache and disable automatic resizing for it.
cargo fmt --all --check
cargo check --workspace
cargo test --workspaceGLM-5.2 was created by Z.ai.
The target Q2 GGUF was produced and published by Antirez in antirez/glm-5.2-gguf.
Inferno is licensed under the MIT License.