Skip to content

Commit 26b9ff0

Browse files
authored
Bump llama.cpp to b10830 (465e49b9c, upstream v0.4.0) (#89)
165 commits past b10665 (ca3d5a3e1). LLAMA_COMMIT in the Makefile moves with the submodule. No NIF source change. API surface - include/llama.h: only llama_tensor_read_lazy -> llama_lazy_mode (tensor_read_lazy -> lazy_mode in llama_model_params, #27969). The NIF neither sets nor reads it. - common/common.h: same rename plus preserve_reasoning_specified. - ggml-backend.h, ggml-rpc.h, chat.h, json-schema-to-grammar.h, speculative.h: unchanged. - llama_model_default_params() / llama_context_default_params(): value-identical to b10665 apart from the rename. Upstream defects we work around: all three still stand (source diff) - ggml_backend_rpc_start_server still returns void. - ggml_backend_cuda_comm_init untouched; RPC set/get_tensor_2d hooks still NULL. RPC diff is #26500 (cross-server buffer serialisation) and #27960 (ggml_op_alloc_size_may_expand). - ggml-cpu/CMakeLists.txt diff adds iqp.cpp and gates SpacemiT IME sources; the -mcpu=native probe is untouched. - :row split mode still throws on CUDA (no ggml_backend_split_buffer_type). Behaviour change worth knowing - TENSOR_READ_LAZY tensors (Gemma 4 per-layer embeddings, Qwen4-exp PLE) now map the file regardless of load_mode (#27837). Verification (macOS, Metal, M1 Max, LLAMA_BACKEND=metal) - mix test (no model): 428 passed, 149 excluded - --include smoke embeddings slow mtp (Qwen3.5-0.8B-UD-Q4_K_XL, Qwen3-Embedding-0.6B-f16, Qwen3.6-35B-A3B-MTP-UD-Q4_K_XL): 569 passed, 8 excluded - --include mtp_sidecar (Qwen3.8-27B-Q4_K_M + mtp-Qwen3.8-27B-Q4_0): 434 passed, 143 excluded - --only mtp_cancel: fails as documented, "verify decode failed: code=-1", no abort - mix compile --warnings-as-errors and mix format --check-formatted clean Not run: rpc_live (needs an RPC build and a worker); the Hex tarball source-build check.
1 parent 2114abe commit 26b9ff0

5 files changed

Lines changed: 48 additions & 11 deletions

File tree

‎CHANGELOG.md‎

Lines changed: 37 additions & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -6,10 +6,10 @@ DGX Spark (GB10) support: a silent ARM code-generation bug fixed, the ggml RPC
66
backend wired up so a model can span two machines, and a measured runbook for
77
both configurations in [docs/dgx-spark.md](docs/dgx-spark.md).
88

9-
llama.cpp bumped to [`b10665`](https://github.com/ggml-org/llama.cpp/releases/tag/b10665)
10-
(`ca3d5a3e1`), by way of b10435 and b10582, which brought Qwen 3.8 in under the
11-
existing `qwen35` architecture, and MTP support for its target/sidecar split
12-
(see Added).
9+
llama.cpp bumped to [`b10830`](https://github.com/ggml-org/llama.cpp/releases/tag/b10830)
10+
(`465e49b9c`, upstream v0.4.0), by way of b10435, b10582 and b10665, which
11+
brought Qwen 3.8 in under the existing `qwen35` architecture, and MTP support
12+
for its target/sidecar split (see Added).
1313

1414
Verified on macOS (Metal) at `e85caa81e`, running **every** tag the suite
1515
excludes by default. Default build: **428 passed, 149 excluded** with no model;
@@ -26,6 +26,13 @@ Re-verified at `ca3d5a3e1` (b10665) on macOS (Metal): default build
2626
**428 passed, 149 excluded** with no model; **434 passed** for
2727
`--include mtp_sidecar` (Qwen3.8-27B-Q4_K_M plus its `mtp-*-Q4_0` head).
2828

29+
Re-verified at `465e49b9c` (b10830) on macOS (Metal), M1 Max: default build
30+
**428 passed, 149 excluded** with no model; **569 passed, 8 excluded** for
31+
`--include smoke --include embeddings --include slow --include mtp`
32+
(Qwen3.5-0.8B-UD-Q4_K_XL, Qwen3-Embedding-0.6B-f16,
33+
Qwen3.6-35B-A3B-MTP-UD-Q4_K_XL); **434 passed, 143 excluded** for
34+
`--include mtp_sidecar` (Qwen3.8-27B-Q4_K_M plus its `mtp-*-Q4_0` head).
35+
2936
The one tag that is not green is `:mtp_cancel`, and it moved: see Changed.
3037

3138
### Fixed
@@ -198,6 +205,27 @@ The one tag that is not green is `:mtp_cancel`, and it moved: see Changed.
198205
`common/speculative.h` only gained functions (`common_speculative_n_max`,
199206
synthetic-acceptance helpers); every `common_speculative_*` call the NIF
200207
makes is signature-identical.
208+
- **llama.cpp bumped to `465e49b9c`** (b10830, upstream v0.4.0), 165 commits
209+
past b10665, and `LLAMA_COMMIT` moved with the submodule. No binding edit:
210+
the only `include/llama.h` change is the rename `llama_tensor_read_lazy` →
211+
`llama_lazy_mode` (`tensor_read_lazy` → `lazy_mode` in
212+
`llama_model_params`, #27969), which the NIF neither sets nor reads;
213+
`llama_model_default_params()` / `llama_context_default_params()` are
214+
otherwise value-identical to b10665, and `ggml-backend.h`, `ggml-rpc.h`,
215+
`common/chat.h`, `common/json-schema-to-grammar.h` and
216+
`common/speculative.h` did not change. One upstream behaviour change worth
217+
knowing: `TENSOR_READ_LAZY` tensors (Gemma 4 per-layer embeddings, Qwen4-exp
218+
PLE) now map the file even under a non-mmap `load_mode` (#27837), so
219+
`:use_mlock` alone no longer guarantees "no mapping" for those two tensor
220+
kinds — every other tensor still behaves as `Model.load/2` documents.
221+
All three defects in [docs/release-guide.md](docs/release-guide.md) still
222+
stand, re-checked as a source diff: `ggml_backend_rpc_start_server` is
223+
unchanged, `ggml_backend_cuda_comm_init` was untouched (the RPC diff is
224+
#26500's cross-server buffer serialisation fix and #27960's
225+
`ggml_op_alloc_size_may_expand`, with the 2-D tensor hooks still `NULL`),
226+
and the `ggml-cpu` CMake diff adds `iqp.cpp` and gates the SpacemiT IME
227+
kernels — nothing near the `-mcpu=native` probe. `:row` split mode still
228+
throws on CUDA (`ggml-cuda` exports no `ggml_backend_split_buffer_type`).
201229
- **The `:mtp_cancel` bug no longer aborts the VM — it returns an error.** The
202230
race is unchanged and unfixed: cancellation is fire-and-forget, so reusing an
203231
`%MTP{}` session immediately after halting a stream can start decoding on
@@ -207,10 +235,11 @@ The one tag that is not green is `:mtp_cancel`, and it moved: see Changed.
207235
`GGML_ASSERT(buf != NULL && "tensor buffer not set")`, three exit 139), at
208236
b10582 it aborted **0 of 4** — three runs failed with
209237
`{:error, "prompt decode failed: code=-1"}` / `"verify decode failed:
210-
code=-1"` and one passed. The decode paths now refuse the half-released
211-
context instead of writing through it. The test stays on its own tag: a flaky
212-
failure still does not belong in a green run, and the real fix is still an
213-
acknowledged-cancellation protocol, not a bump.
238+
code=-1"` and one passed; at b10830 one run failed the same way,
239+
`"verify decode failed: code=-1"`, without aborting. The decode paths now
240+
refuse the half-released context instead of writing through it. The test
241+
stays on its own tag: a flaky failure still does not belong in a green run,
242+
and the real fix is still an acknowledged-cancellation protocol, not a bump.
214243
- **`--include rpc_live` must run on its own**, documented in
215244
`test/test_helper.exs` after it took the VM down here. The live test calls
216245
`RPC.add_server/1`, which mutates the *process-global* ggml device registry,

‎Makefile‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -36,7 +36,7 @@ endif
3636
# Pinned llama.cpp commit, used when vendor/llama.cpp has to be cloned. MUST
3737
# match the vendor/llama.cpp submodule; bump both together, see
3838
# docs/release-guide.md. Override to build the NIF against another revision.
39-
LLAMA_COMMIT ?= ca3d5a3e10d53f7ea672cb9b6178faca3e2807bc
39+
LLAMA_COMMIT ?= 465e49b9cea78a68b9c244ffb48d0ee24a82873d
4040

4141
# The commit actually on disk. A submodule can be bumped without LLAMA_COMMIT
4242
# following it, and the build has to key off what is really there.

‎docs/release-guide.md‎

Lines changed: 8 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -107,6 +107,14 @@ none of it near the `-mcpu=native` probe. A source diff is enough to say a defec
107107
is *still there*; it is not enough to say it is *gone*, so if a diff ever touches
108108
the probe itself, run the command in the last column.
109109

110+
Re-checked at `465e49b9c` (b10830), again as a source diff: still all three.
111+
`ggml_backend_rpc_start_server` and `ggml_backend_cuda_comm_init` are
112+
untouched — the RPC diff is #26500 (do not serialise buffers that belong to
113+
another server) and #27960 (`ggml_op_alloc_size_may_expand`), and the RPC
114+
buffer's `set_tensor_2d`/`get_tensor_2d` hooks are still `NULL`. The
115+
`ggml-cpu/CMakeLists.txt` diff adds `iqp.cpp` and gates the SpacemiT IME
116+
kernel sources; the `-mcpu=native` probe is untouched.
117+
110118
| # | Upstream defect | Our workaround | Still needed? |
111119
|---|---|---|---|
112120
| 1 | `GGML_NATIVE=ON` makes ggml's `-mcpu=native` probe resolve to **base ARMv8-A** on Cortex-X925/A725 with GCC 13.3 — silently, with a soft CMake warning and exit 0. Costs every `sdot`/`smmla`/SVE kernel. | `LLAMA_CPU_ARM_ARCH` + `LLAMA_CUDA_ARCH` in the `Makefile`, which must be set together. See [DGX Spark](dgx-spark.md) and [Cross-Platform Builds](cross-platform-builds.md). | `scripts/spark/verify-build-flags.sh` on an aarch64 host. If a default build (no `LLAMA_CPU_ARM_ARCH`) now reports non-zero `sdot`/`smmla`, upstream fixed the probe. |

‎lib/llama_cpp_ex/model.ex‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -80,7 +80,7 @@ defmodule LlamaCppEx.Model do
8080
`:tensor_split` and `:main_gpu` would index a list you never saw. With
8181
`:devices` set, they index this one.
8282
83-
> #### Split modes at llama.cpp b10362 {: .warning}
83+
> #### Split modes at llama.cpp b10830 (`465e49b9`) {: .warning}
8484
>
8585
> `:layer` splits contiguous layer ranges across devices, one KV cache per
8686
> device, and is the only mode that works across hosts.

‎vendor/llama.cpp‎

Submodule llama.cpp updated 442 files

0 commit comments

Comments
 (0)