Skip to content

perf(mlx): opt-in Gemma 4 expert-QMM tile kernel with parallel descriptor builder - #4

Merged
Gajesh2007 merged 3 commits into
darkbloom-basefrom
codex/gemma4-autoresearch-v0.8.2
Aug 10, 2026
Merged

perf(mlx): opt-in Gemma 4 expert-QMM tile kernel with parallel descriptor builder#4
Gajesh2007 merged 3 commits into
darkbloom-basefrom
codex/gemma4-autoresearch-v0.8.2

Conversation

@Gajesh2007

Copy link
Copy Markdown
Member

What

Adds an opt-in, exactly-scoped Gemma 4 26B-A4B expert-QMM tile path behind MLX_GATHER_QMM_EXPERT_SLICES:

  • A distinctly named qmm_t_expert_impl (BM32 expert tiles, BM16 fallback rows) taking a private/by-value row count. The shared qmm_t_impl constant-address ABI (const constant int& M) and every ordinary gathered / batched / dense QMM route are byte-for-byte unchanged.
  • build_gemma4_sorted_expert_tiles_bm32 — the descriptor builder runs on one 128-thread threadgroup (parallel binary search of each expert's sorted range + Hillis–Steele exclusive scan + strided upper-bound emission) instead of a single GPU thread serially building all descriptors.
  • Selector runs after the existing NAX-first route (gather_qmm_rhs_nax keeps priority — a NAX engagement is recorded as non-engagement, never bypassed) and requires: affine BF16 transposed activations, contiguous uint32 sorted indices, 4-bit gs=64 weights, exactly 128 experts, assignment count ∈ {4096, 8192, 16384}, and one of the two exact production geometries (packed gate/up [128, 1408, 352]; down [128, 2816, 88]). Every miss returns through the unchanged legacy route.
  • Device-level one-shot request resolution, non-throwing dual-symbol AOT probe + prewarm, and relaxed-atomic per-primitive diagnostics: requested, aotAvailable, naxAvailable, hits, and mutually-exclusive fallback classes (nax, outer_route, quantization, topology, assignment_count, geometry, metallib_unavailable), with attempts = hits + fallbacks maintained exactly.

Before / After — dispatch behavior

flowchart LR
  subgraph Before
    A1[sorted gather QMM<br/>128 experts, gs=64 b=4] --> B1[NAX available?]
    B1 -- yes --> C1[gather_qmm_rhs_nax]
    B1 -- no --> D1[gather_qmm_rhs_*<br/>generic 16-row tiling]
  end
  subgraph After
    A2[sorted gather QMM<br/>128 experts, gs=64 b=4] --> B2[NAX available?]
    B2 -- yes --> C2[gather_qmm_rhs_nax<br/>counted: fallback_nax]
    B2 -- no --> S{exact selector<br/>opt-in + AOT + dtype +<br/>128 experts + 4096/8192/16384<br/>+ exact geometry}
    S -- hit --> R1[build_gemma4_sorted_expert_tiles_bm32<br/>+ affine_gather_qmm_gemma4_expert_tiles_...]
    S -- miss --> D2[gather_qmm_rhs_*<br/>counted per-class fallback]
  end
Loading

Before / After — code

flowchart TD
  subgraph Before
    Q1[quantized.cpp dispatch] --> Q2[qmm_t_impl<br/>shared constant-address M]
    Q2 --> Q3[generic BM16/BM32 tiles]
  end
  subgraph After
    P1[quantized.cpp dispatch] --> P2{selector gate}
    P2 --> P3[quantized.metal<br/>build_gemma4_sorted_expert_tiles_bm32<br/>affine_gather_qmm_gemma4_expert_tiles_bm_32]
    P3 --> P4[gemma4_expert_qmm.h<br/>qmm_t_expert_impl private M]
    P2 -- miss --> P5[qmm_t_impl — unchanged]
    P6[device.cpp/h] --> P7[request resolution, AOT probe/prewarm,<br/>diagnostics snapshot/reset]
    P1 -. counts .-> P6
  end
Loading

Testing

  • tests/gpu_tests.cpp: exact-shape arithmetic parity at all three assignment counts, one-field selector misses, NAX priority, AOT-absence fallback, counter invariant/reset, and zero attempts from ordinary routes.
  • Built deterministically (JIT off) through scripts/fetch-metallib.sh in the downstream tree, which refuses artifacts missing either R1 symbol; the resulting mlx.metallib is byte-identical (3f9d85ec…) to the artifact used in the production retention matrix.

Measured standing (honest read)

In the downstream v0.8.2 production retention matrix (gemma-4-26B-A4B-it-qat-4bit, M4 Max):

  • Standalone profile: dropped (prefill geomean −10.2% vs bracket).
  • Paired with direct weighted-unsort (the retained production tuple): prefill +1.8%, TTFT −7.5%, decode +3.3%, arrival E2E +12.0% vs bracket — retained-final.
  • Production engagement is zero on measured paths: 810 armed attempts, all fallbacks (723 outer-route, 87 assignment-count). It ships strictly as an opt-in experiment with hit/fallback observability; orphan/never-hitting claims are not made.

Stack: upstream of Layr-Labs/mlx-swift (Swift bindings + mirrors) → Layr-Labs/mlx-swift-lm → Layr-Labs/d-inference (v0.8.2).

…scriptor builder

Adds a distinctly-named expert QMM implementation for the Gemma 4
26B-A4B MoE production shapes, gated by MLX_GATHER_QMM_EXPERT_SLICES:

- qmm_t_expert_impl: BM32 expert tile body (BM16 fallback rows) taking a
  private/by-value row count; the shared qmm_t_impl constant-address ABI
  and all ordinary gathered/batched/dense QMM routes are unchanged.
- build_gemma4_sorted_expert_tiles_bm32: one 128-thread threadgroup
  replaces the reference design's single-GPU-thread serial builder;
  parallel expert-range binary search, Hillis-Steele scan, and strided
  upper-bound descriptor emission.
- Selector runs after the NAX-first route and requires affine BF16
  transposed inputs, 4-bit gs=64 weights, 128 experts, assignment counts
  of exactly 4096/8192/16384, and the exact gate/up or down rank-3
  shapes; every miss keeps the legacy route. NAX engagement is
  non-engagement, never bypassed.
- device.{h,cpp}: one-shot request resolution, nonthrowing dual-symbol
  AOT probe/prewarm, relaxed-atomic diagnostics (requested, aotAvailable,
  naxAvailable, hits, per-class fallbacks).
- gpu_tests: exact-shape arithmetic parity, fallback, and counter
  invariant probes.

Retention standing (2026-08-09 production matrix): opt-in experiment.
Standalone profile dropped (prefill -10.2% vs bracket); paired
weighted-unsort+R1 profile retained-final (prefill +1.8%, TTFT -7.5%,
decode +3.3%, arrival E2E +12.0%). NOTE: this source post-dates the
benchmarked binaries/metallib (post-measurement kernel-body edit);
rebuild and re-verify before any performance claim.
@Gajesh2007

Copy link
Copy Markdown
Member Author

@codex review

…unter/atomic hygiene

Review-wave fixes for the R1 expert-QMM path:

- N1 (sortedness trust): build_gemma4_sorted_expert_tiles_bm32 now
  verifies each thread's post-binary-search segment boundary against the
  generalized invariant indices[start - 1] < lid <= indices[start]
  (edge threads check their single neighbor), votes per simdgroup via
  simd_or, folds the votes through threadgroup memory, and on any
  violation retracts count[0] to 0 (tile kernel then early-returns) and
  records the violation in count[1]; the buffer ABI is unchanged
  (count index 1 was previously unused). try_gemma4_expert_qmm allocates
  the second count element, drains the encoder after the builder, and
  re-routes a retracted call to the order-agnostic legacy path instead of
  dispatching the tile kernel (zero count is unambiguous: the selector's
  assignment gate guarantees M is 4096/8192/16384).
- N2 (route-condition duplication): the sorted-RHS gate literal that
  appeared (negated) in the diagnostics record and in the dispatch
  decision is now the shared static constexpr predicate
  takes_sorted_rhs_route, so future tuning of the 16/4 thresholds cannot
  desynchronize counter vs route.
- N3 (per-call bias normalization): gather_qmm_rhs no longer spends
  ensure_row_contiguous on biases before classification reads the raw
  tensor's fields; normalization runs only inside the winning-route
  branch (hit semantics unchanged; the legacy block keeps its own
  normalization point and ordering).
- N4 (armed_ data race): Gemma4ExpertQMMCounters::armed_ is now
  std::atomic<bool> with relaxed loads/stores in armed(), snapshot(),
  snapshot_and_disarm() (read-then-write order preserved) and
  clear_and_arm(); the class remains non-copyable, now enforced.
@Gajesh2007

Copy link
Copy Markdown
Member Author

Deep-review round complete — fixes landed as 8d538a08 (pushed). Summary of the adversarial review (agent-verified with a standalone Metal harness, 14/14 GPU experiments bitwise-clean) and what it produced:

Review finding Resolution
F1 in-tree tests never execute the kernels on GPU Documented: ON-path runs (MLX_GATHER_QMM_EXPERT_SLICES=1 swift test --filter SortedGatherQuantizedMMTests) execute real HITs — verified locally on this machine (hits == 1 asserted against the fresh source-matched metallib). Header doc now carries the exact invocation.
F2 descriptor builder trusts global sortedness (legacy was order-agnostic) Kernel now self-checks: per-thread range verification + simd/threadgroup vote; on violation count[0]=0 and the host synchronizes, reads the count, and re-routes to the legacy path — mis-sorted input degrades to correct-but-slow instead of silently wrong.
F3 armed_ plain bool race std::atomic<bool>, relaxed semantics, explicit order preservation.
F4 bias re-normalization on fallback calls normalization now happens only inside route == hit.
F5 route-predicate literal duplicated single static constexpr takes_sorted_rhs_route(...) feeds dispatch and counters.
F6/F7 nits recorded.

Verification: gpu_tests 261/261 (3531 assertions); xcrun metal -Wall -Wextra -fno-fast-math 0 warnings; symbol gates all present.

Disclosure: the fail-safe re-route requires a CommandEncoder::synchronize() (stream drain) on every requested-hit call before the host chooses the tile vs legacy route. R1 hits are production-inert today (exact-shape selector), but this cost is real for future engaged runs and will be re-measured in the clean perf gate before any engagement flip.

@Gajesh2007

Copy link
Copy Markdown
Member Author

@codex review

…ibution

F1: the per-expert boundary vote was a partial detector -- an inversion
inside a segment used by no other expert's boundary could escape, so
"re-route on any violation" overclaimed. build_gemma4_sorted_expert_tiles_bm32
now also runs a strided adjacent-pair scan: thread lid checks
indices[i-1] <= indices[i] for i = lid+1; i < M; i += 128, covering every
adjacent pair in [1, M) exactly once (1..128 iterations at the reachable
M in {4096,8192,16384}). Adjacent-pair monotonicity is transitive, so a
clean scan is a sound and complete sortedness oracle; it folds into the
same simd_or/threadgroup vote and the same retract (count[0]=0, count[1]=1).
The boundary checks stay as cheap, precise diagnostics.

F2: retracts were write-only in count[1] and surfaced as
fallback_metallib_unavailable -- misattribution in the only observable
surface. A dedicated fallback_sortedness_retracted counter now rides the
GemmA4 route counters and the C diagnostics ABI
(sizeof 80 -> 88, new uint64 at offset 80; existing offsets unchanged).
try_gemma4_expert_qmm returns the route class: count[0]==0 with count[1]==1
records fallback_sortedness_retracted, any other unusable build keeps
fallback_metallib_unavailable, then re-routes to the legacy path as before.

F4: new doctest drives the full armed() -> clear_and_arm() ->
snapshot_and_disarm() cycle and the attempts == hits + fallbacks invariant
including the new class; the route-table and counter-invariant tests now
cover fallback_sortedness_retracted.

Verified: cmake tests 262/262 + 3550 assertions pass; metal -Wall -Wextra
-fno-fast-math compile of kernels/quantized.metal is warning-free.
@Gajesh2007

Copy link
Copy Markdown
Member Author

Final review wave (#2) complete. Every prior finding is closed; the final verifier pass signed off each repo:

  • mlx 9b0d1b4c: the sortedness fail-safe is now a sound oracle (strided adjacent-pair scan ⟺ globally non-decreasing — intra-segment inversions can't escape); retracts get a dedicated fallback_sortedness_retracted counter (diagnostics ABI 80→88 appended-only, facades byte-synced, pins updated); new doctests pin the arm/disarm cycle. 262/262 cases, 3550 assertions.
  • mlx-swift 38eaa93: mirrors byte-identical to canonical; generator post-condition guarantees one MLX_API guard or fails; facade pivoted to the 88-byte ABI; OFF (6 tests, 2 conditional skips) and ON (6/6, real R1 hits on this machine) both green.
  • mlx-swift-lm 06046c3: the LAST textual survivor of the k_eq_v-gated head rule is gone (validateAttentionProducts) — rule eradication is now complete across model, CBv2 layer kinds, MTP validation, and sizing. 599+471 green twice.
  • d-inference b0fa08af6: UpdateCommand now falls back ONLY for missing/unreadable config (malformed files fail loudly — final swallow hole); preload-gate margin widened for loaded machines (make provider-test: 1692/181, 0 failures on this box).

Perf-gate status unchanged: attribution rerun deferred to a cooled machine (see the prior comment — the box measured ~1.7× slow during the late-night window). The remaining open perf item is the nested PR's per-hit synchronize() cost on engaged R1 shapes (production-inert today); it gets a dedicated measurement in that gate.

@codex review

@Gajesh2007

Copy link
Copy Markdown
Member Author

Final merge-readiness pass complete:

  • Upstream conflicts resolved by merging current base branches; no history rewrite.
  • Heads: mlx 9b0d1b4c, mlx-swift 38eaa93a, mlx-swift-lm 29ff95bc, d-inference 6aa321a71.
  • Every prior review finding is fixed and every stale thread is resolved.
  • LM preserves upstream paged/CBv2 work and adds focused regressions for all shared-tower compatibility edges. Weighted expert reduction is now limited to scheduled CBv2 prefill, eliminating the direct-prefill regression found by the final ablation.
  • Root benchmark artifacts now record typed effective Gemma settings across all phases, reject malformed baselines/sample-count drift, load explicit configs read-only, pin KV backend posture, and force fixed decode token budgets.
  • Verification: final downstream make provider-test = 2,017 tests / 207 suites plus 82 XCTest; benchmark contracts 59/59; coordinator tests/build green; UI 499 tests + lint (0 errors) + production build green; pre-push hook green.
  • Final same-binary A/B correctness gates: exact output invariance PASS, explicit contiguous backend PASS in every phase, delivered arrival topology PASS. Late timing attribution is explicitly caveated because the host drifted during the bracket; the earlier clean attribution epoch remains the performance record.

@codex review

@Gajesh2007
Gajesh2007 merged commit a4b2b4c into darkbloom-base Aug 10, 2026
davidtai added a commit that referenced this pull request Aug 25, 2026
* Return tuple in meshgrid (ml-explore#4229)

* Add endpoint parameter to linspace (ml-explore#4184)

Co-authored-by: Cheng <git@zcbenz.com>

* Fix vmap of partition/argpartition dropping the kth argument (ml-explore#4116)

* Fix nan_to_num replacing inf with 0 for float16 and bfloat16 (ml-explore#4222)

Co-authored-by: codeAnqiang-ma <273298913+codeAnqiang-ma@users.noreply.github.com>
Co-authored-by: Cheng <git@zcbenz.com>

* Fix einsum not broadcasting batch dimensions in batched tensordot (ml-explore#4125)

Co-authored-by: Cheng <git@zcbenz.com>

* Dequantize in float32 (ml-explore#4241)

* chore: Reject complex in erf and erfinv (ml-explore#4243)

* Fix cpu compilation failure of abs with uint (ml-explore#4240)

Co-authored-by: Cheng <git@zcbenz.com>

* Fix quantize matrix multiplication floor issue (ml-explore#4251)

* Only use MPI backend for world size > 1 (ml-explore#4210)

* chore: Reject complex in expm1, sigmoid and arctan2 (ml-explore#4257)

* Decompose small kernel-depth 3D convs into 2D convs (ml-explore#3785)

Co-authored-by: katlun-lgtm <264247399+katlun-lgtm@users.noreply.github.com>
Co-authored-by: Cheng <git@zcbenz.com>

* Fix Metal sort of a view with a negative stride (ml-explore#4252)

* Mirror the depth axis in the decomposed 3D conv when flipped (ml-explore#4277)

* Fix Metal row reductions on negative-stride views (ml-explore#4267)

Co-authored-by: Fu Xiaonan <214359569+FU-max-boop@users.noreply.github.com>

* [CUDA] Fix custom kernel cache collision for same name, different source (ml-explore#4273)

Co-authored-by: Cheng <git@zcbenz.com>

* Fix ops rejecting integers larger than INT32_MAX (ml-explore#4255)

Co-authored-by: Feli <feli@hnu.edu.cn>
Co-authored-by: Cheng <git@zcbenz.com>

* Fix var/std for complex numbers (ml-explore#4260)

* Fix int32 overflow in conv padded input and pad shapes (ml-explore#4258)

Co-authored-by: Cheng <git@zcbenz.com>

* chore: Reject complex in remainder (ml-explore#4270)

* chore: Compare the macOS SDK version as a version when gating JACCL (ml-explore#4286)

* Clamp ring socket transfers so a payload of 2 GiB or more can be sent (ml-explore#4281)

Co-authored-by: Cheng <git@zcbenz.com>

* chore: Use normalize_axis_index in split/unstack/partition/topk (ml-explore#4288)

* Remove grouped output in CI (ml-explore#4195)

* [CUDA] Fix finding cuda 13 headers in JIT compilation (ml-explore#3995)

* Refactor wheel building script (ml-explore#3818)

* Make mx.compile cache erasing thread safe (ml-explore#4248)

Co-authored-by: yentur <mr.yentur@gmail.com>

* Add builds for free-threaded python (ml-explore#3812)

* Fix int32 overflow in concatenate/repeat/kron (ml-explore#4303)

* python: Widen list elements that do not fit in int32 to int64 (ml-explore#4305)

* Propagate CPU errors to events (ml-explore#3742)

Co-authored-by: Alessio Pollero <alessio.pollero@gmail.com>

* Fix mx.arange dtype inference overflow regression (ml-explore#4324)

* Add workflow to update pull request limit bypass list (ml-explore#4320)

* Support head dimension 72 in Metal full attention (ml-explore#4330)

* Patch bump to 0.32.2 (ml-explore#4333)

* Preserve subnormal float values when casting to bool (ml-explore#4224)

* python: Support assigning through a bare Ellipsis index (ml-explore#4314)

* Fix divmod truncating the quotient for floats (ml-explore#4108)

Co-authored-by: Cheng <git@zcbenz.com>

* Add force_fused option to scaled_dot_product_attention (ml-explore#4185)

* chore: Reject negative eps in the normalization layers (ml-explore#4312)

* Bound GGUF metadata string/array values against the file mapping (ml-explore#4212)

Co-authored-by: x14ngch3n <x14ngch3n@users.noreply.github.com>
Co-authored-by: Cheng <git@zcbenz.com>

* Read each K/V byte once in gqa-8 decode attention (ml-explore#4077)

* Fix fft vmap and jvp for transforms over a subset of axes (ml-explore#4138)

* Fix median dropping NaN (ml-explore#4146)

* Fix the CPU scan over a size one axis with a padded stride (ml-explore#4139)

Co-authored-by: Cheng <git@zcbenz.com>

* chore: Validate the optimizer betas at construction (ml-explore#4310)

Co-authored-by: Cheng <git@zcbenz.com>

* `RMSNormVJP` backward writes a full `{n_rows, D}` `gw_temp` intermediate (ml-explore#4293)

* [Bug]: add default none value to axis parameter of the take_along_axis (ml-explore#4357)

Co-authored-by: Anastasiia Filippova <a_filippova@apple.com>

* Add a fused full-attention path for head_dim 256 on NAX devices (ml-explore#3842)

Co-authored-by: Cheng <git@zcbenz.com>

* Update nanobind to 2.15.0 (ml-explore#4337)

* Skip unnecessary simdgroup computations for quantised MOE matmuls on NAX (ml-explore#4352)

* Add AI usage policy (ml-explore#4331)

Co-authored-by: Jake Bowhay <60778417+j-bowhay@users.noreply.github.com>

* Raise cpu stream errors from synchronize (ml-explore#4338)

Co-authored-by: Cheng <git@zcbenz.com>

* chore: Validate eps in Adam at construction (ml-explore#4361)

Co-authored-by: Anastasiia Filippova <a_filippova@apple.com>

* Bound winograd conv2d working set by tiling the batch (ml-explore#4102)

Co-authored-by: Cheng <git@zcbenz.com>

* Use a 32-row block in qmm_t_nax when one block covers all of M (ml-explore#4171)

* chore: Deduplicate fftshift and ifftshift (ml-explore#4318)

* Fix Log and Equal is_equivalent ignoring primitive state (ml-explore#4266)

Co-authored-by: Cheng <git@zcbenz.com>

* Stabilize reduced-precision InstanceNorm (ml-explore#4230)

* chore: Normalize negative axes in sort and argsort (ml-explore#4332)

* Clean up main thread compile cache before python interpreter shuts down (ml-explore#4373)

* chore: Check malformed jaccl hostfile that miss rdma in pairs (ml-explore#4284)

Co-authored-by: Cheng <git@zcbenz.com>

* Round mxfp8 block scales up to avoid saturation (ml-explore#4353)

Co-authored-by: Daniel Hiltgen <daniel.hiltgen@ollama.com>
Co-authored-by: Cheng <git@zcbenz.com>

* Add support for the __array_namespace_info__  (ml-explore#4334)

* Stop a failed CUDA graph commit from poisoning the encoder (ml-explore#4356)

Co-authored-by: Cheng <git@zcbenz.com>

* Fix quantized kernels in JIT build (ml-explore#4372)

Co-authored-by: Cheng <git@zcbenz.com>

* Avoid zero work in stride-2 ConvTranspose3d (ml-explore#4343)

* [CUDA] Ce fused kernel (ml-explore#3947)

* Fix cpu exclusive scan for complex numbers (ml-explore#4272)

Co-authored-by: Cheng <git@zcbenz.com>

* Support Relocatable CUDA DLLs on Windows (ml-explore#4382)

* Use cast_to for fused AsType in compiled Metal kernels (ml-explore#4351)

Co-authored-by: katlun-lgtm <katlun@windyviews.com>
Co-authored-by: Cheng <zcbenz@gmail.com>

* python: Declare DLPackCompatible protocol members as methods (ml-explore#4384)

* Fix quantizing sliced arrays (ml-explore#4381)

* Fix einsum dropping a trailing empty subscript (ml-explore#4299)

Co-authored-by: Cheng <git@zcbenz.com>

* Add script to run python tests (ml-explore#4393)

* Hold GIL in AttachedData destructor (ml-explore#4391)

* Bound Metal buffer COUNT, not just bytes, in MetalAllocator

The Metal allocator throws `[metal::malloc] Resource limit (N) exceeded`
when num_resources_ (the live+cached Metal buffer COUNT) reaches
resource_limit_ (the iogpu.rsrc_limit sysctl, default ~499000). Freed
buffers are recycled into a size-keyed cache whose only trim is by BYTES
(release_cached_buffers takes a bytes-to-free target, max_pool_size_ ~=
physical RAM). Under churn with many distinct buffer shapes (varied prompt
lengths, growing KV caches, multiple co-resident models) the cache fills
with entries never reused at that exact size, so the COUNT climbs to the
limit while byte usage stays modest and the byte trim never fires — the
process crashes mid-inference on a machine with most of its RAM free.

malloc() now also reclaims by count: when num_resources_ crosses a 90%
high-water mark of resource_limit_, it clears the (pure-reuse) buffer
cache so the count drops back to the live working set. Clearing the cache
only costs re-allocation, never correctness, so the count limit becomes
unreachable by any request mix or batching method while the existing byte
limits keep total memory bounded.

Adds get_num_resources()/get_resource_limit() to the public memory API
(metal + no_gpu + cuda backends) so the count and its ceiling are
observable from callers. Adds an MLX_RESOURCE_LIMIT env override that can
only LOWER the ceiling (clamped to the OS limit, strictly validated) to
exercise the trim deterministically and as an operator safety valve.

* perf(mlx): opt-in Gemma 4 expert-QMM tile kernel with parallel descriptor builder (#4)

* perf(mlx): add opt-in Gemma 4 expert-QMM tile kernel with parallel descriptor builder

Adds a distinctly-named expert QMM implementation for the Gemma 4
26B-A4B MoE production shapes, gated by MLX_GATHER_QMM_EXPERT_SLICES:

- qmm_t_expert_impl: BM32 expert tile body (BM16 fallback rows) taking a
  private/by-value row count; the shared qmm_t_impl constant-address ABI
  and all ordinary gathered/batched/dense QMM routes are unchanged.
- build_gemma4_sorted_expert_tiles_bm32: one 128-thread threadgroup
  replaces the reference design's single-GPU-thread serial builder;
  parallel expert-range binary search, Hillis-Steele scan, and strided
  upper-bound descriptor emission.
- Selector runs after the NAX-first route and requires affine BF16
  transposed inputs, 4-bit gs=64 weights, 128 experts, assignment counts
  of exactly 4096/8192/16384, and the exact gate/up or down rank-3
  shapes; every miss keeps the legacy route. NAX engagement is
  non-engagement, never bypassed.
- device.{h,cpp}: one-shot request resolution, nonthrowing dual-symbol
  AOT probe/prewarm, relaxed-atomic diagnostics (requested, aotAvailable,
  naxAvailable, hits, per-class fallbacks).
- gpu_tests: exact-shape arithmetic parity, fallback, and counter
  invariant probes.

Retention standing (2026-08-09 production matrix): opt-in experiment.
Standalone profile dropped (prefill -10.2% vs bracket); paired
weighted-unsort+R1 profile retained-final (prefill +1.8%, TTFT -7.5%,
decode +3.3%, arrival E2E +12.0%). NOTE: this source post-dates the
benchmarked binaries/metallib (post-measurement kernel-body edit);
rebuild and re-verify before any performance claim.

* fix(mlx): fail-safe sortedness check in gemma expert tile builder; counter/atomic hygiene

Review-wave fixes for the R1 expert-QMM path:

- N1 (sortedness trust): build_gemma4_sorted_expert_tiles_bm32 now
  verifies each thread's post-binary-search segment boundary against the
  generalized invariant indices[start - 1] < lid <= indices[start]
  (edge threads check their single neighbor), votes per simdgroup via
  simd_or, folds the votes through threadgroup memory, and on any
  violation retracts count[0] to 0 (tile kernel then early-returns) and
  records the violation in count[1]; the buffer ABI is unchanged
  (count index 1 was previously unused). try_gemma4_expert_qmm allocates
  the second count element, drains the encoder after the builder, and
  re-routes a retracted call to the order-agnostic legacy path instead of
  dispatching the tile kernel (zero count is unambiguous: the selector's
  assignment gate guarantees M is 4096/8192/16384).
- N2 (route-condition duplication): the sorted-RHS gate literal that
  appeared (negated) in the diagnostics record and in the dispatch
  decision is now the shared static constexpr predicate
  takes_sorted_rhs_route, so future tuning of the 16/4 thresholds cannot
  desynchronize counter vs route.
- N3 (per-call bias normalization): gather_qmm_rhs no longer spends
  ensure_row_contiguous on biases before classification reads the raw
  tensor's fields; normalization runs only inside the winning-route
  branch (hit semantics unchanged; the legacy block keeps its own
  normalization point and ordering).
- N4 (armed_ data race): Gemma4ExpertQMMCounters::armed_ is now
  std::atomic<bool> with relaxed loads/stores in armed(), snapshot(),
  snapshot_and_disarm() (read-then-write order preserved) and
  clear_and_arm(); the class remains non-copyable, now enforced.

* fix(mlx): make the R1 sortedness fail-safe sound; proper retract attribution

F1: the per-expert boundary vote was a partial detector -- an inversion
inside a segment used by no other expert's boundary could escape, so
"re-route on any violation" overclaimed. build_gemma4_sorted_expert_tiles_bm32
now also runs a strided adjacent-pair scan: thread lid checks
indices[i-1] <= indices[i] for i = lid+1; i < M; i += 128, covering every
adjacent pair in [1, M) exactly once (1..128 iterations at the reachable
M in {4096,8192,16384}). Adjacent-pair monotonicity is transitive, so a
clean scan is a sound and complete sortedness oracle; it folds into the
same simd_or/threadgroup vote and the same retract (count[0]=0, count[1]=1).
The boundary checks stay as cheap, precise diagnostics.

F2: retracts were write-only in count[1] and surfaced as
fallback_metallib_unavailable -- misattribution in the only observable
surface. A dedicated fallback_sortedness_retracted counter now rides the
GemmA4 route counters and the C diagnostics ABI
(sizeof 80 -> 88, new uint64 at offset 80; existing offsets unchanged).
try_gemma4_expert_qmm returns the route class: count[0]==0 with count[1]==1
records fallback_sortedness_retracted, any other unusable build keeps
fallback_metallib_unavailable, then re-routes to the legacy path as before.

F4: new doctest drives the full armed() -> clear_and_arm() ->
snapshot_and_disarm() cycle and the attempts == hits + fallbacks invariant
including the new class; the route-table and counter-invariant tests now
cover fallback_sortedness_retracted.

Verified: cmake tests 262/262 + 3550 assertions pass; metal -Wall -Wextra
-fno-fast-math compile of kernels/quantized.metal is warning-free.

* perf(metal): E=256 expert-tile route + trust + gpu::eval UAF fix — darkbloom-base mirror (#7)

* perf(metal): instantiate E=256 expert-tile route for Qwen 3.5/3.6 MoE prefill (mirror of Cmlx/mlx 58fab46)

* fix(metal): use-after-free in gpu::eval for primitives that synchronize mid-eval (mirror)

* perf(metal): trust mode skips retract readback (mirror)

* fix(compile): preserve all-cache binding cleanup

---------

Co-authored-by: JasonHonKL <148705846+JasonHonKL@users.noreply.github.com>
Co-authored-by: AK <144495202+AKnassa@users.noreply.github.com>
Co-authored-by: Cheng <git@zcbenz.com>
Co-authored-by: Adityaj0 <93090622+Adityaj0@users.noreply.github.com>
Co-authored-by: anchor <codeanqiang@gmail.com>
Co-authored-by: codeAnqiang-ma <273298913+codeAnqiang-ma@users.noreply.github.com>
Co-authored-by: Rohan Gautam <rohan1gautam@gmail.com>
Co-authored-by: Ayaan Gazali <ayaangazali.work@gmail.com>
Co-authored-by: Erwin Zhang <59893706+erwinzhang7@users.noreply.github.com>
Co-authored-by: katlun-lgtm <katlun@gmail.com>
Co-authored-by: katlun-lgtm <264247399+katlun-lgtm@users.noreply.github.com>
Co-authored-by: robertomeroni <150194833+robertomeroni@users.noreply.github.com>
Co-authored-by: Fu Xiaonan <ht3fudatou@163.com>
Co-authored-by: Fu Xiaonan <214359569+FU-max-boop@users.noreply.github.com>
Co-authored-by: Hao Xu <hxu44@apple.com>
Co-authored-by: Feli <89400571+FeliGame@users.noreply.github.com>
Co-authored-by: Feli <feli@hnu.edu.cn>
Co-authored-by: Eyüp Can Akman <eyupcanakman@gmail.com>
Co-authored-by: Cheng <zcbenz@gmail.com>
Co-authored-by: yentur <mr.yentur@gmail.com>
Co-authored-by: Alessio Pollero <alessio.pollero@gmail.com>
Co-authored-by: Zhiqi Zhang <zhiqizhangg@gmail.com>
Co-authored-by: Daniel Hiltgen <dhiltgen@users.noreply.github.com>
Co-authored-by: Tanish Jain <recklurker@gmail.com>
Co-authored-by: hojin12312 <hojin12312@gmail.com>
Co-authored-by: Xiang Chen <46052474+x14ngch3n@users.noreply.github.com>
Co-authored-by: x14ngch3n <x14ngch3n@users.noreply.github.com>
Co-authored-by: Duhyeon, Kim <49020301+dudududukim@users.noreply.github.com>
Co-authored-by: rohith <kapellirohith@gmail.com>
Co-authored-by: Ishaan Samantray <devteam.aegis@gmail.com>
Co-authored-by: Aaishwarya Mishra <aaishwarymishra@gmail.com>
Co-authored-by: Anastasiia Filippova <a_filippova@apple.com>
Co-authored-by: Yanzhao Wang <19340816+wyanzhao@users.noreply.github.com>
Co-authored-by: XXXXRT666 <157766680+XXXXRT666@users.noreply.github.com>
Co-authored-by: Jake Bowhay <60778417+j-bowhay@users.noreply.github.com>
Co-authored-by: vraj patel <87225460+vraj00222@users.noreply.github.com>
Co-authored-by: Gusanidas <33495733+Gusanidas@users.noreply.github.com>
Co-authored-by: Dwijen Patel <dwijen@gmail.com>
Co-authored-by: Vladimir Iglovikov <ternaus@users.noreply.github.com>
Co-authored-by: Brian C. <94733710+deBrian07@users.noreply.github.com>
Co-authored-by: Daniel Hiltgen <daniel.hiltgen@ollama.com>
Co-authored-by: YH Yan <strayberry0w0@gmail.com>
Co-authored-by: katlun-lgtm <katlun@windyviews.com>
Co-authored-by: anupsv <6407789+anupsv@users.noreply.github.com>
Co-authored-by: Gajesh Naik <26431906+Gajesh2007@users.noreply.github.com>
Co-authored-by: David Tai <davidtai@Davids-MBP.lan>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant