V3.2 DSA indexer (MODEL_DEEPSEEK_V32)
Port mlx-lm deepseek_v32.py Indexer - a real attention/cache fork, not a boolean on collapsed V4.
Spec: doc/plans/engine-inference-core.md § E2
Part of: #105 (E2) · Epic: #56
Depends on: MLA attention + MLA KV cache; family split (MODEL_DEEPSEEK_V32)
Tasks
Notes
- Decision 11: this fork is why V3.2 must not share
MODEL_DEEPSEEK_V4 / V3 tags
V3.2 DSA indexer (
MODEL_DEEPSEEK_V32)Port mlx-lm
deepseek_v32.pyIndexer- a real attention/cache fork, not a boolean on collapsed V4.Spec:
doc/plans/engine-inference-core.md§ E2Part of: #105 (E2) · Epic: #56
Depends on: MLA attention + MLA KV cache; family split (
MODEL_DEEPSEEK_V32)Tasks
index_n_heads/index_head_dim/index_topk)take_along_axis) on decodeCacheListwq_b,wk,weights_proj, ...)index_head_dim,index_n_heads,index_topk(from family-split/config issue)Notes
MODEL_DEEPSEEK_V4/ V3 tags