Skip to content

docs: publish M5 Max speed and compression research - #2

Open
PhilipJohnBasile wants to merge 1 commit into
codex/moe-dflash-integration-20260730from
codex/mlx-m5-research-20260810
Open

docs: publish M5 Max speed and compression research#2
PhilipJohnBasile wants to merge 1 commit into
codex/moe-dflash-integration-20260730from
codex/mlx-m5-research-20260810

Conversation

@PhilipJohnBasile

Copy link
Copy Markdown
Owner

Summary

  • publishes source- and receipt-backed research on M5 Max speed, compression, exactness, and dispatch behavior
  • adds CPU-verifiable tools for dispatch, provenance, compressed-artifact comparison, and exactness auditing
  • records negative results and non-comparable evidence explicitly; runtime defaults are unchanged
  • keeps the research bundle repository-only and excludes it from sdist/wheel payloads
  • provides a review map: 10 benchmark sources, 7 focused tests, 3 documentation files, 2 provenance files, 84 JSON artifacts including the manifest, and the package-boundary change

Verification

  • 63 focused research tests passed
  • 261 required public-contract tests passed (test_no_mlx_imports.py, test_public_cli.py, test_runtime_kpis.py)
  • Ruff, format checks, Python compilation checks, and git diff --check passed
  • deterministic manifest validated 23 sources, 83 receipts, 10 content-addressed references, and exact coverage of all 107 staged paths
  • exact-index clean sdist and wheel builds passed scripts/fresh_venv_smoke.sh
  • zero research paths were present in the built sdist or wheel
  • immutable historical timing producer retained SHA-256 1578f6f8cb860be2e6dcb27ea08f1253b069555f0e83da3c71a451a0bcf98d33
  • an independent frozen-diff review found no actionable blocker before publication

Evidence limits

  • hardware: Apple M5 Max
  • research dates: 2026-08-10 through 2026-08-11
  • this is a build-selection research bundle, not a runtime-default change, model release, or production speed claim
  • Qwen3.6-35B-A3B compressed-artifact timings use fixed token IDs and forward counts; expert assignments were not observed and may differ
  • historical HumanEval NLL receipts are preserved but marked NOT COMPARABLE and have no decision authority
  • measured, verified, published, and hypothesis claims are labeled separately in the main report
  • the repository research bundle is intentionally excluded from Python distributions

Review entry points

  • docs/research/mlx-m5max-speed-without-quality-loss-20260810.md
  • benchmarks/results/mlx-m5-research-20260810/summary.json
  • benchmarks/results/mlx-m5-research-20260810/bundle_manifest.json
  • docs/research/mlx-lm-1709-exactness-audit-20260810.md

@PhilipJohnBasile
PhilipJohnBasile marked this pull request as ready for review August 11, 2026 12:08
@PhilipJohnBasile

Copy link
Copy Markdown
Owner Author

Public CI follow-up: the fork did not auto-create pull_request runs, so I manually dispatched the existing release workflow against the exact PR head a204c6427292bf8b282b9a342f296f88824cd39e with publish_to_pypi=false. The macOS build-artifacts job passed; the PyPI publication job was skipped as intended: https://github.com/PhilipJohnBasile/MTPLX/actions/runs/31489988994

@PhilipJohnBasile

Copy link
Copy Markdown
Owner Author

@youssofal — GitHub would not retain a formal reviewer request on this fork, so I am requesting review here. The branch is evidence-only and changes no runtime defaults. The fastest review path is the main research note plus the content-addressed bundle manifest; the public macOS build/smoke run is linked above.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant