- Benchmarks — measured M2 performance and tiled-versus-naive matmul results.
- Correctness — CPU-to-Metal parity and end-to-end GPT-2 validation.
- Contributing — local checks to run before proposing a change.
- Test command reference — the smallest useful validation command for each task.
- Troubleshooting — platform, model-asset, and test setup notes.