corridorkey-mrp-mlx/docs
modelbeast dfc2bf4a51 docs: M4 mini (7th) completes the matched-width M3->M4->M5 generation ladder
hdim-56 fallback is textbook here (1.00x x3); padding = 2.0x. Fits the
corrected fused-kernel-quality theory as a prediction, not a fit: 0.50x
fused/unfused -> 2.0x padding win, exactly between M3 Air (0.40x->2.6x) and
M1 Max (0.87x->1.3x). At 10c matched width the ladder is 138.9/48.8/44.0ms at
hdim56 but 54.3/24.4/8.9ms at hdim64 — M5's gains live almost entirely in the
fused kernel, making the head-dim fix most valuable on the newest silicon.
2026-07-17 14:30:57 +10:00
..
benchmarks docs: wave2 ablation benchmarks + optimization plan + brainstorm 2026-03-09 19:32:04 -02:30
brainstorms docs: wave2 ablation benchmarks + optimization plan + brainstorm 2026-03-09 19:32:04 -02:30
plans docs: wave2 ablation benchmarks + optimization plan + brainstorm 2026-03-09 19:32:04 -02:30
2026-07-16-m-series-fleet-ablation-results.md docs: M4 mini (7th) completes the matched-width M3->M4->M5 generation ladder 2026-07-17 14:30:57 +10:00
MLX Vision Transformer Optimization Techniques.md docs: wave2 ablation benchmarks + optimization plan + brainstorm 2026-03-09 19:32:04 -02:30