Padding 56->64 helps everywhere (1.3x-4.9x) and the hdim-56 silent fallback reproduces on all six machines. But the M1 Max (24c) gains LEAST of all, breaking our 'narrower GPU => bigger win' reading (M5 10c gained most). Real predictor: fused-kernel quality vs that chip's own matmul throughput — M1 Max's fused path is only 0.87x unfused at every size, so missing it costs little; M5's is 0.14x, so missing it costs a lot. Corrected in place rather than left as a tidy-but-wrong story. |
||
|---|---|---|
| .. | ||
| benchmarks | ||
| brainstorms | ||
| plans | ||
| 2026-07-16-m-series-fleet-ablation-results.md | ||
| MLX Vision Transformer Optimization Techniques.md | ||