Shape-spy on a real 2048 forward + exact-shape sweeps on M1/M3 Ultra: fused==unfused at hdim 56 (silent fallback) on both machines, so use_sdpa buys nothing and its transpose choreography costs ~4.5x on M1. Padding to hdim 64 engages the fast kernel: 2.2x (M3) / 2.7x (M1) on the dominant global-attention shape despite 14% extra FLOPs. |
||
|---|---|---|
| .. | ||
| benchmarks | ||
| brainstorms | ||
| plans | ||
| 2026-07-16-m-series-fleet-ablation-results.md | ||
| MLX Vision Transformer Optimization Techniques.md | ||