benchmarks: trellis_mac sparse-conv ablation - Metal flex_gemm = 1.72x end-to-end vs pure-torch fallback
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
parent
793f8a88bb
commit
fcb5e80702
@ -177,3 +177,16 @@ the fleet doc's `m1max@` user and "disk-critical" note are both stale — 170GB
|
||||
Notable: 2nd-best bandwidth in the fleet (336 GB/s > M4 Pro's 225) despite mid-pack
|
||||
compute — good fit for memory-bound MLX work (VAE decode, big-tensor elementwise).
|
||||
Fleet total ≈ 66 TFLOPS fp16 across 7 machines.
|
||||
|
||||
## 2026-07-19 — trellis_mac rung 1: sparse-conv backend verified (m3ultra, anatomy image, seed 42, 1024-cascade)
|
||||
|
||||
| SPARSE_CONV_BACKEND | total gen+bake | wall | peak RSS |
|
||||
|---|---|---|---|
|
||||
| flex_gemm (Metal) | **124.8 s** | 248.8 s | 20.8 GB |
|
||||
| none (pure-torch) | 214.4 s | 337.6 s | 20.8 GB |
|
||||
|
||||
Verdicts:
|
||||
- The 124.8s baseline WAS the Metal fast path (log line "[SPARSE] Conv backend: flex_gemm") — no hidden 10x jackpot; metallib loads fine on macOS 26.5.
|
||||
- Metal sparse-conv is worth **1.72x end-to-end** (saves 89.6s/gen). Keep flex_gemm the default; the conv_none fallback is a real 90s regression if the metallib ever breaks silently — worth a GODBOARD/log assertion.
|
||||
- Remaining fat in the fast path: ~46s diffusion sampling (padded SDPA attention — next kernel frontier), ~50s decode/extract (incl. aten::segment_reduce falling back to CPU — patchable), 19s Metal bake.
|
||||
- Decode actually produces 2.79M faces before the forced ~200k simplify — the quality-recovery headroom (rung 2) is bigger than the recon estimated.
|
||||
|
||||
Loading…
Reference in New Issue
Block a user