modelbeast/BENCHMARKS.md
m3ultra 27564e80eb perfcheck: nightly fleet GPU canary + the contention lessons that shaped it
BENCHMARKS.md says what a model cost the day it was measured; nothing noticed if a
macOS/torch update or a thermal fault halved a node. perfcheck runs the whole fleet
in ~35s off godcheck's 03:30 cron on the m4mini and reports drift into
GODCHECK_LATEST.md.

Probes matmul (fp16/fp32), memory bandwidth, and SDPA at head_dim 64 vs 56 — the
latter turning CorridorKey's fast-path cliff into a permanent canary: it confirms
the padding win fleet-wide (1.96x-5.26x) and tells us if a future torch closes it.
Runs on venvs/rmbg/bin/python, already identical fleet-wide, so nothing new is
installed (nothing lands on the disk-tight m1max).

First cross-machine capability table for all 6 nodes. The M3 Ultra is ~2.2x the M1
Ultra on fp16 matmul, but they share ~625 GB/s — so bandwidth-bound stages run alike
while compute-bound ones scale. Both Ultras reach only ~78% of spec bandwidth on a
single kernel; the smaller Macs hit ~88%.

Measuring a fleet that is doing real work is the whole problem, and naive
benchmarking here is off by 11x:

- min, not median: a concurrent trellis_mac job dragged a median-of-5 matmul from
  ~24500 to ~2150 GFLOP/s, which reads exactly like a catastrophic regression.
- n=4096 not 2048: 2048 is dispatch-bound and swung 48% run-to-run; 4096 reproduces
  to 0.1% even while contended.
- sdpa 16x2048 not 8x1024: sub-ms probes are dispatch noise — 8x1024 gave ratios of
  0.79/3.95/2.35 on three runs of one machine, the first "proving" 56 is faster.
- busy nodes are excluded, not blamed: GPU contention is invisible to load average
  (M3 Ultra read load 2.45 with its GPU pinned), so bench.py samples ioreg GPU% before
  it touches the GPU — our own matmul pins the device, so ordering is the trick.
- baselines are the median of recent history, not a saved best: the M4 Pro also serves
  Ollama and is bimodal (~3200 vs ~5500 fp32), so a best-observed baseline pins to a
  lucky outlier and alerts forever.
- a regression must repeat before it is believed ([~] watching -> [!] CONFIRMED).

Validated by re-running the fleet against its own baselines: zero false alarms,
including a sweep where the M3 Ultra read 43 GB/s under load and was correctly
marked BUSY rather than reported as a 93% regression.

Also found: the m4mini is the only node with Tailscale SSH (RunSSH: true) and it does
NOT propagate remote exit codes — `ssh m4mini "exit 7"` returns 0, so any
`if ssh m4mini ...` test silently always passes. Test on output instead, which is
what godcheck already does (and why it is unaffected). It also cannot ssh to itself,
so run_fleet detects its own tailnet IP and benches the local node via the shell.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 18:03:03 +10:00

150 lines
16 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# MODELBEAST Benchmarks (M3 Ultra, 256GB)
First measurements on this machine, recorded 2026-07-12. Fixtures are synthetic
(Blender-rendered Suzanne), so quality numbers are not representative of real
photography — these validate that the pipeline *runs* and how fast.
## Scan track (fully validated, no gated weights)
| Stage | Input | Settings | Result |
|---|---|---|---|
| `colmap_poses` | 48 frames @ 800×600 | sequential matcher, global (GLOMAP) mapper, OPENCV, CPU SIFT | **48/48 images registered, 1479 points, 0.60px mean reprojection error**; global mapper step ~1.0s; full job a few seconds |
| `brush_train` | above colmap_dataset | 1500 steps, max_res 800, sh 2 | ~3060s wall, 465KB splat.ply, renders in the in-app SplatViewer |
A full-quality `brush_train` run is 30000 steps (the default) — expect minutes,
and a much crisper splat than the 1500-step preview above.
## Image generation — local FLUX lineup (M3 Ultra, mflux/MLX, 1024×1024, validated 2026-07-13)
All five installed and generating. **Warm** = weights cached (the real per-image cost):
| Model | Steps | Warm gen | Peak MLX mem | Gated? | Best for |
|---|---|---|---|---|---|
| **FLUX.2 Klein 4B** | 4 | **9.1s** | 18.0 GB | no (Apache) | volume / sprites — the default |
| FLUX.1 schnell-4bit | 4 | 18.5s | 19.1 GB | no (community quant) | fast draft, FLUX.1 look |
| **FLUX.2 Klein 9B** | 4 | **18.7s** | 28.4 GB | yes | **hero assets — best object accuracy of the whole lineup, cloud included** |
| FLUX.1 schnell | 4 | 20.4s | 25.0 GB | yes | fast draft (Klein 9B beats it at same speed) |
| FLUX.1 dev | 25 | 108.1s | 25.1 GB | yes | cinematic mood / DoF when you can wait ~2 min |
First-run downloads (one-time): schnell/dev ~31GB & ~1619 min each, Klein 9B ~32GB, Klein 4B ~15GB. Believed among the first published M3 Ultra mflux FLUX.1/FLUX.2 numbers.
**Sweet spot: Klein 9B.** Same ~19s as schnell but far better object coherence → it dominates schnell. Klein 4B when speed matters (2× faster), dev only when you want the cinematic atmosphere. Note the `hf_xet` chunked downloader fails on these repos ("Unable to parse string as hex hash value") — the operator sets `HF_HUB_DISABLE_XET=1` to force the reliable HTTP path.
**First A/B (same prompt, 2026-07-12):** FLUX.2 Klein 4B local (8s, $0) vs nano-banana via OpenRouter (`google/gemini-2.5-flash-image`, 7.7s, $0.0387 exact-billed). Klein: cleaner product-photo subject. nano-banana: richer scene dressing (books/inkwell/quill, dust motes) + finer engraving detail. Verdict: Klein is the volume workhorse; nano-banana wins on scene storytelling per prompt-adherence expectations (Elo 1154 vs ~1083).
## Mesh-gen (local, validated 2026-07-13)
| Operator | Config | Result |
|---|---|---|
| `sf3d` | image → GLB, MPS, tex 1024 | **~5s**, ~9GB peak, 1.5MB GLB. Fast draft tier — good on solid objects, struggles on thin/open geometry. Needs `OMP_NUM_THREADS=1`+`KMP_DUPLICATE_LIB_OK` (segfaults otherwise). |
| `trellis_mac` | TRELLIS.2-4B, pipeline 1024, tex 2048, MPS | **289s (~4.8 min) generation** + 16s bake, 18.4MB GLB with PBR. SOTA-tier local quality — clean coherent geometry even on a thin-ringed astrolabe (dramatically better than SF3D). First run adds a one-time ~15GB download (~30 min); cached after. Needs `HF_HUB_DISABLE_XET=1` + the OMP guards. |
| `hunyuan3d_mlx` | Hunyuan3D 2.1, native MLX (fp16), shape+PBR, tex 2048, remesh 40k | **260s (~4.3 min) total** (shape 149s + texture 112s), peak 20.2GB, 7.6MB GLB (40k faces, 2048² baseColor+MR PBR). Weights **public — no HF login**. Needs `diffusers`+`fast_simplification` in the venv. **MLX-native → the one local 3D op that runs on M1 Ultra** (trellis_mac's torch-MPS bf16 is unverified there). |
| `bg_remove_local` | RMBG-2.0, MPS, 1024 | seconds; clean transparent cutout. Run before SF3D for a big geometry improvement. Note: keeps original RGB under alpha (upscale before cutout). |
### Head-to-head, same mermaid cutout (M3 Ultra, 2026-07-16)
| | speed | GLB | faces | face/detail quality |
|---|---|---|---|---|
| `trellis_mac` | 318s | 23MB | 175,842 | **sharper** — defined eyes/nose/mouth, individually raised tail scales, vivid colors |
| `hunyuan3d_mlx` | 260s | 7.6MB | 40,000 | softer — melted face, smoothed scales, muted texture |
**Verdict:** at defaults `trellis_mac` wins on quality (crisper face + geometry, richer color); `hunyuan3d_mlx` is faster, ~3× lighter, and the only local 3D op that runs on M1. `sf3d` stays the ~5s draft tier. All free/offline; fal cloud for on-demand SOTA without the local wait.
## Video gen — mlx-video / Wan 2.1 on M3 Ultra (2026-07-17, first numbers)
[Blaizzy/mlx-video](https://github.com/Blaizzy/mlx-video) (native MLX, MIT) at `vendor/mlx-video`. Wan2.1-T2V-1.3B, PyTorch weights → MLX via their converter (56s), raw `.pth` deleted after (MLX copy is what runs; 14GB).
| run | config | time | note |
|---|---|---|---|
| **T2V 3s** | 832×480, 49 frames, 50 steps, `--tiling aggressive` | **809.6s (13.5 min)** | T5 6.5s · **denoise 780.7s (15.6s/step)** · VAE decode 21.7s |
| smoke | 832×480, 5 frames, 2 steps | 11.8s | pipeline sanity |
Quality: genuinely good — photoreal fox walking through snow, coherent across all 49 frames, correct gait, no flicker/melting. Output registered in the library (`wan_first.mp4`).
**⚠️ THE GOTCHA — default `--tiling auto` gets Wan SIGKILLed past ~25 frames.** Verified bisect (2 steps, 832×480): 13f (seq 6240) OK · 25f (seq 10920) OK · **33f (seq 14040) SIGKILL**. It is **not** memory: it dies at **4GB RSS with 195GB free**, and MLX reports `max_buffer_length: 179GB` / working set 239GB on this machine. The `auto` heuristic never enables **temporal** tiling, so the VAE decode explodes. Worse, it dies *after* denoising completes — with defaults you wait 14 min for the expensive part, then lose it at the decode with no error. **Fix: `--tiling aggressive`** (→ spatial=256px, temporal=32f) — 33f then saves fine. `--tiling temporal` does *not* help.
Other mlx-video packaging bugs found (all upstream-PR-worthy): (1) `torch` is required to read the `.pth` T5/VAE but is **undeclared** — conversion dies *halfway*, leaving a plausible-looking broken model dir; (2) `librosa` is a hard dep though only LTX-2's *audio* VAE uses it — drags in ancient numba and **breaks install** on modern Python (should be an optional `[audio]` extra); (3) README documents `python -m mlx_video.wan2.convert` — that module doesn't exist (it's `mlx_video.models.wan_2.convert`).
**Wan2.2-TI2V-5B** (23GB MLX, text+image→video): 704², 49 frames, 40 steps = **337s (5.6 min)***faster than the 1.3B* at 832×480 (fewer steps + better-conditioned I2V). Note `Wan2.2-I2V-A14B` is **126GB** (~250GB with conversion) — not attempted.
### ❌ NEGATIVE RESULT: I2V-synthesized orbits do NOT reconstruct in 3D (2026-07-17)
Tested the appealing shortcut — *single image → I2V "camera orbit" → frames → colmap_poses → brush_train* — i.e. synthesize the turntable instead of filming it. Prompted TI2V-5B explicitly for camera orbit around a static subject (parallax being what COLMAP needs), fed the RMBG mermaid cutout on white.
| source | registered images | **3D points triangulated** |
|---|---|---|
| real iPad room video | 34 | **1006** ✅ |
| **synthesized I2V orbit** | 26 | **0** ❌ |
**Zero triangulated points.** COLMAP matched features and guessed poses but could not place a single consistent point in space (hence `Failed to fix Gauge … insufficient number of fixed points: 0`, and a "reconstruction" in 2.4s vs ~60s for real footage). Brush still emitted a 190MB .ply — trained faithfully on geometry-free input, i.e. garbage. **Bigger splat ≠ better.**
**Why:** the video visibly rotates (back of the head by frame 47) but the subject is **not rigid** — the tail dissolves, an orb materialises, the figure morphs into a different object. Video models optimise **temporal plausibility**, not **multi-view consistency**; 3D needs the latter. Each frame is individually gorgeous and mutually incompatible.
**Conclusion: you cannot synthesize your way out of capture.** For video→3D, film the real turntable (→ `bg_remove_local`/CorridorKey → colmap → brush/hunyuan3d_mlx). Use I2V for motion/B-roll, not as a multi-view source. Don't re-run this experiment — it's a property of the objective, not a tuning problem. (A purpose-built multi-view/NVS model — e.g. TRELLIS-class or a camera-controlled LoRA — is the thing that could work; a general I2V model can't.)
## CorridorKey (neural green-screen keyer, MLX) — Ultra tuning (2026-07-16)
Corridor Digital's keyer (`vendor/corridorkey`, 14.4k★) with the native MLX backend (`corridorkey-mlx`, resolved from git — not on PyPI). Benchmarked on a synthetic 12-frame 2048² green-screen set with **exact ground-truth alpha** (RMBG mermaid cutout + soft-alpha stripes/disk over chroma green); hint = 8× downscaled truth. Scores = alpha MAE / soft-IoU vs truth, steady-state after 2-frame warmup.
| config | M3 s/f | M1 s/f | MAE ↓ | IoU ↑ | peak GB |
|---|---|---|---|---|---|
| full 2048 (stock default) | 3.81 | 4.97 | 0.0091 | 0.925 | **28.2** |
| full 1024 | 0.61 | 0.75 | 0.0111 | 0.908 | 3.7 |
| full 512 | 0.25 | 0.31 | 0.0139 | 0.885 | 2.6 |
| tiled 512 (stock: compile forced off) | 3.64 | 4.65 | **0.0082** | **0.932** | 2.5 |
| **tiled 512 + our compile patch** | **2.48** | 4.19 | **0.0082** | **0.932** | **2.3** |
**Findings:** (1) `compile` gains nothing full-frame at 2048 on Ultras (within noise) — the documented "1.52×" is a small-res/laptop figure. (2) **Tiled-512 is the QUALITY winner**, not just the memory fallback — best alpha accuracy, model at native tile scale over full-res input. (3) Upstream hard-codes `compile=False` in tiled mode; tiles are fixed-shape so compilation applies — our 1-line patch (`vendor/corridorkey-mlx`, branch `modelbeast`, editable-installed into the app venv on M3+M1) makes tiled **1.47× faster on M3** (3.64→2.48 s/f), 1.11× on M1, output bit-identical. (4) tile 1024 tested worse (quality + speed) — 512 is the sweet spot. (5) full-frame 2048 peaks **28 GB** → 32GB Macs (M1 Max) should run **tiled** (2.3 GB) — which is also the best-quality config anyway.
**Fleet verdict:** best-quality config = `tile_size=512, overlap=64` + our patch: M3 ~0.40 fps, M1 ~0.24 fps at 2048², IoU 0.932, 2.3 GB — runs on every node including the M4 24GB. Throughput mode: full-1024 (1.65/1.34 fps, IoU 0.908). MLX gaps: blue-screen checkpoint + despill/despeckle not yet on MLX (torch backend covers those). License: CC BY-NC-SA (non-commercial).
**8-machine matrix (2026-07-17) — tiled keys 2048 on EVERY Mac, same accuracy.** `tile_size=768, overlap=64` + our compile patch, full 2048² input, **peak ~2.2GB and alpha MAE 0.00849 identical on all**: M3 Ultra 1949ms · M1 Ultra 3275ms · M1 Max(24c/32GB) 4062ms · M4 mini(10c/16GB) 7284ms · **OG M1 mini (2020, 8c/8GB) 14269ms**. Full-frame 2048 needs 26GB → unreachable on most Macs ever sold; **tiled isn't the fallback, it's the answer**. Also: hdim-56 fast-path fallback confirmed on all 8 (padding win 1.3×4.9×, tracks fused-kernel quality **not** core count — the M1 Max refuted our width theory, the M4 mini then fit the corrected one as a prediction); M5's MLX gains live almost entirely in the fused kernel; fanless M3 Air does **not** throttle (1.01× at 5 min).
**Fleet ablation follow-up (2026-07-16, full report in the fork: `monster/corridorkey-mrp-mlx` → `docs/2026-07-16-m-series-fleet-ablation-results.md`):** ran upstream's own 6-toggle benchmark matrix on M3+M1+M4. New best config on BOTH Ultras = **tiled 768/64 + compile** (M3 1949ms, M1 3275ms per 2048² raster — vs 2788/5373 full-frame) at 2.4GB. Big discovery: **`sdpa` is a 4.5× regression on M1-class GPUs at 2048 (18.6s vs 4.1s) and free on M3; bf16 is neutral on both** — gate sdpa by GPU generation, the "M1 bf16 danger" assumption is wrong for this workload. `stage_gc` harmful on Ultras (0.530.66×). M4 24GB swaps at full-frame 2048 (~2024s all configs) → tiled mandatory there; healthy ≤1024.
### Phase D — hunyuan Studio tuning (2026-07-16, M3 Ultra) → new operator defaults
Raised the config from the laptop-tuned defaults to `octree_resolution 384` + `remesh_faces 120000` + `texture_size 4096`. **4096² bake works on the Studio GPU — no Metal command-buffer watchdog** (the existing `extract_textiles` tiling handles it; `uv_feature_map` never needed patching). Result: **380s** (shape 160 + tex 221), peak 20.2GB, 21.5MB GLB, 78k verts / **120k faces**, 4096² baseColor+MR. Quality jump is real — the melted face gains defined eyes + structure, tail geometry sharpens, textures crisper; closes most of the gap to trellis (trellis still edges the face). Cost: ~46% slower + ~3× file size vs the 40k/2048 default. **These are now the `hunyuan3d_mlx` operator defaults** (all still param-overridable; drop to `remesh_faces 40000`/`texture_size 2048` for fast drafts). **4096 confirmed watchdog-free on the M1 Ultra too** (751s total — M1 runs it at ~2× M3 time, so speed-critical jobs prefer M3, which is first in the pool). Defaults are fleet-safe on both Ultras.
## Mesh-gen — earlier install notes (superseded by the table above)
| Operator | Install | Runtime status |
|---|---|---|
| `sf3d` | venv + Metal texture_baker/uv_unwrapper kernels compiled OK; torch 2.13 MPS available | Runs end-to-end; **weights gated**`stabilityai/stable-fast-3d` returns `GatedRepoError` until the owner accepts the license + sets an HF token. Then expect seconds-to-a-minute on MPS. |
| `trellis_mac` | setup.sh built .venv (py3.11) + mtl* Metal kernels; torch 2.13 MPS available | Runs end-to-end; **weights gated** — needs HF access to `facebook/dinov3-vitl16-pretrain-lvd1689m` + `briaai/RMBG-2.0`. Expect ~35 min/gen once authed (M4 Pro reference; M3 Ultra should match or beat). |
| `fal_*` (trellis / trellis2 / hunyuan3d / rodin) | none (API) | Gated on `FAL_KEY`. Verified param surfaces; ~1s1min server-side per fal docs. |
## To unblock the gated local operators
1. Accept the model licenses on HuggingFace (one-time, usually instant):
- https://huggingface.co/stabilityai/stable-fast-3d
- https://huggingface.co/facebook/dinov3-vitl16-pretrain-lvd1689m
- https://huggingface.co/briaai/RMBG-2.0
2. Either `huggingface-cli login` on the machine, or paste an HF token into
Settings → "HuggingFace token" (injected as `HF_TOKEN` for the operators).
## Fleet hardware baseline (perfcheck, torch 2.13.0/MPS, 2026-07-17)
What each node's GPU is *capable* of, independent of any model — the denominator for every
number above, and the reference for "is this machine still healthy". Measured nightly by
`scripts/perfcheck/`; see **[PERFCHECK.md](PERFCHECK.md)** for the full story.
| node | chip | RAM | matmul fp16 | bandwidth | sdpa 56/64 |
|---|---|---|---|---|---|
| m3ultra | M3 Ultra | 256GB | 25202 | 627 GB/s | 4.35x |
| m1 | M1 Ultra | 128GB | ~17100 | 639 GB/s | 3.89x |
| studio | M1 Max | 32GB | 6996 | 348 GB/s | 3.09x |
| m4 | M4 Pro | 24GB | 5678 | 241 GB/s | 1.96x |
| m4mini | M4 | 16GB | 3774 | 105 GB/s | 5.10x |
| mini | M1 (2020) | 8GB | 2310 | 61 GB/s | 5.26x |
GFLOP/s. The M3 Ultra is **~2.2x** the M1 Ultra on fp16 matmul but they share the same ~625
GB/s memory bandwidth — so bandwidth-bound stages (VAE decode, big texture bakes) run at
similar speed on both, while compute-bound ones (denoise) scale with the newer silicon. Both
Ultras reach only ~78% of their ~800 GB/s spec on a single kernel; the smaller Macs hit ~88%.
`sdpa 56/64` is the fused-attention cliff at CorridorKey's awkward `head_dim=56` — it's why
padding to 64 wins (`CORRIDORKEY.md`), reproduced independently on all 6 nodes, and it is
watched nightly so we learn if a future torch ever closes it.
## Method
Timings are wall-clock from the job runner (`started_at`→`finished_at`), single
job at a time (gpu lane = 1). Re-run `tests/smoke.sh` for the framework
regression suite (12 checks, ~30s).
**Benchmarking on this fleet is contended.** These are working machines; a neighbouring job
can make a node look 511x slower and it is invisible to load average (the M3 Ultra reads load
2.45 with its GPU pinned). Anything measured here should use min-of-runs and check
`ioreg … IOAccelerator` "Device Utilization %" *before* touching the GPU. See PERFCHECK.md.