modelbeast/BENCHMARKS.md
2026-07-19 00:35:15 +10:00

18 KiB
Raw Blame History

MODELBEAST Benchmarks (M3 Ultra, 256GB)

First measurements on this machine, recorded 2026-07-12. Fixtures are synthetic (Blender-rendered Suzanne), so quality numbers are not representative of real photography — these validate that the pipeline runs and how fast.

Scan track (fully validated, no gated weights)

Stage Input Settings Result
colmap_poses 48 frames @ 800×600 sequential matcher, global (GLOMAP) mapper, OPENCV, CPU SIFT 48/48 images registered, 1479 points, 0.60px mean reprojection error; global mapper step ~1.0s; full job a few seconds
brush_train above colmap_dataset 1500 steps, max_res 800, sh 2 ~3060s wall, 465KB splat.ply, renders in the in-app SplatViewer

A full-quality brush_train run is 30000 steps (the default) — expect minutes, and a much crisper splat than the 1500-step preview above.

Image generation — local FLUX lineup (M3 Ultra, mflux/MLX, 1024×1024, validated 2026-07-13)

All five installed and generating. Warm = weights cached (the real per-image cost):

Model Steps Warm gen Peak MLX mem Gated? Best for
FLUX.2 Klein 4B 4 9.1s 18.0 GB no (Apache) volume / sprites — the default
FLUX.1 schnell-4bit 4 18.5s 19.1 GB no (community quant) fast draft, FLUX.1 look
FLUX.2 Klein 9B 4 18.7s 28.4 GB yes hero assets — best object accuracy of the whole lineup, cloud included
FLUX.1 schnell 4 20.4s 25.0 GB yes fast draft (Klein 9B beats it at same speed)
FLUX.1 dev 25 108.1s 25.1 GB yes cinematic mood / DoF when you can wait ~2 min

First-run downloads (one-time): schnell/dev ~31GB & ~1619 min each, Klein 9B ~32GB, Klein 4B ~15GB. Believed among the first published M3 Ultra mflux FLUX.1/FLUX.2 numbers.

Sweet spot: Klein 9B. Same ~19s as schnell but far better object coherence → it dominates schnell. Klein 4B when speed matters (2× faster), dev only when you want the cinematic atmosphere. Note the hf_xet chunked downloader fails on these repos ("Unable to parse string as hex hash value") — the operator sets HF_HUB_DISABLE_XET=1 to force the reliable HTTP path.

First A/B (same prompt, 2026-07-12): FLUX.2 Klein 4B local (8s, $0) vs nano-banana via OpenRouter (google/gemini-2.5-flash-image, 7.7s, $0.0387 exact-billed). Klein: cleaner product-photo subject. nano-banana: richer scene dressing (books/inkwell/quill, dust motes) + finer engraving detail. Verdict: Klein is the volume workhorse; nano-banana wins on scene storytelling per prompt-adherence expectations (Elo 1154 vs ~1083).

Mesh-gen (local, validated 2026-07-13)

Operator Config Result
sf3d image → GLB, MPS, tex 1024 ~5s, ~9GB peak, 1.5MB GLB. Fast draft tier — good on solid objects, struggles on thin/open geometry. Needs OMP_NUM_THREADS=1+KMP_DUPLICATE_LIB_OK (segfaults otherwise).
trellis_mac TRELLIS.2-4B, pipeline 1024, tex 2048, MPS 289s (~4.8 min) generation + 16s bake, 18.4MB GLB with PBR. SOTA-tier local quality — clean coherent geometry even on a thin-ringed astrolabe (dramatically better than SF3D). First run adds a one-time ~15GB download (~30 min); cached after. Needs HF_HUB_DISABLE_XET=1 + the OMP guards.
hunyuan3d_mlx Hunyuan3D 2.1, native MLX (fp16), shape+PBR, tex 2048, remesh 40k 260s (~4.3 min) total (shape 149s + texture 112s), peak 20.2GB, 7.6MB GLB (40k faces, 2048² baseColor+MR PBR). Weights public — no HF login. Needs diffusers+fast_simplification in the venv. MLX-native → the one local 3D op that runs on M1 Ultra (trellis_mac's torch-MPS bf16 is unverified there).
bg_remove_local RMBG-2.0, MPS, 1024 seconds; clean transparent cutout. Run before SF3D for a big geometry improvement. Note: keeps original RGB under alpha (upscale before cutout).

Head-to-head, same mermaid cutout (M3 Ultra, 2026-07-16)

speed GLB faces face/detail quality
trellis_mac 318s 23MB 175,842 sharper — defined eyes/nose/mouth, individually raised tail scales, vivid colors
hunyuan3d_mlx 260s 7.6MB 40,000 softer — melted face, smoothed scales, muted texture

Verdict: at defaults trellis_mac wins on quality (crisper face + geometry, richer color); hunyuan3d_mlx is faster, ~3× lighter, and the only local 3D op that runs on M1. sf3d stays the ~5s draft tier. All free/offline; fal cloud for on-demand SOTA without the local wait.

Video gen — mlx-video / Wan 2.1 on M3 Ultra (2026-07-17, first numbers)

Blaizzy/mlx-video (native MLX, MIT) at vendor/mlx-video. Wan2.1-T2V-1.3B, PyTorch weights → MLX via their converter (56s), raw .pth deleted after (MLX copy is what runs; 14GB).

run config time note
T2V 3s 832×480, 49 frames, 50 steps, --tiling aggressive 809.6s (13.5 min) T5 6.5s · denoise 780.7s (15.6s/step) · VAE decode 21.7s
smoke 832×480, 5 frames, 2 steps 11.8s pipeline sanity

Quality: genuinely good — photoreal fox walking through snow, coherent across all 49 frames, correct gait, no flicker/melting. Output registered in the library (wan_first.mp4).

⚠️ THE GOTCHA — default --tiling auto gets Wan SIGKILLed past ~25 frames. Verified bisect (2 steps, 832×480): 13f (seq 6240) OK · 25f (seq 10920) OK · 33f (seq 14040) SIGKILL. It is not memory: it dies at 4GB RSS with 195GB free, and MLX reports max_buffer_length: 179GB / working set 239GB on this machine. The auto heuristic never enables temporal tiling, so the VAE decode explodes. Worse, it dies after denoising completes — with defaults you wait 14 min for the expensive part, then lose it at the decode with no error. Fix: --tiling aggressive (→ spatial=256px, temporal=32f) — 33f then saves fine. --tiling temporal does not help.

Other mlx-video packaging bugs found (all upstream-PR-worthy): (1) torch is required to read the .pth T5/VAE but is undeclared — conversion dies halfway, leaving a plausible-looking broken model dir; (2) librosa is a hard dep though only LTX-2's audio VAE uses it — drags in ancient numba and breaks install on modern Python (should be an optional [audio] extra); (3) README documents python -m mlx_video.wan2.convert — that module doesn't exist (it's mlx_video.models.wan_2.convert).

Wan2.2-TI2V-5B (23GB MLX, text+image→video): 704², 49 frames, 40 steps = 337s (5.6 min)faster than the 1.3B at 832×480 (fewer steps + better-conditioned I2V). Note Wan2.2-I2V-A14B is 126GB (~250GB with conversion) — not attempted.

NEGATIVE RESULT: I2V-synthesized orbits do NOT reconstruct in 3D (2026-07-17)

Tested the appealing shortcut — single image → I2V "camera orbit" → frames → colmap_poses → brush_train — i.e. synthesize the turntable instead of filming it. Prompted TI2V-5B explicitly for camera orbit around a static subject (parallax being what COLMAP needs), fed the RMBG mermaid cutout on white.

source registered images 3D points triangulated
real iPad room video 34 1006
synthesized I2V orbit 26 0

Zero triangulated points. COLMAP matched features and guessed poses but could not place a single consistent point in space (hence Failed to fix Gauge … insufficient number of fixed points: 0, and a "reconstruction" in 2.4s vs ~60s for real footage). Brush still emitted a 190MB .ply — trained faithfully on geometry-free input, i.e. garbage. Bigger splat ≠ better.

Why: the video visibly rotates (back of the head by frame 47) but the subject is not rigid — the tail dissolves, an orb materialises, the figure morphs into a different object. Video models optimise temporal plausibility, not multi-view consistency; 3D needs the latter. Each frame is individually gorgeous and mutually incompatible.

Conclusion: you cannot synthesize your way out of capture. For video→3D, film the real turntable (→ bg_remove_local/CorridorKey → colmap → brush/hunyuan3d_mlx). Use I2V for motion/B-roll, not as a multi-view source. Don't re-run this experiment — it's a property of the objective, not a tuning problem. (A purpose-built multi-view/NVS model — e.g. TRELLIS-class or a camera-controlled LoRA — is the thing that could work; a general I2V model can't.)

CorridorKey (neural green-screen keyer, MLX) — Ultra tuning (2026-07-16)

Corridor Digital's keyer (vendor/corridorkey, 14.4k★) with the native MLX backend (corridorkey-mlx, resolved from git — not on PyPI). Benchmarked on a synthetic 12-frame 2048² green-screen set with exact ground-truth alpha (RMBG mermaid cutout + soft-alpha stripes/disk over chroma green); hint = 8× downscaled truth. Scores = alpha MAE / soft-IoU vs truth, steady-state after 2-frame warmup.

config M3 s/f M1 s/f MAE ↓ IoU ↑ peak GB
full 2048 (stock default) 3.81 4.97 0.0091 0.925 28.2
full 1024 0.61 0.75 0.0111 0.908 3.7
full 512 0.25 0.31 0.0139 0.885 2.6
tiled 512 (stock: compile forced off) 3.64 4.65 0.0082 0.932 2.5
tiled 512 + our compile patch 2.48 4.19 0.0082 0.932 2.3

Findings: (1) compile gains nothing full-frame at 2048 on Ultras (within noise) — the documented "1.52×" is a small-res/laptop figure. (2) Tiled-512 is the QUALITY winner, not just the memory fallback — best alpha accuracy, model at native tile scale over full-res input. (3) Upstream hard-codes compile=False in tiled mode; tiles are fixed-shape so compilation applies — our 1-line patch (vendor/corridorkey-mlx, branch modelbeast, editable-installed into the app venv on M3+M1) makes tiled 1.47× faster on M3 (3.64→2.48 s/f), 1.11× on M1, output bit-identical. (4) tile 1024 tested worse (quality + speed) — 512 is the sweet spot. (5) full-frame 2048 peaks 28 GB → 32GB Macs (M1 Max) should run tiled (2.3 GB) — which is also the best-quality config anyway.

Fleet verdict: best-quality config = tile_size=512, overlap=64 + our patch: M3 ~0.40 fps, M1 ~0.24 fps at 2048², IoU 0.932, 2.3 GB — runs on every node including the M4 24GB. Throughput mode: full-1024 (1.65/1.34 fps, IoU 0.908). MLX gaps: blue-screen checkpoint + despill/despeckle not yet on MLX (torch backend covers those). License: CC BY-NC-SA (non-commercial).

8-machine matrix (2026-07-17) — tiled keys 2048 on EVERY Mac, same accuracy. tile_size=768, overlap=64 + our compile patch, full 2048² input, peak ~2.2GB and alpha MAE 0.00849 identical on all: M3 Ultra 1949ms · M1 Ultra 3275ms · M1 Max(24c/32GB) 4062ms · M4 mini(10c/16GB) 7284ms · OG M1 mini (2020, 8c/8GB) 14269ms. Full-frame 2048 needs 26GB → unreachable on most Macs ever sold; tiled isn't the fallback, it's the answer. Also: hdim-56 fast-path fallback confirmed on all 8 (padding win 1.3×4.9×, tracks fused-kernel quality not core count — the M1 Max refuted our width theory, the M4 mini then fit the corrected one as a prediction); M5's MLX gains live almost entirely in the fused kernel; fanless M3 Air does not throttle (1.01× at 5 min).

Fleet ablation follow-up (2026-07-16, full report in the fork: monster/corridorkey-mrp-mlxdocs/2026-07-16-m-series-fleet-ablation-results.md): ran upstream's own 6-toggle benchmark matrix on M3+M1+M4. New best config on BOTH Ultras = tiled 768/64 + compile (M3 1949ms, M1 3275ms per 2048² raster — vs 2788/5373 full-frame) at 2.4GB. Big discovery: sdpa is a 4.5× regression on M1-class GPUs at 2048 (18.6s vs 4.1s) and free on M3; bf16 is neutral on both — gate sdpa by GPU generation, the "M1 bf16 danger" assumption is wrong for this workload. stage_gc harmful on Ultras (0.530.66×). M4 24GB swaps at full-frame 2048 (~2024s all configs) → tiled mandatory there; healthy ≤1024.

Phase D — hunyuan Studio tuning (2026-07-16, M3 Ultra) → new operator defaults

Raised the config from the laptop-tuned defaults to octree_resolution 384 + remesh_faces 120000 + texture_size 4096. 4096² bake works on the Studio GPU — no Metal command-buffer watchdog (the existing extract_textiles tiling handles it; uv_feature_map never needed patching). Result: 380s (shape 160 + tex 221), peak 20.2GB, 21.5MB GLB, 78k verts / 120k faces, 4096² baseColor+MR. Quality jump is real — the melted face gains defined eyes + structure, tail geometry sharpens, textures crisper; closes most of the gap to trellis (trellis still edges the face). Cost: ~46% slower + ~3× file size vs the 40k/2048 default. These are now the hunyuan3d_mlx operator defaults (all still param-overridable; drop to remesh_faces 40000/texture_size 2048 for fast drafts). 4096 confirmed watchdog-free on the M1 Ultra too (751s total — M1 runs it at ~2× M3 time, so speed-critical jobs prefer M3, which is first in the pool). Defaults are fleet-safe on both Ultras.

Mesh-gen — earlier install notes (superseded by the table above)

Operator Install Runtime status
sf3d venv + Metal texture_baker/uv_unwrapper kernels compiled OK; torch 2.13 MPS available Runs end-to-end; weights gatedstabilityai/stable-fast-3d returns GatedRepoError until the owner accepts the license + sets an HF token. Then expect seconds-to-a-minute on MPS.
trellis_mac setup.sh built .venv (py3.11) + mtl* Metal kernels; torch 2.13 MPS available Runs end-to-end; weights gated — needs HF access to facebook/dinov3-vitl16-pretrain-lvd1689m + briaai/RMBG-2.0. Expect ~35 min/gen once authed (M4 Pro reference; M3 Ultra should match or beat).
fal_* (trellis / trellis2 / hunyuan3d / rodin) none (API) Gated on FAL_KEY. Verified param surfaces; ~1s1min server-side per fal docs.

To unblock the gated local operators

  1. Accept the model licenses on HuggingFace (one-time, usually instant):
  2. Either huggingface-cli login on the machine, or paste an HF token into Settings → "HuggingFace token" (injected as HF_TOKEN for the operators).

Fleet hardware baseline (perfcheck, torch 2.13.0/MPS, 2026-07-17)

What each node's GPU is capable of, independent of any model — the denominator for every number above, and the reference for "is this machine still healthy". Measured nightly by scripts/perfcheck/; see PERFCHECK.md for the full story.

node chip RAM matmul fp16 bandwidth sdpa 56/64
m3ultra M3 Ultra 256GB 25202 627 GB/s 4.35x
m1 M1 Ultra 128GB ~17100 639 GB/s 3.89x
studio M1 Max 32GB 6996 348 GB/s 3.09x
m4 M4 Pro 24GB 5678 241 GB/s 1.96x
m4mini M4 16GB 3774 105 GB/s 5.10x
mini M1 (2020) 8GB 2310 61 GB/s 5.26x

GFLOP/s. The M3 Ultra is ~2.2x the M1 Ultra on fp16 matmul but they share the same ~625 GB/s memory bandwidth — so bandwidth-bound stages (VAE decode, big texture bakes) run at similar speed on both, while compute-bound ones (denoise) scale with the newer silicon. Both Ultras reach only ~78% of their ~800 GB/s spec on a single kernel; the smaller Macs hit ~88%.

sdpa 56/64 is the fused-attention cliff at CorridorKey's awkward head_dim=56 — it's why padding to 64 wins (CORRIDORKEY.md), reproduced independently on all 6 nodes, and it is watched nightly so we learn if a future torch ever closes it.

Method

Timings are wall-clock from the job runner (started_atfinished_at), single job at a time (gpu lane = 1). Re-run tests/smoke.sh for the framework regression suite (12 checks, ~30s).

Benchmarking on this fleet is contended. These are working machines; a neighbouring job can make a node look 511x slower and it is invisible to load average (the M3 Ultra reads load 2.45 with its GPU pinned). Anything measured here should use min-of-runs and check ioreg … IOAccelerator "Device Utilization %" before touching the GPU. See PERFCHECK.md.

2026-07-19 — fleet MLX capability matrix (mlx 0.32.0 everywhere)

Probe: 4096² matmul TFLOPS (fp16/bf16), 256MB elementwise GB/s, sdpa head_dim 56 vs 64 (B1·H8·L2048).

host chip RAM fp16 TF bf16 TF ~GB/s sdpa d56/d64
m3ultra M3 Ultra 256 23.48 23.54 400 1.51×
ultra M1 Ultra 128 16.31 13.16 297 0.94×
JING5 M5 16 8.61 8.99 43* 4.72×
m4pro M4 Pro 24 5.24 3.86 225 1.29×
m4mini M4 16 3.65 3.65 98 1.55×
AIRPRO M3 16 1.55 1.64 87 0.95×

*JING5 bandwidth measured under interactive load — retest idle before trusting.

Verdicts:

  • bf16-on-M1 question answered: works, ~20% slower than fp16 (16.3→13.2). Ship fp16 on M1-class, bf16 free on M3 Ultra/M4/M5.
  • head_dim pad-to-64 matters MOST on newest chips (M5 4.7×!, M4-class ~1.3-1.6×, M3 Ultra 1.5×) and ~zero on M1-class (no fused fast path there anyway — matches the corridorkey ablation).
  • Fleet total ≈ 59 TFLOPS fp16; the two Ultras are 2/3 of it. TRELLIS.2-MLX lane assignment: m3ultra primary, ultra second (fp16), M4 Pro light ops.
  • Probe script: reusable one-liner via uv run --with mlx — mlxprobe.py (this bench).

Addendum (same day) — m1max studio joins the matrix (login studio@100.92.78.24; the fleet doc's m1max@ user and "disk-critical" note are both stale — 170GB free):

| M1MAX | M1 Max | 32 | 7.09 | 6.04 | 336 | 1.22× |

Notable: 2nd-best bandwidth in the fleet (336 GB/s > M4 Pro's 225) despite mid-pack compute — good fit for memory-bound MLX work (VAE decode, big-tensor elementwise). Fleet total ≈ 66 TFLOPS fp16 across 7 machines.