modelbeast/BENCHMARKS.md
MODELBEAST c42f066723 flux_local: HF_HUB_DISABLE_XET fix + full local FLUX lineup benchmarks
- run.py: force HF_HUB_DISABLE_XET=1 (xet chunked downloader fails on BFL repos
  with 'Unable to parse string as hex hash value'; HTTP path is reliable)
- BENCHMARKS.md: warm generation times for all 5 local FLUX models on M3 Ultra
  (Klein 4B 9.1s, schnell-4bit 18.5s, Klein 9B 18.7s, schnell 20.4s, dev 108s).
  Verdict: Klein 9B = best hero-asset quality (beats schnell at same speed),
  Klein 4B = volume workhorse.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 00:37:24 +10:00

4.1 KiB
Raw Blame History

MODELBEAST Benchmarks (M3 Ultra, 256GB)

First measurements on this machine, recorded 2026-07-12. Fixtures are synthetic (Blender-rendered Suzanne), so quality numbers are not representative of real photography — these validate that the pipeline runs and how fast.

Scan track (fully validated, no gated weights)

Stage Input Settings Result
colmap_poses 48 frames @ 800×600 sequential matcher, global (GLOMAP) mapper, OPENCV, CPU SIFT 48/48 images registered, 1479 points, 0.60px mean reprojection error; global mapper step ~1.0s; full job a few seconds
brush_train above colmap_dataset 1500 steps, max_res 800, sh 2 ~3060s wall, 465KB splat.ply, renders in the in-app SplatViewer

A full-quality brush_train run is 30000 steps (the default) — expect minutes, and a much crisper splat than the 1500-step preview above.

Image generation — local FLUX lineup (M3 Ultra, mflux/MLX, 1024×1024, validated 2026-07-13)

All five installed and generating. Warm = weights cached (the real per-image cost):

Model Steps Warm gen Peak MLX mem Gated? Best for
FLUX.2 Klein 4B 4 9.1s 18.0 GB no (Apache) volume / sprites — the default
FLUX.1 schnell-4bit 4 18.5s 19.1 GB no (community quant) fast draft, FLUX.1 look
FLUX.2 Klein 9B 4 18.7s 28.4 GB yes hero assets — best object accuracy of the whole lineup, cloud included
FLUX.1 schnell 4 20.4s 25.0 GB yes fast draft (Klein 9B beats it at same speed)
FLUX.1 dev 25 108.1s 25.1 GB yes cinematic mood / DoF when you can wait ~2 min

First-run downloads (one-time): schnell/dev ~31GB & ~1619 min each, Klein 9B ~32GB, Klein 4B ~15GB. Believed among the first published M3 Ultra mflux FLUX.1/FLUX.2 numbers.

Sweet spot: Klein 9B. Same ~19s as schnell but far better object coherence → it dominates schnell. Klein 4B when speed matters (2× faster), dev only when you want the cinematic atmosphere. Note the hf_xet chunked downloader fails on these repos ("Unable to parse string as hex hash value") — the operator sets HF_HUB_DISABLE_XET=1 to force the reliable HTTP path.

First A/B (same prompt, 2026-07-12): FLUX.2 Klein 4B local (8s, $0) vs nano-banana via OpenRouter (google/gemini-2.5-flash-image, 7.7s, $0.0387 exact-billed). Klein: cleaner product-photo subject. nano-banana: richer scene dressing (books/inkwell/quill, dust motes) + finer engraving detail. Verdict: Klein is the volume workhorse; nano-banana wins on scene storytelling per prompt-adherence expectations (Elo 1154 vs ~1083).

Mesh-gen (installed; first real run blocked on owner HuggingFace auth)

Operator Install Runtime status
sf3d venv + Metal texture_baker/uv_unwrapper kernels compiled OK; torch 2.13 MPS available Runs end-to-end; weights gatedstabilityai/stable-fast-3d returns GatedRepoError until the owner accepts the license + sets an HF token. Then expect seconds-to-a-minute on MPS.
trellis_mac setup.sh built .venv (py3.11) + mtl* Metal kernels; torch 2.13 MPS available Runs end-to-end; weights gated — needs HF access to facebook/dinov3-vitl16-pretrain-lvd1689m + briaai/RMBG-2.0. Expect ~35 min/gen once authed (M4 Pro reference; M3 Ultra should match or beat).
fal_* (trellis / trellis2 / hunyuan3d / rodin) none (API) Gated on FAL_KEY. Verified param surfaces; ~1s1min server-side per fal docs.

To unblock the gated local operators

  1. Accept the model licenses on HuggingFace (one-time, usually instant):
  2. Either huggingface-cli login on the machine, or paste an HF token into Settings → "HuggingFace token" (injected as HF_TOKEN for the operators).

Method

Timings are wall-clock from the job runner (started_atfinished_at), single job at a time (gpu lane = 1). Re-run tests/smoke.sh for the framework regression suite (12 checks, ~30s).