12 KiB
MRP-MLX — the MonsterRobotParty MLX stack
The short version: every serious generative model we run — image, layered image, image→3D, video keying — runs locally, on Apple Silicon, in MLX or Metal, from forks we control, at $0 per generation. This document is the canonical reference: what each piece is, where it lives, how fast it goes, how to run it, and the hard-won rules that keep it fast. Written 2026-07-20 after the rung-3 kernel push; every number here was measured, not estimated.
1 · Philosophy (why this stack wins)
- Local-first. Hosted GPU (fal) is for interactive one-offs and A/B references only. Volume work runs on the fleet. (July fal bill was ~$93/mo before this stack; the same workload is now ≈ $0.)
- Adopt-then-patch, never from-scratch. Every port here started as the best community MLX effort, vendored to our Gitea, then patched where measurement said to. Proven 5×: hunyuan3d (dgrauet), corridorkey (cmoyates/Niko), trellis-2 (Jourloy PR#175), qwen-layered (mflux PR#302), flux (mflux upstream).
- Forks live on our Gitea (
ssh://git@100.71.119.27:222/monster/…) with upstream kept asoriginfor rebases and a pin tag on the adopted commit. Our infra, our history, our rollback. - Parity-first. No optimization ships without a numeric gate against the reference implementation (tolerances recorded), and e2e output stats compared at fixed seed. Fixture sets are kept for regression.
- Bench-everything, keep the negatives. BENCHMARKS.md records failures (q4 = noise, mx.compile ≈ 0 on big GEMMs) with the same care as wins, so no future session re-chases a ghost.
- Env-gated changes. Every behavioral change has an env knob and a default that preserves upstream behavior (or our validated best). Prod rollback is an environment variable, not a revert.
2 · The fleet it runs on
| box | chip | RAM | fp16 TFLOPS | bandwidth | role |
|---|---|---|---|---|---|
| m3ultra | M3 Ultra | 256GB | 23.5 | 400 GB/s | primary: TRELLIS.2, queue, heavy gen |
| m1ultra | M1 Ultra | 128GB | 16.3 | 297 GB/s | second lane: everything (post rung-3) |
| m1max | M1 Max | 32GB | 7.1 | 336 GB/s | bandwidth-friendly light ops |
| m4pro | M4 Pro | 24GB | 5.2 | 225 GB/s | MB light ops + Ollama (hands off) |
Full 7-machine MLX matrix + probe script: BENCHMARKS.md (2026-07-19) +
scripts/mlxprobe.py. Rules from it: fp16 on M1-class (bf16 −20%),
bf16 free on M3/M4/M5; pad attention head_dim→64 on M4/M5-class.
3 · The pillars
3.1 TRELLIS.2 — image → full-density 3D · monster/trellis-2-mrp-mlx
- Base: Jourloy/TRELLIS.2 (upstream PR#175 head, pin
pr175-head-754d403) — hybrid MLX (dense stages) + pedronaugusto Metal kernels (sparse ops). - Our commits:
fa972aapure-MLX sampler loop (batched dense CFG, once-per-stage conversions, one eval/step; parity Δ0.03% e2e) ·7860148opt-in Metal spconv kernel + vertex-baker scheduler fix ·0db816bvertex-normal export + COLOR_0 asset class in the validator. - Weights: original
microsoft/TRELLIS.2-4Bsafetensors, loaded directly (mx.load+ key renames; no conversion step exists or should — a converter would double-transpose ss_dec into silent garbage). - Performance (anatomy bench, seed 42, 1024-cascade, vertex baker):
| box | e2e | peak mem |
|---|---|---|
| m3ultra | 72.0 s | 26.5 GB |
| m1ultra | 181.5 s | 17.2 GB |
(old torch-MPS trellis_mac) |
124.8 s gen (248.8 wall) | 20.8 GB |
| (fal, same image) | 298.6 s wall | $0.30 |
- Run:
cd <fork> && .venv/bin/python scripts/generate_asset.py IMG --output-dir OUT --backend mlx-experimental --baker vertex --pipeline-type 1024_cascade --seed N· Farm: operatortrellis2_mlx. - Knobs:
TRELLIS2_MLX_SAMPLER=0(revert to torch sampler) ·TRELLIS2_MLX_COMPILE=1(≈0 on M3U; try on small boxes) ·TRELLIS2_METAL_SPCONV=1(opt-in kernel; slower-but-leaner, reference). - Output class: vertex-colored full-density GLB (albedo only). MR maps
need the UV path (
--baker metal— has the dark-patch sampling bug) or the future fast-UV bake. Game budgets: MESHGOD finish farm downstream. - Checkouts: m3
~/Documents/trellis2-mlx-staging/trellis2-jourloy· m1~/trellis2-mlx· fresh boxes:scripts/install_trellis2_mlx.sh.
3.2 Hunyuan3D 2.1 — image → textured 3D · monster/Hunyuan3D-2.2-mrp-MLX
- Base: dgrauet/Hunyuan3D-2.1-mlx (pure MLX, both stages; 13GB
pre-converted MLX weights on HF). Terminology note: this IS the
v2-generation Hunyuan; fal's
/v2,/v21,/turboare hosted variants. - Our branch
mlx-tune@5a32240(deployed both Ultras): T1 UNet mx.compile + T2 whole-DiT compile (envHY3D_MLX_COMPILE=1; parity 3e-5/7e-6) · T3 StaticMoELayer (bit-identical, kills ~4800 host syncs, makes the DiT compile-legal) · T4 fused-SDPA DINO attention (1.41× its stage, parity 1e-7). - Performance: 260 s baseline → 243 s tuned (m3ultra, defaults). Runs on every box M1-and-up (fp16). ~20GB peak.
- Run:
generate_e2e.py IMG --output OUTinvendor/hunyuan3d-mlx(venv.venv) · Farm: operatorhunyuan3d_mlx.
3.3 FLUX / image generation — mflux (upstream; not forked)
- mflux 0.18.0 (latest) in
venvs/mflux, both Ultras + m4pro. Written from scratch in MLX by Filip Strand — no port debt to carry. - Audit (2026-07-19): our default path (
flux2-klein-4b, 4 steps, ~9 s/image) has mx.compile on the denoise step + a shapeless-compiled scheduler — healthy, keep. Known gap: the FLUX.1 family (dev/schnell/krea) has zero mx.compile — a backport is the same 10-line pattern as Hunyuan T1, expected win 1–4% on Ultras (more on small boxes), and belongs UPSTREAM as an mflux PR rather than a site-packages patch. Queued, not urgent. - Operator
flux_local(klein default) +mflux_image_edit.
3.4 CorridorKey — neural green-screen · monster/corridorkey-mrp-mlx
- Base: nikopueringer/corridorkey-mlx (cmoyates' phased parity-first
port — its
prompts/folder is our template for any future from-scratch port). License CC BY-NC-SA — non-commercial only. - Our patch (
4c660df, branchmodelbeast): honor the compile flag in tiled mode — 8 lines, 1.47× on m3ultra (3.64→2.48 s/frame), output bit-identical. Plus the 8-machine ablation that found the head_dim→64 fast-path rule and the "tiled beats full-frame on quality AND memory" result (2.3GB vs 28GB). Nightly perfcheck watches the fast path.
3.5 Qwen-Image-Layered — image → editable RGBA layers · monster/Qwen-Image-Layered-MRP-MLX
- Base: mflux PR#302 (ZimengXiong; unmerged upstream), pin
pr302-head-a255e4f. 20B, Apache 2.0 (ship-commercial OK). - Validated ladder: q8 is the quality FLOOR — q6 breaks the layer semantics (silhouettes), q4 is noise; speed is FLAT across quants (compute-bound), so never quantize below q8. 20 steps ≈ 50 steps visually (2.9× faster). 1024-res = hero tier.
- Performance (4 layers, 640, baked q8, 20 steps): m3ultra 243 s · m1ultra 366 s · 1024-res hero: 858 s (m3). Peak ~37GB.
- Baked models: m3
~/qwen-layered/qwen-layered-q8· m1~/qwen-layered-staging/qwen-layered-q8(34GB; rebake viamflux-save). - Run:
mflux-generate-qwen-layered --image IMG --layers 4 --resolution 640 --steps 20 --model-path <baked>· Farm: operatorqwen_layered_local. Bonus: the subject layer is matting-grade — the premium tier of image preprocessing for hero assets.
4 · MODELBEAST integration
| operator | backs | speed (m3 / m1) | notes |
|---|---|---|---|
trellis2_mlx |
3.1 | 72 s / 182 s | vertex baker default; preferred 3D lane |
trellis_mac |
torch-MPS port | 125 s gen / n/a | legacy; UV-textured (dark-patch bug) |
hunyuan3d_mlx |
3.2 | 243 s / ~5-6 min | textured GLB; runs everywhere |
flux_local |
3.3 | ~9 s/img | klein-4b 4-step default |
qwen_layered_local |
3.5 | 243 s / 366 s | RGBA layer decomposition |
Routing lives in nodes.json (machine-local, gitignored — edit on the
m3 primary). MESHGOD's local/ lane calls these via the MB REST API.
5 · The lessons ledger (do not re-learn these)
- Profile before optimizing — every time. The rung-3 3× came from a scheduler bug fix + baker swap the profiler exposed; the glamorous kernel work measured ≈ 0. The 75GB "sparse conv" memory monster was actually the Metal texture bake.
- mx.compile ≈ 0 on GEMM-bound 2B+ models on M3 Ultra (measured 4×: Hunyuan T1, T2, TRELLIS step, sampler). It's free insurance, not a win. Compile wins live on small models (corridorkey 1.47×) and maybe small chips.
- Quantization floors are model-specific and cliff-shaped. Qwen- Layered: q8 fine, q6 semantically broken, q4 noise — and quant does NOT buy speed on compute-bound models, only footprint.
- glTF axis + gamma: pipeline space is Z-up in [-0.5,0.5]³; glTF is
Y-up (
(x,y,z)→(x,z,−y)); COLOR_0 is LINEAR (never pre-gamma). A "mangled blob" render is usually just the figure viewed down its own axis after a convention miss. - xatlas hangs (hours) on écorché-class meshes (thousands of disconnected shells). Vertex bake sidesteps UV unwrap entirely.
- macOS TCC blocks launchd from ~/Documents — services live in
~/. - Gated HF weights don't need logins on workers — rsync the model dirs from a cache that has them.
- The fal-outage class of bug: any
while Truepolling a paid API needs a deadline + a cancel; an in-memory queue wedges silently. - Adapters that convert per-call hide O(steps) waste — convert at stage boundaries, hold state in the fast runtime.
- Sentinel fallthroughs bite: a baker/scheduler that silently falls back can burn 5 hours before anyone notices. Fail loudly instead.
6 · Runbooks
New box bring-up (any Apple Silicon, macOS 26+):
- Keys: box's SSH key added to Gitea (
monsteraccount). git clonethe MODELBEAST repo; run the relevantscripts/install_*.sh(trellis2_mlx, qwen_layered, hunyuan…) — each is idempotent, uses uv, and states its weight needs.- Weights: rsync the HF cache model dirs from m3ultra (see §5.7).
- Bench:
uv run --with mlx python scripts/mlxprobe.py→ add the row to BENCHMARKS fleet matrix; set fp16/bf16 per §2 rules. - Register the box + operator list in
nodes.jsonon the m3 primary.
Rollback: every MRP patch is env-gated (§3 knobs) or a pinned-tag fork
— git checkout <pin-tag> restores the adopted upstream exactly.
Where results go: every experiment appends to BENCHMARKS.md (wall
time, peak mem via mx.get_peak_memory(), parity numbers, and a verdict
line). Session-scale context lives in docs/FABLE_STARTUP_PACKET.md.
7 · Roadmap (scoped, in value order)
- DiT-internals deep-dive — TRELLIS.2 sampling runs ≈2.3× above ideal
FLOPs (54 s of the 72). Attention/MLP kernel work inside the forward;
the corridorkey
prompts/phased method is the template. - Fast UV bake — bucket rasterizer (written, in fast_bake_test.py) +
an écorché-safe unwrap → texture+MR parity with fal on the vertex
baker's quality. Unlocks
--baker vertex-quality WITH UV maps. - simdgroup spconv kernel — upgrade the opt-in scalar kernel if the decoders ever matter for time (they're 10 s today).
- Upstreaming — our sampler + baker + validator to Jourloy PR#175; FLUX.1 compile backport as an mflux PR; spconv kernel as reference.
- MESHGOD default flip — route its
local/3D lane totrellis2_mlxonce finish-farm verts→texture baking lands for game budgets.