# MRP-MLX — the MonsterRobotParty MLX stack **The short version:** every serious generative model we run — image, layered image, image→3D, video keying — runs *locally, on Apple Silicon, in MLX or Metal, from forks we control, at $0 per generation.* This document is the canonical reference: what each piece is, where it lives, how fast it goes, how to run it, and the hard-won rules that keep it fast. Written 2026-07-20 after the rung-3 kernel push; every number here was measured, not estimated. --- ## 1 · Philosophy (why this stack wins) 1. **Local-first.** Hosted GPU (fal) is for interactive one-offs and A/B references only. Volume work runs on the fleet. (July fal bill was ~$93/mo before this stack; the same workload is now ≈ $0.) 2. **Adopt-then-patch, never from-scratch.** Every port here started as the best community MLX effort, vendored to our Gitea, then patched where measurement said to. Proven 5×: hunyuan3d (dgrauet), corridorkey (cmoyates/Niko), trellis-2 (Jourloy PR#175), qwen-layered (mflux PR#302), flux (mflux upstream). 3. **Forks live on our Gitea** (`ssh://git@100.71.119.27:222/monster/…`) with upstream kept as `origin` for rebases and a pin tag on the adopted commit. Our infra, our history, our rollback. 4. **Parity-first.** No optimization ships without a numeric gate against the reference implementation (tolerances recorded), and e2e output stats compared at fixed seed. Fixture sets are kept for regression. 5. **Bench-everything, keep the negatives.** BENCHMARKS.md records failures (q4 = noise, mx.compile ≈ 0 on big GEMMs) with the same care as wins, so no future session re-chases a ghost. 6. **Env-gated changes.** Every behavioral change has an env knob and a default that preserves upstream behavior (or our validated best). Prod rollback is an environment variable, not a revert. ## 2 · The fleet it runs on | box | chip | RAM | fp16 TFLOPS | bandwidth | role | |---|---|---|---|---|---| | m3ultra | M3 Ultra | 256GB | 23.5 | 400 GB/s | primary: TRELLIS.2, queue, heavy gen | | m1ultra | M1 Ultra | 128GB | 16.3 | 297 GB/s | second lane: everything (post rung-3) | | m1max | M1 Max | 32GB | 7.1 | 336 GB/s | bandwidth-friendly light ops | | m4pro | M4 Pro | 24GB | 5.2 | 225 GB/s | MB light ops + Ollama (hands off) | Full 7-machine MLX matrix + probe script: `BENCHMARKS.md` (2026-07-19) + `scripts/mlxprobe.py`. Rules from it: **fp16 on M1-class** (bf16 −20%), bf16 free on M3/M4/M5; pad attention head_dim→64 on M4/M5-class. ## 3 · The pillars ### 3.1 TRELLIS.2 — image → full-density 3D · `monster/trellis-2-mrp-mlx` - **Base:** Jourloy/TRELLIS.2 (upstream PR#175 head, pin `pr175-head-754d403`) — hybrid MLX (dense stages) + pedronaugusto Metal kernels (sparse ops). - **Our commits:** `fa972aa` pure-MLX sampler loop (batched dense CFG, once-per-stage conversions, one eval/step; parity Δ0.03% e2e) · `7860148` opt-in Metal spconv kernel + vertex-baker scheduler fix · `0db816b` vertex-normal export + COLOR_0 asset class in the validator. - **Weights:** original `microsoft/TRELLIS.2-4B` safetensors, loaded directly (`mx.load` + key renames; **no conversion step exists or should** — a converter would double-transpose ss_dec into silent garbage). - **Performance (anatomy bench, seed 42, 1024-cascade, vertex baker):** | box | e2e | peak mem | |---|---|---| | m3ultra | **72.0 s** | 26.5 GB | | m1ultra | **181.5 s** | 17.2 GB | | (old torch-MPS `trellis_mac`) | 124.8 s gen (248.8 wall) | 20.8 GB | | (fal, same image) | 298.6 s wall | $0.30 | - **Run:** `cd && .venv/bin/python scripts/generate_asset.py IMG --output-dir OUT --backend mlx-experimental --baker vertex --pipeline-type 1024_cascade --seed N` · Farm: operator `trellis2_mlx`. - **Knobs:** `TRELLIS2_MLX_SAMPLER=0` (revert to torch sampler) · `TRELLIS2_MLX_COMPILE=1` (≈0 on M3U; try on small boxes) · `TRELLIS2_METAL_SPCONV=1` (opt-in kernel; slower-but-leaner, reference). - **Output class:** vertex-colored full-density GLB (albedo only). MR maps need the UV path (`--baker metal` — has the dark-patch sampling bug) or the future fast-UV bake. Game budgets: MESHGOD finish farm downstream. - **Checkouts:** m3 `~/Documents/trellis2-mlx-staging/trellis2-jourloy` · m1 `~/trellis2-mlx` · fresh boxes: `scripts/install_trellis2_mlx.sh`. ### 3.2 Hunyuan3D 2.1 — image → textured 3D · `monster/Hunyuan3D-2.2-mrp-MLX` - **Base:** dgrauet/Hunyuan3D-2.1-mlx (pure MLX, both stages; 13GB pre-converted MLX weights on HF). *Terminology note: this IS the v2-generation Hunyuan; fal's `/v2`,`/v21`,`/turbo` are hosted variants.* - **Our branch `mlx-tune` @ `5a32240`** (deployed both Ultras): T1 UNet mx.compile + T2 whole-DiT compile (env `HY3D_MLX_COMPILE=1`; parity 3e-5/7e-6) · T3 StaticMoELayer (bit-identical, kills ~4800 host syncs, makes the DiT compile-legal) · T4 fused-SDPA DINO attention (1.41× its stage, parity 1e-7). - **Performance:** 260 s baseline → **243 s** tuned (m3ultra, defaults). Runs on every box M1-and-up (fp16). ~20GB peak. - **Run:** `generate_e2e.py IMG --output OUT` in `vendor/hunyuan3d-mlx` (venv `.venv`) · Farm: operator `hunyuan3d_mlx`. ### 3.3 FLUX / image generation — `mflux` (upstream; not forked) - mflux 0.18.0 (latest) in `venvs/mflux`, both Ultras + m4pro. Written from scratch in MLX by Filip Strand — no port debt to carry. - **Audit (2026-07-19):** our default path (`flux2-klein-4b`, 4 steps, ~9 s/image) has mx.compile on the denoise step + a shapeless-compiled scheduler — healthy, keep. **Known gap:** the FLUX.1 family (dev/schnell/krea) has zero mx.compile — a backport is the same 10-line pattern as Hunyuan T1, expected win 1–4% on Ultras (more on small boxes), and belongs UPSTREAM as an mflux PR rather than a site-packages patch. Queued, not urgent. - Operator `flux_local` (klein default) + `mflux_image_edit`. ### 3.4 CorridorKey — neural green-screen · `monster/corridorkey-mrp-mlx` - **Base:** nikopueringer/corridorkey-mlx (cmoyates' phased parity-first port — its `prompts/` folder is our template for any future from-scratch port). License CC BY-NC-SA — non-commercial only. - **Our patch (`4c660df`, branch `modelbeast`):** honor the compile flag in tiled mode — 8 lines, **1.47×** on m3ultra (3.64→2.48 s/frame), output bit-identical. Plus the 8-machine ablation that found the head_dim→64 fast-path rule and the "tiled beats full-frame on quality AND memory" result (2.3GB vs 28GB). Nightly perfcheck watches the fast path. ### 3.5 Qwen-Image-Layered — image → editable RGBA layers · `monster/Qwen-Image-Layered-MRP-MLX` - **Base:** mflux PR#302 (ZimengXiong; unmerged upstream), pin `pr302-head-a255e4f`. 20B, Apache 2.0 (ship-commercial OK). - **Validated ladder:** **q8 is the quality FLOOR** — q6 breaks the layer semantics (silhouettes), q4 is noise; speed is FLAT across quants (compute-bound), so never quantize below q8. 20 steps ≈ 50 steps visually (2.9× faster). 1024-res = hero tier. - **Performance (4 layers, 640, baked q8, 20 steps):** m3ultra **243 s** · m1ultra 366 s · 1024-res hero: 858 s (m3). Peak ~37GB. - **Baked models:** m3 `~/qwen-layered/qwen-layered-q8` · m1 `~/qwen-layered-staging/qwen-layered-q8` (34GB; rebake via `mflux-save`). - **Run:** `mflux-generate-qwen-layered --image IMG --layers 4 --resolution 640 --steps 20 --model-path ` · Farm: operator `qwen_layered_local`. Bonus: the subject layer is matting-grade — the premium tier of image preprocessing for hero assets. ## 4 · MODELBEAST integration | operator | backs | speed (m3 / m1) | notes | |---|---|---|---| | `trellis2_mlx` | 3.1 | 72 s / 182 s | vertex baker default; **preferred 3D lane** | | `trellis_mac` | torch-MPS port | 125 s gen / n/a | legacy; UV-textured (dark-patch bug) | | `hunyuan3d_mlx` | 3.2 | 243 s / ~5-6 min | textured GLB; runs everywhere | | `flux_local` | 3.3 | ~9 s/img | klein-4b 4-step default | | `qwen_layered_local` | 3.5 | 243 s / 366 s | RGBA layer decomposition | Routing lives in `nodes.json` (machine-local, gitignored — edit on the m3 primary). MESHGOD's `local/` lane calls these via the MB REST API. ## 5 · The lessons ledger (do not re-learn these) 1. **Profile before optimizing — every time.** The rung-3 3× came from a scheduler bug fix + baker swap the profiler exposed; the glamorous kernel work measured ≈ 0. The 75GB "sparse conv" memory monster was actually the Metal texture bake. 2. **mx.compile ≈ 0 on GEMM-bound 2B+ models on M3 Ultra** (measured 4×: Hunyuan T1, T2, TRELLIS step, sampler). It's free insurance, not a win. Compile wins live on small models (corridorkey 1.47×) and maybe small chips. 3. **Quantization floors are model-specific and cliff-shaped.** Qwen- Layered: q8 fine, q6 semantically broken, q4 noise — and quant does NOT buy speed on compute-bound models, only footprint. 4. **glTF axis + gamma:** pipeline space is Z-up in [-0.5,0.5]³; glTF is Y-up (`(x,y,z)→(x,z,−y)`); COLOR_0 is LINEAR (never pre-gamma). A "mangled blob" render is usually just the figure viewed down its own axis after a convention miss. 5. **xatlas hangs (hours) on écorché-class meshes** (thousands of disconnected shells). Vertex bake sidesteps UV unwrap entirely. 6. **macOS TCC blocks launchd from ~/Documents** — services live in `~/`. 7. **Gated HF weights don't need logins on workers** — rsync the model dirs from a cache that has them. 8. **The fal-outage class of bug:** any `while True` polling a paid API needs a deadline + a cancel; an in-memory queue wedges silently. 9. **Adapters that convert per-call hide O(steps) waste** — convert at stage boundaries, hold state in the fast runtime. 10. **Sentinel fallthroughs bite:** a baker/scheduler that silently falls back can burn 5 hours before anyone notices. Fail loudly instead. ## 6 · Runbooks **New box bring-up (any Apple Silicon, macOS 26+):** 1. Keys: box's SSH key added to Gitea (`monster` account). 2. `git clone` the MODELBEAST repo; run the relevant `scripts/install_*.sh` (trellis2_mlx, qwen_layered, hunyuan…) — each is idempotent, uses uv, and states its weight needs. 3. Weights: rsync the HF cache model dirs from m3ultra (see §5.7). 4. Bench: `uv run --with mlx python scripts/mlxprobe.py` → add the row to BENCHMARKS fleet matrix; set fp16/bf16 per §2 rules. 5. Register the box + operator list in `nodes.json` on the m3 primary. **Rollback:** every MRP patch is env-gated (§3 knobs) or a pinned-tag fork — `git checkout ` restores the adopted upstream exactly. **Where results go:** every experiment appends to `BENCHMARKS.md` (wall time, peak mem via `mx.get_peak_memory()`, parity numbers, and a verdict line). Session-scale context lives in `docs/FABLE_STARTUP_PACKET.md`. ## 7 · Roadmap (scoped, in value order) 1. **DiT-internals deep-dive** — TRELLIS.2 sampling runs ≈2.3× above ideal FLOPs (54 s of the 72). Attention/MLP kernel work inside the forward; the corridorkey `prompts/` phased method is the template. 2. **Fast UV bake** — bucket rasterizer (written, in fast_bake_test.py) + an écorché-safe unwrap → texture+MR parity with fal on the vertex baker's quality. Unlocks `--baker vertex`-quality WITH UV maps. 3. **simdgroup spconv kernel** — upgrade the opt-in scalar kernel if the decoders ever matter for time (they're 10 s today). 4. **Upstreaming** — our sampler + baker + validator to Jourloy PR#175; FLUX.1 compile backport as an mflux PR; spconv kernel as reference. 5. **MESHGOD default flip** — route its `local/` 3D lane to `trellis2_mlx` once finish-farm verts→texture baking lands for game budgets.