modelbeast/docs/TRELLIS2_MLX_RECON.md
2026-07-19 00:22:41 +10:00

107 lines
6.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# TRELLIS.2 on Apple Silicon — recon (2026-07-18)
**Verdict up front:** we already run TRELLIS.2 locally — `vendor/trellis-mac`
(torch-MPS + Metal kernels) IS TRELLIS.2-4B and is wired as the
`trellis_mac` operator. There is no "MLX interface" shortcut for CUDA code
(CUDA kernels must be *rewritten*, not wrapped), but ~90% of TRELLIS.2 is
standard tensor math that MPS/MLX already runs; only four custom CUDA
pieces matter, and the community has replaced all four. The open work is
**speed** (unfused sparse ops ≈ 10× slower than CUDA) and **quality gaps**
(hole-filling disabled, forced pre-simplification) — not feasibility.
## What we run today (`vendor/trellis-mac`, shivampkumar fork)
- torch-MPS + `PYTORCH_ENABLE_MPS_FALLBACK`, Metal kernels by @pedronaugusto
(mtlgemm / mtldiffrast / mtlbvh / mtlmesh), SDPA attention.
- Full inference: sparse structure → shape SLat → tex SLat → decode →
dual-grid mesh extract → simplify → Metal PBR bake → GLB.
- Our benchmark: **318 s/gen on m3ultra**, ~18 GB peak (M4 Pro 24GB: 5m13s
cold at 512). H100 does 317 s — the gap is almost entirely the unfused
sparse conv + padded attention.
- Known gaps vs fal's trellis-2: CuMesh skipped → **no hole filling**,
meshes pre-simplified ~858K→200K faces before baking, Metal BVH
instability, macOS GPU-watchdog can kill long kernels (detected +
workarounds printed by generate.py). macOS 26 needed for the metallib
(fleet is on 26.5 ✓).
- Licensing: DINOv3 is Meta-gated; **RMBG-2.0 preprocessing is CC BY-NC**.
## What TRELLIS.2 actually needs (upstream: Linux, CUDA 12.4, ≥24 GB)
| CUDA dep | Role | Mac status |
|---|---|---|
| o-voxel ext (hash/convert/rasterize) | the O-Voxel representation | pure-Python reimpl (trellis-mac `backends/`) + CPU fork |
| FlexGEMM (**Triton**) | all sparse conv | **the crux** — Triton ≠ Metal; Metal `mtlgemm` or slow pure-torch gather/scatter |
| flash-attn / xformers | 4B flow transformer attention | SDPA (padded → unfused, big cost) |
| CuMesh | decimate/remesh/**hole-fill**/UV | skipped → `fast_simplification`; hole-fill lost |
| nvdiffrast | texture bake | `mtldiffrast` (Metal) |
| nvdiffrec | preview renders only | not needed |
## The MLX question, answered
- **No shim exists or can exist**: CUDA kernels are NVIDIA-machine code;
"an MLX interface" means rewriting each custom op in MLX/Metal. MLX can
express them (`mx.fast.metal_kernel` JIT-compiles Metal from Python) —
but MLX has **no sparse-tensor type and no sparse-voxel precedent**;
SparseTensor/varlen semantics must be hand-rolled. A TRELLIS.2-MLX would
be a first.
- **Our own playbook (proven 2×) is adopt-then-patch, not from-scratch**:
hunyuan3d-mlx = dgrauet's 8.4k-LOC port + our thin packaging layer;
corridorkey-mlx = cmoyates/Niko's 5.3k-LOC port + our **8-line**
`mx.compile`-in-tiled-mode patch = the 1.47× win. From-scratch dense→MLX
ports of this scale are months of solo work (mflux, mlx-video authors).
- **Starting points already exist**: upstream **PR #175** (Jourloy,
2026-07-17 — MPS + Metal + an *experimental `mlx` backend flag*, 28
tests green on M4 Max) and **pedronaugusto/trellis2-apple** (an
`mlx_backend/` dir, no benchmarks yet).
## Recommended ladder (effort-ordered)
1. **Hours — tune what we have**: benchmark `trellis_mac` 1024_cascade vs
fal trellis-2 on identical inputs (BENCHMARKS.md format); route
MESHGOD's batch/overnight work to the local lane (m3ultra clears 18 GB
~14× over; fal stays for interactive one-offs). $93/mo → mostly $0.
2. **Days — adopt + fleet-patch**: vendor PR #175 / trellis2-apple as
`monster/trellis2-*` Gitea forks (house pattern); run the corridorkey
ablation moves on them: attention head_dim → pad to 64 fast-path,
`mx.compile` on fixed shapes, sdpa gating by GPU generation, tiled-vs-
full sweeps. Prize: **M1 Ultra compatibility** (MLX-native, like
hunyuan3d — today trellis is m3-only in practice) = 2nd free 3D box.
3. **Weeks — the real kernel work** (only if we want fal-class speed):
fused Metal sparse-conv (gather-GEMM-scatter) + varlen attention via
`mx.fast.metal_kernel`; port CuMesh hole-filling. Closes most of the
10× gap; genuinely novel, upstreamable to PR #175.
4. **Quality parity misc**: raise simplification budget on 256 GB boxes,
swap RMBG-2.0 → our licensed bg-remove lane, wire `trellis_mac` into
MESHGOD's `local/` model list next to hunyuan3d-mlx.
## Sources
- github.com/microsoft/TRELLIS.2 (setup.sh = dep manifest) · PR #175 ·
issue #74 · shivampkumar/trellis-mac · pedronaugusto/trellis2-apple
- ml-explore.github.io/mlx custom-metal-kernels docs
- Local: HANDOFF_HY3D_MLX.md, CORRIDORKEY.md, BENCHMARKS.md,
vendor/trellis-mac/README, vendor/corridorkey-mlx/prompts/ (the 6-phase
parity-first port template — the blueprint if we ever do ladder step 3)
## First benchmark (2026-07-19, m3ultra, anatomy écorché test image)
| run | tris | time | peak mem | cost |
|---|---|---|---|---|
| trellis_mac 1024-cascade, tex 2048 | 191,336 | **124.8 s** compute (248.8 s wall, 103 s load) | 20.8 GB | $0 |
| fal trellis-2, same image | 468,049 | 298.6 s wall (queue incl.) | — | $0.30 |
**Local already beats fal on wall-clock.** Quality is the gap, not speed:
the local master shows shattered dark patches — CuMesh hole-fill skipped +
forced 858k→191k pre-simplify before baking (the un-simplified master IS
saved as .obj alongside the GLB). Fix quality first, then fuse kernels.
Tri-budget ladder (Blender decimate on the fal master): 2k = confetti
(ratio floors at 0.01 → 4,674 min), 8k = torn, 30k = good, 50k ≈ master.
**Écorché-class meshes (many disconnected thin shells) cannot pure-decimate
below ~30k** — game budgets need the MESHGOD solidify/remesh route.
Discovered in passing: MESHGOD finish_glb.py's tri budget is silently
ignored on Blender 5 (modifier_apply cancels — fix task spawned 2026-07-19).
Bench artifacts: m3ultra ~/Documents/trellis2-bench/ (masters, sweeps,
renders, logs, sweep2.csv / sweep_fal.csv).