Fork of dgrauet/Hunyuan3D-2.1-mlx + our generate_e2e.py CLI, env-tunable remesh (HY3D_REMESH_FACES), and HARDWARE.md. Upstream training data (mini_trainset) and demo images stripped — inference needs none of it. Full upstream history: github.com/dgrauet/Hunyuan3D-2.1-mlx
3.8 KiB
Hunyuan3D 2.1 MLX — hardware & RAM guide
Which Apple Silicon Mac can run this pipeline, and how fast. It's a two-stage pipeline (Stage 1 shape → Stage 2 PBR texture); Stage 2 is the memory- and time-heavy half. Everything is fp16 MLX (Metal) — no bf16 dependency, so it runs on every Apple Silicon generation (M1 → M5), unlike torch-MPS ports.
Memory by precision
The pre-converted weights (monster/Hunyuan3D-2.2-mrp-MLX / upstream dgrauet/hunyuan3d-2.1-mlx) ship fp16. INT8/INT4 are available by self-converting (mlx-forge convert hunyuan3d-2.1 --quantize --bits 8).
| Precision | DiT | Stage 1 peak | Both stages peak | Min Mac | Recommended |
|---|---|---|---|---|---|
| FP16 (default) | 5.7 GB | ~10 GB | ~16–20 GB (measured 20.2 GB at texture 4096) | 24 GB (tight) | 32 GB |
| INT8 | 3.0 GB | ~6 GB | ~10–12 GB | 16 GB | 16–24 GB |
| INT4 | 1.6 GB | ~4 GB | ~8–10 GB | 16 GB | 16 GB |
- Stage 1 only (
--no-texture, geometry): ~10 GB fp16 — fits a 16 GB Mac comfortably. - Full pipeline (shape + PBR paint UNet + dual-stream ref UNet + VAE + DINO + RealESRGAN + bake): ~20 GB fp16 at our 4096 quality config → wants 32 GB.
- Texture resolution barely moves peak (2048 vs 4096 both ~20 GB) — it's the resident model stack, not the atlas, that dominates.
Measured performance (our fleet, mermaid test image)
| Machine | Config | Shape | Texture | Total | Peak |
|---|---|---|---|---|---|
| M3 Ultra 256 GB | 2048 / remesh 40k | 149 s | 112 s | 260 s | 20.2 GB |
| M3 Ultra 256 GB | 4096 / remesh 120k / octree 384 | 160 s | 221 s | 380 s | 20.2 GB |
| M1 Ultra 128 GB | 4096 / remesh 120k / octree 384 | 349 s | 402 s | 751 s | ~20 GB |
| M2 Pro (upstream ref) | 2048, 6-view @ 512 | — | ~9 min | — | — |
Takeaways: an Ultra does a hero-quality asset in 4–13 min. M-series is ~2× per tier (M1 Ultra ≈ 2× M3 Ultra here). Speed scales with GPU cores + bandwidth (Ultra > Max > Pro > base) and generation (M5 > M4 > M3 …), while fit is purely RAM.
What each Mac can do
| RAM | Hunyuan capability |
|---|---|
| 8 GB | ✗ — even INT4 both-stages (~8 GB) leaves no OS headroom. Not recommended. |
| 16 GB | ✓ INT8/INT4 full pipeline; ✓ fp16 Stage 1 only (--no-texture). Fast-ish on Pro/Max, slow on base. |
| 24 GB | ✓ fp16 full pipeline (tight — close other apps). The practical floor for textured fp16 output. |
| 32 GB | ✓ fp16 full pipeline comfortably, incl. our 4096 / 120k quality default. The recommended target. |
| 48 GB+ | ✓ headroom to keep it resident alongside other work; push remesh_faces 200k, max_num_view 8–12. |
| Ultra (96–256 GB) | ✓ trivially; batch several resident. Studio GPUs bake 4096² with no Metal command-buffer watchdog (verified M3 + M1 Ultra). |
Quality knobs vs cost (from our tuning)
| Knob | Fast draft | Studio default | Max |
|---|---|---|---|
texture_size |
2048 | 4096 | 4096 |
remesh_faces |
40000 | 120000 | 200000 |
octree_resolution |
256 | 384 | 384 |
max_num_view |
6 | 6 | 8–12 |
| → time (M3 Ultra) | ~260 s | ~380 s | longer |
| → GLB / faces | 7.6 MB / 40k | 21.5 MB / 120k | larger |
Raising remesh_faces from the laptop-tuned 40k is the biggest quality lever (it's what un-melts the face); 4096 texture + octree 384 sharpen further. All are CLI flags on generate_e2e.py and params on the MODELBEAST hunyuan3d_mlx operator. On a ≤16 GB Mac, stay at 2048 / 40k (or INT8) and expect the bake to be the tight part.
Bottom line
- Minimum to run textured output: 16 GB (INT8/INT4) or 24 GB (fp16, tight).
- Recommended: 32 GB fp16. Best experience: an M-series Max/Ultra with 48 GB+.
- Any M1 works — this is the fp16-MLX path, no bf16 gotcha.