From 2be37219f483b00d3b164eb29c5e14155c3fe34f Mon Sep 17 00:00:00 2001 From: modelbeast Date: Thu, 16 Jul 2026 14:00:41 +1000 Subject: [PATCH] HARDWARE.md: RAM/precision guide (fp16/int8/int4, min+recommended per Mac) --- HARDWARE.md | 58 +++++++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 58 insertions(+) create mode 100644 HARDWARE.md diff --git a/HARDWARE.md b/HARDWARE.md new file mode 100644 index 0000000..b5515af --- /dev/null +++ b/HARDWARE.md @@ -0,0 +1,58 @@ +# Hunyuan3D 2.1 MLX — hardware & RAM guide + +Which Apple Silicon Mac can run this pipeline, and how fast. It's a **two-stage** pipeline (Stage 1 shape → Stage 2 PBR texture); Stage 2 is the memory- and time-heavy half. Everything is **fp16 MLX** (Metal) — no bf16 dependency, so it runs on **every** Apple Silicon generation (M1 → M5), unlike torch-MPS ports. + +## Memory by precision + +The pre-converted weights (`monster/Hunyuan3D-2.2-mrp-MLX` / upstream `dgrauet/hunyuan3d-2.1-mlx`) ship **fp16**. INT8/INT4 are available by self-converting (`mlx-forge convert hunyuan3d-2.1 --quantize --bits 8`). + +| Precision | DiT | Stage 1 peak | Both stages peak | Min Mac | Recommended | +|---|---|---|---|---|---| +| **FP16** (default) | 5.7 GB | ~10 GB | **~16–20 GB** *(measured 20.2 GB at texture 4096)* | **24 GB** *(tight)* | **32 GB** | +| **INT8** | 3.0 GB | ~6 GB | ~10–12 GB | **16 GB** | 16–24 GB | +| **INT4** | 1.6 GB | ~4 GB | ~8–10 GB | **16 GB** | 16 GB | + +- **Stage 1 only** (`--no-texture`, geometry): ~10 GB fp16 — fits a **16 GB** Mac comfortably. +- **Full pipeline** (shape + PBR paint UNet + dual-stream ref UNet + VAE + DINO + RealESRGAN + bake): ~20 GB fp16 at our 4096 quality config → wants **32 GB**. +- Texture resolution barely moves peak (2048 vs 4096 both ~20 GB) — it's the resident model stack, not the atlas, that dominates. + +## Measured performance (our fleet, mermaid test image) + +| Machine | Config | Shape | Texture | Total | Peak | +|---|---|---|---|---|---| +| **M3 Ultra** 256 GB | 2048 / remesh 40k | 149 s | 112 s | **260 s** | 20.2 GB | +| **M3 Ultra** 256 GB | **4096 / remesh 120k / octree 384** | 160 s | 221 s | **380 s** | 20.2 GB | +| **M1 Ultra** 128 GB | 4096 / remesh 120k / octree 384 | 349 s | 402 s | **751 s** | ~20 GB | +| M2 Pro (upstream ref) | 2048, 6-view @ 512 | — | ~9 min | — | — | + +Takeaways: an **Ultra** does a hero-quality asset in 4–13 min. **M-series is ~2× per tier** (M1 Ultra ≈ 2× M3 Ultra here). Speed scales with GPU cores + bandwidth (Ultra > Max > Pro > base) and generation (M5 > M4 > M3 …), while *fit* is purely RAM. + +## What each Mac can do + +| RAM | Hunyuan capability | +|---|---| +| **8 GB** | ✗ — even INT4 both-stages (~8 GB) leaves no OS headroom. Not recommended. | +| **16 GB** | ✓ **INT8/INT4** full pipeline; ✓ fp16 **Stage 1 only** (`--no-texture`). Fast-ish on Pro/Max, slow on base. | +| **24 GB** | ✓ **fp16 full pipeline** (tight — close other apps). The practical floor for textured fp16 output. | +| **32 GB** | ✓ fp16 full pipeline comfortably, incl. our **4096 / 120k** quality default. The recommended target. | +| **48 GB+** | ✓ headroom to keep it resident alongside other work; push `remesh_faces` 200k, `max_num_view` 8–12. | +| **Ultra (96–256 GB)** | ✓ trivially; batch several resident. Studio GPUs bake 4096² with no Metal command-buffer watchdog (verified M3 + M1 Ultra). | + +## Quality knobs vs cost (from our tuning) + +| Knob | Fast draft | Studio default | Max | +|---|---|---|---| +| `texture_size` | 2048 | **4096** | 4096 | +| `remesh_faces` | 40000 | **120000** | 200000 | +| `octree_resolution` | 256 | **384** | 384 | +| `max_num_view` | 6 | 6 | 8–12 | +| → time (M3 Ultra) | ~260 s | ~380 s | longer | +| → GLB / faces | 7.6 MB / 40k | 21.5 MB / 120k | larger | + +Raising `remesh_faces` from the laptop-tuned 40k is the biggest quality lever (it's what un-melts the face); 4096 texture + octree 384 sharpen further. All are CLI flags on `generate_e2e.py` and params on the MODELBEAST `hunyuan3d_mlx` operator. On a **≤16 GB** Mac, stay at 2048 / 40k (or INT8) and expect the bake to be the tight part. + +## Bottom line + +- **Minimum to run textured output:** 16 GB (INT8/INT4) or 24 GB (fp16, tight). +- **Recommended:** 32 GB fp16. **Best experience:** an M-series Max/Ultra with 48 GB+. +- **Any M1 works** — this is the fp16-MLX path, no bf16 gotcha.