pool cpu lane across nodes; hunyuan Studio-quality defaults

runner: gpu AND cpu are now a node pool. Per-node cpu_slots (primary 3,
helpers 2, nodes.json-overridable); net stays primary-only. Python-less ops
(ffmpeg/ffprobe) now run remotely on the node's system python3. Verified: 9
concurrent ffmpeg_frames distributed 4 local / 2 m1 / 2 m4.

hunyuan3d_mlx: default to Studio-quality (octree 384, texture 4096, remesh
120k) — 4096 bake verified watchdog-free on M3 Ultra; big quality gain
(defined face, 120k faces). remesh_faces now a param. All param-overridable.
This commit is contained in:
m3ultra 2026-07-16 10:59:44 +10:00
parent 9b525fa420
commit 06f33be865
5 changed files with 58 additions and 26 deletions

View File

@ -45,7 +45,10 @@ First-run downloads (one-time): schnell/dev ~31GB & ~1619 min each, Klein 9B
| `trellis_mac` | 318s | 23MB | 175,842 | **sharper** — defined eyes/nose/mouth, individually raised tail scales, vivid colors | | `trellis_mac` | 318s | 23MB | 175,842 | **sharper** — defined eyes/nose/mouth, individually raised tail scales, vivid colors |
| `hunyuan3d_mlx` | 260s | 7.6MB | 40,000 | softer — melted face, smoothed scales, muted texture | | `hunyuan3d_mlx` | 260s | 7.6MB | 40,000 | softer — melted face, smoothed scales, muted texture |
**Verdict:** at defaults `trellis_mac` wins on quality (crisper face + geometry, richer color); `hunyuan3d_mlx` is faster, ~3× lighter, and the only local 3D op that runs on M1. Hunyuan's softness is largely its **laptop-tuned 40k-face decimation** — on a Studio, raise `HY3D_REMESH_FACES` (100200k) and try `texture_size 4096` / `octree_resolution 384` (Phase D) before judging it beaten. `sf3d` stays the ~5s draft tier. All free/offline; fal cloud for on-demand SOTA without the local wait. **Verdict:** at defaults `trellis_mac` wins on quality (crisper face + geometry, richer color); `hunyuan3d_mlx` is faster, ~3× lighter, and the only local 3D op that runs on M1. `sf3d` stays the ~5s draft tier. All free/offline; fal cloud for on-demand SOTA without the local wait.
### Phase D — hunyuan Studio tuning (2026-07-16, M3 Ultra) → new operator defaults
Raised the config from the laptop-tuned defaults to `octree_resolution 384` + `remesh_faces 120000` + `texture_size 4096`. **4096² bake works on the Studio GPU — no Metal command-buffer watchdog** (the existing `extract_textiles` tiling handles it; `uv_feature_map` never needed patching). Result: **380s** (shape 160 + tex 221), peak 20.2GB, 21.5MB GLB, 78k verts / **120k faces**, 4096² baseColor+MR. Quality jump is real — the melted face gains defined eyes + structure, tail geometry sharpens, textures crisper; closes most of the gap to trellis (trellis still edges the face). Cost: ~46% slower + ~3× file size vs the 40k/2048 default. **These are now the `hunyuan3d_mlx` operator defaults** (all still param-overridable; drop to `remesh_faces 40000`/`texture_size 2048` for fast drafts). 4096 on the M1 Ultra: verifying.
## Mesh-gen — earlier install notes (superseded by the table above) ## Mesh-gen — earlier install notes (superseded by the table above)
| Operator | Install | Runtime status | | Operator | Install | Runtime status |

View File

@ -12,7 +12,7 @@ The order of preference for any task:
|---|---|---|---| |---|---|---|---|
| **M3 Ultra** (primary) | `m3ultra.local` · 100.89.131.57 | 256 GB | Runs the server + the central queue on :8777. Handles everything, especially the heavy 3D models (`hunyuan3d_mlx`, `trellis_mac`) and FLUX. | | **M3 Ultra** (primary) | `m3ultra.local` · 100.89.131.57 | 256 GB | Runs the server + the central queue on :8777. Handles everything, especially the heavy 3D models (`hunyuan3d_mlx`, `trellis_mac`) and FLUX. |
| **M1 Ultra** (worker) | `ultra.local` · 100.91.239.7 | 128 GB | GPU worker in the pool **and** its own standalone instance. MLX-native — the right node for `hunyuan3d_mlx` (trellis's torch-MPS is unverified here). Also FLUX, `sf3d`, `bg_remove_local`. | | **M1 Ultra** (worker) | `ultra.local` · 100.91.239.7 | 128 GB | GPU worker in the pool **and** its own standalone instance. MLX-native — the right node for `hunyuan3d_mlx` (trellis's torch-MPS is unverified here). Also FLUX, `sf3d`, `bg_remove_local`. |
| **M4 Pro** (helper) | `m4pro.local` · 100.69.21.128 | 24 GB | GPU-pool helper for light ops (`bg_remove_local`, verified) + a standalone **Ollama LLM** (`qwen2.5:7b` at `http://100.69.21.128:11434`). **Not** the big 3D/diffusion models — too little RAM. | | **M4 Pro** (helper) | `m4pro.local` · 100.69.21.128 | 24 GB | Helper for light gpu ops (`bg_remove_local`) + **cpu ops** (`ffmpeg_frames`, `ffprobe`; 2 slots) + a standalone **Ollama LLM** (`qwen2.5:7b` at `http://100.69.21.128:11434`). **Not** the big 3D/diffusion models — too little RAM. |
Git origin (shared by all nodes): `ssh://git@100.71.119.27:222/monster/modelbeast.git` Git origin (shared by all nodes): `ssh://git@100.71.119.27:222/monster/modelbeast.git`
@ -20,7 +20,8 @@ Git origin (shared by all nodes): `ssh://git@100.71.119.27:222/monster/modelbeas
The M3 primary runs one server with per-lane concurrency: The M3 primary runs one server with per-lane concurrency:
- **`gpu` lane = a pool of nodes.** Each Mac has one Metal GPU, so each node runs **one** gpu job at a time. When gpu jobs queue up, the primary hands each to the first *free* node that supports that operator — rsyncing inputs over SSH, running the operator's `run.py` on that node, and rsyncing outputs back. This is the "unified queue": submit to M3, it load-balances across M3 + M1 (+ M4 for light ops). - **`gpu` lane = a pool of nodes.** Each Mac has one Metal GPU, so each node runs **one** gpu job at a time. When gpu jobs queue up, the primary hands each to the first *free* node that supports that operator — rsyncing inputs over SSH, running the operator's `run.py` on that node, and rsyncing outputs back. This is the "unified queue": submit to M3, it load-balances across M3 + M1 (+ M4 for light ops).
- **`cpu` / `net` lanes** run on the primary (M3) with simple concurrency limits (`cpu` 3, `net` 6). **Note:** only the gpu lane is pooled today — cpu-lane ops (`ffmpeg_frames`, `colmap_poses`, `ffprobe`) run on M3 only, so a helper node can offload *gpu* work but not CPU work yet. Pooling the cpu lane is a worthwhile follow-up. - **`cpu` lane is also pooled** (since 2026-07-16): cpu ops distribute across nodes too — each node runs `cpu_slots` at once (primary 3, helpers 2 by default; set per node in nodes.json). A cpu op routes to a node only if it's in that node's allowlist **and** the node has the needed binary — `ffmpeg_frames`/`ffprobe` need ffmpeg (M1+M4 have it); `colmap_poses`/`blender_convert` stay on nodes with COLMAP/Blender. Python-less ops run on the node's system `python3`.
- **`net` lane** stays primary-only (the cloud API keys live only on M3), concurrency 6.
**Ollama LLM (M4):** a standalone Qwen2.5-7B server runs on M4 at `http://100.69.21.128:11434` (OpenAI-compatible at `/v1`), reachable across the tailnet. It is not a MODELBEAST operator — point any copilot/captioning client straight at that endpoint. Swap models with `ollama pull <model>` on M4. **Ollama LLM (M4):** a standalone Qwen2.5-7B server runs on M4 at `http://100.69.21.128:11434` (OpenAI-compatible at `/v1`), reachable across the tailnet. It is not a MODELBEAST operator — point any copilot/captioning client straight at that endpoint. Swap models with `ollama pull <model>` on M4.

View File

@ -13,8 +13,9 @@
"properties": { "properties": {
"steps": {"type": "integer", "default": 50, "description": "Stage 1 shape denoising steps"}, "steps": {"type": "integer", "default": 50, "description": "Stage 1 shape denoising steps"},
"guidance_scale": {"type": "number", "default": 7.5, "description": "Stage 1 CFG scale"}, "guidance_scale": {"type": "number", "default": 7.5, "description": "Stage 1 CFG scale"},
"octree_resolution": {"type": "integer", "enum": [256, 384], "default": 256, "description": "Stage 1 marching-cubes grid (384 = finer geometry, slower)"}, "octree_resolution": {"type": "integer", "enum": [256, 384], "default": 384, "description": "Stage 1 marching-cubes grid (384 = finer geometry)"},
"texture_size": {"type": "integer", "enum": [1024, 2048, 4096], "default": 2048, "description": "Stage 2 baked UV atlas resolution"}, "texture_size": {"type": "integer", "enum": [1024, 2048, 4096], "default": 4096, "description": "Stage 2 baked UV atlas resolution (4096 verified on Studio GPUs)"},
"remesh_faces": {"type": "integer", "default": 120000, "description": "Stage 2 bake-mesh face budget (Studio default 120k; 40k = laptop/fast)"},
"max_num_view": {"type": "integer", "default": 6, "description": "Stage 2 number of baked views"}, "max_num_view": {"type": "integer", "default": 6, "description": "Stage 2 number of baked views"},
"seed": {"type": "integer", "default": 42, "description": "Random seed"}, "seed": {"type": "integer", "default": 42, "description": "Random seed"},
"no_texture": {"type": "boolean", "default": false, "description": "Geometry only (skip Stage 2, much faster)"} "no_texture": {"type": "boolean", "default": false, "description": "Geometry only (skip Stage 2, much faster)"}

View File

@ -30,8 +30,9 @@ cmd = [
"--output", str(out_stem.resolve()), "--output", str(out_stem.resolve()),
"--steps", str(p.get("steps", 50)), "--steps", str(p.get("steps", 50)),
"--guidance", str(p.get("guidance_scale", 7.5)), "--guidance", str(p.get("guidance_scale", 7.5)),
"--octree-resolution", str(p.get("octree_resolution", 256)), "--octree-resolution", str(p.get("octree_resolution", 384)),
"--texture-size", str(p.get("texture_size", 2048)), "--texture-size", str(p.get("texture_size", 4096)),
"--remesh-faces", str(p.get("remesh_faces", 120000)),
"--max-num-view", str(p.get("max_num_view", 6)), "--max-num-view", str(p.get("max_num_view", 6)),
"--seed", str(p.get("seed", 42)), "--seed", str(p.get("seed", 42)),
] ]

View File

@ -8,8 +8,9 @@ Operator contract:
{"outputs": [{"path": "rel/or/abs", "name": "...", "meta": {...}}, ...], {"outputs": [{"path": "rel/or/abs", "name": "...", "meta": {...}}, ...],
"summary": {...}} "summary": {...}}
- if result.json is absent, every top-level file in outdir is registered - if result.json is absent, every top-level file in outdir is registered
Concurrency lanes (from manifest "resources"): gpu=1 (Metal contention), Concurrency lanes (from manifest "resources"): gpu and cpu are a node pool
cpu=3, net=6. Env vars from settings are injected; "requires_env" gates a job. (distributed across the primary + remote workers gpu 1/node, cpu N/node);
net=6 stays primary-only. Env from settings is injected; "requires_env" gates.
""" """
import asyncio import asyncio
import json import json
@ -21,8 +22,10 @@ from pathlib import Path
from . import db, remote, settings as settings_mod, store from . import db, remote, settings as settings_mod, store
JOBS_DIR = db.DATA / "jobs" JOBS_DIR = db.DATA / "jobs"
# gpu is a NODE POOL (this Mac + remote workers), not a plain semaphore; cpu/net # gpu AND cpu are a NODE POOL (this Mac + remote workers): gpu runs 1 job/node
# are semaphores. Each Mac has ONE Metal device, so each gpu node runs one job. # (one Metal device each), cpu runs per-node slots — LANE_LIMITS["cpu"] on the
# primary, fewer on helpers (nodes.json "cpu_slots" overrides). net stays a
# primary-only semaphore because the cloud API keys live only on the primary.
LANE_LIMITS = {"cpu": 3, "net": 6} LANE_LIMITS = {"cpu": 3, "net": 6}
@ -141,8 +144,10 @@ class Runner:
for key in LANE_LIMITS: for key in LANE_LIMITS:
self.lanes[key] = asyncio.Semaphore(LANE_LIMITS[key]) self.lanes[key] = asyncio.Semaphore(LANE_LIMITS[key])
self.gpu_nodes = remote.load_gpu_nodes() self.gpu_nodes = remote.load_gpu_nodes()
for n in self.gpu_nodes:
n.setdefault("inflight", {"gpu": 0, "cpu": 0})
names = ", ".join(n.get("name", n.get("ssh", "?")) for n in self.gpu_nodes) names = ", ".join(n.get("name", n.get("ssh", "?")) for n in self.gpu_nodes)
print(f"[runner] gpu pool: {names}", flush=True) print(f"[runner] node pool (gpu 1/node + cpu distributed): {names}", flush=True)
# re-queue jobs left running/queued by a previous process # re-queue jobs left running/queued by a previous process
for row in con.execute("SELECT id FROM jobs WHERE status IN ('queued','running')"): for row in con.execute("SELECT id FROM jobs WHERE status IN ('queued','running')"):
con.execute("UPDATE jobs SET status='queued' WHERE id = ?", (row["id"],)) con.execute("UPDATE jobs SET status='queued' WHERE id = ?", (row["id"],))
@ -164,8 +169,8 @@ class Runner:
await self._update(con, job_id, status="cancelled", finished_at=db.now()) await self._update(con, job_id, status="cancelled", finished_at=db.now())
return return
lane = self.lane_of(job["operator"]) lane = self.lane_of(job["operator"])
if lane == "gpu": if lane in ("gpu", "cpu"): # pooled lanes — distribute across the node pool
node = await self._acquire_gpu_node(job["operator"]) node = await self._acquire_node(job["operator"], lane)
try: try:
if job_id in self.cancelled: if job_id in self.cancelled:
self.cancelled.discard(job_id) self.cancelled.discard(job_id)
@ -175,9 +180,9 @@ class Runner:
except Exception as e: except Exception as e:
await self._update(con, job_id, status="error", error=str(e), finished_at=db.now()) await self._update(con, job_id, status="error", error=str(e), finished_at=db.now())
finally: finally:
node["busy"] = False self._release_node(node, lane)
return return
async with self.lanes[lane]: async with self.lanes[lane]: # net: cloud keys live on the primary — not pooled
if job_id in self.cancelled: if job_id in self.cancelled:
self.cancelled.discard(job_id) self.cancelled.discard(job_id)
await self._update(con, job_id, status="cancelled", finished_at=db.now()) await self._update(con, job_id, status="cancelled", finished_at=db.now())
@ -187,21 +192,39 @@ class Runner:
except Exception as e: except Exception as e:
await self._update(con, job_id, status="error", error=str(e), finished_at=db.now()) await self._update(con, job_id, status="error", error=str(e), finished_at=db.now())
async def _acquire_gpu_node(self, op_id: str) -> dict: def _node_capacity(self, node: dict, lane: str) -> int:
"""Return the first free gpu node that supports this operator (local always """Max concurrent jobs of this lane a node may run. gpu=1 (one Metal
does; a remote must list it + be reachable). Polls gpu jobs are coarse.""" device per Mac); cpu=per-node slots (small helper nodes get fewer),
overridable per node in nodes.json via "cpu_slots"; net is not pooled
(cloud keys live only on the primary)."""
if lane == "gpu":
return 1
if lane == "cpu":
default = LANE_LIMITS["cpu"] if node.get("local") else 2
return int(node.get("cpu_slots", default))
return 0
async def _acquire_node(self, op_id: str, lane: str) -> dict:
"""Return the first node with free capacity in this lane that supports the
operator (local always does; a remote must list it + be reachable). Polls."""
while True: while True:
for n in self.gpu_nodes: for n in self.gpu_nodes:
if n["busy"] or not remote.node_supports(n, op_id): if not remote.node_supports(n, op_id):
continue continue
n["busy"] = True # reserve synchronously — the health check below cap = self._node_capacity(n, lane)
# awaits, and without reserving first two jobs could grab one node if cap <= 0 or n["inflight"].get(lane, 0) >= cap:
continue
n["inflight"][lane] = n["inflight"].get(lane, 0) + 1 # reserve before the await
if not n.get("local") and not await remote.healthy(n): if not n.get("local") and not await remote.healthy(n):
n["busy"] = False # unreachable right now; release and keep looking n["inflight"][lane] -= 1 # unreachable right now; release and keep looking
continue continue
return n return n
await asyncio.sleep(0.25) await asyncio.sleep(0.25)
def _release_node(self, node: dict, lane: str) -> None:
if node is not None and "inflight" in node:
node["inflight"][lane] = max(0, node["inflight"].get(lane, 0) - 1)
async def _run_job(self, con, job_id: str, node: dict | None = None): async def _run_job(self, con, job_id: str, node: dict | None = None):
job = self.get_job(con, job_id) job = self.get_job(con, job_id)
if not job: if not job:
@ -249,10 +272,13 @@ class Runner:
cmd += ["--input", p] cmd += ["--input", p]
run_env = env run_env = env
else: # remote node: rsync inputs out, run over ssh, rsync outputs back else: # remote node: rsync inputs out, run over ssh, rsync outputs back
if not op_py_rel or Path(op_py_rel).is_absolute(): if op_py_rel and Path(op_py_rel).is_absolute():
await self._update(con, job_id, status="error", finished_at=db.now(), await self._update(con, job_id, status="error", finished_at=db.now(),
error="operator has no repo-relative python — can't run remotely") error="operator has an absolute python path — can't run remotely")
return return
# python-less ops (ffmpeg/ffprobe wrappers) run on the node's system
# python3, resolved via the remote login shell (bash -lc in remote_cmd).
remote_python = op_py_rel or "python3"
try: try:
remote_inputs = await remote.prepare(node, job_id, local_inputs) remote_inputs = await remote.prepare(node, job_id, local_inputs)
except Exception as e: except Exception as e:
@ -260,7 +286,7 @@ class Runner:
error=f"remote prepare failed: {e}") error=f"remote prepare failed: {e}")
return return
entry_rel = os.path.relpath(str(entry), str(db.ROOT)) entry_rel = os.path.relpath(str(entry), str(db.ROOT))
cmd = remote.remote_cmd(node, op["id"], op_py_rel, entry_rel, cmd = remote.remote_cmd(node, op["id"], remote_python, entry_rel,
job_id, params_json, remote_inputs) job_id, params_json, remote_inputs)
run_env = os.environ.copy() run_env = os.environ.copy()