modelbeast/CLUSTER.md
m3ultra 9b525fa420 hunyuan install clones self-owned MLX fork; CLUSTER fleet final state
- install_hunyuan3d_mlx.sh: auto-clone monster/Hunyuan3D-2.2-mrp-MLX (our
  Gitea fork) into vendor/ when missing, so any node bootstraps identically.
- CLUSTER.md: M1 hunyuan verified, M4 = bg_remove + Ollama qwen2.5:7b endpoint,
  cpu-lane-not-pooled caveat, self-owned repo reference.
2026-07-16 10:35:11 +10:00

4.5 KiB

MODELBEAST cluster — read this first

Policy: LOCAL FIRST. This is a multi-Mac render farm on Tailscale. Every generation should run on our own silicon ($0, private, unlimited) unless a cloud model is genuinely better for the job. Reach for fal_* cloud operators only when no local operator covers the need (e.g. Rodin hero-grade meshes, the multi-view volume bargain) — never as the default. Cost of a local gen: nothing. Cost of a fal gen: real money (see BENCHMARKS.md for per-call prices).

The order of preference for any task:

  1. Local operator on whichever node is free (the queue picks it — see below).
  2. fal cloud only for capabilities local can't match yet.

The fleet

Node Host / Tailscale IP RAM Role
M3 Ultra (primary) m3ultra.local · 100.89.131.57 256 GB Runs the server + the central queue on :8777. Handles everything, especially the heavy 3D models (hunyuan3d_mlx, trellis_mac) and FLUX.
M1 Ultra (worker) ultra.local · 100.91.239.7 128 GB GPU worker in the pool and its own standalone instance. MLX-native — the right node for hunyuan3d_mlx (trellis's torch-MPS is unverified here). Also FLUX, sf3d, bg_remove_local.
M4 Pro (helper) m4pro.local · 100.69.21.128 24 GB GPU-pool helper for light ops (bg_remove_local, verified) + a standalone Ollama LLM (qwen2.5:7b at http://100.69.21.128:11434). Not the big 3D/diffusion models — too little RAM.

Git origin (shared by all nodes): ssh://git@100.71.119.27:222/monster/modelbeast.git

The central queue (how tasks route to the right machine)

The M3 primary runs one server with per-lane concurrency:

  • gpu lane = a pool of nodes. Each Mac has one Metal GPU, so each node runs one gpu job at a time. When gpu jobs queue up, the primary hands each to the first free node that supports that operator — rsyncing inputs over SSH, running the operator's run.py on that node, and rsyncing outputs back. This is the "unified queue": submit to M3, it load-balances across M3 + M1 (+ M4 for light ops).
  • cpu / net lanes run on the primary (M3) with simple concurrency limits (cpu 3, net 6). Note: only the gpu lane is pooled today — cpu-lane ops (ffmpeg_frames, colmap_poses, ffprobe) run on M3 only, so a helper node can offload gpu work but not CPU work yet. Pooling the cpu lane is a worthwhile follow-up.

Ollama LLM (M4): a standalone Qwen2.5-7B server runs on M4 at http://100.69.21.128:11434 (OpenAI-compatible at /v1), reachable across the tailnet. It is not a MODELBEAST operator — point any copilot/captioning client straight at that endpoint. Swap models with ollama pull <model> on M4.

Which node may run which operator is set in nodes.json (repo root, gitignored, primary-only) via each node's operators allowlist. Omit the allowlist to let a node run every gpu op. nodes.json is read at server startup — restart the server (scripts/serve.sh) after editing it.

Adding a machine to the fleet

  1. git clone ssh://git@100.71.119.27:222/monster/modelbeast.git ~/MODELBEAST
  2. Install uv (curl -LsSf https://astral.sh/uv/install.sh | sh) — the venvs use it.
  3. Run the per-operator install scripts you want that node to serve (scripts/install_*.sh). Each builds its own vendor/*/.venv. (install_hunyuan3d_mlx.sh auto-clones our self-owned MLX fork — monster/Hunyuan3D-2.2-mrp-MLX on the Gitea — into vendor/, so every node runs the same patched code.)
  4. Model weights: copy from a node that already has them (rsync -a ~/.cache/huggingface/hub/models--… newnode:~/.cache/huggingface/hub/) rather than re-downloading — faster and dodges HF rate limits. We have an HF token in Settings if a fresh download is unavoidable.
  5. On the primary, add the node to nodes.json with an operators allowlist matched to its RAM, then restart the server.

Operator cheat-sheet (local first!)

  • image → 3D mesh: local hunyuan3d_mlx (MLX, M1-capable, no HF login) or trellis_mac (sharper at defaults) → sf3d for fast drafts. fal only for Rodin hero-grade / multi-view. Always bg_remove_local (or fal_bg_remove) the input first.
  • prompt/image → image: local flux_local (FLUX.2 Klein) → mflux_image_edit. fal/OpenRouter for nano-banana scene-storytelling.
  • video → frames / poses: ffmpeg_frames, colmap_poses (CPU — great M4 work).
  • Full matrix + prices: AGENTS.md · BENCHMARKS.md.