Go to file
type-two d5217575eb wip(corridorkey_local): operator scaffold — BLOCKED on GVM inference crash
Full staging/zip/cleanup contract works; generate-alphas dies in GVM with
'MPSNDArray buffer is not large enough. Must be 59768832 bytes' on BOTH
m1ultra and m3ultra, any input dims, torch or mlx backend, gpu or cpu post.
GVM --device cpu wrote 0 frames in 3.5h (hung). Weights installed both nodes
(gvm_core/weights, 6GB). Next: either debug GVM's torch-MPS path or replace
the alpha-hint stage with the corridorkey-mlx fork / bg_remove_local hints.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-18 22:30:39 +10:00
.claude Phase 0 baseline: FastAPI+SQLite operator pipeline, React/three.js UI, ffprobe/ffmpeg_frames/blender_convert operators 2026-07-12 21:05:22 +10:00
docs runner: resolve relative operator python paths against repo root (multi-node portability) 2026-07-14 11:47:39 +10:00
scripts install_acestep: torchcodec for torchaudio>=2.9 save path 2026-07-17 19:47:03 +10:00
server wip(corridorkey_local): operator scaffold — BLOCKED on GVM inference crash 2026-07-18 22:30:39 +10:00
tests Security hardening from adversarial review (5-agent workflow) 2026-07-13 12:25:13 +10:00
web Unified worker pool: gpu lane dispatches across M3 + M1 nodes (HANDOFF2 phase B) 2026-07-14 18:18:09 +10:00
.gitignore README: local-first rewrite + ignore the .env.* family 2026-07-17 22:25:08 +10:00
AGENTS.md hunyuan3d_mlx operator (local MLX image→3D) + cluster docs 2026-07-16 10:10:43 +10:00
BENCHMARKS.md perfcheck: nightly fleet GPU canary + the contention lessons that shaped it 2026-07-17 18:03:03 +10:00
CLUSTER.md docs: register the M1 Max / M4 mini / M1 mini nodes in CLUSTER + HARDWARE 2026-07-17 22:24:36 +10:00
CORRIDORKEY.md CORRIDORKEY.md: status, Ultra tuning findings, video->3D use-case map 2026-07-16 21:50:56 +10:00
HANDOFF2.md HANDOFF2: execution brief — multi-user auth (guests local-only), dashboard, VPS hosting 2026-07-13 11:10:27 +10:00
HANDOFF_HY3D_MLX.md hunyuan3d_mlx operator (local MLX image→3D) + cluster docs 2026-07-16 10:10:43 +10:00
HANDOFF.md fal panel: 5 new operators, image outputs, grouped UI, no-input ops + review fixes 2026-07-12 21:51:16 +10:00
HARDWARE.md docs: register the M1 Max / M4 mini / M1 mini nodes in CLUSTER + HARDWARE 2026-07-17 22:24:36 +10:00
mb Security hardening from adversarial review (5-agent workflow) 2026-07-13 12:25:13 +10:00
PERFCHECK.md perfcheck: nightly fleet GPU canary + the contention lessons that shaped it 2026-07-17 18:03:03 +10:00
PLAN.md Phase 0 baseline: FastAPI+SQLite operator pipeline, React/three.js UI, ffprobe/ffmpeg_frames/blender_convert operators 2026-07-12 21:05:22 +10:00
pyproject.toml Auth + dashboard (HANDOFF2 phases 1-2): guests local-only, system snapshot 2026-07-13 12:06:52 +10:00
README.md README: local-first rewrite + ignore the .env.* family 2026-07-17 22:25:08 +10:00
uv.lock Auth + dashboard (HANDOFF2 phases 1-2): guests local-only, system snapshot 2026-07-13 12:06:52 +10:00

MODELBEAST

Local-first generation farm for a fleet of Macs. Video, image, audio, text or 3D in → meshes, splats, images, video, speech, music, motion and text out. Everything below runs on your own machines — nothing leaves the tailnet unless you explicitly pick a fal_* or openrouter_* operator, and guests can't pick those at all.

🖥️ Multi-Mac fleet? Read CLUSTER.md first — the local-first policy, node roles, and how the central queue routes jobs across machines. 💾 What runs on which Mac? HARDWARE.md — min/recommended RAM per operator across every Apple Silicon tier (8GB → 256GB). 📊 Is the fleet healthy / how fast is it? PERFCHECK.md — nightly GPU canary + per-node capability table. BENCHMARKS.md — what each model actually costs. 🎬 Green-screen keying → 3D: CORRIDORKEY.md — CorridorKey status, our Ultra tuning, and the video→cutout→mesh/splat use-case map.

See PLAN.md for the verified tool matrix and roadmap, HANDOFF.md for the agent build brief.

Run

/opt/homebrew/bin/uv run uvicorn server.main:app --host 0.0.0.0 --port 8777

Open http://localhost:8777 (or http://<tailscale-ip>:8777 from any device on the tailnet). Prefer scripts/serve.sh — it runs headless and prints the one-time owner password on first start.

Drag/drop/paste any video, image, audio or 3D file; or drop files into data/inbox/ for auto-ingest. Pick an operator, tune its parameters, run. Several operators (flux_local, comfyui_sd, tts_local, llm_local, music_local, motion_local) need no input asset — just a prompt. Use ⊞ Compare to view outputs side by side, and 📊 Dashboard for a live CPU/RAM/GPU/queue snapshot.

Accounts & sharing

Login is required. The first scripts/serve.sh creates an owner (monster) and prints a one-time password to data/server.log — change it in ⚙ Settings → Users (or uv run python scripts/users.py passwd monster). The owner can add guest accounts for friends: guests are local-only — they can run only on-device operators (never the owner's paid keys, enforced server-side), see only their own assets/jobs, and have a per-user concurrent-job cap. So mates cost you electricity, nothing else. Public hosting via a VPS + Caddy is in docs/VPS.md; the CUDA-worker offload plan is in docs/CUDA-WORKER.md.

Headless / agents

Everything is driveable without the UI via the zero-dependency mb CLI (works from any tailnet machine) or raw REST. AGENTS.md is the complete handover brief — endpoints, mb reference, operator catalog, recipes, job semantics, etiquette. Server ops: scripts/serve.sh (start/restart headless), scripts/install_launchagent.sh (optional boot persistence, owner-run).

export MB_HOST=http://100.89.131.57:8777 MB_TOKEN=mbt_...   # token from ⚙ Settings → Users
./mb run flux_local -p prompt="a brass astrolabe on velvet" --wait --download out/
./mb run bg_remove_local --file photo.jpg --wait --download out/
./mb run hunyuan3d_mlx --file cutout.png --wait --download out/

Operators

Each subfolder of server/operators/ with a manifest.json is an operator. The UI auto-renders its parameter form from params_schema (JSON Schema) and filters by the selected asset's kind (accepts).

Local — free, offline, on-fleet

This is the default. Reach for cloud only when local genuinely can't do it.

id lane what
image → 3D
hunyuan3d_mlx gpu image → GLB + PBR (Hunyuan3D 2.1, native MLX). Weights public — no HF login. The only local 3D op that runs on M1. ~4.3 min
trellis_mac gpu image → GLB + PBR (TRELLIS.2 MPS port). Best local quality — sharper than hunyuan at defaults. ~5 min [HF-gated]
sf3d gpu image → GLB in ~5s. Draft tier — good on solid objects, struggles on thin/open geometry [HF-gated]
image gen / edit
flux_local gpu prompt → image via mflux/MLX. FLUX.2 Klein 4B is Apache + ungated, ~9s/image
comfyui_sd gpu prompt → image via local SD/SDXL checkpoints + LoRAs (ComfyUI on Metal). flux_local cannot load SD/SDXL — this is the one that can
mflux_image_edit gpu image + instruction → edited image (Qwen-Image-Edit, ungated)
seedvr2_upscale gpu image → faithful upscale (SeedVR2, MIT) — sharpens without inventing detail
bg_remove_local gpu image → subject cutout (RMBG-2.0). The single biggest quality lever before any image→3D
video
wan_video gpu prompt → MP4, or image + prompt → MP4 (Wan 2.2 TI2V-5B on the resident ComfyUI). 24fps, 1280×704
audio / speech
tts_local gpu text → speech (Kokoro-82M, ~50 voices, Apache, ~350MB)
stt_local gpu audio → transcript + timestamps (Whisper large-v3-turbo on MLX)
voice_clone_local gpu ~10s reference voice + text → that voice speaking it (Chatterbox, MIT)
music_local gpu style tags (+ optional lyrics) → music WAV (ACE-Step 3.5B, Apache)
lipsync_local cpu speech WAV → viseme timing JSON (Rhubarb, fully offline)
text / motion
llm_local gpu prompt → text (Qwen3-30B-A3B-4bit MoE on MLX)
motion_local cpu text → character motion as BVH + mp4 preview (MoMask)
scan → splat
ffmpeg_frames cpu video → frames (fps sampling, mpdecimate dedupe, blur cull)
colmap_poses cpu frames → camera poses + sparse cloud (COLMAP 4.x + GLOMAP)
brush_train gpu colmap dataset → gaussian splat .ply (Brush, native Metal)
utility
ffprobe cpu media inspection
blender_convert cpu GLB/GLTF/OBJ/FBX/USD/STL/PLY/BLEND ⇄ conversion via headless Blender

Recommended local image→3D chain: bg_remove_local → (seedvr2_upscale if thin/blurry) → trellis_mac (quality) or hunyuan3d_mlx (faster, ~3× lighter, runs on M1).

Cloud — costs real money, opt-in, owner-only

Use when local can't, or for on-demand SOTA without the local wait.

id lane what
fal_trellis2 / fal_hunyuan3d_v21 / fal_hunyuan3d / fal_trellis / fal_rodin net image → mesh via fal.ai [FAL_KEY]
fal_bg_remove / fal_upscale / fal_image_edit / fal_text_image net cutout / upscale / nano-banana edit / Ideogram v3 (readable text) [FAL_KEY]
openrouter_image net prompt → image via OpenRouter (nano-banana family; exact cost reported per job) [OPENROUTER_API_KEY]

Note: fal's Hunyuan v21 multi-view is broken (verified 2026-07) — v21 is single-image only; use v2 for multi-view. fal_rodin is the film/hero tier (quad topology, ~$0.40+).

Setup for the gated local models (one-time)

Most local operators work out of the box. Only these need anything:

  • sf3d, trellis_mac — weights are HuggingFace-gated. Accept the licences (stabilityai/stable-fast-3d, facebook/dinov3-vitl16-pretrain-lvd1689m), then huggingface-cli login or paste an HF token in ⚙ Settings.
  • bg_remove_local — briaai/RMBG-2.0 licence, same deal.
  • hunyuan3d_mlxnothing. Weights are public.
  • Cloud operators — FAL_KEY / OPENROUTER_API_KEY in ⚙ Settings.

Heavy tools live in vendor/ (repos) + venvs/ (per-tool envs), both gitignored — reinstall with scripts/install_*.sh.

Develop

  • Backend: server/ — FastAPI + SQLite (data/modelbeast.db), job runner runs operators as subprocesses across concurrency lanes. gpu and cpu are a node pool: gpu = 1 job per node, cpu = per-node slots (cpu_slots in nodes.json; default 3 local / 2 remote), net = 6 on the primary only. Settings/secrets in server/settings.py (env-injected, log-redacted).
  • Frontend: web/ — React + Vite + three.js + @mkkellogg/gaussian-splats-3d. After editing: cd web && npm run build (the server serves web/dist).
  • Data: data/assets/ (store), data/jobs/ (job workdirs), data/inbox/ (watch folder). Delete data/ to reset. Tests use MODELBEAST_DATA=<tmp>.
  • Tests: ./tests/smoke.sh (12 framework checks).

Writing an operator

Contract: the runner invokes <python> run.py --input <asset> --outdir <jobdir> --params '<json>'. Write outputs into the outdir; optionally write result.json ({"outputs": [{"path": ..., "name": ..., "meta": ...}]}) to control what gets registered as assets. stdout/stderr become the job log; the runner stamps [node: <name>] as log line 1.

Manifest fields: id, name, category, description, accepts, produces, resources (gpu/cpu/net lane — defaults to cpu if omitted), requires_env (gates the operator until the key is set), python (venv path for heavy tools — repo-relative, so remote nodes resolve it), and params_schema. Operators can tag output asset kind via result.json output meta.kind (e.g. splat, colmap_dataset).

Remaining roadmap (object_capture, freemocap, retargeting, character APIs, workflow presets, LLM copilot) is in HANDOFF.md §58.