- flux_local: prompt->image fully on-device via mflux (schnell/dev, steps, quantize, seed, size). Both FLUX repos are HF-gated as of 2026-07 (schnell included) — clean GatedRepoError surfaces with a hint; needs owner HF token. - openrouter_image: prompt->image via OpenRouter chat/completions with modalities [image,text]; parses data-URL images from the response; model picker for nano-banana / nano-banana-pro. Gated on OPENROUTER_API_KEY. - settings: openrouter_key added to vault (masked/redacted/env-injected) - scripts/install_mflux.sh; venv installed - A/B flow: run both with the same prompt, judge in Compare mode Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|---|---|---|
| .claude | ||
| scripts | ||
| server | ||
| tests | ||
| web | ||
| .gitignore | ||
| BENCHMARKS.md | ||
| HANDOFF.md | ||
| PLAN.md | ||
| pyproject.toml | ||
| README.md | ||
| uv.lock | ||
MODELBEAST
Local-first web app that turns videos, images, and 3D files into meshes, splats, mocap, and rigged characters on the M3 Ultra. See PLAN.md for the full verified tool matrix and roadmap, and HANDOFF.md for the agent build brief (phases 1–4 instructions).
Run
/opt/homebrew/bin/uv run uvicorn server.main:app --host 0.0.0.0 --port 8777
Open http://localhost:8777 (or http://<tailscale-ip>:8777 from any device on the tailnet).
Drag/drop/paste any video, image, or 3D file; or drop files into data/inbox/ for auto-ingest. Pick an operator, tune its parameters, run. Add API keys / HuggingFace token under ⚙ Settings. Use ⊞ Compare to view several outputs side by side.
Develop
- Backend:
server/— FastAPI + SQLite (data/modelbeast.db), job runner runs operators as subprocesses across concurrency lanes (gpu=1, cpu=3, net=6). Settings/secrets inserver/settings.py(env-injected, log-redacted). - Frontend:
web/— React + Vite + three.js +@mkkellogg/gaussian-splats-3d. After editing:cd web && npm run build(the server servesweb/dist). - Data:
data/assets/(store),data/jobs/(job workdirs),data/inbox/(watch folder). Deletedata/to reset. Tests useMODELBEAST_DATA=<tmp>. - Tests:
./tests/smoke.sh(12 framework checks). Benchmarks in BENCHMARKS.md. - Heavy tools live in
vendor/(repos) +venvs/(per-tool envs), both gitignored. Reinstall withscripts/install_*.sh.
Setup for the local mesh generators (one-time)
sf3d and trellis_mac are installed but their weights are HuggingFace-gated. Accept the licenses (stabilityai/stable-fast-3d, facebook/dinov3-vitl16-pretrain-lvd1689m, briaai/RMBG-2.0), then huggingface-cli login or paste an HF token in Settings. Cloud fal_* operators need FAL_KEY in Settings.
Operators
Each subfolder of server/operators/ with a manifest.json is an operator. The UI auto-renders its parameter form from params_schema (JSON Schema) and filters by the selected asset's kind (accepts).
Contract: the runner invokes <python> run.py --input <asset> --outdir <jobdir> --params '<json>'. Write outputs into the outdir; optionally write result.json ({"outputs": [{"path": ..., "name": ..., "meta": ...}]}) to control what gets registered as assets. stdout/stderr become the job log.
Manifest fields: id, name, category, description, accepts, produces, resources (gpu/cpu/net lane), requires_env (gates the operator until the key is set), python (absolute venv path for heavy tools), params_schema. Operators can tag output asset kind via result.json output meta.kind (e.g. splat, colmap_dataset).
Current operators:
| id | lane | what |
|---|---|---|
ffprobe |
cpu | media inspection |
ffmpeg_frames |
cpu | video → frames (fps, mpdecimate dedupe, blur cull, max cap) |
blender_convert |
cpu | universal 3D format conversion via headless Blender |
colmap_poses |
cpu | frames → camera poses + sparse cloud (COLMAP 4.x + GLOMAP) |
brush_train |
gpu | colmap dataset → 3D gaussian splat (Brush, native Metal) |
sf3d |
gpu | image → GLB locally (Stable-Fast-3D, MPS) [HF-gated] |
trellis_mac |
gpu | image → GLB+PBR locally (TRELLIS.2 MPS port) [HF-gated] |
fal_trellis / fal_trellis2 / fal_hunyuan3d / fal_hunyuan3d_v21 / fal_rodin |
net | image → mesh via fal.ai API [needs FAL_KEY] |
fal_bg_remove |
net | image → subject cutout (BiRefNet v2) — run before any image→3D for a big quality jump |
fal_upscale |
net | image → faithful upscale (SeedVR) |
fal_image_edit |
net | image + instruction → edited image (nano-banana) |
fal_text_image |
net | prompt → image (Ideogram v3, readable text) — no input asset needed |
Recommended fal chain for best image→3D: fal_bg_remove → (fal_upscale if thin) → fal_trellis2 or fal_hunyuan3d_v21. Note: Hunyuan v21 multi-view is broken on fal (verified 2026-07) — v21 is single-image only; use v2 for multi-view.
Heavy operators get their own uv venv ("python": "<abs venv path>" in the manifest). Remaining roadmap (object_capture, freemocap, retargeting, tripo/meshy character APIs, workflow presets, LLM copilot) is in HANDOFF.md §5–8.