vidgod/README.md
type-two 8c84243b61 phase 2: vg-remove (ProPainter), vg-interp (RIFE), vg-cutie, farm ops, zoo mirror
- vg-remove: object/logo/watermark removal via ProPainter on MPS; static --box,
  SAM2-tracked --point for moving objects, or user --mask. Output always scaled
  back to source dims (imageio macro-block-pads ProPainter output).
- vg-interp: RIFE frame interpolation via rife-ncnn-vulkan (universal binary,
  native Metal/MoltenVK, rife-v4.6); smooth (fps x N) or --slowmo.
- vg-cutie: Cutie interactive segmentation GUI launcher (local GUI session).
- setup/fetch_phase2.sh: idempotent clones + weights + deps + patches.
- patches: propainter-cv2-reader (torchvision >= 0.23 removed read_video),
  cutie-device (get_default_model hard-coded .cuda(); now cuda->mps->cpu).
- smoke_test.sh: adds the RIFE lane (skips when not fetched).
- Farm: vidgod_roto/vidgod_index operators live in MODELBEAST (8965d22),
  verified from JING5; weights mirrored to NAS modelzoo/vidgod-weights.

All lanes verified on ultra 2026-08-24: de-logo reconstruction eyeballed clean,
24->48fps interp, Cutie headless propagation PASS, smoke test 4/4.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-24 15:42:52 +10:00

85 lines
5.1 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# VIDGOD
Open-source AI video production toolkit for music-video editing around DaVinci Resolve.
All inference is local on Apple Silicon (PyTorch MPS + MLX) — no cloud, no subscriptions.
Built for the workflow: hoard old footage → normalize it → index every shot and every
spoken word → search it → arrange to a beat grid in Resolve → cut actors out with clicks
instead of rotoscoping splines.
## The tools
| Command | What it does |
|---|---|
| `bin/vg-roto CLIP --point x,y` | Click an actor → tracked cutout as **ProRes 4444 with alpha** (SAM 2.1 + MatAnyone). `--grab-frame` saves the prompt frame so you can find coords. `--mode mask` for hard binary masks via SAM2 video propagation. |
| `bin/vg-transcode DIR` | Batch-convert old AVI/WMV/MPEG/etc → ProRes LT, mirrored dir tree, auto-deinterlace, skips already-done files. `--h264` for small proxies. |
| `bin/vg-index PATHS` | Shot-detect (PySceneDetect) + transcribe (mlx-whisper large-v3-turbo) + thumbnail everything into `library/clips.sqlite` (FTS5). Resumable. |
| `bin/vg-find WORDS` | Full-text search the dialogue; shows the enclosing shot. `--cut DIR` exports each matching shot as a ProRes clip. `--shots VIDEO` lists a video's shots. |
| `bin/vg-beats SONG` | Beat-track a music file (librosa) → `.beats.csv` + a **Resolve marker EDL** (import: right-click timeline in media pool → Timelines → Import → Timeline Markers from EDL). |
| `bin/vg-remove CLIP --box x1,y1,x2,y2` | Erase logos/watermarks/objects (ProPainter inpainting). `--point x,y` makes SAM2 track a *moving* object first; `--mask` takes your own. |
| `bin/vg-interp CLIP --factor 2` | RIFE frame interpolation (native Metal binary): smooth motion at fps×factor, or `--slowmo` for factor× slow motion. |
| `bin/vg-cutie --video CLIP` | Cutie interactive segmentation GUI — scribble on the actor, propagate, correct frame by frame; exports per-frame masks (run from a GUI session, not SSH). |
## Install
Needs: Apple Silicon Mac, `brew install ffmpeg uv`, ~4GB disk for venvs + models.
```bash
git clone ssh://git@100.71.119.27:222/monster/vidgod.git && cd vidgod
./setup/setup_venvs.sh # venvs, SAM2 + MatAnyone clones, checkpoints, patches (~10 min)
./setup/fetch_phase2.sh # ProPainter + Cutie + RIFE: clones, weights, deps, patches
./setup/smoke_test.sh # fast lanes; add --roto for the full cutout pipeline
```
Everything heavy (venvs/, tools/, models/, library/, work/, tests/) lives untracked inside
the working copy; the repo is just the scripts. Long jobs heartbeat to `~/.jobs/vidgod-*.status`.
## Typical session
```bash
bin/vg-transcode ~/old-tv-rips # → ~/old-tv-rips-prores
bin/vg-index ~/old-tv-rips-prores # overnight for a big archive
bin/vg-find "radical" --cut ~/mv/pulls # grab every shot where someone says it
bin/vg-beats ~/mv/song.mp3 --fps 25 # markers for the Resolve timeline
bin/vg-roto ~/mv/pulls/003_*.mov --grab-frame # find your click coords
bin/vg-roto ~/mv/pulls/003_*.mov --point 970,220 # → ProRes 4444 alpha cutout
```
## Notes & gotchas
- **MatAnyone patch**: torchvision ≥ 0.23 removed `torchvision.io.read_video`; setup applies
`patches/matanyone-cv2-reader.patch` to its clone automatically. If you re-clone
`tools/MatAnyone` by hand, re-run `setup/setup_venvs.sh`.
- MatAnyone's checkpoint comes from its GitHub release (the HF repo doesn't carry the
file at the root path); its inference script insists on `pretrained_models/` inside its
repo dir, so setup symlinks it there.
- MatAnyone propagates **forward from a first-frame mask**`vg-roto --frame N` trims the
clip at N first. Prompt on a frame where the target is clearly visible.
- Scripts set `/opt/homebrew/bin` in PATH themselves (non-interactive SSH doesn't) and
`PYTORCH_ENABLE_MPS_FALLBACK=1` for stray unsupported MPS kernels.
- Rough speed on an M1 Ultra: MatAnyone ≈ 3.4 fr/s at 1920×804; whisper large-v3-turbo
transcribes far faster than realtime; ProRes transcodes are ffmpeg-bound.
## Farm integration (MODELBEAST)
Two operators live in the MODELBEAST repo (`server/operators/`): **`vidgod_roto`**
(gpu lane — video + click points → the three roto outputs as job assets) and
**`vidgod_index`** (cpu lane — index a clip into the node's library; the video is
copied to `library/ingest/` so entries persist). They're thin stdlib wrappers that
shell out to this repo's `bin/` CLIs, gated by manifest `requires_path` so only nodes
with `~/Documents/VIDGOD` set up will take the jobs (allowlisted on ultra). Submit
from any fleet box with its own MB token; outputs come back through the assets API.
## Weights mirror
All checkpoints (SAM2.1, MatAnyone, Cutie, ProPainter, the RIFE bundle) are mirrored
to the NAS at `/Volumes/projects/modelzoo/vidgod-weights/` — restore instead of
re-downloading when standing up a new box.
## Maybe later
Practical-RIFE v4.25+ (newer models than the ncnn bundle), Depth Anything V2 depth-matte
lane, DDColor colorization, a vg-roto mode that takes Cutie mask dirs end-to-end.
Upscaling already exists fleet-side (`seedvr2_upscale` on the farm); Resolve's Optical
Flow retime covers basic slow-mo without RIFE.