SCENEGOD/PLAN.md
type-two 86b76f6bdc [m6] Flow video plates spec + flow_intake.py; M6 queued behind M5
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-19 01:25:14 +10:00

247 lines
13 KiB
Markdown

# SCENEGOD master plan
Owner: John. Orchestrator: Fable (reviews `logs/`, updates `lanes/`).
Executors: three Opus 4.8 lane agents working in parallel on `main`.
## 1. What we're building
A browser page where John can:
- drag characters (rigged GLB/FBX), props (MeshGod gallery GLBs), and image
backdrops onto a lit 3D stage;
- drop Mixamo-style animation clips onto characters (client-side retarget,
ported from MESHGOD rigroom);
- keyframe entity transforms, camera moves + cuts, and lighting on a
timeline with a scrubber;
- save/load scenes as JSON;
- render the timeline to mp4 (draft: MediaRecorder webm; final: per-frame
PNG → server ffmpeg);
- later: audio tracks, MODELBEAST tts/voice-clone/lipsync, viseme lip sync
for blendshape characters, and a headless-Blender limb-graft script.
Non-goals (v1): multiplayer, physics, in-browser modeling, VPS deploy,
game-engine export. It renders video, it is not a game engine.
## 2. Repo layout (create as you go — no empty scaffolding)
```
scenegod/
server.py # FastAPI app (Lane C)
web/
index.html # page skeleton + importmap + mount divs (Lane A)
style.css # (Lane A; B may append a clearly-marked timeline block)
room3d.js # loaders + retarget, ported from rigroom (Lane A)
stage.js # Stage class: scene/entities/gizmos (Lane A)
dock.js # asset dock UI (Lane A)
timeline.js # Timeline class: clock, tracks, evaluation (Lane B)
tlui.js # timeline DOM/canvas UI (Lane B)
render.js # draft+final render client (Lane C)
scripts/
graft_limb.py # headless Blender limb graft (Lane C, M4)
test_server.py # endpoint smoke tests (Lane C)
requirements.txt # fastapi, uvicorn (Lane C)
lanes/ logs/ # orchestrator + agents
```
## 3. File ownership (hard rule)
| Lane | Owns | May read everything, edits ONLY its own files |
|---|---|---|
| A — Stage | `web/index.html`, `web/style.css`, `web/room3d.js`, `web/stage.js`, `web/dock.js` | |
| B — Timeline | `web/timeline.js`, `web/tlui.js` | appends to `style.css` only inside `/* === LANE B === */` block |
| C — Server | `scenegod/server.py`, `web/render.js`, `scripts/*`, `requirements.txt`, `.gitignore` | |
`PLAN.md` + `lanes/*` = orchestrator only. `logs/lane<X>.md` = its agent only.
Schema changes (§4): propose in your log, orchestrator amends PLAN.md, then
everyone follows the new version.
## 4. Contracts (the reason three lanes can run in parallel)
### 4.1 Scene JSON — schema v1
```jsonc
{
"version": 1,
"name": "street-scene",
"fps": 30,
"duration": 20.0, // seconds
"entities": [
{
"id": "e1", // unique, stable
"kind": "character", // character|prop|backdrop|camera|light
"label": "lady",
"source": { "type": "assets", // assets|upload|none (upload = client-side only, warn on
"path": "characters/pack01/lady.glb" }, // save; none = needs no bytes: camera/light)
"params": {}, // kind-specific, see below
"transform": { "pos": [0,0,0], "rot": [0,0,0], "scale": 1 }, // rest pose (used when no keys)
"tracks": {
"transform": [ // sorted by t; euler rad; ease: linear|in|out|inout|step
{ "t": 0, "pos": [0,0,0], "rot": [0,0,0], "scale": 1, "ease": "linear" }
],
"params": [ // keyable scalars/colors, e.g. camera fov, light intensity
{ "t": 0, "key": "fov", "value": 45, "ease": "inout" }
],
"clips": [ // characters only; non-overlapping except fade regions
{ "path": "animations/pack01/walk.fbx", "clipIndex": 0,
"start": 1.0, // timeline seconds where block begins
"in": 0.0, "out": 2.4, // trim within source clip
"loop": 2, // repeat count (1 = once)
"fade": 0.25 } // crossfade seconds into NEXT block
]
}
}
],
"cameraCuts": [ { "t": 0, "camera": "e7" } ], // active camera entity id from t onward
"audio": [ { "path": "audio/vo1.wav", "start": 0.0, "gain": 1.0 } ] // M4
}
```
`params` by kind — `camera`: `{fov}`; `light`: `{type:"key"|"ambient", color:"#fff",
intensity, castShadow}`; `backdrop`: `{mode:"plane"|"corner"|"dome", image:"backdrops/street.jpg",
width}`. Exactly one `light` entity of type `ambient` per scene (hemisphere);
`key` lights are directional, multiple allowed.
### 4.2 JS module contracts
`room3d.js` (Lane A) exports:
```js
parseAny(arrayBuffer, filename) -> Promise<{root, anims}> // glb|fbx|obj|bvh
canon(name) -> string
boneMap(root) -> {canonName: Bone}
captureRest(root) -> {rest: Map<Object3D,{wq,parent}>, order: Object3D[]}
bakeRetarget(clip, srcRoot, srcRest, tgtRoot, tgtRest) -> AnimationClip // parameterized target (rigroom closes over a global — fix that in the port)
disposeRoot(root); stats(root) -> {tris, meshes, bones}
```
`stage.js` (Lane A) exports `class Stage`:
```js
new Stage(viewportEl) // renderer, scene, IBL env, grid, orbit "director cam"
async addEntity(desc) -> entity // desc = scene-JSON entity; loads asset, builds object
removeEntity(id); getEntity(id); entities() -> entity[]
select(id|null); onSelect(cb); onChange(cb) // gizmo edits fire onChange(entity)
entityTransform(id) -> {pos,rot,scale} // current gizmo state (for "capture keyframe")
setTransform(id, {pos,rot,scale}) // timeline drives this every tick
setParam(id, key, value) // fov, light color/intensity, ...
prepareClip(id, path, clipIndex) -> Promise<AnimationClip> // fetch+parse+retarget, CACHED
entityMixer(id) -> AnimationMixer
setActiveCamera(id|null) // null = director orbit cam
renderActiveCamera(canvasOrNull) // PiP preview / final-render target
captureState() -> sceneJsonEntities; async applyState(sceneJson)
```
`timeline.js` (Lane B) exports `class Timeline`:
```js
new Timeline(stage) // talks to Stage ONLY via the API above
load(sceneJson); toJSON() -> sceneJson // owns duration/fps/tracks/cameraCuts
time; playing; play(); pause(); seek(t); onTick(cb(t))
evaluate(t) // interpolate transform+param keys → stage.setTransform/…
// activate clip blocks → mixer weights/time, crossfades
// camera cuts → stage.setActiveCamera
step(frame) // deterministic: seek(frame/fps)+evaluate — used by final render
```
Until Lane A ships Stage (M1), Lane B develops against `stagestub.js`
(Lane B writes it, same API, logs to console) — delete at integration.
### 4.3 HTTP endpoints (Lane C)
```
GET / stageroom page (no-cache) + all /web statics
GET /assets/tree {characters:[...], animations:[...], props:[...], backdrops:[...], audio:[...]}
scans SCENEGOD_ASSETS folder tree live (rigroom-library style:
dir+stem grouping, no index file); props may also merge the
MESHGOD gallery if MESHGOD_GALLERY_URL is set (later, M3)
GET /assets/file?path= streams one asset (path-traversal-guarded)
GET /scenes list; GET /scenes/{name} one
POST /scenes/{name} save scene JSON (validated: version, ids unique, sorted tracks)
POST /render/begin {fps,width,height,name} -> {renderId}
POST /render/{id}/frame/{n} raw PNG body
POST /render/{id}/end {audio?: [...]}: ffmpeg PNG seq (+audio M4) -> mp4
GET /render/{id}/status ; GET /render/{id}/out.mp4
```
Env: `SCENEGOD_ASSETS` (default `./assets`), `SCENEGOD_SCENES` (default
`./scenes`), `SCENEGOD_PORT` (8020). John symlinks/rsyncs his existing
banks into `assets/` — the server never writes into it except `/render`.
## 5. Milestones & sync points
**M1 — walking skeleton** (all lanes, independent):
- A: index.html + Stage with IBL viewer, load a GLB character + an image
backdrop from the dock, gizmo move/rotate/scale, retarget port proven
(load fbx clip onto character, plays once via a temp button).
- B: Timeline UI against stagestub — tracks list, scrubber, add/drag/delete
transform keyframes, play/pause with interpolation visible in console.
- C: server serving the page + /assets/tree + /assets/file + /scenes CRUD,
with scripts/test_server.py passing.
- **SYNC 1** (orchestrator): wire B to real Stage, delete stub. Definition of
done: load scene JSON → character walks between two keyframed positions
in front of a street backdrop, scrubbed by the timeline.
**M2 — direction**: A: multiple cameras + PiP + light entities + param
plumbing. B: clip blocks with trim/loop/crossfade, params track, camera-cut
track, snapping, undo (simple command stack). C: scene validation hardening,
asset thumbnails endpoint. **SYNC 2**: the "lady + man + street + 2 cameras +
sunset" demo scene plays end to end.
**M3 — render**: C: render session endpoints + ffmpeg + render.js client
(draft MediaRecorder + final frame-step using `timeline.step`). A/B: whatever
SYNC 2 revealed + polish pass from orchestrator notes. **SYNC 3**: 10-second
1080p mp4 rendered deterministically.
**M4 — sound & flesh**: C: audio in scene JSON + ffmpeg mux + MODELBEAST
tts/voice-clone submit (mb contract from meshgod/ops.py, token from
`~/Documents/backnforth/.env`, NEVER logged) + `scripts/graft_limb.py`.
B: audio tracks on the timeline UI, Web Audio sync. A: viseme lane — detect
morph targets, Rhubarb-JSON → morph keyframes, greyed out when absent.
**M5 — DIRECTOR MODE (ease of use)**: classic cinema grammar as one-click
presets, LLM strictly optional on top.
- Shared vocabulary in `web/grammar.js` (Lane A owns, data only): shot sizes
ECU/CU/MCU/MS/MWS/WS/EWS (target vertical span as a fraction of subject
height + frame-center height), angles (eye/low/high/dutch/OTS-L/OTS-R),
focal lengths 18/24/35/50/85mm (vertical fov = 2·atan(12/f), full-frame),
lighting presets (day, golden-hour, night, noir, horror, neon, overcast),
blocking marks (center, two-shot L/R, walk-to-mark, face-camera, face-other).
- Lane A `web/presets.js`: `frameSubject(camId, subjectId, shot, angle,
focalmm)` — deterministic solver from the subject's bbox (no LLM), light/
blocking preset appliers, and the DIRECT panel UI (shot buttons on the
selected subject, light preset buttons, mark buttons).
- Lane B: preset actions land as normal keys/cuts AT THE PLAYHEAD through
existing mutators (undo-able); "walk to mark" = 2 transform keys.
- Lane C: `POST /director` — optional NL box ("medium two-shot, golden hour,
lady walks to camera") → tailnet Ollama (OpenAI-compat, env
`SCENEGOD_LLM_URL` + `SCENEGOD_LLM_MODEL`, 503 when unset) → strict JSON
ops validated server-side → client applies via the SAME preset functions.
Buttons must never depend on the LLM lane.
**M6 — FLOW VIDEO PLATES (rear-projection era)**: Google-Flow-generated mp4s
as first-class stage media. Feed: John's flowext rinse drops mp4s in
`~/Downloads/flowrinse/<category>/`; `scripts/flow_intake.py` copies them to
`assets/backdrops/video/<category>/` with poster jpgs.
- Lane A: backdrop `params.mode` gains `"video"` (and existing modes accept
.mp4 sources): `THREE.VideoTexture` on plane/corner/dome, muted, `loop`,
`params.videoStart` offset. NEW entity kind `screen`: a video-textured
plane with a slight emissive boost + optional CRT curve, for TVs/club
screens/jumbotrons INSIDE the set (djsim_screens bank is made for these).
Poster jpg shown until play.
- Lane B: video backdrops/screens follow the master clock — on evaluate,
`video.currentTime = (t - videoStart) % duration` only when drift >50ms
(cheap scrub-follow); play/pause with the clock.
- Lane C: render path must be deterministic — finalRender already steps
frame-by-frame; stage exposes `await stage.syncVideos(t)` (seek + wait
for 'seeked') that render.js awaits per frame before capture. Video audio
is NEVER used (scene audio[] is the only audio source).
## 6. How John launches the lanes
Three Claude Code sessions (Opus 4.8), cwd `~/Documents/SCENEGOD`, one prompt
each — see README-LANES section at the bottom of each lane file, or just:
> You are Lane A (or B / C). Read CLAUDE.md, PLAN.md, lanes/A-stage.md and
> logs/laneA.md, then continue the work from where the log leaves off. Obey
> the file-ownership map and the log protocol. Stop at your current
> milestone boundary.
Orchestrator loop (Fable, this session or any MESHGOD/SCENEGOD session):
read `logs/*.md` → answer BLOCKED/REQUESTS by editing `lanes/*.md` (append a
dated **ORCHESTRATOR UPDATE** section) → run the SYNC integration when all
three logs report their milestone DONE.