22 KiB
Festival 4D — Build Instructions
Audience: Claude Opus 4.8 executing autonomously in this repo (/Users/m3ultra/Documents/festifun).
Goal: Turn multiple fan-shot smartphone videos of the same concert into a synchronized, explorable 4D experience: time-aligned multi-video playback, 3D scene reconstruction with camera poses, a free-roam "god's eye" viewer, AR-style overlays projected onto each video, and AI-tagged moments on a shared timeline.
This is an entertainment/replay project (same family as sports "bullet time" replay). Keep all naming, prompts, and UI copy in the concert domain.
Parallel execution: this spec is the canonical reference; the work is organized into lanes for multiple concurrent agents — see
plan/README.mdfor the phase map, branch protocol, file ownership, and your lane brief. If you were given a lane brief, it overrides the serial ordering note below (your lane's milestones still apply verbatim).
0. Ground rules
- Build milestones in dependency order. Each milestone has acceptance criteria; do not start the next until the current one passes. When executing serially (single agent), go M0→M9; when executing in lanes, follow
plan/. - Work without real footage. The user may not have multi-camera concert videos yet. Every milestone must be verifiable with the synthetic fixtures described in M0. When real footage arrives, the same pipeline runs on it unchanged.
- Prefer boring tech. SQLite (not PostgreSQL — this is a local, single-user, offline-first tool; do not add a DB server dependency). Vanilla JS + Three.js via Vite (no React). FastAPI + uvicorn. Python 3.11+, managed with
uvif available, else venv + pip. - COLMAP is an external binary. Detect it at runtime (
shutil.which("colmap")); if missing, print install instructions (brew install colmapon macOS) and — critically — the whole app must still run using the synthetic-poses fixture. Never make COLMAP a hard import-time dependency. - Commit per milestone with a message describing what now works. Initialize git first (
git init).
1. Repository layout
festifun/
├── README.md # setup + usage, written last (M9)
├── pyproject.toml # backend deps: fastapi, uvicorn, numpy, scipy, librosa,
│ # soundfile, sqlalchemy, pydantic, opencv-python-headless,
│ # google-genai, anthropic, openai (last three = classifier providers)
├── backend/
│ ├── festival4d/
│ │ ├── __init__.py
│ │ ├── cli.py # `python -m festival4d <cmd>`: ingest | sync | reconstruct | events | serve
│ │ ├── config.py # paths, constants
│ │ ├── db.py # SQLAlchemy engine/session, schema
│ │ ├── ingest.py # probe videos (ffprobe), extract audio, register in DB
│ │ ├── audio_sync.py # GCC-PHAT pairwise offsets + global solve
│ │ ├── frames.py # sharpness-aware frame sampling for SfM
│ │ ├── sfm.py # COLMAP orchestration, model export, normalization, interpolation
│ │ ├── geometry.py # quaternion slerp, ray casting, triangulation, coord conversions
│ │ ├── events_ai.py # audio candidate detection + Claude classification
│ │ ├── api.py # FastAPI app
│ │ └── synthetic.py # test-fixture generator (see M0)
│ └── tests/ # pytest; every numeric module gets a test
├── frontend/
│ ├── index.html
│ ├── package.json # vite + three
│ └── src/
│ ├── main.js
│ ├── state.js # global store: manifest, playhead, playing, selected camera
│ ├── transport.js # master clock + per-video sync (M4)
│ ├── videoGrid.js # <video> grid + overlay canvases
│ ├── scene3d.js # Three.js scene: point cloud, frusta, orbit controls
│ ├── overlays.js # 3D→2D anchor projection (M6)
│ ├── timeline.js # scrubber + event markers
│ ├── annotate.js # bbox drawing + submission (M8)
│ └── camPath.js # keyframed camera paths (M9)
└── data/
├── raw/ # user drops source videos here (gitignored)
├── work/ # extracted wavs, frames, colmap workspace (gitignored)
└── project.db # SQLite (gitignored)
2. Data model (SQLite via SQLAlchemy)
videos(id, filename, duration_s, fps, width, height, offset_ms REAL, drift_ppm REAL,
sync_confidence REAL, created_at)
camera_poses(id, video_id FK, frame_idx INT, t_video_s REAL,
qw,qx,qy,qz, tx,ty,tz, -- COLMAP convention: world→camera
fx,fy,cx,cy, -- pinhole intrinsics after undistortion
registered BOOL) -- False = interpolated, not solved
anchors(id, label, x,y,z, color) -- 3D points to overlay (stage corners etc.)
events(id, t_global_s REAL, duration_s REAL, event_type TEXT, confidence REAL,
description TEXT, source TEXT) -- source: 'audio_auto' | 'ai' | 'user'
annotations(id, video_id FK, t_video_s REAL, x0,y0,x1,y1, -- normalized [0,1] bbox
resolved_anchor_id FK NULL)
Timebase convention (write this in config.py docstring and stick to it everywhere):
t_global is the master timeline in seconds. For video v: t_video = (t_global - offset_ms/1000) * (1 + drift_ppm * 1e-6). The reference video has offset 0.
3. Milestones
M0 — Scaffold + synthetic fixtures
- Repo layout above, deps installed,
python -m festival4d --helpworks,pytestruns, Vite dev server serves a hello page. synthetic.pygenerates a full fake project so every later layer is testable without footage:- Fake videos: use ffmpeg to render 3 short (20 s, 640×360, 30 fps) clips of the same procedurally animated scene from different angles — simplest: a
testsrc2/drawtextcomposite is fine, but each clip must share a common audio track (a music-like WAV you synthesize with numpy: beat pulses + tones) shifted by known offsets (e.g. 0 ms, +1370 ms, −842 ms) and with distinct visuals. Store ground-truth offsets in a JSON next to them. - Fake poses: 3 camera trajectories on arcs around origin looking at a fake "stage" (a box of 3D points at the origin) + a sparse point cloud (random points on the box + ground plane). Write them into
camera_posesand apoints.plyin the exact same format the real COLMAP path will produce.
- Fake videos: use ffmpeg to render 3 short (20 s, 640×360, 30 fps) clips of the same procedurally animated scene from different angles — simplest: a
- Accept:
python -m festival4d syntheticpopulatesdata/raw+ DB; tests pass.
M1 — Ingest + audio sync
ingest: ffprobe each file indata/raw, insert intovideos, extract mono 16 kHz WAV per video (ffmpeg) intodata/work/audio/.sync(inaudio_sync.py):- Use GCC-PHAT, not raw cross-correlation. Concert audio is loud, clipped, and reverberant; PHAT weighting (whiten the cross-spectrum:
R = X·conj(Y); R /= |R|+eps; r = irfft(R)) is far more robust. Implement with numpy/scipy FFTs; librosa only for loading. - Compute all pairwise offsets, not just first-vs-rest. Score each pair by peak-to-second-peak ratio. Then solve globally: offsets are edges in a graph, per-video offsets are node values; least-squares solve the overdetermined system, weight by confidence, and reject pairs that violate cycle consistency (
off_ab + off_bc − off_ac> 50 ms). - Drift check: re-run GCC-PHAT on 10 s windows every 30 s along the overlap; fit a line through window-offsets. Store slope as
drift_ppm. Phone clocks drift tens of ppm — over a 10-minute set that's audible/visible. If |drift| < 5 ppm store 0. - Persist
offset_ms,drift_ppm,sync_confidenceto DB and exportdata/work/sync.json.
- Use GCC-PHAT, not raw cross-correlation. Concert audio is loud, clipped, and reverberant; PHAT weighting (whiten the cross-spectrum:
- Tests: synthesize one signal, shift copies by known amounts (+ add noise + small resample to simulate drift), assert recovered offsets within ±10 ms and drift within ±3 ppm. Also run against the M0 fixture and assert the known ground-truth offsets.
- Accept: sync on synthetic clips recovers ground truth.
M2 — Reconstruction (COLMAP) + pose export
frames.py: sample candidate frames ~2/s per video; within each 0.5 s window keep the sharpest frame (variance of Laplacian via OpenCV). Write JPEGs named{video_id}_{frame_idx}.jpg.sfm.py: drive the COLMAP CLI via subprocess:feature_extractor— one camera per video (--ImageReader.single_camera_per_folder 1, put each video's frames in its own subfolder),OPENCVcamera model.sequential_matcherwithin each video +exhaustive_matcheracross (frame counts here are small; exhaustive is fine at this scale).mapper(incremental reconstruction), thenimage_undistorter/model_converterto get a TXT model.- Parse
images.txt/cameras.txt/points3D.txtyourself (formats are documented in COLMAP docs; simple line parsing — no pycolmap dependency needed).
- Normalize the scene before storing: translate so the point-cloud centroid is at origin; scale so the camera-position bounding sphere has radius 10; rotate so world-up = the average of camera up-vectors (phones are held roughly upright — this heuristic works well; write it as such in a comment). Apply the same similarity transform to poses and points.
- Interpolate poses for unregistered frames between registered ones: slerp rotation, lerp translation, mark
registered=False; do not extrapolate beyond first/last registered frame of each video. - Export
data/work/points.ply(binary little-endian PLY) and fillcamera_poses. - Failure handling: if COLMAP registers < 60 % of frames or < 2 videos into one model, print a clear diagnostic (likely causes: too-dark footage, motion blur, no view overlap) and leave the synthetic/previous poses untouched. Concert footage is genuinely hard for SfM (moving lights, crowds, blur) — the app must degrade gracefully to "synced videos, no 3D".
- Tests: parser unit tests on hand-written COLMAP TXT snippets; normalization test (centroid≈0, radius≈10, up≈+Y); slerp test.
- Accept: on the synthetic fixture (which skips COLMAP and injects known poses) the whole export path runs; if COLMAP is installed, the pipeline runs end-to-end on any real folder of images without crashing.
M3 — API
FastAPI app (api.py), CORS enabled for the Vite origin, mounted static serving of data/raw with HTTP Range support (Starlette's StaticFiles supports Range — required or <video> seeking breaks).
GET /api/manifest → {videos:[{id,filename,url,duration_s,fps,width,height,
offset_ms,drift_ppm}], t_global_max, has_poses}
GET /api/videos/{id}/poses → [{frame_idx,t_video_s,q:[w,x,y,z],t:[x,y,z],
intrinsics:{fx,fy,cx,cy},registered}]
GET /api/pointcloud → points.ply file
GET /api/anchors → [...] POST /api/anchors
GET /api/events → [...]
POST /api/events/detect {t_global_s?, window_s?} → runs M7 detection (or on one window)
POST /api/annotations {video_id, t_video_s, bbox:[x0,y0,x1,y1]} → {anchor_id?, point?} (M8)
cli.py serve runs uvicorn. Accept: with the synthetic project, curl of each endpoint returns sane data; a browser can seek the served videos.
M4 — Synchronized multi-video playback (the heart of the app)
- Layout: left = 3D canvas placeholder, right = grid of
<video>elements (one per video), bottom = timeline scrubber + play/pause + speed control. - Master clock in
transport.js:t_global = t0 + (performance.now() - wall0)/1000 * ratewhile playing. Never trust any single video element as the clock. - Per video, each animation frame (use
requestVideoFrameCallbackwhere available, else rAF): compute targett_videofrom the timebase formula; compare tovideo.currentTime:- error > 150 ms → hard seek (
fastSeekif available); - 20–150 ms → nudge
playbackRate(clamp to master rate ×[0.95, 1.05]) until converged; - < 20 ms →
playbackRate = rate.
- error > 150 ms → hard seek (
- Videos whose
t_videofalls outside[0, duration]pause and dim (they weren't recording at that global time). - Audio: only one video unmuted at a time (click to select "audio source"); all others muted — otherwise phasing chaos.
- Accept: with the 3 synthetic clips (known offsets), scrubbing and playback keep all clips visibly locked to the shared audio events; drift error indicator (dev overlay showing per-video sync error in ms) stays < 50 ms during continuous playback.
M5 — 3D viewer
scene3d.js: load/api/pointcloud(PLY via three/examples PLYLoader) as aPointscloud; draw each video's camera path as a line and current pose as a frustum wireframe (colored per video); OrbitControls for free roam.- Coordinate conversion — do this exactly, it's the classic pitfall. COLMAP stores world→camera:
x_cam = R x_world + t, camera axes = (+x right, +y down, +z forward). Three.js cameras look down −z with +y up. Per pose:
Implement once in a shared helper with a unit test in the backend mirror (R = quat_to_mat(q) // world→cam C = -Rᵀ t // camera center in world coords R_c2w = Rᵀ R_three = R_c2w * diag(1, -1, -1) // flip y,z of camera-local axes camera.position = C; camera.setRotationFromMatrix(R_three)geometry.py) so both sides agree. - On each playhead tick, update frusta from the pose at the nearest
t_video(poses are keyed by frame; binary-search + use interpolated entries). - Snap-to-camera: clicking a frustum (or a per-video button) tweens the viewer camera to that video's pose and keeps following it while playing; intrinsics →
PerspectiveCamera:fov = 2·atan(H/(2·fy))·180/π,aspect = W/H(assume centered principal point for the prototype). "Free roam" button detaches. - Accept: on synthetic data, orbiting shows 3 camera arcs around the stage box; snap-to-camera view of camera A roughly matches what video A shows (verifiable exactly on synthetic data since it's generated from those poses).
M6 — Anchor overlays ("x-ray" HUD)
- Each
<video>gets a transparent<canvas>overlay, resolution-matched to the displayed video size (handle devicePixelRatio + object-fit letterboxing math carefully — compute the video's content rect inside the element). - For each anchor and each visible video: build the Three.js camera for that video's current pose (M5 helper),
anchor.project(camera), convert NDC → canvas px, skip if behind camera (vec.z > 1after project, or check the w-divide sign properly by transforming to camera space first and requiringz_cam < 0). - Draw a dot + label; anchors render regardless of real-world occlusion — that is the x-ray feature. Style: small circle, 1 px white halo, label text.
- Seed default anchors on the synthetic stage (4 box corners named "Stage FL/FR/BL/BR").
- Accept: on synthetic data, the projected corners track the drawn stage box across all 3 videos while playing (exact on synthetic since geometry is ground truth).
M7 — Moment detection + AI classification
Two stages — cheap audio candidates first, then vision LLM on candidates only:
-
Candidates (
events_ai.py): on the reference video's audio compute RMS energy and spectral-flux onset strength (librosa); find peaks that are local maxima over a 10 s neighborhood and above the 90th percentile → candidatet_globalmoments (bass drops, pyro bangs, crowd roars). Insert asevents(source='audio_auto', event_type='candidate'). -
Classification — pluggable provider (
events_ai.py). The shared contract is a Pydantic model and a tiny protocol; providers are interchangeable and selected by env var:class MomentClassification(BaseModel): event_type: Literal["bass_drop","pyro","confetti","crowd_wave", "artist_moment","light_show","quiet_moment","other"] confidence: float # 0..1 description: str # one sentence, human-readable class MomentClassifier(Protocol): def classify(self, clip_path: Path, frames: list[Path]) -> MomentClassification: ...Selection:
FESTIVAL4D_CLASSIFIER=gemini|claude|local(defaultgemini). Inputs prepared once per candidate: a 3 s MP4 clip aroundt_globalfrom the best-covering video (ffmpeg, ≤720p, keep it under ~15 MB) and 6 sampled JPEGs (≤768 px long edge) for providers without video input.GeminiClassifier(default — cheapest, native video).google-genaiSDK, modelgemini-2.5-flash, needsGEMINI_API_KEY. Send the clip inline (video bytes are supported directly for small files) and enforce structured output:from google import genai client = genai.Client() # reads GEMINI_API_KEY resp = client.models.generate_content( model="gemini-2.5-flash", contents=[ genai.types.Part.from_bytes(data=clip_bytes, mime_type="video/mp4"), "This is a 3-second concert crowd clip around a detected audio spike. Classify the moment.", ], config={"response_mime_type": "application/json", "response_schema": MomentClassification}, ) result = MomentClassification.model_validate_json(resp.text)ClaudeClassifier(optional,ANTHROPIC_API_KEY): Anthropic Python SDK, modelclaude-opus-4-8, the 6 frames as base64 image blocks, structured output viaclient.messages.parse(..., thinking={"type":"adaptive"}, output_format=MomentClassification)→resp.parsed_output. Do not passtemperature(rejected on Opus 4.8).LocalClassifier(offline/free): any OpenAI-compatible endpoint — Ollama/LM Studio with a vision model (e.g.qwen2.5-vl), or OpenRouter by just pointing the same class athttps://openrouter.ai/api/v1. Config:FESTIVAL4D_OPENAI_BASE_URL,FESTIVAL4D_OPENAI_KEY,FESTIVAL4D_OPENAI_MODEL. Send the 6 frames as image_url/base64 parts, request JSON, validate withMomentClassification.model_validate_jsonand one retry on validation failure (local models are sloppier about JSON).
Common behavior for all providers: catch exceptions per candidate — one failure must not abort the batch; log and continue. If the selected provider's key/endpoint is unconfigured, run candidate detection only, store
event_type='candidate', and log that classification was skipped. Add one pytest with a stub classifier to verify the orchestration loop independent of any API. -
Frontend: colored markers on the timeline (color by event_type, legend in a corner); hover → description tooltip; click → jump playhead there.
-
Accept: on synthetic audio (which has synthesized "bangs"), candidates land within ±0.5 s of the ground-truth pulse times; with an API key present, classification runs and markers render.
M8 — Correction / annotation → 3D
- Clicking a timeline event opens a small panel: shown classification, a dropdown to correct
event_type(writessource='user'), and an "annotate location" mode. - Annotate mode: user drags a rectangle on any video's overlay canvas → normalized bbox →
POST /api/annotations. - Backend resolution (
geometry.py):- Cast a ray from that camera through the bbox center: unproject with the pose+intrinsics at
t_video. - If another annotation exists for the same event from a different video: triangulate (closest point between the two rays; reject if rays are near-parallel or midpoint gap > 0.5 scene units).
- Else: fall back to the nearest point-cloud point to the ray (within a 0.3-unit cylinder); if none, return the ray at depth = distance to point-cloud centroid.
- Store the resulting 3D point as a new anchor linked to the event → it immediately appears in all overlays and the 3D scene (label = event description).
- Cast a ray from that camera through the bbox center: unproject with the pose+intrinsics at
- Accept: on synthetic data, annotating the same stage corner from two videos produces an anchor within 0.2 units of the true corner.
M9 — God's-eye extras + README
- Keyframed camera paths (
camPath.js): "add keyframe" captures current free-roam pose at currentt_global; keyframes render as gizmos; playback interpolates position with Catmull-Rom and orientation with slerp while the transport plays; export/import path as JSON. - Polish: loading states, per-video enable/disable, keyboard shortcuts (space = play/pause, ←/→ = ±1 s, 1..9 = snap to camera N).
- README.md: prerequisites (ffmpeg required; COLMAP optional; a classifier key optional —
GEMINI_API_KEYdefault, or Claude/local/OpenRouter viaFESTIVAL4D_CLASSIFIER), quickstart (synthetic→serve→npm run dev), then the real-footage workflow (drop files in data/raw→ingest→sync→reconstruct→events), architecture sketch, and a short "shooting tips for good reconstructions" section (overlap between viewpoints, avoid pure zoom, keep some static structure in frame, brighter footage works better). - Accept: a fresh clone following README reaches the working synthetic demo with no undocumented steps.
4. Known pitfalls (read before coding)
- COLMAP↔Three.js axes — the diag(1,−1,−1) flip in M5. Get the shared helper + test done before any rendering.
- HTML5 video is not frame-accurate. Design the transport around continuous correction (M4), never around "seek then trust currentTime".
- Range requests — without them, seeking silently breaks. Verify with
curl -H "Range: bytes=0-100"→ expect 206. - Letterboxing math in overlays —
object-fit: containmeans canvas px ≠ video px; compute the content rect. - Only sync what overlaps. Two clips with no common audio can't be aligned; the pairwise-graph design (M1) handles this — videos in disconnected components get
offset_ms=NULLand are excluded from the timeline with a UI notice. .gitignoredata/raw data/work *.db node_modulesfrom M0 — media files must never be committed.- Classifier providers: default is Gemini 2.5 Flash (
GEMINI_API_KEY) with native video input; Claude option usesclaude-opus-4-8with adaptive thinking, no sampling params,messages.parse(model IDs never get date suffixes); local/OpenRouter option is any OpenAI-compatible vision endpoint. All must degrade to candidates-only when unconfigured.
5. Explicit non-goals for this prototype
No user accounts/auth, no cloud storage, no realtime ingest, no NeRF/Gaussian splatting (leave a docs/ideas.md note), no mobile UI, no Docker. Single local user, single project at a time.