festifun/README.md
m3ultra c24af8029e Audio is the master clock + spatial audio free-roam (fresh-eyes opts 1 & feature 1)
Master clock: the selected soundtrack (new GET /api/audio/{id} —
listening-quality stereo AAC, lazily ffmpeg-extracted and cached in
data/work/audio_hq/) plays through WebAudio, and t_global derives from
AudioContext.currentTime. Every <video> is muted picture-only, so hard
seeks and playbackRate corrections are inaudible (the unmuted video
used to warble on every nudge). Graceful fallback to the old
performance.now() clock + <video> audio when a track can't load, with
an in-place upgrade if the decode lands mid-play. Drift is compensated
in the track's playbackRate; _resync() now recomputes from the live
clock since the soundtrack keeps playing while rAF is throttled.

Spatial audio (🎧 3D audio, shown when poses exist): every enabled
camera's track plays through an HRTF PannerNode at its live
reconstructed pose, gains normalized 1/sqrt(N); the listener follows
the viewer camera each frame (free roam, follow-cam, and M9 paths).
Tracks that finish decoding mid-play join the mix in place; toggling,
seeking, rate and enable changes rebuild the graph at the current
t_global with no timeline jump.

The window.__f4d debug hook is now also available in production builds
via localStorage f4dDebug=1 (headless panes throttle rAF; the pump is
the only way to drive the loop there).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 17:02:48 +10:00

246 lines
12 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Festival 4D
Turn multiple fan-shot smartphone videos of the same concert into a synchronized,
explorable **4D experience**: time-aligned multi-video playback, 3D scene reconstruction
with camera poses, a free-roam "god's eye" viewer, AR-style overlays projected onto each
video, AI-tagged moments on a shared timeline, click-to-place 3D annotations, and keyframed
cinematic fly-throughs.
Everything works **from pixels and audio alone** — no depth sensors, no IMU logs. It runs
fully offline and local; the only optional cloud piece is moment classification.
---
## What you can do
- **Scrub a shared timeline** and watch every camera stay locked to the same instant, aligned
by their audio (no manual frame-matching).
- **Fly through the reconstructed scene** in 3D — orbit freely, snap to any real camera, or
play a keyframed cinematic path.
- **See an "x-ray" HUD**: 3D anchor points (stage corners, tagged objects) projected onto every
video, tracking even when occluded by crowd or scaffolding.
- **Jump between AI-tagged moments** (bass drops, pyro, crowd waves…) on the timeline, and
correct any misclassification.
- **Click to locate things in 3D**: draw a box around an object in one or two videos and the
app triangulates its 3D position, dropping an anchor that appears everywhere at once.
- **Hear the scene in 3D** 🎧: toggle spatial audio and every camera becomes a positional
sound source at its reconstructed pose — fly toward a camera and you hear *that* spot in
the crowd get closer.
---
## Prerequisites
| Tool | Required? | Notes |
|---|---|---|
| **Python 3.11+** + [`uv`](https://docs.astral.sh/uv/) (or venv+pip) | yes | backend |
| **ffmpeg** / **ffprobe** on `PATH` | yes | audio extraction, frame sampling, clip cutting |
| **Node 18+** | yes | frontend (Vite + Three.js) |
| **COLMAP** on `PATH` | optional | 3D reconstruction. Without it the app still runs — you get synced videos + timeline, just no 3D scene. `brew install colmap` on macOS. |
| A classifier API key | optional | AI moment labels. Default **Gemini flash** (`GEMINI_API_KEY`; model `gemini-3.1-flash-lite`, override with `FESTIVAL4D_GEMINI_MODEL`); also Claude or any OpenAI-compatible/local/OpenRouter endpoint via `FESTIVAL4D_CLASSIFIER`. Without a key you still get audio-detected candidate moments, just unlabeled. |
---
## Quickstart (synthetic demo — no footage needed)
The synthetic fixture generates three fake camera videos of the same fake stage, with known
audio offsets and known 3D geometry — so you can see the whole app working end-to-end before
you have any real footage.
```bash
# 1. backend env
uv venv --python 3.12
uv pip install -e ".[dev]"
# 2. generate the synthetic project (fake videos + poses + point cloud + seeded events)
uv run python -m festival4d synthetic
# 3. serve the API (http://127.0.0.1:8000)
uv run python -m festival4d serve
# 4. in another terminal, run the frontend
cd frontend
npm install
npm run dev # opens http://localhost:5173 (or the next free port)
```
Open the printed URL. You should see the 3D scene with three camera frusta around a stage box,
the three synced video players, stage-corner anchors projected onto each video, and colored
moment markers on the timeline.
---
## Capture: record straight from webcams & phones
Don't have footage yet? Record it with whatever cameras you have — USB webcams, the laptop cam,
and any phone with a browser. **No app install.** Cameras *record* clips that feed the normal
pipeline (this isn't live streaming — see `docs/ideas.md` for that).
This works because of two properties of the design: browser recordings (`.webm`, or `.mp4` on
iOS) are already ingestible formats, and **sync is audio-based** — so the devices need **no clock
sync and don't even have to start together**. They only have to record the same moment and hear
the same sound.
```bash
# capture is OPT-IN (it accepts uploads) — enable it explicitly:
FESTIVAL4D_CAPTURE=1 uv run python -m festival4d serve
```
Then open **`/capture`** on each device (a 📹 Capture link also appears in the viewer): name the
device, pick a camera + mic, hit record, hit stop — the clip uploads into `data/raw/`. The page
shows a live "who's shooting what" monitor of every connected camera. When you're done:
```bash
uv run python -m festival4d ingest && uv run python -m festival4d sync # then reconstruct / events
```
### Getting a phone connected (the HTTPS bit)
Browsers **only grant camera access on a secure origin**. A phone opening `http://<lan-ip>:8000`
gets *no camera*, silently. The easiest fix if you use Tailscale — it issues a real cert, so
there's nothing to trust manually:
```bash
tailscale serve --bg 8000 # prints https://<machine>.<tailnet>.ts.net
```
Join the phone to the tailnet and open `https://<machine>.<tailnet>.ts.net/capture`. (Without
Tailscale: any HTTPS reverse proxy works; `localhost` is also treated as secure, which is why the
laptop's own cameras work with no setup.)
### Capture tips
- **Every clip needs audio**, and all cameras must hear the same sound — that's the only thing
aligning them. Testing indoors? Play music.
- **USB bandwidth is the real limit** on multiple webcams: two 1080p cams on one controller often
fails. Drop to 720p or use separate ports/hubs.
- Follow the *Shooting tips* below for reconstruction-friendly angles.
- Capture routes stay **unmounted** unless `FESTIVAL4D_CAPTURE=1` — a public deployment must
never expose an open upload endpoint.
## Real-footage workflow
1. **Shoot / gather** 24 videos of the same performance from different positions (see
*Shooting tips* below) and drop the files into `data/raw/`.
2. **Run the pipeline:**
```bash
uv run python -m festival4d ingest # probe videos, extract 16 kHz mono audio
uv run python -m festival4d sync # GCC-PHAT audio alignment (offsets + drift)
uv run python -m festival4d reconstruct # COLMAP structure-from-motion → camera poses + point cloud
uv run python -m festival4d events # detect audio moments + AI-classify them
uv run python -m festival4d serve # serve it
```
Each step is independent and re-runnable. `sync` alone already gives you locked multi-video
playback; `reconstruct` adds the 3D scene and overlays; `events` adds the tagged timeline.
3. **Explore** in the browser (`cd frontend && npm run dev`).
If COLMAP can't reconstruct your footage (common with dark, motion-blurred, or low-overlap
clips), the pipeline says so and leaves you with the synced-video experience — it never
corrupts existing data.
To enable AI moment labels, set a key before `events`:
```bash
export GEMINI_API_KEY=... # default provider (Gemini flash, native video)
# or: export FESTIVAL4D_CLASSIFIER=claude ANTHROPIC_API_KEY=...
# or: export FESTIVAL4D_CLASSIFIER=local FESTIVAL4D_OPENAI_BASE_URL=... FESTIVAL4D_OPENAI_MODEL=...
```
---
## Using the viewer
| Action | How |
|---|---|
| Play / pause | `Space` or the ⏵ button |
| Nudge time ±1 s (±5 s) | `←` / `→` (hold `Shift`) |
| Scrub | drag the timeline |
| Pick the audio you hear | click a video (one soundtrack at a time) |
| Spatial audio | 🎧 **3D audio** (needs poses): every camera plays from its position in the scene; the mix follows the viewer camera as you fly |
| Enable/disable a video | ◉ on the video |
| Orbit the 3D scene | drag in the 3D pane |
| Snap the 3D view to a camera | click its frustum, its ⛶ button, or press `1``9` |
| Free roam again | `Esc` / `0` / **Free roam** |
| Correct a moment / annotate | click a timeline marker → panel |
| Place a 3D anchor | in the panel, **Annotate location**, then drag a box on a video (repeat on a 2nd video to triangulate) |
| Add a camera keyframe | ** Key** or `K` (captures the current free-roam view at the current time) |
| Play the keyframed path | **▶ Path** · export/import with ⤓ / ⤒ |
The top-right **sync error** panel shows each video's alignment error in ms while playing —
a dev aid; it should stay green (<50 ms) once footage is well-synced.
---
## Architecture
```
data/raw/ ──ingest──▶ data/work/audio ──sync────▶ videos.offset_ms / drift_ppm ┐
──frames──▶ data/work/frames ──reconstruct──▶ camera_poses + points.ply
reference audio ──events──▶ events (audio candidates → AI labels)
SQLite (data/project.db) ◀─────────┘
FastAPI (backend/festival4d/api.py)
│ /api/manifest, /poses, /pointcloud,
│ /events, /anchors, /annotations
Vite + Three.js SPA (frontend/src)
transport (master clock) · videoGrid + overlays · scene3d · timeline
annotate (M8) · camPath (M9)
```
- **Backend** Python: `ingest`/`audio_sync` (media + GCC-PHAT sync), `frames`/`sfm`/`geometry`
(COLMAP orchestration + pose math + triangulation), `events_ai` (moment detection + pluggable
classifiers), `resolve` (annotation3D), `api` (FastAPI, Range-capable video serving), `db`
(SQLAlchemy/SQLite), `synthetic` (the test fixture).
- **Frontend** vanilla JS + Three.js. A single master clock (`transport.js`) derives each
video's local time from the shared timeline and continuously corrects playback. When possible
the clock is **audio**: the selected soundtrack (`GET /api/audio/{id}`, extracted+cached
server-side) plays through WebAudio and `t_global` derives from `AudioContext.currentTime`
so every video is muted picture-only and can be seeked/rate-trimmed with zero audible
artifacts (falls back to a `performance.now()` clock + `<video>` audio if the track can't
load). `lib/pose.js` is the one COLMAPThree.js pose conversion, mirrored from `geometry.py`
and locked by a shared test.
- **Timebase contract:** `t_video = (t_global offset_ms/1000) · (1 + drift_ppm·1e6)`; the
reference video has offset 0.
Full design + milestone history: [`OPUS_BUILD_INSTRUCTIONS.md`](OPUS_BUILD_INSTRUCTIONS.md);
the parallel build plan and per-lane status live in [`plan/`](plan/).
---
## Shooting tips (for good reconstructions)
Concert footage is genuinely hard for structure-from-motion. To give COLMAP a chance:
- **Spread the cameras out** but keep **overlapping views** each pair of cameras should see
some of the same stage/structure. No overlap they can't be related in 3D.
- **Keep some static structure in frame** (stage edges, truss, speaker stacks). A frame that's
all moving crowd and lights has nothing stable to triangulate.
- **Avoid pure zoom** physically moving parallax reconstructs far better than zooming.
- **Brighter, sharper is better** motion blur and near-dark frames are the main failure cause.
- **24 phones is the sweet spot** for a first reconstruction; more is fine but slower.
Audio sync is far more forgiving: any clips that share audible sound (the same music/claps)
will align, even across otherwise-unrelated angles.
---
## Tests
```bash
uv run pytest # backend
cd frontend && npm run build # frontend typecheck/build
```
---
## Non-goals (this prototype)
No accounts/auth, no cloud storage, no realtime ingest, no NeRF/Gaussian splatting, no mobile
UI, no Docker. Single local user, one project at a time.