5.1 KiB
Festival 4D — Future Ideas
Parking lot for post-v0.1.0 extensions (spec §5 future-extensions). Not committed work — a place to capture direction so the prototype's scope stays honest. Roughly ordered by payoff-to-effort.
Status 2026-07-17 (evening): splatting ✅ shipped (see
modelbeast-crossover.md); spatial audio ✅ shipped (🎧, WebAudio master clock). Cinematic export ✅ M12, multi-modal audio analysis ✅ M10/M11, anchor manager ✅ M13, path persistence ✅ M16 — all SHIPPED in phase 5 (plan/20-phase5.md, suite floor 189). Still parked here: realtime ingest, object/artist tracking, per-project workspaces, sync-graph visualization, per-moment splats (crossover doc §3), and the new hardware ideas below.
Neural rendering — Gaussian Splatting
Replace the sparse COLMAP point cloud with a 3D Gaussian Splatting model for photorealistic
free-roam instead of a dot cloud. COLMAP poses + images already feed splatting trainers
(e.g. gsplat, Inria 3DGS) directly, so the reconstruction step's output is reusable. The viewer
would swap Points for a splat renderer (a WebGL splat viewer, or bake to a mesh). Biggest visual
upgrade; heaviest compute (needs a GPU training pass per project).
Cinematic export — render the camera path to MP4
M9 already produces keyframed god's-eye paths. Add a headless render: step t_global at a fixed
fps, drive the Three.js camera along the path, capture frames (offscreen canvas → captureStream
or server-side puppeteer + ffmpeg), and mux to MP4. Turns the app from a viewer into a
"director's cut" clip exporter. Path JSON export/import is already in place, so this is additive.
Real-time / streaming ingest
Instead of pre-recorded files, ingest live phone feeds (WebRTC/RTMP). Requires: rolling audio sync on a sliding window (GCC-PHAT already windowed for drift — reuse it), incremental SfM or pose tracking against a pre-built map, and a streaming transport in the frontend. Large effort; the offline pipeline is the right foundation.
Multi-modal audio analysis for moment detection
Today's events step finds candidates from RMS + spectral flux, then a VLM classifies video.
Add audio-native analysis in parallel: beat/tempo tracking, drop detection from
spectrograms, song-boundary segmentation. Fuse audio-derived and vision-derived labels for
higher-confidence events and beat-aligned auto camera cuts. Cheap; complements the existing
classifier providers.
Object / artist tracking
Extend M8 (manual bbox → 3D anchor) with automatic per-frame tracking (e.g. a detector + tracker), so an anchor follows the lead performer across time instead of being static. Enables a "follow artist" auto-camera mode in the 3D viewer.
Stereo / spatial-video ingest (split the eyes, pin the scale)
Spatial video (iPhone 15 Pro+ MV-HEVC, QooCam, Canon dual-fisheye) is two synced views with a factory-known baseline. Ingest tweak: detect multi-view files and split L/R into two pipeline videos (ffmpeg can demux the streams; identical offset, shared audio). Payoff beyond one extra viewpoint: COLMAP reconstructions are scale-free, and a known stereo baseline anchors metric scale for the whole scene — real-meter units for anchors, paths, and spatial audio distances. Cheap ingest change, high leverage.
Depth-camera seeding (small venues only)
RGB-D cams (ZED 2i, OAK-D, LiDAR phones) can seed/regularize splat training (depth-supervised 3DGS → fewer floaters, faster convergence) and densify the point cloud without SfM guessing. Honest constraint: consumer depth dies past ~10–15 m (LiDAR ~5 m), so this only pays at club / small-venue range, not a festival main stage. Ingest: accept a per-video depth track or per-frame depth PNGs alongside the RGB.
Microcontroller rig kit (Arduino as crew, not camera)
An ESP32-CAM is not concert-grade — but a microcontroller on the rig earns its place three ways:
- Sync/calibration beacon: an LED blinking a known pattern visible to all cameras = visual ground truth to validate the audio sync (and a free calibration target).
- Servo pan/tilt or slider mount: one slowly-sweeping fixed camera contributes rich parallax from a single device; the per-frame pose track already handles moving cameras.
- Record start/stop trigger for the whole rig.
DIY camera nodes should be Raspberry Pi + Global Shutter camera module (kills
rolling-shutter wobble under fast stage lights) pushing into the existing
/captureendpoint (FESTIVAL4D_CAPTURE=1) — which was built for exactly this.
Quality-of-life
- Standalone anchor manager — the delete UI currently lives in the event correction panel; a dedicated always-available anchor list would decouple anchor management from events.
- Persist camera paths server-side — paths are JSON export/import only; a
pathstable + endpoints would let them live with the project. - Per-project workspaces — the app assumes one project at a time (SQLite at
data/project.db). - Sync-graph visualization — show which videos aligned into which connected component when some clips share no audio.