Commit Graph

2 Commits

Author SHA1 Message Date
m3ultra
06a080b18d SLAT stage: image -> occupancy -> sparse latents -> mesh, running end to end
The whole geometry chain now runs on a real photograph:

  [1] occupancy   12948 voxels                     19.7s
  [2] cond        proj (1, 262144, 2048) @ 64^3     2.6s
      gathered    proj (12948, 2048)
  [3] SLAT        (12948, 32)                      88.7s
  [4] MESH        3556515 verts, 7071196 faces     11.0s   grid 1024^3
      peak 22.6 GB, bounds inside the unit cube

Two bugs fixed on the way:

1. "global" must be FLAT [M,C] for the sparse blocks, not the dense stage's [B,T,C].
   The sparse cross-attention takes a token stack plus an explicit layout, so the
   dense shape dies inside to_kv's reshape rather than anywhere informative. Gathering
   now reshapes it, and refuses batch > 1 rather than silently mislabelling a layout.
2. o_voxel needs the decoder OUTPUT grid, not its configured resolution. The shape
   decoder applies four 2x upsamples, so a res-64 latent decodes into 1024^3, while
   the config says 256 (upstream overrides it per run via set_resolution). Passing 256
   raised an opaque out-of-bounds inside o_voxel's hashmap insert. Added
   output_resolution() and a guard that names the real cause.

HONEST LIMITATION - this is NOT yet the shipped cascade. Upstream's
sample_shape_slat_cascade runs the 512 flow (res 32) first, denormalises, UPSAMPLES
THE COORDINATE SET through the shape decoder, then runs the 1024 flow on the refined
coords. Running the HR flow straight off the 64^3 occupancy set yields a complete,
exportable mesh whose silhouette IoU is 0.639 - against 0.842 for the occupancy grid
that seeded it. The gap is a halo of geometry outside the true silhouette, exactly
what the missing coordinate refinement would prune. Do not read the current mesh
quality as the model's; wiring the cascade is the next step.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-03 14:53:32 +10:00
m3ultra
5484d59cb3 Mesh export: Flexible Dual Grid -> triangles -> GLB, via o_voxel
The shape decoder's 7 channels are not an SDF — O-Voxel solves a QEF over a Flexible
Dual Grid, which is what lets it carry open and non-manifold surfaces:

  0:3  vertex offset in-voxel, (1+2m)*sigmoid(v)-m so it may sit OUTSIDE its own cell
  3:6  per-axis intersection logits, thresholded at 0
  6:7  quad split weight through softplus

mesh.py is the MLX->torch boundary for export. o_voxel's convert/postprocess are
native (C++/Metal) and deliberately NOT ported: o-voxel builds a CPU CppExtension when
CUDA is absent, and the trellis-2 lane on this fleet already runs it with a Metal
baker, so reusing that build beats reimplementing a QEF solver in MLX. Installed into
the shared venv from ~/Documents/trellis-2-mrp-mlx/o-voxel; it needs cv2 and xatlas,
and NOT utils3d (which drags in open3d, with no cp312 wheel).

Verified against the REAL shape_dec (292/292 params, resolution 256):

  decoded 5954 voxels x 7ch   ->   5954 vertices, 6886 faces   ->   GLB written

Not watertight, correctly: the input was a random latent, and FlexiDualGrid represents
open surfaces by design. Vertices land inside the octant of the unit cube matching the
sparse coords fed in, which is the check that the grid indexing is right.

Still to wire: the SLAT stage itself (sparse latents seeded from the occupancy coords,
with proj features gathered at those coords).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-03 14:44:23 +10:00