Add SOMA and G1 model support

This commit is contained in:
Richard Palethorpe 2026-08-26 15:59:58 +01:00
parent dedbe1066a
commit 92341f3194
35 changed files with 872 additions and 327 deletions

View File

@ -4,14 +4,24 @@ GGML/C++ implementation of NVIDIA's Kimodo text-to-motion model.
## Status
`Kimodo-SMPLX-RP-v1` accepts either a UTF-8 prompt or a precomputed LLM2Vec
embedding and generates unconstrained SMPL-X22 local rotations and root
translations on CPU or Vulkan. The text encoder uses eight-layer Vulkan chunks
by default; set `KIMODO_TEXT_LAYER_CHUNK=1..32` to tune VRAM use.
The five released Kimodo motion checkpoints accept either a UTF-8 prompt or a
precomputed LLM2Vec embedding and generate local rotations plus root
translations on CPU or Vulkan:
- SMPL-X RP v1: 22 joints
- SOMA RP/SEED v1.1: the predicted compact 30-joint control skeleton
- G1 RP/SEED v1: 34 Unitree G1 joints
NVIDIA's Python API expands SOMA's predicted 30 joints to a relaxed-hand
77-joint presentation skeleton. The native API currently returns the 30 joints
the model actually predicts. The text encoder uses eight-layer Vulkan chunks by
default; set `KIMODO_TEXT_LAYER_CHUNK=1..32` to tune VRAM use.
Included: checked GGUF loading, safetensors conversion, DDIM sampling, C/C++
APIs, CPU/Vulkan parity tests, and a local text-to-motion demo. Constraints,
SOMA, G1, GLB export, and quantised models are not implemented yet.
APIs, conditioned multi-prompt transitions, CPU/Vulkan parity tests,
skeleton-only GLB export, and a local text-to-motion demo. General constraint
input, 77-joint SOMA expansion, skinned-mesh GLB export, and quantised models
are not implemented yet.
## Build and test on Linux
@ -21,7 +31,7 @@ support. GGML is a pinned Git submodule:
```sh
git submodule update --init --recursive
scripts/download_gguf_weights.sh --output "$PWD"
scripts/download_gguf_weights.sh --output "$PWD" --model soma-rp-v1.1
cmake --preset debug
cmake --build --preset debug
ctest --preset debug
@ -56,8 +66,9 @@ to the process. The GGUF parser fuzzer requires Clang.
`include/kimodo/kimodo_capi.h` is the C API. Model loading checks the motion
GGUF and text bundle before inference. Use `kimodo_generate_embedding` for
4096 F32 values or `kimodo_generate` for text. Both return SMPL-X22 root
translations and local XYZW rotations.
4096 F32 values or `kimodo_generate` for text. Both return the selected model's
root translations and local XYZW rotations; query the joint count from the
result rather than assuming a fixed skeleton.
## Demo
@ -69,27 +80,43 @@ go run ./demo -addr 0.0.0.0:8094
Open `http://localhost:8094`. The left sidebar contains the prompt and a
persistent history; choosing a previous animation restores its prompt for a
new generation.
new generation. Every successful animation also writes a standalone
`animation.glb` beside its raw streams, for example
`demo-output/<animation-id>/animation.glb`. It contains the selected animated
node hierarchy (no mesh), ready to copy into a Three.js project. It is also
available from `/api/animations/<animation-id>/animation.glb` while the demo
is running.
## Weights
Ready-to-run native GGML weights are published under the Hugging Face
`LocalAI-io` organisation (not GitHub's `localai-org`). The reusable
[Llama-3-Kimodo-GGML](https://huggingface.co/LocalAI-io/Llama-3-Kimodo-GGML)
text encoder and the upstream-linked
[Kimodo-SMPLX-RP-v1-GGML](https://huggingface.co/LocalAI-io/Kimodo-SMPLX-RP-v1-GGML)
diffusion model are separate, so users download rather than recreate them:
text encoder is separate from the four redistributable motion repositories,
each of which preserves a one-to-one relationship to its NVIDIA upstream:
- [Kimodo-SOMA-RP-v1.1-GGML](https://huggingface.co/LocalAI-io/Kimodo-SOMA-RP-v1.1-GGML)
- [Kimodo-SOMA-SEED-v1.1-GGML](https://huggingface.co/LocalAI-io/Kimodo-SOMA-SEED-v1.1-GGML)
- [Kimodo-G1-RP-v1-GGML](https://huggingface.co/LocalAI-io/Kimodo-G1-RP-v1-GGML)
- [Kimodo-G1-SEED-v1-GGML](https://huggingface.co/LocalAI-io/Kimodo-G1-SEED-v1-GGML)
Download one or repeat `--model` to install several:
```sh
scripts/download_gguf_weights.sh --output "$PWD"
scripts/download_gguf_weights.sh --output "$PWD" \
--model soma-rp-v1.1 --model g1-rp-v1
```
The installer verifies each published manifest and SHA-256 hashes. Use
`--motion-only` when supplying a precomputed 4096-float LLM2Vec embedding.
SMPL-X RP is deliberately absent from the published-weight installer: its
internal-R&D licence prohibits distributing derivative models, so it must be
converted locally after the user obtains the upstream checkpoint under its
gated terms.
The GGUF bundle includes converted Meta Llama 3 material and Kimodo is
non-commercial research-only. Review the published model card and upstream
licences before downloading or redistributing.
The text bundle includes converted Meta Llama 3 material and retains its
separate terms. Review every selected model card before downloading or
redistributing.
## License
@ -97,6 +124,22 @@ The C++ port and its original tooling are licensed under Apache-2.0; see
[LICENSE](LICENSE). GGML and the model weights retain their respective
licences.
| Motion checkpoint | Upstream terms | Commercial use |
| --- | --- | --- |
| Kimodo-SMPLX-RP-v1 | [NVIDIA Internal Scientific Research and Development Model License](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-internal-scientific-research-and-development-model-license/) | No; internal, non-production R&D only; derivative model redistribution is prohibited |
| SOMA RP/SEED v1.1 | [NVIDIA Open Model License](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/) | Permitted by the model licence |
| G1 RP/SEED v1 | [NVIDIA Open Model License](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/) | Permitted by the model licence |
The SMPL-X warning is about NVIDIA's trained Kimodo checkpoint, not the mere
fact that its output uses an SMPL-X-shaped 22-joint hierarchy. Converting that
checkpoint to GGUF is a new runtime representation of the same weights and does
not replace its licence. Skeleton names, parent links, and the Apache-2.0 port
source do not by themselves make the SOMA or G1 checkpoints non-commercial.
The SMPL-X Hugging Face metadata, model card, and access terms identify the
internal-R&D licence; treat those restrictive terms as controlling even though
an apparently inconsistent `LICENSE` file has also appeared in that upstream
repository.
### Regenerating the bundle
This is only needed to reproduce a conversion. The SMPL-X checkpoint and Llama
@ -105,8 +148,9 @@ authenticating, download the exact revisions and hash manifests with:
```sh
nix develop path:. --command hf auth login
nix develop path:. --command scripts/download_weights.sh \
--output "$PWD/models" --with-text
scripts/download_weights.sh --output "$PWD/models" --with-text \
--model smplx-rp-v1 --model soma-rp-v1.1 --model soma-seed-v1.1 \
--model g1-rp-v1 --model g1-seed-v1
```
Convert the local LLM2Vec model to the native component bundle with:
@ -120,9 +164,10 @@ Validate a prospective release without network access, then explicitly upload
it from an account allowed to publish to `LocalAI-io`:
```sh
nix develop path:. --command python scripts/publish_gguf.py --component motion
nix develop path:. --command python scripts/publish_gguf.py --component motion \
--upload --confirm-upstream-licences
--motion-model soma-rp-v1.1
nix develop path:. --command python scripts/publish_gguf.py --component motion \
--motion-model soma-rp-v1.1 --upload --confirm-upstream-licences
nix develop path:. --command python scripts/publish_gguf.py --component text \
--upload --confirm-upstream-licences
```

View File

@ -1,20 +1,21 @@
<!doctype html><html lang="en"><meta charset="utf-8"><meta name="viewport" content="width=device-width, initial-scale=1"><title>Kimodo text to motion</title>
<style>
:root{color-scheme:dark;font-family:Inter,system-ui,sans-serif;background:#10131b;color:#f2f5fb}body{margin:0;overflow:hidden;background:#10131b}main{width:100vw;height:100vh;display:grid;grid-template-columns:360px minmax(0,1fr)}.sidebar{padding:24px 18px;display:flex;flex-direction:column;gap:18px;overflow-y:auto;overflow-x:hidden;background:radial-gradient(circle at 0 0,#293b62,transparent 32rem),#141a27;border-right:1px solid #2b354b}h1{font-size:2rem;letter-spacing:-.05em;margin:0}.eyebrow{font-size:.7rem;text-transform:uppercase;letter-spacing:.14em;color:#9bb8ff;margin-bottom:8px}.card{background:#171c28dd;border:1px solid #2b354b;border-radius:14px;overflow:hidden}.stage{min-width:0;min-height:0;display:flex;flex-direction:column;background:#0c1019}canvas{display:block;width:100%;height:100%;flex:1;min-height:0;background:linear-gradient(#111b31,#0c1019);cursor:grab;touch-action:none}canvas.dragging{cursor:grabbing}.controls{padding:14px 18px;display:flex;align-items:center;gap:10px;flex-wrap:wrap;border-top:1px solid #2b354b}.promptbox{padding:16px;display:grid;gap:10px}.promptbox textarea{min-height:110px;resize:none;overflow:hidden}textarea,input{box-sizing:border-box;width:100%;border:1px solid #35415b;border-radius:10px;background:#0d121d;color:#f2f5fb;padding:10px;font:inherit}button{border:0;border-radius:999px;padding:10px 16px;background:#9bb8ff;color:#10131b;font:inherit;font-weight:700;cursor:pointer}button:disabled{opacity:.5;cursor:wait}.readout,.hint{color:#a9b3c8;font-size:.88rem}.gallery{padding:4px;display:grid;gap:8px;align-content:start}.history{flex:none;min-height:0;overflow:visible}.gallery h2{font-size:1rem;margin:8px}.item{width:100%;background:#101622;border:1px solid #2c3750;border-radius:12px;padding:11px;text-align:left;color:#e8eefc}.item:hover,.item.active{border-color:#9bb8ff;background:#19233a}.item p{margin:0 0 6px;font-size:.88rem;line-height:1.35}.status{font-size:.75rem;color:#a9b3c8}.error{color:#ffacac;display:-webkit-box;-webkit-line-clamp:3;-webkit-box-orient:vertical;overflow:hidden}.camera{flex:1 1 100%;color:#a9b3c8;font-size:.82rem}.progress{display:inline-flex;align-items:center;gap:7px;color:#d9e5ff;font-weight:600}.progress::before{content:'';width:8px;height:8px;border-radius:50%;background:#9bb8ff;box-shadow:0 0 0 0 #9bb8ff;animation:pulse 1.25s infinite}@keyframes pulse{70%{box-shadow:0 0 0 8px #9bb8ff00}}.item.running{border-color:#739cf7}@media(max-width:850px){body{overflow:auto}main{height:auto;min-height:100vh;grid-template-columns:1fr}.sidebar{overflow:visible}.stage{height:min(70vh,700px)}}
:root{color-scheme:dark;font-family:Inter,system-ui,sans-serif;background:#10131b;color:#f2f5fb}body{margin:0;overflow:hidden;background:#10131b}main{width:100vw;height:100vh;display:grid;grid-template-columns:360px minmax(0,1fr)}.sidebar{padding:24px 18px;display:flex;flex-direction:column;gap:18px;overflow-y:auto;overflow-x:hidden;background:radial-gradient(circle at 0 0,#293b62,transparent 32rem),#141a27;border-right:1px solid #2b354b}h1{font-size:2rem;letter-spacing:-.05em;margin:0}.eyebrow{font-size:.7rem;text-transform:uppercase;letter-spacing:.14em;color:#9bb8ff;margin-bottom:8px}.card{background:#171c28dd;border:1px solid #2b354b;border-radius:14px;overflow:hidden}.stage{min-width:0;min-height:0;display:flex;flex-direction:column;background:#0c1019}canvas{display:block;width:100%;height:100%;flex:1;min-height:0;background:linear-gradient(#111b31,#0c1019);cursor:grab;touch-action:none}canvas.dragging{cursor:grabbing}.controls{padding:14px 18px;display:flex;align-items:center;gap:10px;flex-wrap:wrap;border-top:1px solid #2b354b}.promptbox{padding:16px;display:grid;gap:10px}.promptbox textarea{min-height:110px;resize:none;overflow:hidden}textarea,input{box-sizing:border-box;width:100%;border:1px solid #35415b;border-radius:10px;background:#0d121d;color:#f2f5fb;padding:10px;font:inherit}button,.download{border:0;border-radius:999px;padding:10px 16px;background:#9bb8ff;color:#10131b;font:inherit;font-weight:700;cursor:pointer;text-decoration:none}button:disabled{opacity:.5;cursor:wait}.download{background:#24313a;color:#dce9e8}.readout,.hint{color:#a9b3c8;font-size:.88rem}.gallery{padding:4px;display:grid;gap:8px;align-content:start}.history{flex:none;min-height:0;overflow:visible}.gallery h2{font-size:1rem;margin:8px}.item{width:100%;background:#101622;border:1px solid #2c3750;border-radius:12px;padding:11px;text-align:left;color:#e8eefc}.item:hover,.item.active{border-color:#9bb8ff;background:#19233a}.item p{margin:0 0 6px;font-size:.88rem;line-height:1.35}.status{font-size:.75rem;color:#a9b3c8}.error{color:#ffacac;display:-webkit-box;-webkit-line-clamp:3;-webkit-box-orient:vertical;overflow:hidden}.camera{flex:1 1 100%;color:#a9b3c8;font-size:.82rem}.progress{display:inline-flex;align-items:center;gap:7px;color:#d9e5ff;font-weight:600}.progress::before{content:'';width:8px;height:8px;border-radius:50%;background:#9bb8ff;box-shadow:0 0 0 0 #9bb8ff;animation:pulse 1.25s infinite}@keyframes pulse{70%{box-shadow:0 0 0 8px #9bb8ff00}}.item.running{border-color:#739cf7}@media(max-width:850px){body{overflow:auto}main{height:auto;min-height:100vh;grid-template-columns:1fr}.sidebar{overflow:visible}.stage{height:min(70vh,700px)}}
/* LocalAI house style: dark graphite surfaces with a cyan/violet edge. */
:root{--la-bg:#080b0f;--la-panel:#10151b;--la-panel-2:#151c23;--la-line:#25313b;--la-text:#f3f6f4;--la-muted:#8b9aa5;--la-cyan:#3bd3d0;--la-cyan-2:#5ee6be;--la-violet:#9f7aea}body{background:var(--la-bg);color:var(--la-text)}.sidebar{background:radial-gradient(circle at 18% 0,rgba(59,211,208,.12),transparent 26rem),var(--la-panel);border-color:var(--la-line);box-sizing:border-box;min-height:0;min-width:0}.card{background:var(--la-panel-2);border-color:var(--la-line)}.stage{background:#090d12}canvas{background:radial-gradient(circle at 50% 25%,#102632 0,#090d12 56%)}.eyebrow{color:var(--la-cyan);font-family:ui-monospace,SFMono-Regular,Menlo,monospace}.product-title{display:flex;align-items:center;gap:11px}.product-mark{display:grid;place-items:center;width:36px;height:36px;border-radius:11px;background:linear-gradient(135deg,var(--la-cyan),#218eac 55%,var(--la-violet));color:#051013;font-size:20px;font-weight:900;box-shadow:0 0 28px rgba(59,211,208,.2)}.localai-logo{position:fixed;top:18px;right:18px;z-index:2;width:56px;height:56px;object-fit:contain;filter:drop-shadow(0 5px 16px #0009)}.promptbox label,.gallery h2{color:var(--la-text)}textarea,input{border-color:var(--la-line);background:#0b1015}button{background:linear-gradient(135deg,var(--la-cyan),var(--la-cyan-2));color:#051013;border-radius:8px;box-shadow:0 8px 22px rgba(59,211,208,.12)}.history,.gallery,#items{min-width:0}.item{background:#0b1015;border-color:var(--la-line);box-sizing:border-box;min-width:0;white-space:normal;overflow-wrap:anywhere}.item:hover,.item.active{border-color:var(--la-cyan);background:#111d24}.status,.readout,.hint,.camera{color:var(--la-muted)}.progress{color:var(--la-cyan-2)}.progress::before{background:var(--la-cyan);box-shadow:0 0 0 0 var(--la-cyan)}.item.running{border-color:var(--la-cyan)}
</style><main><aside class="sidebar"><header><div class="eyebrow">Kimodo-SMPLX-RP-v1 · Vulkan</div><div class="product-title"><span class="product-mark">K</span><h1>Text to motion</h1></div></header><section class="card promptbox"><label for="prompt">Describe a motion</label><textarea id="prompt">A person runs forward and then leaps over an obstacle in front of them.</textarea><button id="generate">Generate motion</button><span id="status" class="readout">Ready</span><div class="hint">One generation runs at a time. Select an animation below to restore and edit its prompt.</div></section><section class="card history"><div class="gallery"><h2>Past animations</h2><div id="items" class="readout">Loading…</div></div></section></aside><section class="stage"><canvas id="view" width="1280" height="720" aria-label="Animated SMPL-X skeleton"></canvas><div class="controls"><button id="play">Pause</button><button id="reset">Reset view</button><input id="frame" type="range" min="0" value="0" step="1"><span id="frameText" class="readout"></span><span class="camera">Camera tracks motion · Drag: rotate · wheel: zoom · Shift/right-drag: pan · double-click: reset</span></div></section></main><img class="localai-logo" src="/localai.png" alt="LocalAI">
<style>.promptbox,.history{flex:0 0 auto}</style>
</style><main><aside class="sidebar"><header><div class="eyebrow">Kimodo-SMPLX-RP-v1 · Vulkan</div><div class="product-title"><span class="product-mark">K</span><h1>Text to motion</h1></div></header><section class="card promptbox"><label for="prompt">Describe a motion</label><textarea id="prompt">A person runs forward and then leaps over an obstacle in front of them.</textarea><button id="generate">Generate motion</button><span id="status" class="readout">Ready</span><div class="hint">One generation runs at a time. Select an animation below to restore and edit its prompt.</div></section><section class="card history"><div class="gallery"><h2>Past animations</h2><div id="items" class="readout">Loading…</div></div></section></aside><section class="stage"><canvas id="view" width="1280" height="720" aria-label="Animated SMPL-X skeleton"></canvas><div class="controls"><button id="play">Pause</button><button id="reset">Reset view</button><a id="download" class="download" hidden download>Download GLB</a><input id="frame" type="range" min="0" value="0" step="1"><span id="frameText" class="readout"></span><span class="camera">Camera tracks motion · Drag: rotate · wheel: zoom · Shift/right-drag: pan · double-click: reset</span></div></section></main><img class="localai-logo" src="/localai.png" alt="LocalAI">
<style>.promptbox,.history{flex:0 0 auto}.license-warning{color:#ffbf78!important;font-weight:650}</style>
<script src="/models.js"></script><script>
const parents=[-1,0,0,0,1,2,3,4,5,6,7,8,9,9,9,12,13,14,16,17,18,19];
let parents=[-1,0,0,0,1,2,3,4,5,6,7,8,9,9,9,12,13,14,16,17,18,19];
// SMPL-X22 rest offsets calibrated from the captured upstream posed-joint /
// global-rotation fixture. The generator persists local rotations and root
// translations; these fixed parent-local bone vectors make them viewable here.
const offsets=[[0,0,0],[.052299,-.093936,-.027607],[-.057193,-.106548,-.022218],[-.001496,.11293,-.024981],[.058867,-.416442,-.006557],[-.048074,-.39756,-.014061],[.0069,.145636,-.006859],[-.041738,-.437584,-.029512],[.014489,-.446853,-.01803],[-.010334,.056082,.021116],[.049294,-.065279,.126259],[-.040575,-.065287,.127076],[-.011026,.171365,-.028827],[.047725,.087643,-.008375],[-.046636,.086612,-.014864],[.024654,.175391,.024463],[.126285,.05768,-.013885],[-.109342,.053674,-.009118],[.272907,-.069853,-.039094],[-.292029,-.03544,-.024565],[.276174,.021254,-.002478],[-.271878,-.004835,-.016445]];
const canvas=document.querySelector('#view'),ctx=canvas.getContext('2d'),slider=document.querySelector('#frame'),promptBox=document.querySelector('#prompt'),generate=document.querySelector('#generate'),status=document.querySelector('#status'),items=document.querySelector('#items');let selected,root,rotations,frame=0,playbackBlend=0,playing=true,last=0,cameraLast=0,animations=[],activeRequest,activeStarted;const defaultView={yaw:-.68,pitch:-.20,zoom:2700,panX:0,panY:0,cameraX:0,cameraZ:0,cameraZoom:2700,cameraSpeed:0},view={...defaultView};window.addEventListener('kimodo:sequence-controls-ready',()=>{if(selected)window.dispatchEvent(new CustomEvent('kimodo:restore-sequence',{detail:{segments:selected.segments,model:selected.model}}))});
let offsets=[[0,0,0],[.052299,-.093936,-.027607],[-.057193,-.106548,-.022218],[-.001496,.11293,-.024981],[.058867,-.416442,-.006557],[-.048074,-.39756,-.014061],[.0069,.145636,-.006859],[-.041738,-.437584,-.029512],[.014489,-.446853,-.01803],[-.010334,.056082,.021116],[.049294,-.065279,.126259],[-.040575,-.065287,.127076],[-.011026,.171365,-.028827],[.047725,.087643,-.008375],[-.046636,.086612,-.014864],[.024654,.175391,.024463],[.126285,.05768,-.013885],[-.109342,.053674,-.009118],[.272907,-.069853,-.039094],[-.292029,-.03544,-.024565],[.276174,.021254,-.002478],[-.271878,-.004835,-.016445]];
window.addEventListener('kimodo:restore-sequence',event=>{const id=event.detail?.model||'smplx-rp-v1',model=window.kimodoModels?.find(item=>item.id===id);if(model?.parents?.length&&model.offsets?.length===model.parents.length){parents=model.parents;offsets=model.offsets}});
const canvas=document.querySelector('#view'),ctx=canvas.getContext('2d'),slider=document.querySelector('#frame'),promptBox=document.querySelector('#prompt'),generate=document.querySelector('#generate'),status=document.querySelector('#status'),items=document.querySelector('#items'),download=document.querySelector('#download');let selected,root,rotations,frame=0,playbackBlend=0,playing=true,last=0,cameraLast=0,animations=[],activeRequest,activeStarted;const defaultView={yaw:-.68,pitch:-.20,zoom:2700,panX:0,panY:0,cameraX:0,cameraZ:0,cameraZoom:2700,cameraSpeed:0},view={...defaultView};window.addEventListener('kimodo:sequence-controls-ready',()=>{if(selected)window.dispatchEvent(new CustomEvent('kimodo:restore-sequence',{detail:{segments:selected.segments,model:selected.model}}))});
function renderedRoot(){if(!selected||!root)return[0,0,0];const at=frame*3,next=((frame+1)%selected.frames)*3,b=playbackBlend;return[root[at]*(1-b)+root[next]*b,root[at+1]*(1-b)+root[next+1]*b,root[at+2]*(1-b)+root[next+2]*b]}function snapCamera(){const[x,,z]=renderedRoot();view.cameraX=x;view.cameraZ=z;view.cameraZoom=view.zoom;view.cameraSpeed=0;cameraLast=performance.now()}function reset(){Object.assign(view,defaultView);snapCamera();draw()}function updateCamera(now){if(!selected||!root)return;const dt=Math.min(.1,Math.max(0,(now-cameraLast)/1000));cameraLast=now;const at=frame*3,next=((frame+1)%selected.frames)*3,[rootX,,rootZ]=renderedRoot(),speed=playing?Math.hypot(root[next]-root[at],root[next+2]-root[at+2])*30:0,dx=rootX-view.cameraX,dz=rootZ-view.cameraZ,distance=Math.hypot(dx,dz),deadZone=.18;const targetX=distance>deadZone?rootX-dx/distance*deadZone:view.cameraX,targetZ=distance>deadZone?rootZ-dz/distance*deadZone:view.cameraZ,follow=1-Math.exp(-3*dt),speedFollow=1-Math.exp(-2.4*dt),zoomFollow=1-Math.exp(-2.6*dt);view.cameraSpeed+=(speed-view.cameraSpeed)*speedFollow;const targetZoom=view.zoom*(1-Math.min(.22,view.cameraSpeed*.035));view.cameraX+=(targetX-view.cameraX)*follow;view.cameraZ+=(targetZ-view.cameraZ)*follow;view.cameraZoom+=(targetZoom-view.cameraZoom)*zoomFollow}function rotate(q,v){const[x,y,z,w]=q,[vx,vy,vz]=v,tx=2*(y*vz-z*vy),ty=2*(z*vx-x*vz),tz=2*(x*vy-y*vx);return[vx+w*tx+y*tz-z*ty,vy+w*ty+z*tx-x*tz,vz+w*tz+x*ty-y*tx]};function add(a,b){return[a[0]+b[0],a[1]+b[1],a[2]+b[2]]}function multiply(a,b){const[x,y,z,w]=a,[X,Y,Z,W]=b;return[x*W+w*X+y*Z-z*Y,y*W+w*Y+z*X-x*Z,z*W+w*Z+x*Y-y*X,w*W-x*X-y*Y-z*Z]}
function pose(){if(!selected||!root||!rotations)return[];const positions=[],global=[],rootPosition=renderedRoot();for(let j=0;j<22;j++){const q=Array.from(rotations.subarray((frame*22+j)*4,(frame*22+j+1)*4)),p=parents[j];if(p<0){global[j]=q;positions[j]=rootPosition}else{global[j]=multiply(global[p],q);positions[j]=add(positions[p],rotate(global[p],offsets[j]))}}return positions}function project([x,y,z]){x-=view.cameraX;z-=view.cameraZ;const rx=x*Math.cos(view.yaw)-z*Math.sin(view.yaw),rz=x*Math.sin(view.yaw)+z*Math.cos(view.yaw),ry=y*Math.cos(view.pitch)-rz*Math.sin(view.pitch),dz=y*Math.sin(view.pitch)+rz*Math.cos(view.pitch)+8;return[canvas.width/2+view.panX+view.cameraZoom*rx/dz,canvas.height*.78+view.panY-view.cameraZoom*ry/dz]}
function ground(){const cx=0,cz=0,step=.5,initialExtent=6,drawGrid=(extent,alpha,colour,width)=>{ctx.globalAlpha=alpha;ctx.strokeStyle=colour;ctx.lineWidth=width;for(let i=-extent/step;i<=extent/step;i++){let a=project([cx-extent,0,cz+i*step]),b=project([cx+extent,0,cz+i*step]);ctx.beginPath();ctx.moveTo(a[0],a[1]);ctx.lineTo(b[0],b[1]);ctx.stroke();a=project([cx+i*step,0,cz-extent]),b=project([cx+i*step,0,cz+extent]);ctx.beginPath();ctx.moveTo(a[0],a[1]);ctx.lineTo(b[0],b[1]);ctx.stroke()}};const[x,,z]=renderedRoot(),rootDistance=Math.max(Math.abs(x),Math.abs(z)),outerAlpha=.72*Math.max(0,Math.min(1,(rootDistance-3.5)/2.5));if(outerAlpha)drawGrid(18,outerAlpha,'#173943',1);drawGrid(initialExtent,1,'#173943',2);ctx.globalAlpha=1;ctx.strokeStyle='#23606b';ctx.lineWidth=3;const xAxis=project([cx-initialExtent,0,cz]),zAxis=project([cx,0,cz-initialExtent]),xEnd=project([cx+initialExtent,0,cz]),zEnd=project([cx,0,cz+initialExtent]);ctx.beginPath();ctx.moveTo(xAxis[0],xAxis[1]);ctx.lineTo(xEnd[0],xEnd[1]);ctx.moveTo(zAxis[0],zAxis[1]);ctx.lineTo(zEnd[0],zEnd[1]);ctx.stroke()}function draw(){ctx.clearRect(0,0,canvas.width,canvas.height);const p=pose();ground();if(p.length){ctx.lineCap='round';for(let pass=0;pass<2;pass++){ctx.strokeStyle=pass?'#65eee1':'#071014';ctx.lineWidth=pass?7:13;for(let i=1;i<22;i++){const a=project(p[i]),b=project(p[parents[i]]);ctx.beginPath();ctx.moveTo(a[0],a[1]);ctx.lineTo(b[0],b[1]);ctx.stroke()}}for(const x of p){const a=project(x);ctx.fillStyle='#071014';ctx.beginPath();ctx.arc(a[0],a[1],7,0,Math.PI*2);ctx.fill();ctx.fillStyle='#ecfffd';ctx.beginPath();ctx.arc(a[0],a[1],4,0,Math.PI*2);ctx.fill()}}slider.value=frame;document.querySelector('#frameText').textContent=selected?`frame ${frame+1} / ${selected.frames}`:'No animation selected'}function tick(t){if(selected){const interval=1000/30;if(!last)last=t;if(playing){while(t-last>=interval){frame=(frame+1)%selected.frames;last+=interval}playbackBlend=Math.min(1,(t-last)/interval)}else playbackBlend=0;updateCamera(t);draw()}requestAnimationFrame(tick)}
async function select(a){if(a.status!=='ready')return;selected=a;promptBox.value=a.prompt;window.dispatchEvent(new CustomEvent('kimodo:restore-sequence',{detail:{segments:a.segments,model:a.model}}));status.className='readout';status.textContent=`Selected ${a.id.slice(0,8)} · sequence restored`;[root,rotations]=await Promise.all([fetch(`/api/animations/${a.id}/root.f32`).then(r=>r.arrayBuffer()).then(b=>new Float32Array(b)),fetch(`/api/animations/${a.id}/rotations.f32`).then(r=>r.arrayBuffer()).then(b=>new Float32Array(b))]);frame=0;playbackBlend=0;last=performance.now();snapCamera();slider.max=a.frames-1;renderGallery();draw()}function renderGallery(){if(!animations.length){items.textContent='No animations yet.';return}items.replaceChildren(...animations.map(a=>{const b=document.createElement('button');b.className='item '+a.status+(selected?.id===a.id?' active':'');b.disabled=a.status!=='ready';b.innerHTML=`<p>${a.prompt}</p><span class="status">${a.status} · ${a.frames} frames · ${a.diffusion_steps} steps</span>${a.error?`<div class="error">${a.error}</div>`:''}`;b.onclick=()=>select(a);return b}))}function showProgress(){const a=animations.find(a=>a.id===activeRequest);if(!a)return;if(a.status==='ready'){status.className='readout';status.textContent='Generation complete — select it from the gallery to play it.';activeRequest=undefined;generate.disabled=false;return}if(a.status==='failed'){status.className='error';status.textContent=`Generation failed: ${a.error}`;activeRequest=undefined;generate.disabled=false;return}const seconds=Math.max(0,Math.floor((Date.now()-activeStarted)/1000));status.className='readout progress';status.textContent=a.status==='running'?`Generating motion… ${seconds}s elapsed`:`Queued for generation… ${seconds}s elapsed`;generate.disabled=true}async function refresh(){animations=await fetch('/api/animations').then(r=>r.json());renderGallery();showProgress()}generate.onclick=async()=>{const prompt=promptBox.value.trim();if(!prompt)return;generate.disabled=true;status.className='readout progress';status.textContent='Submitting generation…';try{const r=await fetch('/api/generate',{method:'POST',headers:{'Content-Type':'application/json'},body:JSON.stringify({prompt,frames:150,steps:100,seed:0})});if(!r.ok)throw new Error(await r.text());const a=await r.json();activeRequest=a.id;activeStarted=Date.now();await refresh()}catch(e){status.className='error';status.textContent=e.message;generate.disabled=false}};
function pose(){if(!selected||!root||!rotations)return[];const positions=[],global=[],rootPosition=renderedRoot(),joints=parents.length;for(let j=0;j<joints;j++){const q=Array.from(rotations.subarray((frame*joints+j)*4,(frame*joints+j+1)*4)),p=parents[j];if(p<0){global[j]=q;positions[j]=rootPosition}else{global[j]=multiply(global[p],q);positions[j]=add(positions[p],rotate(global[p],offsets[j]))}}return positions}function project([x,y,z]){x-=view.cameraX;z-=view.cameraZ;const rx=x*Math.cos(view.yaw)-z*Math.sin(view.yaw),rz=x*Math.sin(view.yaw)+z*Math.cos(view.yaw),ry=y*Math.cos(view.pitch)-rz*Math.sin(view.pitch),dz=y*Math.sin(view.pitch)+rz*Math.cos(view.pitch)+8;return[canvas.width/2+view.panX+view.cameraZoom*rx/dz,canvas.height*.78+view.panY-view.cameraZoom*ry/dz]}
function ground(){const cx=0,cz=0,step=.5,initialExtent=6,drawGrid=(extent,alpha,colour,width)=>{ctx.globalAlpha=alpha;ctx.strokeStyle=colour;ctx.lineWidth=width;for(let i=-extent/step;i<=extent/step;i++){let a=project([cx-extent,0,cz+i*step]),b=project([cx+extent,0,cz+i*step]);ctx.beginPath();ctx.moveTo(a[0],a[1]);ctx.lineTo(b[0],b[1]);ctx.stroke();a=project([cx+i*step,0,cz-extent]),b=project([cx+i*step,0,cz+extent]);ctx.beginPath();ctx.moveTo(a[0],a[1]);ctx.lineTo(b[0],b[1]);ctx.stroke()}};const[x,,z]=renderedRoot(),rootDistance=Math.max(Math.abs(x),Math.abs(z)),outerAlpha=.72*Math.max(0,Math.min(1,(rootDistance-3.5)/2.5));if(outerAlpha)drawGrid(18,outerAlpha,'#173943',1);drawGrid(initialExtent,1,'#173943',2);ctx.globalAlpha=1;ctx.strokeStyle='#23606b';ctx.lineWidth=3;const xAxis=project([cx-initialExtent,0,cz]),zAxis=project([cx,0,cz-initialExtent]),xEnd=project([cx+initialExtent,0,cz]),zEnd=project([cx,0,cz+initialExtent]);ctx.beginPath();ctx.moveTo(xAxis[0],xAxis[1]);ctx.lineTo(xEnd[0],xEnd[1]);ctx.moveTo(zAxis[0],zAxis[1]);ctx.lineTo(zEnd[0],zEnd[1]);ctx.stroke()}function draw(){ctx.clearRect(0,0,canvas.width,canvas.height);const p=pose();ground();if(p.length){ctx.lineCap='round';for(let pass=0;pass<2;pass++){ctx.strokeStyle=pass?'#65eee1':'#071014';ctx.lineWidth=pass?7:13;for(let i=1;i<parents.length;i++){const a=project(p[i]),b=project(p[parents[i]]);ctx.beginPath();ctx.moveTo(a[0],a[1]);ctx.lineTo(b[0],b[1]);ctx.stroke()}}for(const x of p){const a=project(x);ctx.fillStyle='#071014';ctx.beginPath();ctx.arc(a[0],a[1],7,0,Math.PI*2);ctx.fill();ctx.fillStyle='#ecfffd';ctx.beginPath();ctx.arc(a[0],a[1],4,0,Math.PI*2);ctx.fill()}}slider.value=frame;document.querySelector('#frameText').textContent=selected?`frame ${frame+1} / ${selected.frames}`:'No animation selected'}function tick(t){if(selected){const interval=1000/30;if(!last)last=t;if(playing){while(t-last>=interval){frame=(frame+1)%selected.frames;last+=interval}playbackBlend=Math.min(1,(t-last)/interval)}else playbackBlend=0;updateCamera(t);draw()}requestAnimationFrame(tick)}
async function select(a){if(a.status!=='ready')return;selected=a;download.href=`/api/animations/${a.id}/animation.glb`;download.download=`kimodo-${a.id}.glb`;download.hidden=false;promptBox.value=a.prompt;window.dispatchEvent(new CustomEvent('kimodo:restore-sequence',{detail:{segments:a.segments,model:a.model}}));status.className='readout';status.textContent=`Selected ${a.id.slice(0,8)} · sequence restored`;[root,rotations]=await Promise.all([fetch(`/api/animations/${a.id}/root.f32`).then(r=>r.arrayBuffer()).then(b=>new Float32Array(b)),fetch(`/api/animations/${a.id}/rotations.f32`).then(r=>r.arrayBuffer()).then(b=>new Float32Array(b))]);frame=0;playbackBlend=0;last=performance.now();snapCamera();slider.max=a.frames-1;renderGallery();draw()}function renderGallery(){if(!animations.length){items.textContent='No animations yet.';return}items.replaceChildren(...animations.map(a=>{const b=document.createElement('button');b.className='item '+a.status+(selected?.id===a.id?' active':'');b.disabled=a.status!=='ready';b.innerHTML=`<p>${a.prompt}</p><span class="status">${a.status} · ${a.frames} frames · ${a.diffusion_steps} steps</span>${a.error?`<div class="error">${a.error}</div>`:''}`;b.onclick=()=>select(a);return b}))}function showProgress(){const a=animations.find(a=>a.id===activeRequest);if(!a)return;if(a.status==='ready'){status.className='readout';status.textContent='Generation complete — select it from the gallery to play it.';activeRequest=undefined;generate.disabled=false;return}if(a.status==='failed'){status.className='error';status.textContent=`Generation failed: ${a.error}`;activeRequest=undefined;generate.disabled=false;return}const seconds=Math.max(0,Math.floor((Date.now()-activeStarted)/1000));status.className='readout progress';status.textContent=a.status==='running'?`Generating motion… ${seconds}s elapsed`:`Queued for generation… ${seconds}s elapsed`;generate.disabled=true}async function refresh(){animations=await fetch('/api/animations').then(r=>r.json());renderGallery();showProgress()}generate.onclick=async()=>{const prompt=promptBox.value.trim();if(!prompt)return;generate.disabled=true;status.className='readout progress';status.textContent='Submitting generation…';try{const r=await fetch('/api/generate',{method:'POST',headers:{'Content-Type':'application/json'},body:JSON.stringify({prompt,frames:150,steps:100,seed:0})});if(!r.ok)throw new Error(await r.text());const a=await r.json();activeRequest=a.id;activeStarted=Date.now();await refresh()}catch(e){status.className='error';status.textContent=e.message;generate.disabled=false}};
document.querySelector('#play').onclick=e=>{playing=!playing;last=performance.now();playbackBlend=0;e.target.textContent=playing?'Pause':'Play'};slider.oninput=()=>{frame=Number(slider.value);playbackBlend=0;last=performance.now();draw()};document.querySelector('#reset').onclick=reset;canvas.addEventListener('dblclick',reset);canvas.addEventListener('contextmenu',e=>e.preventDefault());let drag;canvas.addEventListener('pointerdown',e=>{canvas.setPointerCapture(e.pointerId);drag={x:e.clientX,y:e.clientY,pan:e.button===2||e.shiftKey};canvas.classList.add('dragging')});canvas.addEventListener('pointermove',e=>{if(!drag)return;const dx=e.clientX-drag.x,dy=e.clientY-drag.y;drag.x=e.clientX;e.clientY;drag.y=e.clientY;if(drag.pan){view.panX+=dx;view.panY+=dy}else{view.yaw+=dx*.008;view.pitch=Math.max(-1.25,Math.min(1.25,view.pitch+dy*.008))}draw()});function end(){drag=undefined;canvas.classList.remove('dragging')}canvas.addEventListener('pointerup',end);canvas.addEventListener('pointercancel',end);canvas.addEventListener('wheel',e=>{e.preventDefault();view.zoom=Math.max(350,Math.min(3600,view.zoom*Math.exp(-e.deltaY*.001)));draw()},{passive:false});refresh().then(()=>{const a=animations.find(a=>a.status==='ready');if(a)return select(a);draw()});setInterval(refresh,2500);requestAnimationFrame(tick);
</script></html>

View File

@ -51,13 +51,19 @@ type promptSegment struct {
Frames int `json:"frames"`
}
type motionModel struct {
ID string `json:"id"`
Label string `json:"label"`
Skeleton string `json:"skeleton"`
Upstream string `json:"upstream"`
Available bool `json:"available"`
Reason string `json:"reason,omitempty"`
Motion string `json:"-"`
ID string `json:"id"`
Label string `json:"label"`
Skeleton string `json:"skeleton"`
SkeletonKey string `json:"skeleton_key"`
Upstream string `json:"upstream"`
License string `json:"license"`
LicenseURL string `json:"license_url"`
Commercial bool `json:"commercial"`
Available bool `json:"available"`
Reason string `json:"reason,omitempty"`
Parents []int `json:"parents"`
Offsets [][3]float32 `json:"offsets"`
Motion string `json:"-"`
}
type gallery struct {
mu sync.RWMutex
@ -94,14 +100,6 @@ func (g *gallery) list() []*animation {
return result
}
func copyFile(dst, src string) error {
b, err := os.ReadFile(src)
if err != nil {
return err
}
return os.WriteFile(dst, b, 0600)
}
func readF32(path string) ([]float32, error) {
b, err := os.ReadFile(path)
if err != nil {
@ -125,79 +123,159 @@ func writeF32(path string, values []float32) error {
return os.WriteFile(path, b, 0600)
}
func blendQuaternion(a, b []float32, alpha float32) {
dot := a[0]*b[0] + a[1]*b[1] + a[2]*b[2] + a[3]*b[3]
if dot < 0 {
for i := range b {
b[i] = -b[i]
}
}
length := float32(0)
for i := range a {
a[i] = alpha*a[i] + (1-alpha)*b[i]
length += a[i] * a[i]
}
if length > 0 {
length = 1 / float32(math.Sqrt(float64(length)))
for i := range a {
a[i] *= length
}
}
// Each motion is exported as a node-only GLB: it deliberately has no mesh or
// skin, so consumers can attach their own Three.js geometry to the named
// joints. Kimodo stores root translations and local XYZW rotations.
var smplx22Parents = [...]int{-1, 0, 0, 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 9, 9, 12, 13, 14, 16, 17, 18, 19}
var smplx22Names = [...]string{"pelvis", "left_hip", "right_hip", "spine1", "left_knee", "right_knee", "spine2", "left_ankle", "right_ankle", "spine3", "left_foot", "right_foot", "neck", "left_collar", "right_collar", "head", "left_shoulder", "right_shoulder", "left_elbow", "right_elbow", "left_wrist", "right_wrist"}
var smplx22Offsets = [...][3]float32{{}, {.052299, -.093936, -.027607}, {-.057193, -.106548, -.022218}, {-.001496, .11293, -.024981}, {.058867, -.416442, -.006557}, {-.048074, -.39756, -.014061}, {.0069, .145636, -.006859}, {-.041738, -.437584, -.029512}, {.014489, -.446853, -.01803}, {-.010334, .056082, .021116}, {.049294, -.065279, .126259}, {-.040575, -.065287, .127076}, {-.011026, .171365, -.028827}, {.047725, .087643, -.008375}, {-.046636, .086612, -.014864}, {.024654, .175391, .024463}, {.126285, .05768, -.013885}, {-.109342, .053674, -.009118}, {.272907, -.069853, -.039094}, {-.292029, -.03544, -.024565}, {.276174, .021254, -.002478}, {-.271878, -.004835, -.016445}}
type skeletonDefinition struct {
key string
names []string
parents []int
offsets [][3]float32
}
// stitchSegments joins independently sampled demo segments. The overlap is
// blended in root space and by normalized linear interpolation for quaternions.
// Native observed-motion conditioning is deliberately a later parity step.
func stitchSegments(output string, dirs []string, overlap int) error {
var roots, rotations []float32
for index, dir := range dirs {
root, err := readF32(filepath.Join(dir, "root_positions.f32"))
if err != nil {
return err
}
rot, err := readF32(filepath.Join(dir, "local_rotations_xyzw.f32"))
if err != nil {
return err
}
frames := len(root) / 3
if frames == 0 || len(rot) != frames*22*4 {
return fmt.Errorf("invalid motion segment %d", index+1)
}
if index == 0 {
roots, rotations = root, rot
continue
}
n := overlap
if n > frames {
n = frames
}
if n > len(roots)/3 {
n = len(roots) / 3
}
for frame := 0; frame < n; frame++ {
alpha := float32(0.5)
if n > 1 {
alpha = 1 - float32(frame)/float32(n-1)
}
old := (len(roots)/3 - n + frame) * 3
newest := frame * 3
for axis := 0; axis < 3; axis++ {
roots[old+axis] = alpha*roots[old+axis] + (1-alpha)*root[newest+axis]
}
for joint := 0; joint < 22; joint++ {
oldQ := (len(rotations)/4 - n*22 + frame*22 + joint) * 4
newQ := (frame*22 + joint) * 4
blendQuaternion(rotations[oldQ:oldQ+4], append([]float32(nil), rot[newQ:newQ+4]...), alpha)
}
}
roots = append(roots, root[n*3:]...)
rotations = append(rotations, rot[n*22*4:]...)
var skeletonDefinitions = map[string]skeletonDefinition{
"smplx22": {key: "smplx22", names: smplx22Names[:], parents: smplx22Parents[:], offsets: smplx22Offsets[:]},
}
type gltfBufferView struct {
Buffer int `json:"buffer"`
ByteOffset int `json:"byteOffset,omitempty"`
ByteLength int `json:"byteLength"`
}
type gltfAccessor struct {
BufferView int `json:"bufferView"`
ComponentType int `json:"componentType"`
Count int `json:"count"`
Type string `json:"type"`
}
func appendF32(dst []byte, values []float32) []byte {
for _, value := range values {
var b [4]byte
binary.LittleEndian.PutUint32(b[:], math.Float32bits(value))
dst = append(dst, b[:]...)
}
if err := writeF32(filepath.Join(output, "root_positions.f32"), roots); err != nil {
return dst
}
func writeSkeletonGLB(path string, roots, rotations []float32, skeleton skeletonDefinition) error {
frames := len(roots) / 3
joints := len(skeleton.parents)
if frames < 1 || joints < 1 || len(skeleton.names) != joints || len(skeleton.offsets) != joints || len(roots) != frames*3 || len(rotations) != frames*joints*4 {
return fmt.Errorf("invalid %s motion for GLB export", skeleton.key)
}
times := make([]float32, frames)
for i := range times {
times[i] = float32(i) / 30
}
bin := make([]byte, 0, (frames+frames*3+frames*22*4)*4)
views := make([]gltfBufferView, 0, 24)
addView := func(values []float32) int {
offset := len(bin)
bin = appendF32(bin, values)
views = append(views, gltfBufferView{Buffer: 0, ByteOffset: offset, ByteLength: len(bin) - offset})
return len(views) - 1
}
timeView, rootView := addView(times), addView(roots)
rotationViews := make([]int, joints)
for joint := range rotationViews {
track := make([]float32, frames*4)
for frame := 0; frame < frames; frame++ {
copy(track[frame*4:], rotations[(frame*joints+joint)*4:(frame*joints+joint+1)*4])
}
rotationViews[joint] = addView(track)
}
accessors := []gltfAccessor{{BufferView: timeView, ComponentType: 5126, Count: frames, Type: "SCALAR"}, {BufferView: rootView, ComponentType: 5126, Count: frames, Type: "VEC3"}}
for _, view := range rotationViews {
accessors = append(accessors, gltfAccessor{BufferView: view, ComponentType: 5126, Count: frames, Type: "VEC4"})
}
nodes := make([]map[string]any, joints)
for joint := range nodes {
node := map[string]any{"name": skeleton.names[joint]}
if joint != 0 {
node["translation"] = skeleton.offsets[joint]
}
children := make([]int, 0, 3)
for child, parent := range skeleton.parents {
if parent == joint {
children = append(children, child)
}
}
if len(children) != 0 {
node["children"] = children
}
nodes[joint] = node
}
samplers := make([]map[string]any, 0, 23)
channels := make([]map[string]any, 0, 23)
addChannel := func(node, output int, path string) {
samplers = append(samplers, map[string]any{"input": 0, "output": output, "interpolation": "LINEAR"})
channels = append(channels, map[string]any{"sampler": len(samplers) - 1, "target": map[string]any{"node": node, "path": path}})
}
addChannel(0, 1, "translation")
for joint := 0; joint < joints; joint++ {
addChannel(joint, joint+2, "rotation")
}
document := map[string]any{
"asset": map[string]string{"version": "2.0", "generator": "kimodo.cpp skeleton exporter"},
"scene": 0,
"scenes": []map[string]any{{"nodes": []int{0}}},
"nodes": nodes,
"buffers": []map[string]int{{"byteLength": len(bin)}},
"bufferViews": views,
"accessors": accessors,
"animations": []map[string]any{{"name": "KimodoMotion", "samplers": samplers, "channels": channels}},
"extras": map[string]any{"skeleton": skeleton.key, "fps": 30, "rotation_order": "xyzw"},
}
jsonChunk, err := json.Marshal(document)
if err != nil {
return err
}
return writeF32(filepath.Join(output, "local_rotations_xyzw.f32"), rotations)
for len(jsonChunk)%4 != 0 {
jsonChunk = append(jsonChunk, ' ')
}
for len(bin)%4 != 0 {
bin = append(bin, 0)
}
total := 12 + 8 + len(jsonChunk) + 8 + len(bin)
out := make([]byte, 0, total)
putU32 := func(value uint32) {
var b [4]byte
binary.LittleEndian.PutUint32(b[:], value)
out = append(out, b[:]...)
}
putU32(0x46546c67)
putU32(2)
putU32(uint32(total))
putU32(uint32(len(jsonChunk)))
putU32(0x4e4f534a)
out = append(out, jsonChunk...)
putU32(uint32(len(bin)))
putU32(0x004e4942)
out = append(out, bin...)
return os.WriteFile(path, out, 0600)
}
func exportSkeletonGLB(dir, skeletonKey string) error {
skeleton, ok := skeletonDefinitions[skeletonKey]
if !ok {
return fmt.Errorf("unsupported skeleton %q", skeletonKey)
}
roots, err := readF32(filepath.Join(dir, "root_positions.f32"))
if err != nil {
return err
}
rotations, err := readF32(filepath.Join(dir, "local_rotations_xyzw.f32"))
if err != nil {
return err
}
return writeSkeletonGLB(filepath.Join(dir, "animation.glb"), roots, rotations, skeleton)
}
func (g *gallery) worker() {
for id := range g.queue {
g.mu.Lock()
@ -239,6 +317,9 @@ func (g *gallery) worker() {
err = fmt.Errorf("sequence: %w: %s", runErr, strings.TrimSpace(string(output)))
}
}
if err == nil {
err = exportSkeletonGLB(dir, model.SkeletonKey)
}
}
}
g.mu.Lock()
@ -259,6 +340,10 @@ func (g *gallery) worker() {
func main() {
addr := flag.String("addr", "127.0.0.1:8090", "listen address")
motion := flag.String("motion-model", "models/kimodo-smplx-rp-v1-f32.gguf", "motion GGUF")
somaRP := flag.String("soma-rp-model", "models/kimodo-soma-rp-v1.1-f32.gguf", "SOMA RP v1.1 motion GGUF")
somaSEED := flag.String("soma-seed-model", "models/kimodo-soma-seed-v1.1-f32.gguf", "SOMA SEED v1.1 motion GGUF")
g1RP := flag.String("g1-rp-model", "models/kimodo-g1-rp-v1-f32.gguf", "G1 RP v1 motion GGUF")
g1SEED := flag.String("g1-seed-model", "models/kimodo-g1-seed-v1-f32.gguf", "G1 SEED v1 motion GGUF")
text := flag.String("text-bundle", "generated/llm2vec-text-bundle", "native LLM2Vec component directory")
generator := flag.String("generator", "build/debug/kmd-generate", "native text-to-motion command")
output := flag.String("output", "demo-output", "persistent gallery directory")
@ -266,12 +351,24 @@ func main() {
if err := os.MkdirAll(*output, 0755); err != nil {
log.Fatal(err)
}
makeModel := func(id, label, skeletonLabel, skeletonKey, upstream, license, licenseURL, path string, commercial bool) motionModel {
definition := skeletonDefinitions[skeletonKey]
model := motionModel{ID: id, Label: label, Skeleton: skeletonLabel, SkeletonKey: skeletonKey, Upstream: upstream, License: license, LicenseURL: licenseURL, Commercial: commercial, Parents: definition.parents, Offsets: definition.offsets, Motion: path}
if info, err := os.Stat(path); err == nil && info.Mode().IsRegular() {
model.Available = true
} else {
model.Reason = "GGUF not found at " + path
}
return model
}
const internalLicense = "https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-internal-scientific-research-and-development-model-license/"
const openLicense = "https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/"
models := map[string]motionModel{
"smplx-rp-v1": {ID: "smplx-rp-v1", Label: "SMPL-X RP v1", Skeleton: "SMPL-X 22 joints", Upstream: "nvidia/Kimodo-SMPLX-RP-v1", Available: true, Motion: *motion},
"soma-rp-v1.1": {ID: "soma-rp-v1.1", Label: "SOMA RP v1.1", Skeleton: "SOMA 30 joints", Upstream: "nvidia/Kimodo-SOMA-RP-v1.1", Reason: "SOMA decoder and GGML conversion are being added"},
"soma-seed-v1.1": {ID: "soma-seed-v1.1", Label: "SOMA SEED v1.1", Skeleton: "SOMA 30 joints", Upstream: "nvidia/Kimodo-SOMA-SEED-v1.1", Reason: "SOMA decoder and GGML conversion are being added"},
"g1-rp-v1": {ID: "g1-rp-v1", Label: "G1 RP v1", Skeleton: "Unitree G1 34 joints", Upstream: "nvidia/Kimodo-G1-RP-v1", Reason: "G1 decoder and GGML conversion are being added"},
"g1-seed-v1": {ID: "g1-seed-v1", Label: "G1 SEED v1", Skeleton: "Unitree G1 34 joints", Upstream: "nvidia/Kimodo-G1-SEED-v1", Reason: "G1 decoder and GGML conversion are being added"},
"smplx-rp-v1": makeModel("smplx-rp-v1", "SMPL-X RP v1", "SMPL-X 22 joints", "smplx22", "nvidia/Kimodo-SMPLX-RP-v1", "NVIDIA Internal Scientific R&D (non-commercial)", internalLicense, *motion, false),
"soma-rp-v1.1": makeModel("soma-rp-v1.1", "SOMA RP v1.1", "SOMA compact 30-joint control skeleton", "soma30", "nvidia/Kimodo-SOMA-RP-v1.1", "NVIDIA Open Model License", openLicense, *somaRP, true),
"soma-seed-v1.1": makeModel("soma-seed-v1.1", "SOMA SEED v1.1", "SOMA compact 30-joint control skeleton", "soma30", "nvidia/Kimodo-SOMA-SEED-v1.1", "NVIDIA Open Model License", openLicense, *somaSEED, true),
"g1-rp-v1": makeModel("g1-rp-v1", "G1 RP v1", "Unitree G1 34 joints", "g1skel34", "nvidia/Kimodo-G1-RP-v1", "NVIDIA Open Model License", openLicense, *g1RP, true),
"g1-seed-v1": makeModel("g1-seed-v1", "G1 SEED v1", "Unitree G1 34 joints", "g1skel34", "nvidia/Kimodo-G1-SEED-v1", "NVIDIA Open Model License", openLicense, *g1SEED, true),
}
g := &gallery{items: map[string]*animation{}, output: *output, queue: make(chan string, 32), generator: *generator, motion: *motion, text: *text, models: models}
entries, _ := filepath.Glob(filepath.Join(*output, "*.json"))
@ -283,6 +380,15 @@ func main() {
var a animation
if json.Unmarshal(b, &a) == nil {
g.items[a.ID] = &a
if a.Status == "ready" {
model, ok := models[a.Model]
if !ok {
model = models["smplx-rp-v1"]
}
if err := exportSkeletonGLB(filepath.Join(*output, a.ID), model.SkeletonKey); err != nil && !os.IsNotExist(err) {
log.Printf("export existing animation %s: %v", a.ID, err)
}
}
}
}
go g.worker()
@ -400,7 +506,7 @@ func main() {
})
mux.HandleFunc("/api/animations/", func(w http.ResponseWriter, r *http.Request) {
parts := strings.Split(strings.TrimPrefix(r.URL.Path, "/api/animations/"), "/")
if len(parts) != 2 || (parts[1] != "root.f32" && parts[1] != "rotations.f32") {
if len(parts) != 2 || (parts[1] != "root.f32" && parts[1] != "rotations.f32" && parts[1] != "animation.glb") {
http.NotFound(w, r)
return
}
@ -415,7 +521,23 @@ func main() {
if parts[1] == "rotations.f32" {
name = "local_rotations_xyzw.f32"
}
w.Header().Set("Content-Type", "application/octet-stream")
if parts[1] == "animation.glb" {
name = "animation.glb"
w.Header().Set("Content-Type", "model/gltf-binary")
w.Header().Set("Content-Disposition", "attachment; filename=kimodo-"+a.ID+".glb")
// A GLB is a compact asset; read it directly so browsers always receive
// it as a download rather than invoking any path-cleaning redirects.
data, err := os.ReadFile(filepath.Join(g.output, a.ID, name))
if err != nil {
http.NotFound(w, r)
return
}
w.Header().Set("Content-Length", fmt.Sprint(len(data)))
_, _ = w.Write(data)
return
} else {
w.Header().Set("Content-Type", "application/octet-stream")
}
w.Header().Set("Cache-Control", "no-store")
http.ServeFile(w, r, filepath.Join(g.output, a.ID, name))
})

55
demo/main_test.go Normal file
View File

@ -0,0 +1,55 @@
package main
import (
"encoding/binary"
"encoding/json"
"os"
"path/filepath"
"testing"
)
func TestWriteSkeletonGLB(t *testing.T) {
for _, skeletonKey := range []string{"smplx22", "soma30", "g1skel34"} {
t.Run(skeletonKey, func(t *testing.T) {
skeleton := skeletonDefinitions[skeletonKey]
path := filepath.Join(t.TempDir(), "animation.glb")
roots := []float32{0, 0, 0, 1, 0, 0}
rotations := make([]float32, 2*len(skeleton.parents)*4)
for frame := 0; frame < 2; frame++ {
for joint := range skeleton.parents {
rotations[(frame*len(skeleton.parents)+joint)*4+3] = 1
}
}
if err := writeSkeletonGLB(path, roots, rotations, skeleton); err != nil {
t.Fatal(err)
}
assertSkeletonGLB(t, path, len(skeleton.parents))
})
}
}
func assertSkeletonGLB(t *testing.T, path string, expectedJoints int) {
t.Helper()
b, err := os.ReadFile(path)
if err != nil {
t.Fatal(err)
}
if len(b) < 20 || binary.LittleEndian.Uint32(b) != 0x46546c67 || binary.LittleEndian.Uint32(b[4:]) != 2 || int(binary.LittleEndian.Uint32(b[8:])) != len(b) {
t.Fatalf("invalid GLB header")
}
jsonLength := int(binary.LittleEndian.Uint32(b[12:]))
if binary.LittleEndian.Uint32(b[16:]) != 0x4e4f534a || 20+jsonLength > len(b) {
t.Fatalf("invalid GLB JSON chunk")
}
var document struct {
Asset map[string]string `json:"asset"`
Nodes []json.RawMessage `json:"nodes"`
Animations []json.RawMessage `json:"animations"`
}
if err := json.Unmarshal(b[20:20+jsonLength], &document); err != nil {
t.Fatal(err)
}
if document.Asset["version"] != "2.0" || len(document.Nodes) != expectedJoints || len(document.Animations) != 1 {
t.Fatalf("unexpected GLB document: %s", b[20:20+jsonLength])
}
}

View File

@ -5,6 +5,7 @@ window.addEventListener('load', async () => {
const generate = document.querySelector('#generate');
if (!prompt || !form || !generate) return;
const models = await fetch('/api/models').then(r => r.json());
window.kimodoModels = models;
const modelLabel = document.createElement('label');
modelLabel.htmlFor = 'motionModel'; modelLabel.textContent = 'Motion model';
@ -19,7 +20,14 @@ window.addEventListener('load', async () => {
const modelHint = document.createElement('div'); modelHint.className = 'hint';
const updateModel = () => {
const model = models.find(item => item.id === select.value);
modelHint.textContent = model.available ? `${model.skeleton} · ${model.upstream}` : `${model.skeleton} · ${model.reason}`;
if (!model) return;
const title = document.querySelector('.eyebrow');
if (title) title.textContent = `${model.label} · Vulkan`;
modelHint.classList.toggle('license-warning', !model.commercial);
const terms = model.commercial ? 'commercial use permitted under NVIDIA Open Model License' : '⚠ non-commercial research use only';
const detail = model.available ? `${model.skeleton} · ${model.upstream} · ` : `${model.skeleton} · ${model.reason} · `;
const link = document.createElement('a'); link.href = model.license_url; link.target = '_blank'; link.rel = 'noreferrer'; link.textContent = terms;
modelHint.replaceChildren(document.createTextNode(detail), link);
};
select.onchange = updateModel;
form.insertBefore(modelLabel, prompt); form.insertBefore(select, prompt); form.insertBefore(modelHint, prompt); updateModel();

19
demo/skeletons_extra.go Normal file
View File

@ -0,0 +1,19 @@
package main
// Copied from NVIDIA Kimodo's Apache-2.0 skeleton definitions. Offsets are
// the parent-local differences extracted from the official neutral-joint
// assets in the trusted reference container.
func init() {
skeletonDefinitions["soma30"] = skeletonDefinition{
key: "soma30",
names: []string{"Hips", "Spine1", "Spine2", "Chest", "Neck1", "Neck2", "Head", "Jaw", "LeftEye", "RightEye", "LeftShoulder", "LeftArm", "LeftForeArm", "LeftHand", "LeftHandThumbEnd", "LeftHandMiddleEnd", "RightShoulder", "RightArm", "RightForeArm", "RightHand", "RightHandThumbEnd", "RightHandMiddleEnd", "LeftLeg", "LeftShin", "LeftFoot", "LeftToeBase", "RightLeg", "RightShin", "RightFoot", "RightToeBase"},
parents: []int{-1, 0, 1, 2, 3, 4, 5, 6, 6, 6, 3, 10, 11, 12, 13, 13, 3, 16, 17, 18, 19, 19, 0, 22, 23, 24, 0, 26, 27, 28},
offsets: [][3]float32{{0, 0, 0}, {-.00013727, .0500376256, -.00053726669}, {-1.86574103e-9, .0712530139, -.000298248546}, {-5.75188398e-9, .0755006305, -.00815970992}, {-.00181676517, .263112953, -.00553348292}, {-2.85102231e-8, .0770939664, .0230258546}, {-4.5975437e-8, .0612891595, .0195370861}, {2.63687901e-5, .0047559225, .0309494062}, {.0320638079, .0538020513, .0758688308}, {-.0322244017, .05361869, .0755823359}, {.0162165175, .232371641, .0511341324}, {.149198457, 2.19397873e-8, -.0550232576}, {.287393078, 2.50268389e-9, -2.58787737e-5}, {.270939812, -7.06625108e-9, 2.60897248e-5}, {.122686267, -.0322017573, .0483306876}, {.190119595, -.00312878387, -.000339570373}, {-.0138011824, .231803086, .0521415786}, {-.150371962, 1.17387901e-7, -.0554560437}, {-.287366393, 1.87628082e-8, -2.59709359e-5}, {-.271336198, -1.16767401e-9, 2.61269368e-5}, {-.122642483, -.0321145448, .0480403904}, {-.190005945, -.00306615542, -.0003157343}, {.10043214, -.0843452671, .0259565473}, {-1e-8, -.432217537, -.00802912805}, {1e-8, -.421550959, -.0348152298}, {0, -.0505947206, .132315294}, {-.10047278, -.0829525995, .0262031695}, {1e-8, -.433622059, -.00805555828}, {2e-8, -.421173943, -.0347839785}, {-3.42907669e-9, -.0507960932, .132841956}},
}
skeletonDefinitions["g1skel34"] = skeletonDefinition{
key: "g1skel34",
names: []string{"pelvis_skel", "left_hip_pitch_skel", "left_hip_roll_skel", "left_hip_yaw_skel", "left_knee_skel", "left_ankle_pitch_skel", "left_ankle_roll_skel", "left_toe_base", "right_hip_pitch_skel", "right_hip_roll_skel", "right_hip_yaw_skel", "right_knee_skel", "right_ankle_pitch_skel", "right_ankle_roll_skel", "right_toe_base", "waist_yaw_skel", "waist_roll_skel", "waist_pitch_skel", "left_shoulder_pitch_skel", "left_shoulder_roll_skel", "left_shoulder_yaw_skel", "left_elbow_skel", "left_wrist_roll_skel", "left_wrist_pitch_skel", "left_wrist_yaw_skel", "left_hand_roll_skel", "right_shoulder_pitch_skel", "right_shoulder_roll_skel", "right_shoulder_yaw_skel", "right_elbow_skel", "right_wrist_roll_skel", "right_wrist_pitch_skel", "right_wrist_yaw_skel", "right_hand_roll_skel"},
parents: []int{-1, 0, 1, 2, 3, 4, 5, 6, 0, 8, 9, 10, 11, 12, 13, 0, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 17, 26, 27, 28, 29, 30, 31, 32},
offsets: [][3]float32{{0, 0, 0}, {.064452, -.1027, 0}, {.052, -.030465, 0}, {0, -.12412, .025001}, {.0021489, -.17734, -.078273}, {-.000094445, -.30001, 0}, {0, -.017558, 0}, {0, -.035, .14}, {-.064452, -.1027, 0}, {-.052, -.030465, 0}, {0, -.12412, .025001}, {-.0021489, -.17734, -.078273}, {.000094445, -.30001, 0}, {0, -.017558, 0}, {0, -.035, .14}, {0, 0, 0}, {0, .044, -.0039635}, {0, 0, 0}, {.10022, .24778, .0039563}, {.038, -.013831, 0}, {.00624, -.1032, 0}, {0, -.080518, .015783}, {.00188791, -.01, .1}, {0, 0, .038}, {0, 0, .046}, {0, 0, .1}, {-.10021, .24778, .0039563}, {-.038, -.013831, 0}, {-.00624, -.1032, 0}, {0, -.080518, .015783}, {-.00188791, -.01, .1}, {0, 0, .038}, {0, 0, .046}, {0, 0, .1}},
}
}

View File

@ -32,7 +32,7 @@ def parse_args() -> argparse.Namespace:
parser.add_argument("--seed", required=True, type=int)
parser.add_argument("--device", default="cuda")
parser.add_argument("--checkpoint-dir", type=Path,
help="local directory containing Kimodo-SMPLX-RP-v1; never download at capture time")
help="local directory containing the selected Kimodo checkpoint; never download at capture time")
parser.add_argument("--zero-embedding", action="store_true",
help="use a deterministic [1,1,4096] zero embedding; enables motion-only fixtures")
parser.add_argument("--text-base", type=Path,
@ -113,8 +113,21 @@ def main() -> None:
raise SystemExit("--zero-embedding cannot be combined with real text model paths")
if args.checkpoint_dir:
checkpoint = args.checkpoint_dir.resolve()
if not (checkpoint / "Kimodo-SMPLX-RP-v1" / "config.yaml").is_file():
raise SystemExit("--checkpoint-dir must contain Kimodo-SMPLX-RP-v1/config.yaml")
model_folders = {
"kimodo-smplx-rp": "Kimodo-SMPLX-RP-v1",
"kimodo-smplx-rp-v1": "Kimodo-SMPLX-RP-v1",
"kimodo-soma-rp": "Kimodo-SOMA-RP-v1.1",
"kimodo-soma-rp-v1.1": "Kimodo-SOMA-RP-v1.1",
"kimodo-soma-seed": "Kimodo-SOMA-SEED-v1.1",
"kimodo-soma-seed-v1.1": "Kimodo-SOMA-SEED-v1.1",
"kimodo-g1-rp": "Kimodo-G1-RP-v1",
"kimodo-g1-rp-v1": "Kimodo-G1-RP-v1",
"kimodo-g1-seed": "Kimodo-G1-SEED-v1",
"kimodo-g1-seed-v1": "Kimodo-G1-SEED-v1",
}
folder = model_folders.get(args.model)
if folder is None or not (checkpoint / folder / "config.yaml").is_file():
raise SystemExit("--checkpoint-dir does not contain the selected official Kimodo model")
# This is deliberately set only for the reference subprocess. It
# prevents a missing local model from silently falling back to HF.
import os

View File

@ -1,5 +1,5 @@
#!/usr/bin/env python3
"""Convert the Kimodo SMPL-X safetensors checkpoint to a self-describing GGUF.
"""Convert a supported Kimodo motion safetensors checkpoint to GGUF.
This converter deliberately implements only the safe safetensors and NPY
formats. It never imports torch, never deserializes pickle, and writes to a
@ -22,6 +22,15 @@ GGUF_MAGIC, GGUF_VERSION, GGML_TYPE_F32 = 0x46554747, 3, 0
TYPE_UINT64, TYPE_STRING, TYPE_FLOAT32 = 10, 8, 6
TYPE_UINT32 = 4
MODEL_SPECS = {
"nvidia/Kimodo-SMPLX-RP-v1": ("smplx22", "SMPLXSkeleton22", 22, False,
"nvidia-internal-scientific-research-and-development-model-license"),
"nvidia/Kimodo-SOMA-RP-v1.1": ("soma30", "SOMASkeleton30", 30, True, "nvidia-open-model-license"),
"nvidia/Kimodo-SOMA-SEED-v1.1": ("soma30", "SOMASkeleton30", 30, True, "nvidia-open-model-license"),
"nvidia/Kimodo-G1-RP-v1": ("g1skel34", "G1Skeleton34", 34, True, "nvidia-open-model-license"),
"nvidia/Kimodo-G1-SEED-v1": ("g1skel34", "G1Skeleton34", 34, True, "nvidia-open-model-license"),
}
@dataclass(frozen=True)
class Tensor:
name: str
@ -183,7 +192,7 @@ def copy_range(dst, tensor: Tensor) -> None:
def main() -> None:
p = argparse.ArgumentParser(description=__doc__)
p.add_argument("--input", required=True, type=Path, help="downloaded Kimodo-SMPLX-RP-v1 directory")
p.add_argument("--input", required=True, type=Path, help="downloaded Kimodo model directory")
p.add_argument("--output", required=True, type=Path)
args = p.parse_args()
root = args.input.resolve()
@ -195,22 +204,33 @@ def main() -> None:
for part in ("global_root", "local_root", "body"):
for stat in ("mean", "std"):
tensors.append(read_npy(root / "stats" / "motion" / part / f"{stat}.npy", f"stats.{part}.{stat}"))
revision = (root / "REVISION").read_text(encoding="utf-8").split()[0]
revision_fields = (root / "REVISION").read_text(encoding="utf-8").split()
if len(revision_fields) != 2 or revision_fields[1] not in MODEL_SPECS:
raise SystemExit("REVISION does not identify a supported official Kimodo model")
revision, model_id = revision_fields
skeleton, skeleton_class, joints, commercial, license_name = MODEL_SPECS[model_id]
config = (root / "config.yaml").read_text(encoding="utf-8")
if f"_target_: kimodo.skeleton.{skeleton_class}" not in config:
raise SystemExit("config.yaml skeleton does not match REVISION model identity")
motion_dim = 9 + 12 * joints
body_dim = motion_dim - 5
meta = [
metadata_string("general.architecture", "kimodo-motion"),
metadata_string("general.name", "Kimodo-SMPLX-RP-v1"),
metadata_string("general.name", model_id.removeprefix("nvidia/")),
# GGML's own loader requires general.alignment to be UINT32.
metadata_uint32("general.alignment", ALIGNMENT),
metadata_uint("kimodo.format_version", 1),
metadata_string("kimodo.skeleton", "smplx22"),
metadata_string("kimodo.model_identity", f"nvidia/Kimodo-SMPLX-RP-v1@{revision}"),
metadata_string("kimodo.skeleton", skeleton),
metadata_string("kimodo.model_identity", f"{model_id}@{revision}"),
metadata_string("kimodo.license", license_name),
metadata_uint("kimodo.commercial_use", int(commercial)),
metadata_string("kimodo.source_revision", revision),
metadata_string("kimodo.source_sha256", sha256(ckpt)),
metadata_uint("kimodo.text_embedding_width", 4096),
metadata_uint("kimodo.motion_dim", 273),
metadata_uint("kimodo.motion_dim", motion_dim),
metadata_uint("kimodo.global_root_dim", 5),
metadata_uint("kimodo.local_root_dim", 4),
metadata_uint("kimodo.body_dim", 268),
metadata_uint("kimodo.body_dim", body_dim),
metadata_uint("kimodo.hidden_size", 1024),
metadata_uint("kimodo.layers", 16),
metadata_uint("kimodo.heads", 8),

View File

@ -4,19 +4,20 @@ set -euo pipefail
export HF_HUB_DISABLE_PROGRESS_BARS=1
ORG="${GGUF_ORG:-LocalAI-io}"
MOTION_REPO_DEFAULT="$ORG/Kimodo-SMPLX-RP-v1-GGML"
TEXT_REPO_DEFAULT="$ORG/Llama-3-Kimodo-GGML"
usage() {
printf '%s\n' "usage: $0 --output DIR [--motion-repo HF_REPO] [--text-repo HF_REPO] [--revision REVISION] [--motion-only]" >&2
printf '%s\n' "usage: $0 --output DIR [--model MODEL]... [--motion-repo HF_REPO] [--text-repo HF_REPO] [--revision REVISION] [--motion-only]" >&2
printf '%s\n' "models: soma-rp-v1.1, soma-seed-v1.1, g1-rp-v1, g1-seed-v1" >&2
exit 2
}
output='' motion_repo="$MOTION_REPO_DEFAULT" text_repo="$TEXT_REPO_DEFAULT" revision='main' motion_only=0
output='' motion_repo_override='' text_repo="$TEXT_REPO_DEFAULT" revision='main' motion_only=0 models=()
while [ "$#" -gt 0 ]; do
case "$1" in
--output) [ "$#" -ge 2 ] || usage; output=$2; shift 2 ;;
--motion-repo) [ "$#" -ge 2 ] || usage; motion_repo=$2; shift 2 ;;
--model) [ "$#" -ge 2 ] || usage; models+=("$2"); shift 2 ;;
--motion-repo) [ "$#" -ge 2 ] || usage; motion_repo_override=$2; shift 2 ;;
--text-repo) [ "$#" -ge 2 ] || usage; text_repo=$2; shift 2 ;;
--revision) [ "$#" -ge 2 ] || usage; revision=$2; shift 2 ;;
--motion-only) motion_only=1; shift ;;
@ -25,6 +26,8 @@ while [ "$#" -gt 0 ]; do
done
[ -n "$output" ] || usage
command -v hf >/dev/null || { echo "hf not found; enter the Nix shell first" >&2; exit 1; }
[ "${#models[@]}" -gt 0 ] || models=(soma-rp-v1.1)
[ -z "$motion_repo_override" ] || [ "${#models[@]}" -eq 1 ] || { echo "--motion-repo requires exactly one --model" >&2; exit 2; }
mkdir -p "$output"
@ -70,7 +73,17 @@ print("verified native Kimodo GGUF bundle")
PY
}
download_and_verify "$motion_repo" "models/kimodo-smplx-rp-v1-f32.gguf"
for model in "${models[@]}"; do
case "$model" in
soma-rp-v1.1) motion_repo="$ORG/Kimodo-SOMA-RP-v1.1-GGML"; motion_file='models/kimodo-soma-rp-v1.1-f32.gguf' ;;
soma-seed-v1.1) motion_repo="$ORG/Kimodo-SOMA-SEED-v1.1-GGML"; motion_file='models/kimodo-soma-seed-v1.1-f32.gguf' ;;
g1-rp-v1) motion_repo="$ORG/Kimodo-G1-RP-v1-GGML"; motion_file='models/kimodo-g1-rp-v1-f32.gguf' ;;
g1-seed-v1) motion_repo="$ORG/Kimodo-G1-SEED-v1-GGML"; motion_file='models/kimodo-g1-seed-v1-f32.gguf' ;;
*) echo "unknown motion model: $model" >&2; usage ;;
esac
[ -z "$motion_repo_override" ] || motion_repo=$motion_repo_override
download_and_verify "$motion_repo" "$motion_file"
done
if [ "$motion_only" -eq 0 ]; then
download_and_verify "$text_repo" "generated/llm2vec-text-bundle/*"
fi

View File

@ -6,15 +6,17 @@ set -euo pipefail
export HF_HUB_DISABLE_PROGRESS_BARS=1
usage() {
printf '%s\n' "usage: $0 --output DIR [--revision REVISION] [--with-text]" >&2
printf '%s\n' "usage: $0 --output DIR [--revision REVISION] [--model NAME]... [--with-text]" >&2
exit 2
}
output='' revision='main' with_text=0
models=()
while [ "$#" -gt 0 ]; do
case "$1" in
--output) [ "$#" -ge 2 ] || usage; output=$2; shift 2 ;;
--revision) [ "$#" -ge 2 ] || usage; revision=$2; shift 2 ;;
--model) [ "$#" -ge 2 ] || usage; models+=("$2"); shift 2 ;;
--with-text) with_text=1; shift ;;
*) usage ;;
esac
@ -52,7 +54,18 @@ download() {
printf '%s %s\n' "$sha" "$repo" > "$target/REVISION"
}
download nvidia/Kimodo-SMPLX-RP-v1 Kimodo-SMPLX-RP-v1
if [ "${#models[@]}" -eq 0 ]; then models=(smplx-rp-v1); fi
for model in "${models[@]}"; do
case "$model" in
smplx-rp-v1) repo=nvidia/Kimodo-SMPLX-RP-v1; folder=Kimodo-SMPLX-RP-v1 ;;
soma-rp-v1.1) repo=nvidia/Kimodo-SOMA-RP-v1.1; folder=Kimodo-SOMA-RP-v1.1 ;;
soma-seed-v1.1) repo=nvidia/Kimodo-SOMA-SEED-v1.1; folder=Kimodo-SOMA-SEED-v1.1 ;;
g1-rp-v1) repo=nvidia/Kimodo-G1-RP-v1; folder=Kimodo-G1-RP-v1 ;;
g1-seed-v1) repo=nvidia/Kimodo-G1-SEED-v1; folder=Kimodo-G1-SEED-v1 ;;
*) echo "Unknown Kimodo model: $model" >&2; usage ;;
esac
download "$repo" "$folder"
done
if [ "$with_text" -eq 1 ]; then
# The MNTP repo is a LoRA adapter, not the Llama base checkpoint. Keep all
# three identities separately so converter provenance cannot confuse them.

View File

@ -0,0 +1,4 @@
Licensed by NVIDIA Corporation under the NVIDIA Open Model License.
Upstream model: nvidia/Kimodo-G1-RP-v1. This GGUF conversion grants no
additional rights.

View File

@ -0,0 +1,25 @@
---
license: other
library_name: ggml
base_model: nvidia/Kimodo-G1-RP-v1
base_model_relation: quantized
tags: [gguf, ggml, text-to-motion, unitree-g1, kimodo]
---
# Kimodo-G1-RP-v1-GGML
Native F32 GGML/GGUF conversion of
[nvidia/Kimodo-G1-RP-v1](https://huggingface.co/nvidia/Kimodo-G1-RP-v1),
targeting the 34-joint Unitree G1 skeleton. Its reusable Llama-derived text
encoder is distributed separately as
[`Llama-3-Kimodo-GGML`](https://huggingface.co/LocalAI-io/Llama-3-Kimodo-GGML).
The model is installed at `models/kimodo-g1-rp-v1-f32.gguf` by
`scripts/download_gguf_weights.sh --output "$PWD" --model g1-rp-v1`.
## Provenance and licence
Converted from upstream revision `3020ad8c419c244e0429d360163730c63c4ed011`.
`MANIFEST.json` records the source revision and SHA-256 of the GGUF. The model
remains subject to the [NVIDIA Open Model License](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/).
This conversion grants no additional rights.

View File

@ -0,0 +1,4 @@
Licensed by NVIDIA Corporation under the NVIDIA Open Model License.
Upstream model: nvidia/Kimodo-G1-SEED-v1. This GGUF conversion grants no
additional rights.

View File

@ -0,0 +1,25 @@
---
license: other
library_name: ggml
base_model: nvidia/Kimodo-G1-SEED-v1
base_model_relation: quantized
tags: [gguf, ggml, text-to-motion, unitree-g1, kimodo]
---
# Kimodo-G1-SEED-v1-GGML
Native F32 GGML/GGUF conversion of
[nvidia/Kimodo-G1-SEED-v1](https://huggingface.co/nvidia/Kimodo-G1-SEED-v1),
targeting the 34-joint Unitree G1 skeleton. Its reusable Llama-derived text
encoder is distributed separately as
[`Llama-3-Kimodo-GGML`](https://huggingface.co/LocalAI-io/Llama-3-Kimodo-GGML).
The model is installed at `models/kimodo-g1-seed-v1-f32.gguf` by
`scripts/download_gguf_weights.sh --output "$PWD" --model g1-seed-v1`.
## Provenance and licence
Converted from upstream revision `5e6f2c7e18c2ab834c8d7983b9dcce701e5c6097`.
`MANIFEST.json` records the source revision and SHA-256 of the GGUF. The model
remains subject to the [NVIDIA Open Model License](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/).
This conversion grants no additional rights.

View File

@ -1,3 +1,4 @@
Kimodo-SMPLX-RP-v1 source model: NVIDIA. This converted distribution remains
subject to the NVIDIA Internal Scientific Research and Development Model License
and is for non-commercial research use only.
Kimodo-SMPLX-RP-v1 source model: NVIDIA. The NVIDIA Internal Scientific
Research and Development Model License limits derivative models to internal,
non-production R&D and prohibits their distribution. This notice is retained
for local conversion records; it does not authorize a converted distribution.

View File

@ -6,30 +6,31 @@ base_model_relation: quantized
tags: [gguf, ggml, text-to-motion, smplx, kimodo]
---
# Kimodo-SMPLX-RP-v1-GGML
# Oops — Kimodo-SMPLX-RP-v1 is local-conversion only
Native F32 GGML/GGUF conversion of
The kimodo.cpp converter can produce a local F32 GGML/GGUF representation of
[nvidia/Kimodo-SMPLX-RP-v1](https://huggingface.co/nvidia/Kimodo-SMPLX-RP-v1),
the SMPL-X 22-joint text-and-constraint conditioned motion diffusion model.
This repository contains only the diffusion model; its reusable Llama-derived
text encoder is distributed separately as
[`Llama-3-Kimodo-GGML`](https://huggingface.co/LocalAI-io/Llama-3-Kimodo-GGML).
It is not a redistributable GGUF release.
From a kimodo.cpp checkout with the Hugging Face CLI installed, install both with:
We originally published the converted weights here, then noticed that the
upstream NVIDIA Internal Scientific Research and Development Model License
explicitly prohibits distributing derivative models. Oops. The GGUF, manifest,
and checksums have therefore been removed; this card remains so that existing
links explain what happened instead of becoming a mysterious 404.
```sh
scripts/download_gguf_weights.sh --output "$PWD"
```
The model is installed at `models/kimodo-smplx-rp-v1-f32.gguf`. Use
`--motion-only` when supplying a precomputed LLM2Vec embedding.
If someone at NVIDIA is willing to give LocalAI-io written permission to
redistribute this checkpoint as GGML/GGUF, that would be very welcome. We would
be happy to restore the conversion with its upstream provenance and verified
parity results.
## Provenance and licence
Converted by kimodo.cpp from upstream commit
`1419ba56b734c48bbafb41fefa84088ca94583b5`. `MANIFEST.json` records the
source revision and SHA-256 of the GGUF.
The local converter targets upstream revision
`1419ba56b734c48bbafb41fefa84088ca94583b5`.
Kimodo-SMPLX-RP-v1 is for non-commercial research use only and remains subject
to the [NVIDIA Internal Scientific Research and Development Model License](https://huggingface.co/nvidia/Kimodo-SMPLX-RP-v1).
This conversion grants no additional rights.
Kimodo-SMPLX-RP-v1 remains subject to the
[NVIDIA Internal Scientific Research and Development Model License](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-internal-scientific-research-and-development-model-license/).
That licence limits the checkpoint and derivative models to internal,
non-production R&D and prohibits their distribution. Converting the weights to
GGUF does not grant additional rights.

View File

@ -0,0 +1,4 @@
Licensed by NVIDIA Corporation under the NVIDIA Open Model License.
Upstream model: nvidia/Kimodo-SOMA-RP-v1.1. This GGUF conversion grants no
additional rights.

View File

@ -0,0 +1,25 @@
---
license: other
library_name: ggml
base_model: nvidia/Kimodo-SOMA-RP-v1.1
base_model_relation: quantized
tags: [gguf, ggml, text-to-motion, soma, kimodo]
---
# Kimodo-SOMA-RP-v1.1-GGML
Native F32 GGML/GGUF conversion of
[nvidia/Kimodo-SOMA-RP-v1.1](https://huggingface.co/nvidia/Kimodo-SOMA-RP-v1.1).
The model predicts the compact SOMA 30-joint control skeleton. Its reusable
Llama-derived text encoder is distributed separately as
[`Llama-3-Kimodo-GGML`](https://huggingface.co/LocalAI-io/Llama-3-Kimodo-GGML).
The model is installed at `models/kimodo-soma-rp-v1.1-f32.gguf` by
`scripts/download_gguf_weights.sh --output "$PWD" --model soma-rp-v1.1`.
## Provenance and licence
Converted from upstream revision `6c9233af1180b8151e3c4703477104af5dce9dd5`.
`MANIFEST.json` records the source revision and SHA-256 of the GGUF. The model
remains subject to the [NVIDIA Open Model License](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/).
This conversion grants no additional rights.

View File

@ -0,0 +1,4 @@
Licensed by NVIDIA Corporation under the NVIDIA Open Model License.
Upstream model: nvidia/Kimodo-SOMA-SEED-v1.1. This GGUF conversion grants no
additional rights.

View File

@ -0,0 +1,25 @@
---
license: other
library_name: ggml
base_model: nvidia/Kimodo-SOMA-SEED-v1.1
base_model_relation: quantized
tags: [gguf, ggml, text-to-motion, soma, kimodo]
---
# Kimodo-SOMA-SEED-v1.1-GGML
Native F32 GGML/GGUF conversion of
[nvidia/Kimodo-SOMA-SEED-v1.1](https://huggingface.co/nvidia/Kimodo-SOMA-SEED-v1.1).
The model predicts the compact SOMA 30-joint control skeleton. Its reusable
Llama-derived text encoder is distributed separately as
[`Llama-3-Kimodo-GGML`](https://huggingface.co/LocalAI-io/Llama-3-Kimodo-GGML).
The model is installed at `models/kimodo-soma-seed-v1.1-f32.gguf` by
`scripts/download_gguf_weights.sh --output "$PWD" --model soma-seed-v1.1`.
## Provenance and licence
Converted from upstream revision `aae3af194322c60d21bc44062b64c3fec912be50`.
`MANIFEST.json` records the source revision and SHA-256 of the GGUF. The model
remains subject to the [NVIDIA Open Model License](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/).
This conversion grants no additional rights.

View File

@ -23,15 +23,54 @@ ROOT = Path(__file__).resolve().parent.parent
HF_ORG = "LocalAI-io" # Hugging Face organisation; GitHub is localai-org.
DEFAULT_REPOS = {
"text": f"{HF_ORG}/Llama-3-Kimodo-GGML",
"motion": f"{HF_ORG}/Kimodo-SMPLX-RP-v1-GGML",
}
MOTION_NAME = "kimodo-smplx-rp-v1-f32.gguf"
MOTION_MODELS = {
"smplx-rp-v1": {
"repo": f"{HF_ORG}/Kimodo-SMPLX-RP-v1-GGML",
"source": "nvidia/Kimodo-SMPLX-RP-v1",
"folder": "Kimodo-SMPLX-RP-v1",
"revision": "1419ba56b734c48bbafb41fefa84088ca94583b5",
"file": "kimodo-smplx-rp-v1-f32.gguf",
"redistributable": False,
},
"soma-rp-v1.1": {
"repo": f"{HF_ORG}/Kimodo-SOMA-RP-v1.1-GGML",
"source": "nvidia/Kimodo-SOMA-RP-v1.1",
"folder": "Kimodo-SOMA-RP-v1.1",
"revision": "6c9233af1180b8151e3c4703477104af5dce9dd5",
"file": "kimodo-soma-rp-v1.1-f32.gguf",
"redistributable": True,
},
"soma-seed-v1.1": {
"repo": f"{HF_ORG}/Kimodo-SOMA-SEED-v1.1-GGML",
"source": "nvidia/Kimodo-SOMA-SEED-v1.1",
"folder": "Kimodo-SOMA-SEED-v1.1",
"revision": "aae3af194322c60d21bc44062b64c3fec912be50",
"file": "kimodo-soma-seed-v1.1-f32.gguf",
"redistributable": True,
},
"g1-rp-v1": {
"repo": f"{HF_ORG}/Kimodo-G1-RP-v1-GGML",
"source": "nvidia/Kimodo-G1-RP-v1",
"folder": "Kimodo-G1-RP-v1",
"revision": "3020ad8c419c244e0429d360163730c63c4ed011",
"file": "kimodo-g1-rp-v1-f32.gguf",
"redistributable": True,
},
"g1-seed-v1": {
"repo": f"{HF_ORG}/Kimodo-G1-SEED-v1-GGML",
"source": "nvidia/Kimodo-G1-SEED-v1",
"folder": "Kimodo-G1-SEED-v1",
"revision": "5e6f2c7e18c2ab834c8d7983b9dcce701e5c6097",
"file": "kimodo-g1-seed-v1-f32.gguf",
"redistributable": True,
},
}
TEXT_NAMES = (
"tokenizer.gguf", "embedding.gguf", "final-norm.gguf",
*(f"layer-{index:02d}.gguf" for index in range(32)),
)
SOURCE_REVISIONS = {
"nvidia/Kimodo-SMPLX-RP-v1": "1419ba56b734c48bbafb41fefa84088ca94583b5",
"meta-llama/Meta-Llama-3-8B-Instruct": "8afb486c1db24fe5011ec46dfbe5b5dccdb575c2",
"McGill-NLP/LLM2Vec-Meta-Llama-3-8B-Instruct-mntp": "31474e395ada192e8ed1586db6be79fb3b70c9c0",
"McGill-NLP/LLM2Vec-Meta-Llama-3-8B-Instruct-mntp-supervised": "baa8ebf04a1c2500e61288e7dad65e8ae42601a7",
@ -47,24 +86,23 @@ def digest(path: Path) -> str:
return value.hexdigest()
def require_revision(repo: str) -> None:
revision = ROOT / "models" / {
"nvidia/Kimodo-SMPLX-RP-v1": "Kimodo-SMPLX-RP-v1",
def require_revision(repo: str, expected: str, folder: str | None = None) -> None:
revision = ROOT / "models" / (folder or {
"meta-llama/Meta-Llama-3-8B-Instruct": "llama3-8b-instruct-base",
"McGill-NLP/LLM2Vec-Meta-Llama-3-8B-Instruct-mntp": "llm2vec-mntp-adapter",
"McGill-NLP/LLM2Vec-Meta-Llama-3-8B-Instruct-mntp-supervised": "llm2vec-adapter",
}[repo] / "REVISION"
}[repo]) / "REVISION"
if not revision.is_file():
raise ValueError(f"missing provenance file: {revision}")
actual = revision.read_text(encoding="utf-8").split()[0]
if actual != SOURCE_REVISIONS[repo]:
raise ValueError(f"unexpected {repo} revision: {actual} (expected {SOURCE_REVISIONS[repo]})")
if actual != expected:
raise ValueError(f"unexpected {repo} revision: {actual} (expected {expected})")
def artifacts(component: str, motion: Path, bundle: Path) -> list[tuple[Path, str]]:
def artifacts(component: str, motion: Path, motion_name: str, bundle: Path) -> list[tuple[Path, str]]:
result: list[tuple[Path, str]] = []
if component == "motion":
result.append((motion, f"models/{MOTION_NAME}"))
result.append((motion, f"models/{motion_name}"))
else:
result.extend((bundle / name, f"generated/llm2vec-text-bundle/{name}") for name in TEXT_NAMES)
for source, destination in result:
@ -82,7 +120,9 @@ def artifacts(component: str, motion: Path, bundle: Path) -> list[tuple[Path, st
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--motion", type=Path, default=ROOT / "models" / MOTION_NAME)
parser.add_argument("--motion", type=Path, default=None,
help="motion GGUF (defaults to the selected model's converted output)")
parser.add_argument("--motion-model", choices=tuple(MOTION_MODELS), default="soma-rp-v1.1")
parser.add_argument("--text-bundle", type=Path,
default=ROOT / "generated/llm2vec-text-bundle")
parser.add_argument("--component", choices=("text", "motion"), required=True,
@ -93,21 +133,28 @@ def main() -> int:
parser.add_argument("--confirm-upstream-licences", action="store_true",
help="required with --upload; confirms authority to redistribute all inputs")
args = parser.parse_args()
repo = args.repo or DEFAULT_REPOS[args.component]
card_dir = ROOT / "scripts/hf" / ("Llama-3-Kimodo-GGML" if args.component == "text" else "Kimodo-SMPLX-RP-v1-GGML")
motion_spec = MOTION_MODELS[args.motion_model]
if args.component == "motion" and args.upload and not motion_spec["redistributable"]:
print("error: the SMPL-X checkpoint licence prohibits distributing Derivative Models; local conversion only",
file=sys.stderr)
return 2
motion = args.motion or ROOT / "models" / motion_spec["file"]
repo = args.repo or (DEFAULT_REPOS["text"] if args.component == "text" else motion_spec["repo"])
card_dir = ROOT / "scripts/hf" / ("Llama-3-Kimodo-GGML" if args.component == "text" else repo.rsplit("/", 1)[1])
card = card_dir / "README.md"
notice = card_dir / "NOTICE"
relevant_sources = (SOURCE_REVISIONS if args.component == "text"
else {"nvidia/Kimodo-SMPLX-RP-v1": SOURCE_REVISIONS["nvidia/Kimodo-SMPLX-RP-v1"]})
else {motion_spec["source"]: motion_spec["revision"]})
try:
if not card.is_file() or not notice.is_file():
raise ValueError("version-controlled model card or NOTICE is missing")
if args.component == "text" and not LLAMA_LICENSE.is_file():
raise ValueError("Meta Llama 3 licence is missing")
for source_repo in relevant_sources:
require_revision(source_repo)
files = artifacts(args.component, args.motion, args.text_bundle)
for source_repo, revision in relevant_sources.items():
folder = motion_spec["folder"] if args.component == "motion" else None
require_revision(source_repo, revision, folder)
files = artifacts(args.component, motion, motion_spec["file"], args.text_bundle)
except ValueError as error:
print(f"error: {error}", file=sys.stderr)
return 1

View File

@ -90,9 +90,10 @@ std::expected<std::vector<float>, std::string> run_two_stage_denoiser(
std::span<const float> embedding, std::span<const float> timesteps,
std::span<const float> headings, std::span<const float> mask,
std::size_t batch, std::size_t frames) {
if (!batch || !frames || x.size()!=batch*frames*546 || mask.size()!=batch*frames)
const size_t dim=weights.motion_dim(), root_input_dim=2*dim, body_input_dim=2*dim-1;
if (!batch || !frames || !dim || x.size()!=batch*frames*root_input_dim || mask.size()!=batch*frames)
return std::unexpected("invalid two-stage denoiser input dimensions");
auto root=run_motion_transformer(weights,"root_model.",x,546,embedding,timesteps,headings,batch,frames);
auto root=run_motion_transformer(weights,"root_model.",x,root_input_dim,embedding,timesteps,headings,batch,frames);
if(!root)return std::unexpected(root.error());
auto gm=weights.f32_values("stats.global_root.mean"), gs=weights.f32_values("stats.global_root.std"), lm=weights.f32_values("stats.local_root.mean"), ls=weights.f32_values("stats.local_root.std");
if(!gm)return std::unexpected(gm.error());
@ -101,16 +102,16 @@ std::expected<std::vector<float>, std::string> run_two_stage_denoiser(
if(!ls)return std::unexpected(ls.error());
auto local=global_root_to_local_root(*root,mask,batch,frames,*gm,*gs,*lm,*ls);
if(!local)return std::unexpected(local.error());
std::vector<float> body_input(batch*frames*545);
std::vector<float> body_input(batch*frames*body_input_dim);
for(std::size_t b=0;b<batch;++b) for(std::size_t t=0;t<frames;++t) {
const auto src=(b*frames+t)*546, dst=(b*frames+t)*545;
const auto src=(b*frames+t)*root_input_dim, dst=(b*frames+t)*body_input_dim;
std::memcpy(body_input.data()+dst,local->data()+(b*frames+t)*4,4*sizeof(float));
std::memcpy(body_input.data()+dst+4,x.data()+src+5,541*sizeof(float));
std::memcpy(body_input.data()+dst+4,x.data()+src+5,(root_input_dim-5)*sizeof(float));
}
auto body=run_motion_transformer(weights,"body_model.",body_input,545,embedding,timesteps,headings,batch,frames);
auto body=run_motion_transformer(weights,"body_model.",body_input,body_input_dim,embedding,timesteps,headings,batch,frames);
if(!body)return std::unexpected(body.error());
std::vector<float> output(batch*frames*273);
for(std::size_t b=0;b<batch;++b)for(std::size_t t=0;t<frames;++t){const auto r=(b*frames+t)*5, q=(b*frames+t)*268, o=(b*frames+t)*273;std::memcpy(output.data()+o,root->data()+r,5*sizeof(float));std::memcpy(output.data()+o+5,body->data()+q,268*sizeof(float));}
std::vector<float> output(batch*frames*dim);
for(std::size_t b=0;b<batch;++b)for(std::size_t t=0;t<frames;++t){const auto r=(b*frames+t)*5, q=(b*frames+t)*(dim-5), o=(b*frames+t)*dim;std::memcpy(output.data()+o,root->data()+r,5*sizeof(float));std::memcpy(output.data()+o+5,body->data()+q,(dim-5)*sizeof(float));}
return output;
}
@ -118,7 +119,7 @@ std::expected<std::vector<float>, std::string> run_separated_cfg_denoiser(
const ggml_motion_weights &weights, std::span<const float> motion,
std::span<const float> embedding, float timestep, float text_weight,
float constraint_weight, std::size_t frames) {
const std::vector<float> empty(frames*273, 0.f);
const std::vector<float> empty(frames*weights.motion_dim(), 0.f);
return run_separated_cfg_denoiser_conditioned(weights, motion, embedding, empty, empty,
timestep, 0.f, text_weight, constraint_weight, frames);
}
@ -128,25 +129,26 @@ std::expected<std::vector<float>, std::string> run_separated_cfg_denoiser_condit
std::span<const float> embedding, std::span<const float> observed,
std::span<const float> observed_mask, float timestep, float heading,
float text_weight, float constraint_weight, std::size_t frames) {
if (motion.size()!=frames*273 || embedding.size()!=4096 || !std::isfinite(timestep) || !std::isfinite(text_weight) || !std::isfinite(constraint_weight))
const size_t dim=weights.motion_dim();
if (!dim || motion.size()!=frames*dim || embedding.size()!=4096 || !std::isfinite(timestep) || !std::isfinite(text_weight) || !std::isfinite(constraint_weight))
return std::unexpected("invalid separated CFG denoiser input");
if (observed.size()!=frames*273 || observed_mask.size()!=frames*273 || !std::isfinite(heading))
if (observed.size()!=frames*dim || observed_mask.size()!=frames*dim || !std::isfinite(heading))
return std::unexpected("invalid separated CFG condition dimensions");
constexpr size_t cfg_batch=3; std::vector<float> extended(cfg_batch*frames*546), text(cfg_batch*4096), times(cfg_batch,timestep), headings(cfg_batch,heading), mask(cfg_batch*frames,1.f);
constexpr size_t cfg_batch=3; std::vector<float> extended(cfg_batch*frames*2*dim), text(cfg_batch*4096), times(cfg_batch,timestep), headings(cfg_batch,heading), mask(cfg_batch*frames,1.f);
for(size_t b=0;b<cfg_batch;++b) for(size_t t=0;t<frames;++t) {
auto *dst=extended.data()+(b*frames+t)*546;
std::memcpy(dst,motion.data()+t*273,273*sizeof(float));
auto *dst=extended.data()+(b*frames+t)*2*dim;
std::memcpy(dst,motion.data()+t*dim,dim*sizeof(float));
// Upstream separated CFG is [text, constraint, unconditional]. Only
// the constraint branch receives observed motion and its feature mask.
if (b==1) for (size_t d=0;d<273;++d) dst[d]=motion[t*273+d]*(1.f-observed_mask[t*273+d])+observed[t*273+d]*observed_mask[t*273+d];
if (b==1) std::memcpy(dst+273,observed_mask.data()+t*273,273*sizeof(float));
if (b==1) for (size_t d=0;d<dim;++d) dst[d]=motion[t*dim+d]*(1.f-observed_mask[t*dim+d])+observed[t*dim+d]*observed_mask[t*dim+d];
if (b==1) std::memcpy(dst+dim,observed_mask.data()+t*dim,dim*sizeof(float));
}
// Only branch zero has text. Branch one is constraint-only; branch two is
// unconditional. This is the upstream separated-CFG batch order.
std::memcpy(text.data(),embedding.data(),4096*sizeof(float));
auto all=run_two_stage_denoiser(weights,extended,text,times,headings,mask,cfg_batch,frames);
if(!all)return std::unexpected(all.error());
std::vector<float> result(frames*273);
std::vector<float> result(frames*dim);
for(size_t i=0;i<result.size();++i) result[i]=(*all)[2*result.size()+i]+text_weight*((*all)[i]-(*all)[2*result.size()+i])+constraint_weight*((*all)[result.size()+i]-(*all)[2*result.size()+i]);
return result;
}
@ -155,7 +157,7 @@ std::expected<std::vector<float>, std::string> sample_motion_from_noise(
const ggml_motion_weights &weights, std::span<const float> initial,
std::span<const float> embedding, std::size_t frames, unsigned steps,
float text_weight, float constraint_weight) {
if(initial.size()!=frames*273) return std::unexpected("invalid initial motion noise dimensions");
if(initial.size()!=frames*weights.motion_dim()) return std::unexpected("invalid initial motion noise dimensions");
auto schedule=make_cosine_schedule(1000,steps); if(!schedule)return std::unexpected(schedule.error());
std::vector<float> state(initial.begin(),initial.end()), next(state.size());
for(unsigned i=steps;i-->0;) {
@ -173,7 +175,7 @@ std::expected<std::vector<float>, std::string> sample_motion_from_noise_conditio
std::span<const float> embedding, std::span<const float> observed,
std::span<const float> observed_mask, float heading, std::size_t frames,
unsigned steps, float text_weight, float constraint_weight) {
if(initial.size()!=frames*273 || observed.size()!=initial.size() || observed_mask.size()!=initial.size())
if(initial.size()!=frames*weights.motion_dim() || observed.size()!=initial.size() || observed_mask.size()!=initial.size())
return std::unexpected("invalid conditioned motion noise dimensions");
auto schedule=make_cosine_schedule(1000,steps); if(!schedule)return std::unexpected(schedule.error());
std::vector<float> state(initial.begin(),initial.end()), next(state.size());

View File

@ -34,15 +34,15 @@ std::expected<std::vector<float>, std::string> run_motion_transformer(
std::span<const float> headings, std::size_t batch, std::size_t frames);
// Exact two-stage Kimodo denoiser for concatenated motion/mask inputs
// [B,T,546]. Returned clean prediction is [B,T,273].
// [B,T,2*motion_dim]. Returned clean prediction is [B,T,motion_dim].
std::expected<std::vector<float>, std::string> run_two_stage_denoiser(
const ggml_motion_weights &weights, std::span<const float> motion_and_mask,
std::span<const float> text_embedding, std::span<const float> timesteps,
std::span<const float> headings, std::span<const float> motion_mask,
std::size_t batch, std::size_t frames);
// Unconstrained separated CFG wrapper. `motion` is [T,273], embedding is
// [4096], and the result is one clean [T,273] prediction.
// Unconstrained separated CFG wrapper. `motion` is [T,motion_dim], embedding
// is [4096], and the result is one clean prediction.
std::expected<std::vector<float>, std::string> run_separated_cfg_denoiser(
const ggml_motion_weights &weights, std::span<const float> motion,
std::span<const float> embedding, float timestep, float text_weight,
@ -54,7 +54,7 @@ std::expected<std::vector<float>, std::string> sample_motion_from_noise(
std::span<const float> embedding, std::size_t frames, unsigned steps,
float text_weight, float constraint_weight);
// Multi-prompt transition sampler. `observed` and `observed_mask` are [T,273]
// Multi-prompt transition sampler. `observed` and `observed_mask` are [T,motion_dim]
// normalized motion-representation values/masks. This mirrors the upstream
// concat-mask denoiser: text, constraint, and unconditional CFG branches.
std::expected<std::vector<float>, std::string> sample_motion_from_noise_conditioned(

View File

@ -41,7 +41,7 @@ int main(int argc, char **argv) try {
const std::filesystem::path output(argv[7]); std::filesystem::create_directories(output);
write_f32(output / "root_positions.f32", motion->root_positions);
write_f32(output / "local_rotations_xyzw.f32", motion->local_rotations_xyzw);
std::cout << "generated " << motion->frames << " SMPL-X22 sequence frames\n";
std::cout << "generated " << motion->frames << " frames with " << motion->joints << " joints\n";
return 0;
}
if (argc != 8) {
@ -63,7 +63,7 @@ int main(int argc, char **argv) try {
std::filesystem::create_directories(output);
write_f32(output / "root_positions.f32", motion->root_positions);
write_f32(output / "local_rotations_xyzw.f32", motion->local_rotations_xyzw);
std::cout << "generated " << motion->frames << " SMPL-X22 frames\n";
std::cout << "generated " << motion->frames << " frames with " << motion->joints << " joints\n";
return 0;
} catch (const std::exception &error) {
std::cerr << error.what() << '\n';

View File

@ -61,6 +61,9 @@ std::expected<std::unique_ptr<ggml_motion_weights>, std::string> ggml_motion_wei
if (!checked) return std::unexpected(checked.error());
if (auto valid = validate_motion_gguf(*checked); !valid) return std::unexpected(valid.error());
auto result = std::unique_ptr<ggml_motion_weights>(new ggml_motion_weights);
result->skeleton_ = checked->strings.at("kimodo.skeleton");
result->motion_dim_ = static_cast<size_t>(checked->uints.at("kimodo.motion_dim"));
result->body_dim_ = static_cast<size_t>(checked->uints.at("kimodo.body_dim"));
gguf_init_params params{true, &result->context_};
result->gguf_ = gguf_init_from_file(std::string(path).c_str(), params);
if (!result->gguf_ || !result->context_) return std::unexpected("GGML could not load checked motion GGUF");

View File

@ -1,5 +1,6 @@
#pragma once
#include <cstddef>
#include <expected>
#include <memory>
#include <string>
@ -27,6 +28,9 @@ public:
ggml_tensor *tensor(std::string_view name) const;
std::expected<std::vector<float>, std::string> f32_values(std::string_view name) const;
ggml_backend *backend() const noexcept { return backend_; }
std::string_view skeleton_key() const noexcept { return skeleton_; }
std::size_t motion_dim() const noexcept { return motion_dim_; }
std::size_t body_dim() const noexcept { return body_dim_; }
private:
ggml_motion_weights() = default;
@ -34,6 +38,9 @@ private:
gguf_context *gguf_ = nullptr;
ggml_backend *backend_ = nullptr;
ggml_backend_buffer *buffer_ = nullptr;
std::string skeleton_;
std::size_t motion_dim_ = 0;
std::size_t body_dim_ = 0;
};
} // namespace kimodo::detail

View File

@ -1,4 +1,5 @@
#include "gguf.hpp"
#include "skeleton.hpp"
#include <algorithm>
#include <array>
@ -106,8 +107,14 @@ std::expected<void, std::string> validate_motion_gguf(const gguf_file &file) {
if (format == file.uints.end() || format->second != 1)
return std::unexpected("unsupported Kimodo motion GGUF format");
const auto skeleton = file.strings.find("kimodo.skeleton");
if (skeleton == file.strings.end() || skeleton->second != "smplx22")
return std::unexpected("first runtime supports only smplx22 skeletons");
if (skeleton == file.strings.end() || !find_skeleton(skeleton->second))
return std::unexpected("motion GGUF has an unsupported skeleton");
const auto &spec = *find_skeleton(skeleton->second);
const auto motion_dim = file.uints.find("kimodo.motion_dim");
const auto body_dim = file.uints.find("kimodo.body_dim");
if (motion_dim == file.uints.end() || motion_dim->second != spec.motion_dim() ||
body_dim == file.uints.end() || body_dim->second != spec.body_dim())
return std::unexpected("motion GGUF dimensions do not match its skeleton");
const auto width = file.uints.find("kimodo.text_embedding_width");
if (width == file.uints.end() || width->second != 4096)
return std::unexpected("motion GGUF has incompatible text embedding width");

View File

@ -22,5 +22,5 @@ int main(int argc, char **argv) {
}
gguf_free(file); ggml_free(tensor_context);
#endif
std::puts("Kimodo SMPL-X motion GGUF: valid (414 F32 tensors)");
std::puts("Kimodo motion GGUF: valid (414 F32 tensors)");
}

View File

@ -1,5 +1,6 @@
#include <kimodo/kimodo.hpp>
#include "gguf.hpp"
#include "skeleton.hpp"
#ifdef KIMODO_HAVE_GGML
#include "ggml_weights.hpp"
#include "denoiser.hpp"
@ -15,6 +16,7 @@ namespace kimodo {
struct model::impl {
detail::gguf_file motion;
std::string motion_path;
const detail::skeleton_spec *skeleton = nullptr;
#ifdef KIMODO_HAVE_GGML
mutable std::unique_ptr<detail::ggml_motion_weights> weights;
std::unique_ptr<detail::llm_text_encoder> text;
@ -30,6 +32,7 @@ std::expected<std::unique_ptr<model>, std::string> model::load(std::string_view
auto state = std::make_unique<impl>();
state->motion = std::move(*file);
state->motion_path = std::string(motion_path);
state->skeleton = detail::find_skeleton(state->motion.strings.at("kimodo.skeleton"));
#ifdef KIMODO_HAVE_GGML
if (!text_path.empty()) {
auto text = detail::llm_text_encoder::load(text_path);
@ -74,7 +77,8 @@ std::expected<motion_data, std::string> model::generate_embedding(
}
std::mt19937_64 rng(seed);
std::normal_distribution<float> normal(0.f, 1.f);
std::vector<float> noise(static_cast<size_t>(frames)*273);
const size_t motion_dim=impl_->skeleton->motion_dim();
std::vector<float> noise(static_cast<size_t>(frames)*motion_dim);
for (float &value : noise) value = normal(rng);
auto sampled = detail::sample_motion_from_noise(*impl_->weights, noise, embedding, frames, steps, text_cfg, constraint_cfg);
if (!sampled) return std::unexpected(sampled.error());
@ -84,10 +88,10 @@ std::expected<motion_data, std::string> model::generate_embedding(
if (!global_std) return std::unexpected(global_std.error());
if (!body_mean) return std::unexpected(body_mean.error());
if (!body_std) return std::unexpected(body_std.error());
auto decoded=detail::decode_smplx22(*sampled,frames,*global_mean,*global_std,*body_mean,*body_std);
auto decoded=detail::decode_motion(*sampled,frames,*impl_->skeleton,*global_mean,*global_std,*body_mean,*body_std);
if (!decoded) return std::unexpected(decoded.error());
motion_data result;
result.frames=frames; result.joints=22;
result.frames=frames; result.joints=static_cast<unsigned>(impl_->skeleton->joints());
result.local_rotations_xyzw=std::move(decoded->local_xyzw);
result.root_positions=std::move(decoded->root_positions);
return result;
@ -127,22 +131,23 @@ std::expected<motion_data, std::string> model::generate_text_sequence(
const auto sampled_frames = static_cast<size_t>(segment.frames) +
(index == 0 ? 0 : transition_frames);
embeddings.push_back(*embedding);
noise.emplace_back(sampled_frames*273);
noise.emplace_back(sampled_frames*impl_->skeleton->motion_dim());
for (float &value : noise.back()) value=normal(rng);
sampled.push_back({embeddings.back(), noise.back(), segment.frames});
}
auto joined=detail::sample_motion_sequence_from_noise(*impl_->weights,sampled,transition_frames,steps,text_cfg,constraint_cfg);
if (!joined) return std::unexpected(joined.error());
const auto frames=static_cast<unsigned>(joined->size()/273);
const size_t motion_dim=impl_->skeleton->motion_dim(), body_dim=impl_->skeleton->body_dim();
const auto frames=static_cast<unsigned>(joined->size()/motion_dim);
auto normalized=*joined;
for (size_t row=0; row<frames; ++row) {
auto *value=normalized.data()+row*273;
auto *value=normalized.data()+row*motion_dim;
for (size_t d=0; d<5; ++d) value[d]=(value[d]-(*gm)[d])/std::sqrt((*gs)[d]*(*gs)[d]+1.e-5F);
for (size_t d=0; d<268; ++d) value[5+d]=(value[5+d]-(*bm)[d])/std::sqrt((*bs)[d]*(*bs)[d]+1.e-5F);
for (size_t d=0; d<body_dim; ++d) value[5+d]=(value[5+d]-(*bm)[d])/std::sqrt((*bs)[d]*(*bs)[d]+1.e-5F);
}
auto decoded=detail::decode_smplx22(normalized,frames,*gm,*gs,*bm,*bs);
auto decoded=detail::decode_motion(normalized,frames,*impl_->skeleton,*gm,*gs,*bm,*bs);
if (!decoded) return std::unexpected(decoded.error());
motion_data result; result.frames=frames; result.joints=22;
motion_data result; result.frames=frames; result.joints=static_cast<unsigned>(impl_->skeleton->joints());
result.local_rotations_xyzw=std::move(decoded->local_xyzw); result.root_positions=std::move(decoded->root_positions);
return result;
#else

View File

@ -1,13 +1,14 @@
#include "motion_decode.hpp"
#include <array>
#include "skeleton.hpp"
#include <cmath>
#include <vector>
namespace kimodo::detail { namespace {
constexpr int parent[22]={-1,0,0,0,1,2,3,4,5,6,7,8,9,9,9,12,13,14,16,17,18,19};
struct M{float v[9];};
M mul(const M&a,const M&b){M r{};for(int i=0;i<3;++i)for(int j=0;j<3;++j)for(int k=0;k<3;++k)r.v[i*3+j]+=a.v[i*3+k]*b.v[k*3+j];return r;}
M tr(const M&a){M r{};for(int i=0;i<3;++i)for(int j=0;j<3;++j)r.v[i*3+j]=a.v[j*3+i];return r;}
M six(const float*x){float n=std::sqrt(x[0]*x[0]+x[1]*x[1]+x[2]*x[2]);float a[3]={x[0]/n,x[1]/n,x[2]/n};float z[3]={a[1]*x[5]-a[2]*x[4],a[2]*x[3]-a[0]*x[5],a[0]*x[4]-a[1]*x[3]};n=std::sqrt(z[0]*z[0]+z[1]*z[1]+z[2]*z[2]);for(float&v:z)v/=n;float b[3]={z[1]*a[2]-z[2]*a[1],z[2]*a[0]-z[0]*a[2],z[0]*a[1]-z[1]*a[0]};return M{{a[0],b[0],z[0],a[1],b[1],z[1],a[2],b[2],z[2]}};}
void quat(const M&m,float*q){float w,x,y,z,t=m.v[0]+m.v[4]+m.v[8];if(t>0){float s=2*std::sqrt(t+1);w=.25f*s;x=(m.v[7]-m.v[5])/s;y=(m.v[2]-m.v[6])/s;z=(m.v[3]-m.v[1])/s;}else if(m.v[0]>m.v[4]&&m.v[0]>m.v[8]){float s=2*std::sqrt(1+m.v[0]-m.v[4]-m.v[8]);w=(m.v[7]-m.v[5])/s;x=.25f*s;y=(m.v[1]+m.v[3])/s;z=(m.v[2]+m.v[6])/s;}else if(m.v[4]>m.v[8]){float s=2*std::sqrt(1+m.v[4]-m.v[0]-m.v[8]);w=(m.v[2]-m.v[6])/s;x=(m.v[1]+m.v[3])/s;y=.25f*s;z=(m.v[5]+m.v[7])/s;}else{float s=2*std::sqrt(1+m.v[8]-m.v[0]-m.v[4]);w=(m.v[3]-m.v[1])/s;x=(m.v[2]+m.v[6])/s;y=(m.v[5]+m.v[7])/s;z=.25f*s;}q[0]=x;q[1]=y;q[2]=z;q[3]=w;}
}
std::expected<decoded_motion,std::string> decode_smplx22(std::span<const float>x,size_t T,std::span<const float>gm,std::span<const float>gs,std::span<const float>bm,std::span<const float>bs){if(x.size()!=T*273||gm.size()!=5||gs.size()!=5||bm.size()!=268||bs.size()!=268)return std::unexpected("invalid SMPL-X decode inputs");decoded_motion o;o.local_xyzw.resize(T*22*4);o.root_positions.resize(T*3);auto scale=[](float s){return std::sqrt(s*s+1.e-5f);};for(size_t t=0;t<T;++t){const float*in=x.data()+t*273;std::array<float,273> f{};for(int i=0;i<5;++i)f[i]=in[i]*scale(gs[i])+gm[i];for(int i=0;i<268;++i)f[5+i]=in[5+i]*scale(bs[i])+bm[i];o.root_positions[t*3]=f[0]+f[5];o.root_positions[t*3+1]=f[6];o.root_positions[t*3+2]=f[2]+f[7];M g[22],l[22];for(int j=0;j<22;++j)g[j]=six(f.data()+71+j*6);for(int j=0;j<22;++j)l[j]=parent[j]<0?g[j]:mul(tr(g[parent[j]]),g[j]);for(int j=0;j<22;++j)quat(l[j],o.local_xyzw.data()+(t*22+j)*4);}return o;}
std::expected<decoded_motion,std::string> decode_motion(std::span<const float>x,size_t T,const skeleton_spec&s,std::span<const float>gm,std::span<const float>gs,std::span<const float>bm,std::span<const float>bs){const size_t D=s.motion_dim(),J=s.joints(),body=D-5,rotation=5+3*J;if(x.size()!=T*D||gm.size()!=5||gs.size()!=5||bm.size()!=body||bs.size()!=body)return std::unexpected("invalid "+std::string(s.key)+" decode inputs");decoded_motion o;o.local_xyzw.resize(T*J*4);o.root_positions.resize(T*3);auto scale=[](float v){return std::sqrt(v*v+1.e-5f);};std::vector<float>f(D);std::vector<M>g(J),l(J);for(size_t t=0;t<T;++t){const float*in=x.data()+t*D;for(size_t i=0;i<5;++i)f[i]=in[i]*scale(gs[i])+gm[i];for(size_t i=0;i<body;++i)f[5+i]=in[5+i]*scale(bs[i])+bm[i];o.root_positions[t*3]=f[0]+f[5];o.root_positions[t*3+1]=f[6];o.root_positions[t*3+2]=f[2]+f[7];for(size_t j=0;j<J;++j)g[j]=six(f.data()+rotation+j*6);for(size_t j=0;j<J;++j)l[j]=s.parents[j]<0?g[j]:mul(tr(g[static_cast<size_t>(s.parents[j])]),g[j]);for(size_t j=0;j<J;++j)quat(l[j],o.local_xyzw.data()+(t*J+j)*4);}return o;}
std::expected<decoded_motion,std::string> decode_smplx22(std::span<const float>x,size_t T,std::span<const float>gm,std::span<const float>gs,std::span<const float>bm,std::span<const float>bs){return decode_motion(x,T,smplx22_spec,gm,gs,bm,bs);}
}

View File

@ -4,6 +4,8 @@
#include <string>
#include <vector>
namespace kimodo::detail {
struct skeleton_spec;
struct decoded_motion { std::vector<float> local_xyzw, root_positions; };
std::expected<decoded_motion,std::string> decode_motion(std::span<const float> normalized, std::size_t frames, const skeleton_spec &skeleton, std::span<const float> global_mean, std::span<const float> global_std, std::span<const float> body_mean, std::span<const float> body_std);
std::expected<decoded_motion,std::string> decode_smplx22(std::span<const float> normalized, std::size_t frames, std::span<const float> global_mean, std::span<const float> global_std, std::span<const float> body_mean, std::span<const float> body_std);
}

View File

@ -3,6 +3,7 @@
#include "denoiser.hpp"
#include "ggml_weights.hpp"
#include "motion_decode.hpp"
#include "skeleton.hpp"
#include <algorithm>
#include <cmath>
@ -54,10 +55,11 @@ int main(int argc, char **argv) try {
const std::filesystem::path fixture(argv[2]), output(argv[5]);
const auto embedding = read_f32(fixture / "text_features.f32");
const auto noise = read_f32(fixture / "sampling_initial_noise.f32");
if (embedding.size() != 4096 || noise.size() != frames * 273)
throw std::runtime_error("fixture does not match requested [1,1,4096] embedding and [1,T,273] noise");
auto weights = kimodo::detail::ggml_motion_weights::load(argv[1]);
if (!weights) throw std::runtime_error(weights.error());
const auto *skeleton=kimodo::detail::find_skeleton((*weights)->skeleton_key());
if (!skeleton || embedding.size() != 4096 || noise.size() != frames * skeleton->motion_dim())
throw std::runtime_error("fixture does not match the requested embedding and model motion dimensions");
auto sampled = kimodo::detail::sample_motion_from_noise(**weights, noise, embedding, frames, steps, 2.f, 2.f);
if (!sampled) throw std::runtime_error(sampled.error());
auto gm = (**weights).f32_values("stats.global_root.mean");
@ -65,7 +67,7 @@ int main(int argc, char **argv) try {
auto bm = (**weights).f32_values("stats.body.mean");
auto bs = (**weights).f32_values("stats.body.std");
if (!gm || !gs || !bm || !bs) throw std::runtime_error("missing motion normalisation tensors");
auto decoded = kimodo::detail::decode_smplx22(*sampled, frames, *gm, *gs, *bm, *bs);
auto decoded = kimodo::detail::decode_motion(*sampled, frames, *skeleton, *gm, *gs, *bm, *bs);
if (!decoded) throw std::runtime_error(decoded.error());
std::filesystem::create_directories(output);
write_f32(output / "sampling_final_state.f32", *sampled);

View File

@ -1,132 +1,72 @@
#include "denoiser.hpp"
#include "ggml_weights.hpp"
#include "skeleton.hpp"
#include <algorithm>
#include <array>
#include <cmath>
#include <vector>
namespace kimodo::detail {
namespace {
constexpr int parent[22]={-1,0,0,0,1,2,3,4,5,6,7,8,9,9,9,12,13,14,16,17,18,19};
constexpr float offset[22][3]={{0,0,0},{.052299179F,-.093935639F,-.027606763F},{-.057192899F,-.106548190F,-.022217851F},{-.001495834F,.112929940F,-.024981268F},{.058866613F,-.416441321F,-.006556974F},{-.048074268F,-.397559673F,-.014061437F},{.006900469F,.145636231F,-.006858510F},{-.041737989F,-.437583506F,-.029511765F},{.014489345F,-.446852267F,-.018029511F},{-.010334037F,.056081813F,.021115851F},{.049293540F,-.065279245F,.126259089F},{-.040575184F,-.065286517F,.127075911F},{-.011025756F,.171365142F,-.028827066F},{.047724526F,.087643057F,-.008375450F},{-.046636276F,.086612143F,-.014864366F},{.024654359F,.175390735F,.024463326F},{.126284808F,.057680372F,-.013885141F},{-.109341696F,.053674292F,-.009117880F},{.272907287F,-.069853373F,-.039094493F},{-.292028785F,-.035440356F,-.024564851F},{.276173830F,.021254137F,-.002478220F},{-.271878421F,-.004834589F,-.016445294F}};
struct mat { double v[9]; };
mat mul(const mat&a,const mat&b){mat r{};for(int i=0;i<3;++i)for(int j=0;j<3;++j)for(int k=0;k<3;++k)r.v[i*3+j]+=a.v[i*3+k]*b.v[k*3+j];return r;}
mat trans(const mat&a){mat r{};for(int i=0;i<3;++i)for(int j=0;j<3;++j)r.v[i*3+j]=a.v[j*3+i];return r;}
mat cont6(const float *x) { double a[3]={x[0],x[1],x[2]}, n=std::sqrt(a[0]*a[0]+a[1]*a[1]+a[2]*a[2]); for(double &q:a)q/=n; double z[3]={a[1]*x[5]-a[2]*x[4],a[2]*x[3]-a[0]*x[5],a[0]*x[4]-a[1]*x[3]}; n=std::sqrt(z[0]*z[0]+z[1]*z[1]+z[2]*z[2]);for(double&q:z)q/=n; double b[3]={z[1]*a[2]-z[2]*a[1],z[2]*a[0]-z[0]*a[2],z[0]*a[1]-z[1]*a[0]};return {{a[0],b[0],z[0],a[1],b[1],z[1],a[2],b[2],z[2]}}; }
void rotate(const mat&m,const float *x,float *o){for(int i=0;i<3;++i)o[i]=static_cast<float>(m.v[i*3]*x[0]+m.v[i*3+1]*x[1]+m.v[i*3+2]*x[2]);}
mat cont6(const float*x){double a[3]={x[0],x[1],x[2]},n=std::sqrt(a[0]*a[0]+a[1]*a[1]+a[2]*a[2]);for(double&q:a)q/=n;double z[3]={a[1]*x[5]-a[2]*x[4],a[2]*x[3]-a[0]*x[5],a[0]*x[4]-a[1]*x[3]};n=std::sqrt(z[0]*z[0]+z[1]*z[1]+z[2]*z[2]);for(double&q:z)q/=n;double b[3]={z[1]*a[2]-z[2]*a[1],z[2]*a[0]-z[0]*a[2],z[0]*a[1]-z[1]*a[0]};return{{a[0],b[0],z[0],a[1],b[1],z[1],a[2],b[2],z[2]}};}
void rotate(const mat&m,const std::array<float,3>&x,float*o){for(int i=0;i<3;++i)o[i]=static_cast<float>(m.v[i*3]*x[0]+m.v[i*3+1]*x[1]+m.v[i*3+2]*x[2]);}
// Reconstruct the full-body/end-effector condition used by upstream
// `_multiprompt`, generalized over the three released skeleton layouts.
float condition_row(const float *raw, const skeleton_spec &s, float *value) {
const size_t D=s.motion_dim(),J=s.joints(),rotation_begin=5+3*J;
std::copy_n(raw,D,value);
std::vector<mat> decoded(J),local(J),global(J);
for(size_t j=0;j<J;++j)decoded[j]=cont6(value+rotation_begin+j*6);
for(size_t j=0;j<J;++j)local[j]=s.parents[j]<0?decoded[j]:mul(trans(decoded[static_cast<size_t>(s.parents[j])]),decoded[j]);
const float root[3]={value[0]+value[5],value[6],value[2]+value[7]};
std::vector<std::array<float,3>> posed(J);
for(size_t j=0;j<J;++j){const int parent=s.parents[j];if(parent<0){global[j]=local[j];posed[j]={root[0],root[1],root[2]};}else{global[j]=mul(global[static_cast<size_t>(parent)],local[j]);float offset[3];rotate(global[static_cast<size_t>(parent)],s.offsets[j],offset);for(int k=0;k<3;++k)posed[j][k]=posed[static_cast<size_t>(parent)][k]+offset[k];}}
const auto right=s.hips[0],left=s.hips[1];
const float angle=std::atan2(posed[right][2]-posed[left][2],-(posed[right][0]-posed[left][0]));
value[1]=root[1];value[3]=std::cos(angle);value[4]=std::sin(angle);
for(size_t j=0;j<J;++j){value[5+j*3]=posed[j][0]-value[0];value[6+j*3]=posed[j][1];value[7+j*3]=posed[j][2]-value[2];}
for(unsigned joint:s.end_effectors)for(int d=0;d<6;++d)value[rotation_begin+joint*6+static_cast<size_t>(d)]=static_cast<float>(global[joint].v[(d%3)*3+d/3]);
return angle;
}
}
std::expected<sequence_transition, std::string> prepare_sequence_transition(
const ggml_motion_weights &, std::span<const float> previous,
const ggml_motion_weights &weights, std::span<const float> previous,
std::size_t continuation_frames, unsigned transition_frames) {
constexpr size_t features = 273;
const size_t overlap=transition_frames;
if (!overlap || overlap>=continuation_frames || previous.size()<=(overlap*features))
const auto *s=find_skeleton(weights.skeleton_key());
if(!s)return std::unexpected("unsupported sequence skeleton");
const size_t D=s->motion_dim(),J=s->joints(),rotation_begin=5+3*J,rotation_end=rotation_begin+6*J,overlap=transition_frames;
if(!overlap||overlap>=continuation_frames||previous.size()<=overlap*D||previous.size()%D)
return std::unexpected("invalid sequence transition");
sequence_transition result;
result.observed.resize((continuation_frames+overlap)*features);
result.observed.resize((continuation_frames+overlap)*D);
result.observed_mask.resize(result.observed.size());
const size_t previous_start=previous.size()-overlap*features;
constexpr std::array<std::pair<size_t, size_t>, 3> constrained = {{{0, 71}, {113, 125}, {191, 203}}};
for (size_t frame=0; frame<overlap; ++frame) {
const size_t base=frame*features;
std::array<float, features> value{};
const float *raw=previous.data()+previous_start+base;
std::copy_n(raw,features,value.data());
mat decoded[22],local[22],global[22]; for(int j=0;j<22;++j)decoded[j]=cont6(value.data()+71+j*6);
for(int j=0;j<22;++j) local[j]=parent[j]<0?decoded[j]:mul(trans(decoded[parent[j]]),decoded[j]);
float root[3]={value[0]+value[5],value[6],value[2]+value[7]}, posed[22][3]{};
for(int j=0;j<22;++j){if(parent[j]<0){global[j]=local[j];posed[j][0]=root[0];posed[j][1]=root[1];posed[j][2]=root[2];}else{global[j]=mul(global[parent[j]],local[j]);float d[3];rotate(global[parent[j]],offset[j],d);for(int k=0;k<3;++k)posed[j][k]=posed[parent[j]][k]+d[k];}}
// FullBodyConstraintSet: smooth root, root Y, heading, all joint
// positions; EndEffectorConstraintSet adds its four rotation blocks.
value[0]=value[0]; value[1]=root[1]; value[2]=value[2];
// `compute_heading_angle`: right hip minus left hip.
const float dx=posed[2][0]-posed[1][0], dz=posed[2][2]-posed[1][2], angle=std::atan2(dz,-dx);
value[3]=std::cos(angle); value[4]=std::sin(angle);
for(int j=0;j<22;++j){value[5+j*3]=posed[j][0]-value[0];value[6+j*3]=posed[j][1];value[7+j*3]=posed[j][2]-value[2];}
for(int j: {7,8,20,21}) for(int d=0;d<6;++d)
value[71+j*6+d]=static_cast<float>(global[j].v[(d%3)*3+d/3]);
std::copy_n(value.data(),203,result.observed.data()+base);
for (const auto &[first,last] : constrained)
std::fill(result.observed_mask.begin()+static_cast<std::ptrdiff_t>(base+first),
result.observed_mask.begin()+static_cast<std::ptrdiff_t>(base+last),1.F);
}
result.origin_x=result.observed[0];
result.origin_z=result.observed[2];
for (size_t frame=0; frame<overlap; ++frame) {
auto *row=result.observed.data()+frame*features;
row[0]-=result.origin_x;
row[2]-=result.origin_z;
}
// First heading comes from the first retained full-body constraint.
const size_t first=(previous.size()-overlap*features);
const float *raw=previous.data()+first;
std::array<float, features> value{}; std::copy_n(raw,features,value.data());
mat decoded[22],local[22],global[22]; for(int j=0;j<22;++j)decoded[j]=cont6(value.data()+71+j*6);
for(int j=0;j<22;++j)local[j]=parent[j]<0?decoded[j]:mul(trans(decoded[parent[j]]),decoded[j]);
float root[3]={value[0]+value[5],value[6],value[2]+value[7]}, posed[22][3]{};
for(int j=0;j<22;++j){if(parent[j]<0){global[j]=local[j];for(int k=0;k<3;++k)posed[j][k]=root[k];}else{global[j]=mul(global[parent[j]],local[j]);float d[3];rotate(global[parent[j]],offset[j],d);for(int k=0;k<3;++k)posed[j][k]=posed[parent[j]][k]+d[k];}}
result.first_heading=std::atan2(posed[2][2]-posed[1][2],-(posed[2][0]-posed[1][0]));
const size_t previous_start=previous.size()-overlap*D;
std::vector<float> value(D);
for(size_t frame=0;frame<overlap;++frame){const size_t base=frame*D;condition_row(previous.data()+previous_start+base,*s,value.data());std::copy_n(value.data(),rotation_end,result.observed.data()+base);std::fill(result.observed_mask.begin()+static_cast<std::ptrdiff_t>(base),result.observed_mask.begin()+static_cast<std::ptrdiff_t>(base+rotation_begin),1.F);for(unsigned joint:s->end_effectors){const size_t first=base+rotation_begin+joint*6;std::fill(result.observed_mask.begin()+static_cast<std::ptrdiff_t>(first),result.observed_mask.begin()+static_cast<std::ptrdiff_t>(first+6),1.F);}}
result.origin_x=result.observed[0];result.origin_z=result.observed[2];
for(size_t frame=0;frame<overlap;++frame){auto*row=result.observed.data()+frame*D;row[0]-=result.origin_x;row[2]-=result.origin_z;}
result.first_heading=condition_row(previous.data()+previous_start,*s,value.data());
return result;
}
std::expected<std::vector<float>, std::string> sample_motion_sequence_from_noise(
const ggml_motion_weights &weights, std::span<const sampled_sequence_segment> segments,
unsigned transition_frames, unsigned steps, float text_weight, float constraint_weight) {
constexpr size_t features = 273;
if (segments.empty() || !transition_frames)
return std::unexpected("sequence requires segments and a transition");
auto gm=weights.f32_values("stats.global_root.mean"), gs=weights.f32_values("stats.global_root.std");
auto bm=weights.f32_values("stats.body.mean"), bs=weights.f32_values("stats.body.std");
if (!gm || !gs || !bm || !bs) return std::unexpected("motion GGUF lacks motion statistics");
// Upstream Stats normalizes with sqrt(std^2 + 1e-5), rather than raw std.
auto scale = [](float stddev) { return std::sqrt(stddev * stddev + 1.e-5F); };
auto unnormalize = [&](std::vector<float> &motion) { for(size_t row=0;row<motion.size()/features;++row) { auto *v=motion.data()+row*features; for(size_t d=0;d<5;++d)v[d]=v[d]*scale((*gs)[d])+(*gm)[d]; for(size_t d=0;d<268;++d)v[5+d]=v[5+d]*scale((*bs)[d])+(*bm)[d]; } };
auto normalize = [&](std::vector<float> &motion) { for(size_t row=0;row<motion.size()/features;++row) { auto *v=motion.data()+row*features; for(size_t d=0;d<5;++d)v[d]=(v[d]-(*gm)[d])/scale((*gs)[d]); for(size_t d=0;d<268;++d)v[5+d]=(v[5+d]-(*bm)[d])/scale((*bs)[d]); } };
std::vector<float> joined, previous;
for (size_t index=0; index<segments.size(); ++index) {
const auto &segment=segments[index];
const size_t sampled_frames=segment.frames+(index ? transition_frames : 0);
if (segment.frames < 2 || segment.embedding.size()!=4096 ||
segment.initial_noise.size()!=sampled_frames*features)
return std::unexpected("invalid sampled sequence segment");
std::vector<float> current;
if (!index) {
auto sampled=sample_motion_from_noise(weights,segment.initial_noise,segment.embedding,
sampled_frames,steps,text_weight,constraint_weight);
if (!sampled) return std::unexpected(sampled.error());
current=std::move(*sampled);
unnormalize(current);
} else {
const size_t overlap=transition_frames;
if (overlap >= segment.frames || previous.size()<overlap*features)
return std::unexpected("transition must be shorter than every following segment");
auto transition=prepare_sequence_transition(weights,previous,segment.frames,transition_frames);
if (!transition) return std::unexpected(transition.error());
const float origin_x=transition->origin_x;
const float origin_z=transition->origin_z;
normalize(transition->observed);
auto sampled=sample_motion_from_noise_conditioned(weights,segment.initial_noise,segment.embedding,
transition->observed,transition->observed_mask,transition->first_heading,sampled_frames,steps,text_weight,constraint_weight);
if (!sampled) return std::unexpected(sampled.error());
current=std::move(*sampled);
unnormalize(current);
for (size_t frame=0; frame<sampled_frames; ++frame) {
auto *row=current.data()+frame*features;
row[0]+=origin_x;
row[2]+=origin_z;
}
const size_t start=joined.size()-overlap*features;
for (size_t frame=0; frame<overlap; ++frame) {
const float alpha=overlap==1?.5F:1.F-float(frame)/float(overlap-1);
for (size_t d=0; d<features; ++d)
joined[start+frame*features+d]=alpha*joined[start+frame*features+d]+(1.F-alpha)*current[frame*features+d];
}
joined.insert(joined.end(),current.begin()+static_cast<std::ptrdiff_t>(overlap*features),current.end());
}
if (!index) joined=current;
previous=std::move(current);
}
const size_t D=weights.motion_dim(),body=D-5;
if(segments.empty()||!transition_frames||!D)return std::unexpected("sequence requires segments and a transition");
auto gm=weights.f32_values("stats.global_root.mean"),gs=weights.f32_values("stats.global_root.std");
auto bm=weights.f32_values("stats.body.mean"),bs=weights.f32_values("stats.body.std");
if(!gm||!gs||!bm||!bs||gm->size()!=5||gs->size()!=5||bm->size()!=body||bs->size()!=body)return std::unexpected("motion GGUF lacks compatible motion statistics");
auto scale=[](float stddev){return std::sqrt(stddev*stddev+1.e-5F);};
auto unnormalize=[&](std::vector<float>&motion){for(size_t row=0;row<motion.size()/D;++row){auto*v=motion.data()+row*D;for(size_t d=0;d<5;++d)v[d]=v[d]*scale((*gs)[d])+(*gm)[d];for(size_t d=0;d<body;++d)v[5+d]=v[5+d]*scale((*bs)[d])+(*bm)[d];}};
auto normalize=[&](std::vector<float>&motion){for(size_t row=0;row<motion.size()/D;++row){auto*v=motion.data()+row*D;for(size_t d=0;d<5;++d)v[d]=(v[d]-(*gm)[d])/scale((*gs)[d]);for(size_t d=0;d<body;++d)v[5+d]=(v[5+d]-(*bm)[d])/scale((*bs)[d]);}};
std::vector<float> joined,previous;
for(size_t index=0;index<segments.size();++index){const auto&segment=segments[index];const size_t sampled_frames=segment.frames+(index?transition_frames:0);if(segment.frames<2||segment.embedding.size()!=4096||segment.initial_noise.size()!=sampled_frames*D)return std::unexpected("invalid sampled sequence segment");std::vector<float>current;if(!index){auto sampled=sample_motion_from_noise(weights,segment.initial_noise,segment.embedding,sampled_frames,steps,text_weight,constraint_weight);if(!sampled)return std::unexpected(sampled.error());current=std::move(*sampled);unnormalize(current);}else{const size_t overlap=transition_frames;if(overlap>=segment.frames||previous.size()<overlap*D)return std::unexpected("transition must be shorter than every following segment");auto transition=prepare_sequence_transition(weights,previous,segment.frames,transition_frames);if(!transition)return std::unexpected(transition.error());const float origin_x=transition->origin_x,origin_z=transition->origin_z;normalize(transition->observed);auto sampled=sample_motion_from_noise_conditioned(weights,segment.initial_noise,segment.embedding,transition->observed,transition->observed_mask,transition->first_heading,sampled_frames,steps,text_weight,constraint_weight);if(!sampled)return std::unexpected(sampled.error());current=std::move(*sampled);unnormalize(current);for(size_t frame=0;frame<sampled_frames;++frame){auto*row=current.data()+frame*D;row[0]+=origin_x;row[2]+=origin_z;}const size_t start=joined.size()-overlap*D;for(size_t frame=0;frame<overlap;++frame){const float alpha=overlap==1?.5F:1.F-float(frame)/float(overlap-1);for(size_t d=0;d<D;++d)joined[start+frame*D+d]=alpha*joined[start+frame*D+d]+(1.F-alpha)*current[frame*D+d];}joined.insert(joined.end(),current.begin()+static_cast<std::ptrdiff_t>(overlap*D),current.end());}if(!index)joined=current;previous=std::move(current);}
return joined;
}
} // namespace kimodo::detail

101
src/skeleton.hpp Normal file
View File

@ -0,0 +1,101 @@
#pragma once
#include <array>
#include <cstddef>
#include <span>
#include <string_view>
namespace kimodo::detail {
// Joint names and parent graphs are copied from NVIDIA Kimodo's Apache-2.0
// kimodo/skeleton/definitions.py. Parent-local offsets were extracted from
// the accompanying joints.p assets in the trusted reference container.
struct skeleton_spec {
std::string_view key;
std::span<const std::string_view> names;
std::span<const int> parents;
std::span<const std::array<float, 3>> offsets;
std::array<unsigned, 2> hips; // right, left
std::array<unsigned, 4> end_effectors; // left foot, right foot, left hand, right hand
[[nodiscard]] constexpr std::size_t joints() const noexcept { return names.size(); }
[[nodiscard]] constexpr std::size_t motion_dim() const noexcept { return 9 + 12 * joints(); }
[[nodiscard]] constexpr std::size_t body_dim() const noexcept { return motion_dim() - 5; }
};
inline constexpr std::array<std::string_view,22> smplx22_names{
std::string_view{"pelvis"}, "left_hip", "right_hip", "spine1", "left_knee", "right_knee",
"spine2", "left_ankle", "right_ankle", "spine3", "left_foot", "right_foot", "neck",
"left_collar", "right_collar", "head", "left_shoulder", "right_shoulder", "left_elbow",
"right_elbow", "left_wrist", "right_wrist"};
inline constexpr std::array smplx22_parents{-1,0,0,0,1,2,3,4,5,6,7,8,9,9,9,12,13,14,16,17,18,19};
inline constexpr std::array<std::array<float,3>,22> smplx22_offsets{{
{0,0,0},{.052299179F,-.093935639F,-.027606763F},{-.057192899F,-.106548190F,-.022217851F},
{-.001495834F,.112929940F,-.024981268F},{.058866613F,-.416441321F,-.006556974F},
{-.048074268F,-.397559673F,-.014061437F},{.006900469F,.145636231F,-.006858510F},
{-.041737989F,-.437583506F,-.029511765F},{.014489345F,-.446852267F,-.018029511F},
{-.010334037F,.056081813F,.021115851F},{.049293540F,-.065279245F,.126259089F},
{-.040575184F,-.065286517F,.127075911F},{-.011025756F,.171365142F,-.028827066F},
{.047724526F,.087643057F,-.008375450F},{-.046636276F,.086612143F,-.014864366F},
{.024654359F,.175390735F,.024463326F},{.126284808F,.057680372F,-.013885141F},
{-.109341696F,.053674292F,-.009117880F},{.272907287F,-.069853373F,-.039094493F},
{-.292028785F,-.035440356F,-.024564851F},{.276173830F,.021254137F,-.002478220F},
{-.271878421F,-.004834589F,-.016445294F}}};
inline constexpr std::array<std::string_view,30> soma30_names{
std::string_view{"Hips"}, "Spine1", "Spine2", "Chest", "Neck1", "Neck2", "Head", "Jaw",
"LeftEye", "RightEye", "LeftShoulder", "LeftArm", "LeftForeArm", "LeftHand",
"LeftHandThumbEnd", "LeftHandMiddleEnd", "RightShoulder", "RightArm", "RightForeArm",
"RightHand", "RightHandThumbEnd", "RightHandMiddleEnd", "LeftLeg", "LeftShin", "LeftFoot",
"LeftToeBase", "RightLeg", "RightShin", "RightFoot", "RightToeBase"};
inline constexpr std::array soma30_parents{-1,0,1,2,3,4,5,6,6,6,3,10,11,12,13,13,3,16,17,18,19,19,0,22,23,24,0,26,27,28};
inline constexpr std::array<std::array<float,3>,30> soma30_offsets{{
{0,0,0},{-.00013727F,.0500376256F,-.00053726669F},{-1.86574103e-9F,.0712530139F,-.000298248546F},
{-5.75188398e-9F,.0755006305F,-.00815970992F},{-.00181676517F,.263112953F,-.00553348292F},
{-2.85102231e-8F,.0770939664F,.0230258546F},{-4.5975437e-8F,.0612891595F,.0195370861F},
{2.63687901e-5F,.0047559225F,.0309494062F},{.0320638079F,.0538020513F,.0758688308F},
{-.0322244017F,.05361869F,.0755823359F},{.0162165175F,.232371641F,.0511341324F},
{.149198457F,2.19397873e-8F,-.0550232576F},{.287393078F,2.50268389e-9F,-2.58787737e-5F},
{.270939812F,-7.06625108e-9F,2.60897248e-5F},{.122686267F,-.0322017573F,.0483306876F},
{.190119595F,-.00312878387F,-.000339570373F},{-.0138011824F,.231803086F,.0521415786F},
{-.150371962F,1.17387901e-7F,-.0554560437F},{-.287366393F,1.87628082e-8F,-2.59709359e-5F},
{-.271336198F,-1.16767401e-9F,2.61269368e-5F},{-.122642483F,-.0321145448F,.0480403904F},
{-.190005945F,-.00306615542F,-.0003157343F},{.10043214F,-.0843452671F,.0259565473F},
{-1e-8F,-.432217537F,-.00802912805F},{1e-8F,-.421550959F,-.0348152298F},
{0,-.0505947206F,.132315294F},{-.10047278F,-.0829525995F,.0262031695F},
{1e-8F,-.433622059F,-.00805555828F},{2e-8F,-.421173943F,-.0347839785F},
{-3.42907669e-9F,-.0507960932F,.132841956F}}};
inline constexpr std::array<std::string_view,34> g1skel34_names{
std::string_view{"pelvis_skel"}, "left_hip_pitch_skel", "left_hip_roll_skel", "left_hip_yaw_skel",
"left_knee_skel", "left_ankle_pitch_skel", "left_ankle_roll_skel", "left_toe_base",
"right_hip_pitch_skel", "right_hip_roll_skel", "right_hip_yaw_skel", "right_knee_skel",
"right_ankle_pitch_skel", "right_ankle_roll_skel", "right_toe_base", "waist_yaw_skel",
"waist_roll_skel", "waist_pitch_skel", "left_shoulder_pitch_skel", "left_shoulder_roll_skel",
"left_shoulder_yaw_skel", "left_elbow_skel", "left_wrist_roll_skel", "left_wrist_pitch_skel",
"left_wrist_yaw_skel", "left_hand_roll_skel", "right_shoulder_pitch_skel",
"right_shoulder_roll_skel", "right_shoulder_yaw_skel", "right_elbow_skel",
"right_wrist_roll_skel", "right_wrist_pitch_skel", "right_wrist_yaw_skel", "right_hand_roll_skel"};
inline constexpr std::array g1skel34_parents{-1,0,1,2,3,4,5,6,0,8,9,10,11,12,13,0,15,16,17,18,19,20,21,22,23,24,17,26,27,28,29,30,31,32};
inline constexpr std::array<std::array<float,3>,34> g1skel34_offsets{{
{0,0,0},{.064452F,-.1027F,0},{.052F,-.030465F,0},{0,-.12412F,.025001F},
{.0021489F,-.17734F,-.078273F},{-.000094445F,-.30001F,0},{0,-.017558F,0},{0,-.035F,.14F},
{-.064452F,-.1027F,0},{-.052F,-.030465F,0},{0,-.12412F,.025001F},{-.0021489F,-.17734F,-.078273F},
{.000094445F,-.30001F,0},{0,-.017558F,0},{0,-.035F,.14F},{0,0,0},{0,.044F,-.0039635F},
{0,0,0},{.10022F,.24778F,.0039563F},{.038F,-.013831F,0},{.00624F,-.1032F,0},
{0,-.080518F,.015783F},{.00188791F,-.01F,.1F},{0,0,.038F},{0,0,.046F},{0,0,.1F},
{-.10021F,.24778F,.0039563F},{-.038F,-.013831F,0},{-.00624F,-.1032F,0},
{0,-.080518F,.015783F},{-.00188791F,-.01F,.1F},{0,0,.038F},{0,0,.046F},{0,0,.1F}}};
inline constexpr skeleton_spec smplx22_spec{"smplx22", smplx22_names, smplx22_parents, smplx22_offsets, {2,1}, {7,8,20,21}};
inline constexpr skeleton_spec soma30_spec{"soma30", soma30_names, soma30_parents, soma30_offsets, {26,22}, {24,28,13,19}};
inline constexpr skeleton_spec g1skel34_spec{"g1skel34", g1skel34_names, g1skel34_parents, g1skel34_offsets, {8,1}, {6,13,24,32}};
inline constexpr const skeleton_spec *find_skeleton(std::string_view key) noexcept {
if (key == smplx22_spec.key) return &smplx22_spec;
if (key == soma30_spec.key) return &soma30_spec;
if (key == g1skel34_spec.key) return &g1skel34_spec;
return nullptr;
}
} // namespace kimodo::detail

View File

@ -1,4 +1,5 @@
#include <kimodo/kimodo.hpp>
#include <array>
#include <cstdio>
int main(int argc,char**argv){if(argc!=2)return 2;auto m=kimodo::model::load(argv[1]);if(!m){std::fprintf(stderr,"%s\n",m.error().c_str());return 1;}std::array<float,kimodo::embedding_width> e{};auto r=(*m)->generate_embedding(e,2,1,42,2.f,2.f);if(!r){std::fprintf(stderr,"%s\n",r.error().c_str());return 1;}if(r->frames!=2||r->joints!=22||r->root_positions.size()!=6||r->local_rotations_xyzw.size()!=176)return 1;return 0;}
#include <fstream>
int main(int argc,char**argv){if(argc<2||argc>4)return 2;const unsigned joints=argc>=3?static_cast<unsigned>(std::stoul(argv[2])):22;auto m=kimodo::model::load(argv[1]);if(!m){std::fprintf(stderr,"%s\n",m.error().c_str());return 1;}std::array<float,kimodo::embedding_width> e{};auto r=(*m)->generate_embedding(e,2,1,42,2.f,2.f);if(!r){std::fprintf(stderr,"%s\n",r.error().c_str());return 1;}if(r->frames!=2||r->joints!=joints||r->root_positions.size()!=6||r->local_rotations_xyzw.size()!=2*joints*4)return 1;if(argc==4){std::ofstream out(argv[3],std::ios::binary);out.write(reinterpret_cast<const char*>(r->root_positions.data()),static_cast<std::streamsize>(r->root_positions.size()*sizeof(float)));out.write(reinterpret_cast<const char*>(r->local_rotations_xyzw.data()),static_cast<std::streamsize>(r->local_rotations_xyzw.size()*sizeof(float)));if(!out)return 1;}return 0;}