load_model() now defaults to dtype=bf16, fused_decode=True. Both are free (zero parity regression, bit-exact fused path). Backbone/sigmoid stay fp32. Plan updated with benchmark results: tiled+GC = 12x peak memory reduction at 2048x2048 (27.6GB → 2.3GB), all acceptance criteria met. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| 2026-03-01-feat-2048-smoke-test-plan.md | ||
| 2026-03-01-feat-corridorkey-mlx-inference-port-plan.md | ||
| 2026-03-01-feat-engine-integration-surface-plan.md | ||
| 2026-03-01-fix-converter-review-feedback-plan.md | ||
| 2026-03-01-phase4-hiera-backbone-plan.md | ||
| 2026-03-03-refactor-deep-modules-plan.md | ||
| 2026-03-08-feat-mlx-memory-optimizations-plan.md | ||