corridorkey-mrp-mlx/docs
cmoyates f73aabf5af
chore: default bf16+fused decode on, add benchmark results to plan
load_model() now defaults to dtype=bf16, fused_decode=True. Both are
free (zero parity regression, bit-exact fused path). Backbone/sigmoid
stay fp32.

Plan updated with benchmark results: tiled+GC = 12x peak memory
reduction at 2048x2048 (27.6GB → 2.3GB), all acceptance criteria met.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-08 21:27:56 -02:30
..
brainstorms docs: add brainstorm + plan for MLX memory optimizations 2026-03-08 21:12:59 -02:30
plans chore: default bf16+fused decode on, add benchmark results to plan 2026-03-08 21:27:56 -02:30