- Cache nn.Upsample instances in DecoderHead/GreenFormer __init__ (eliminated ~7 allocations per forward pass) - Add mx.compile() support via load_model(compile=True) - Benchmark harness: eager vs compiled, multi-resolution, parity checks - Tiled inference with overlap blending for large images - Profiling utilities with forced mx.eval for accurate timing - Reference comparison script (scripts/compare_reference.py) - 12 new tests (compiled consistency + tiling) - README performance section with Apple Silicon guidance Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| plans | ||