use_sdpa, stage_gc flags propagated through backbone/model/pipeline. bench_optimizations.py runs exhaustive/ablation/key sweep of all 6 toggles (slim, stage_gc, sdpa, bf16, fused_decode, gpu_preprocess) with latency + peak memory reporting. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| bench_mlx.py | ||
| bench_optimizations.py | ||
| compare_reference.py | ||
| convert_weights.py | ||
| dump_pytorch_reference.py | ||
| infer_pytorch.py | ||
| infer.py | ||
| smoke_2048.py | ||
| smoke_engine.py | ||