modelbeast/scripts/perfcheck/run_fleet.py
m3ultra 27564e80eb perfcheck: nightly fleet GPU canary + the contention lessons that shaped it
BENCHMARKS.md says what a model cost the day it was measured; nothing noticed if a
macOS/torch update or a thermal fault halved a node. perfcheck runs the whole fleet
in ~35s off godcheck's 03:30 cron on the m4mini and reports drift into
GODCHECK_LATEST.md.

Probes matmul (fp16/fp32), memory bandwidth, and SDPA at head_dim 64 vs 56 — the
latter turning CorridorKey's fast-path cliff into a permanent canary: it confirms
the padding win fleet-wide (1.96x-5.26x) and tells us if a future torch closes it.
Runs on venvs/rmbg/bin/python, already identical fleet-wide, so nothing new is
installed (nothing lands on the disk-tight m1max).

First cross-machine capability table for all 6 nodes. The M3 Ultra is ~2.2x the M1
Ultra on fp16 matmul, but they share ~625 GB/s — so bandwidth-bound stages run alike
while compute-bound ones scale. Both Ultras reach only ~78% of spec bandwidth on a
single kernel; the smaller Macs hit ~88%.

Measuring a fleet that is doing real work is the whole problem, and naive
benchmarking here is off by 11x:

- min, not median: a concurrent trellis_mac job dragged a median-of-5 matmul from
  ~24500 to ~2150 GFLOP/s, which reads exactly like a catastrophic regression.
- n=4096 not 2048: 2048 is dispatch-bound and swung 48% run-to-run; 4096 reproduces
  to 0.1% even while contended.
- sdpa 16x2048 not 8x1024: sub-ms probes are dispatch noise — 8x1024 gave ratios of
  0.79/3.95/2.35 on three runs of one machine, the first "proving" 56 is faster.
- busy nodes are excluded, not blamed: GPU contention is invisible to load average
  (M3 Ultra read load 2.45 with its GPU pinned), so bench.py samples ioreg GPU% before
  it touches the GPU — our own matmul pins the device, so ordering is the trick.
- baselines are the median of recent history, not a saved best: the M4 Pro also serves
  Ollama and is bimodal (~3200 vs ~5500 fp32), so a best-observed baseline pins to a
  lucky outlier and alerts forever.
- a regression must repeat before it is believed ([~] watching -> [!] CONFIRMED).

Validated by re-running the fleet against its own baselines: zero false alarms,
including a sweep where the M3 Ultra read 43 GB/s under load and was correctly
marked BUSY rather than reported as a 93% regression.

Also found: the m4mini is the only node with Tailscale SSH (RunSSH: true) and it does
NOT propagate remote exit codes — `ssh m4mini "exit 7"` returns 0, so any
`if ssh m4mini ...` test silently always passes. Test on output instead, which is
what godcheck already does (and why it is unaffected). It also cannot ssh to itself,
so run_fleet detects its own tailnet IP and benches the local node via the shell.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 18:03:03 +10:00

342 lines
15 KiB
Python
Executable File

#!/usr/bin/env python3
"""PERFCHECK fleet runner — benchmark every MODELBEAST node, detect drift, write a report.
Stdlib only, so it runs on any node's system python3 (notably the m4mini, which hosts the
godcheck cron). ``bench.py`` is copied to each node on every run, so a node can never drift
onto a stale copy of the benchmark itself.
./run_fleet.py # bench the fleet, compare against recent history, report
./run_fleet.py --only m1,m4 # subset
./run_fleet.py --show # print the current baselines and exit (no benching)
Exit 1 if any node shows a CONFIRMED regression, so cron/CI can react.
WHY IT LOOKS LIKE THIS — the fleet serves real work, and that dominates every design choice:
* **Baselines are the median of recent history, not a saved "best".** These machines have
neighbours. The M4 Pro also serves Ollama and its fp32 is bimodal (~3200 vs ~5500 GFLOP/s)
depending on whether a request is in flight; the M1 Ultra runs `mrpgi` on its GPU and swings
8301-15891 (1.9x). A best-observed baseline pins to a lucky outlier the node rarely reaches
again, which means alerting every night forever. A median is what the machine actually does.
* **A busy node is excluded, not blamed.** GPU contention is invisible to load average (the M3
Ultra read load 2.45 with its GPU pinned by trellis_mac), so bench.py samples GPU% before it
touches the GPU. Above BUSY_PCT the sample is noise: no alert, and kept out of the baseline.
* **A regression must repeat before it is believed.** Transient contention is far more common
than real regression and does not recur on schedule. First sighting is [~] watching; only a
second consecutive one is [!] CONFIRMED and fails the run.
* **Suspicious nodes are retried and the best attempt kept.** Contention only ever makes a node
look slower, so best-of-attempts is the least-interfered estimate. Only nodes that look off
pay for this, so the common case stays one pass.
"""
from __future__ import annotations
import argparse
import json
import os
import subprocess
import sys
from datetime import datetime, timezone
from pathlib import Path
HERE = Path(__file__).resolve().parent
BENCH = HERE / "bench.py"
HISTORY = HERE / "history.jsonl" # every sample ever taken, append-only
BASELINES = HERE / "baselines.json" # derived view of HISTORY, rewritten each run
PENDING = HERE / "pending.json" # breaches seen last run — see confirm()
# name -> (ssh target, MODELBEAST root). Mirrors nodes.json + the local primary. Embedded
# rather than read from nodes.json because that file is gitignored and primary-only, so the
# m4mini cron would not see it; godcheck hardcodes its hosts the same way.
NODES = [
("m3ultra", "m3ultra@100.89.131.57", "/Users/m3ultra/Documents/MODELBEAST"),
("m1", "johnking@100.91.239.7", "/Users/johnking/MODELBEAST"),
("m4", "m4pro@100.69.21.128", "/Users/m4pro/MODELBEAST"),
("studio", "studio@100.92.78.24", "/Users/studio/MODELBEAST"),
("m4mini", "m4mini@100.124.220.31", "/Users/m4mini/MODELBEAST"),
("mini", "mini@100.79.229.50", "/Users/mini/MODELBEAST"),
]
HIGHER_BETTER = {"matmul_fp16_gflops", "matmul_fp32_gflops", "bandwidth_gbs"}
# sdpa_ratio is diagnostic: a change means the fused-attention landscape moved (e.g. a torch
# release finally handles head_dim=56), which is a headline to read, not a regression to fail.
DIAGNOSTIC = {"sdpa_ratio"}
THRESHOLD = 0.20 # flag a >20% move vs baseline
BUSY_PCT = 50 # GPU% (pre-bench) above which a node is unmeasurable
HISTORY_WINDOW = 9 # samples per node the median is drawn from
MIN_SAMPLES = 3 # below this a node has no baseline yet and cannot alert
# Numbers are only comparable within a bench schema. bench.py bumps SCHEMA when a probe's
# shape changes; older history is then ignored rather than being median'd against the new
# shape and mistaken for a regression. Keep in step with bench.py's SCHEMA.
SCHEMA = 2
SSH = ["ssh", "-o", "ConnectTimeout=10", "-o", "BatchMode=yes"]
# ---------------------------------------------------------------------------
# Running
# ---------------------------------------------------------------------------
def local_ip() -> str:
"""This machine's tailnet IPv4, or "" — used to spot our own entry in NODES."""
for cmd in (["tailscale", "ip", "-4"],
["/Applications/Tailscale.app/Contents/MacOS/Tailscale", "ip", "-4"]):
try:
out = subprocess.run(cmd, capture_output=True, text=True, timeout=8)
if out.returncode == 0 and out.stdout.strip():
return out.stdout.strip().splitlines()[0].strip()
except Exception:
continue
return ""
def run_node(name: str, host: str, root: str, self_ip: str = "") -> dict:
"""Run bench.py on a node (locally if that node is us) and return parsed JSON.
Benching ourselves goes through the shell, not ssh: a host looping back to its own tailnet
address needs its own host key trusted and its own pubkey authorized, which fails out of the
box ("Host key verification failed" on the m4mini watchdog, which must bench itself as part
of the fleet). Running local also skips a pointless copy and round-trip.
"""
remote = "/tmp/perfcheck_bench.py"
is_self = bool(self_ip) and host.split("@")[-1] == self_ip
try:
if is_self:
proc = subprocess.run(
[f"{root}/venvs/rmbg/bin/python", str(BENCH)],
capture_output=True, text=True, timeout=600,
)
else:
subprocess.run(
["scp", "-q", "-o", "ConnectTimeout=10", "-o", "BatchMode=yes",
str(BENCH), f"{host}:{remote}"],
check=True, capture_output=True, timeout=60,
)
proc = subprocess.run(
SSH + [host, f"{root}/venvs/rmbg/bin/python {remote}"],
capture_output=True, text=True, timeout=600,
)
except subprocess.TimeoutExpired:
return {"node": name, "error": "timeout"}
except subprocess.CalledProcessError as e:
return {"node": name, "error": f"scp failed: {e.stderr.decode()[:120].strip()}"}
if proc.returncode != 0:
tail = (proc.stderr or proc.stdout).strip().splitlines()
return {"node": name, "error": (tail[-1][:160] if tail else f"exit {proc.returncode}")}
try:
# The bench prints one JSON object; take the last line in case a lib warns first.
return {"node": name, **json.loads(proc.stdout.strip().splitlines()[-1])}
except (ValueError, IndexError):
return {"node": name, "error": f"unparseable: {proc.stdout[:120]!r}"}
def usable(r: dict) -> bool:
"""True if a sample is trustworthy enough to compare or to feed the baseline."""
return (
"error" not in r
and r.get("schema") == SCHEMA
and r.get("gpu_busy_pre", 0) < BUSY_PCT
and bool(r.get("metrics"))
)
def merge_best(runs: list[dict]) -> dict:
"""Fold repeat attempts at one node into its least-contended sample, per metric.
Contention can only make a node look slower, never faster, so the best value across
attempts is the closest estimate of what the machine can do right now. Same argument as
min-of-runs inside bench.py, lifted one level.
"""
ok = [r for r in runs if r and "error" not in r and r.get("metrics")]
if not ok:
return runs[0]
out = dict(ok[-1])
m = {}
for key in ok[0]["metrics"]:
vals = [r["metrics"][key] for r in ok if r["metrics"].get(key) is not None]
if vals:
m[key] = max(vals) if key in HIGHER_BETTER else min(vals)
# Recompute so the ratio stays consistent with the sdpa numbers reported beside it.
if m.get("sdpa_64_ms"):
m["sdpa_ratio"] = round(m["sdpa_56_ms"] / m["sdpa_64_ms"], 2)
out["metrics"] = m
out["attempts"] = len(ok)
# Least-contended attempt wins, so carry the quietest GPU reading with it.
out["gpu_busy_pre"] = min(r.get("gpu_busy_pre", 0) for r in ok)
return out
# ---------------------------------------------------------------------------
# Baselines (derived from history)
# ---------------------------------------------------------------------------
def _median(xs: list[float]) -> float:
xs = sorted(xs)
n = len(xs)
return xs[n // 2] if n % 2 else (xs[n // 2 - 1] + xs[n // 2]) / 2.0
def baselines_from_history() -> dict:
"""node -> {metrics: {k: median-of-recent}, n: samples}. Only clean samples count."""
per_node: dict[str, list[dict]] = {}
if HISTORY.exists():
for line in HISTORY.read_text().splitlines():
try:
r = json.loads(line)
except ValueError:
continue
if usable(r):
per_node.setdefault(r["node"], []).append(r["metrics"])
out = {}
for node, samples in per_node.items():
recent = samples[-HISTORY_WINDOW:]
if len(recent) < MIN_SAMPLES:
out[node] = {"metrics": {}, "n": len(recent)}
continue
keys = {k for s in recent for k in s}
med = {}
for k in keys:
vals = [s[k] for s in recent if s.get(k) is not None]
if vals:
med[k] = round(_median(vals), 3)
out[node] = {"metrics": med, "n": len(recent)}
return out
def compare(results: list[dict], baselines: dict) -> dict[str, str]:
"""Breaches keyed ``node/metric`` -> readable line. ``{}`` when all within threshold."""
out = {}
for r in results:
if not usable(r):
continue
base = baselines.get(r["node"], {}).get("metrics")
if not base:
continue # not enough history yet
for key, now in r["metrics"].items():
was = base.get(key)
if was in (None, 0) or now is None or key in DIAGNOSTIC:
continue
change = (now - was) / was
slower = -change if key in HIGHER_BETTER else change
if slower > THRESHOLD:
out[f"{r['node']}/{key}"] = f"{r['node']} {key}: {was} -> {now} ({slower:+.0%} SLOWER)"
elif -slower > THRESHOLD:
out[f"+{r['node']}/{key}"] = f"{r['node']} {key}: {was} -> {now} ({-slower:+.0%} faster)"
return out
def confirm(breaches: dict[str, str]) -> tuple[list[str], list[str]]:
"""Split breaches into (confirmed, provisional) by whether they also fired last run.
A regression is believed only once it repeats: transient contention is the overwhelmingly
common cause of one slow sample and it does not recur on schedule, while a real regression
does. A node getting *faster* is reported immediately — it is never an alarm.
"""
try:
prev = set(json.loads(PENDING.read_text()))
except Exception:
prev = set()
confirmed, provisional = [], []
for key, line in sorted(breaches.items()):
if key.startswith("+"):
confirmed.append(f"[+] {line}")
elif key in prev:
confirmed.append(f"[!] {line} — CONFIRMED (2 runs)")
else:
provisional.append(f"[~] {line} — unconfirmed, watching")
PENDING.write_text(json.dumps([k for k in breaches if not k.startswith("+")]) + "\n")
return confirmed, provisional
# ---------------------------------------------------------------------------
# Reporting
# ---------------------------------------------------------------------------
def report(results: list[dict], confirmed: list[str], provisional: list[str],
baselines: dict, stamp: str) -> str:
out = [f"# PERFCHECK {stamp}", ""]
out.append("| node | chip | RAM | matmul fp16 | matmul fp32 | bandwidth | sdpa 64 | sdpa 56 | 56/64 |")
out.append("|---|---|---|---|---|---|---|---|---|")
for r in sorted(results, key=lambda x: x["node"]):
if "error" in r:
out.append(f"| **{r['node']}** | [X] {r['error']} | | | | | | | |")
continue
m, chip = r["metrics"], (r.get("chip") or "").replace("Apple ", "")
busy = r.get("gpu_busy_pre", 0)
if busy >= BUSY_PCT:
chip += f" [~] BUSY {busy}%"
out.append(
f"| **{r['node']}** | {chip} | {r.get('ram_gb')}GB | "
f"{m['matmul_fp16_gflops']:.0f} | {m['matmul_fp32_gflops']:.0f} | "
f"{m['bandwidth_gbs']:.0f} GB/s | {m['sdpa_64_ms']:.2f}ms | "
f"{m['sdpa_56_ms']:.2f}ms | **{m['sdpa_ratio']}x** |"
)
thin = [n for n, b in baselines.items() if b["n"] < MIN_SAMPLES]
out.append("")
out.append("## drift vs baseline (median of last %d clean runs)" % HISTORY_WINDOW)
out.extend(confirmed or ["- no confirmed regressions"])
if provisional:
out.append("")
out.extend(provisional)
if thin:
out.append(f"- (warming up, <{MIN_SAMPLES} clean samples: {', '.join(sorted(thin))})")
return "\n".join(out) + "\n"
def main() -> int:
ap = argparse.ArgumentParser()
ap.add_argument("--only", help="comma-separated node names")
ap.add_argument("--quiet", action="store_true", help="print only the report")
ap.add_argument("--show", action="store_true", help="print current baselines and exit")
ap.add_argument("--attempts", type=int, default=2,
help="extra attempts for a node that looks regressed (default 2)")
a = ap.parse_args()
if a.show:
print(json.dumps(baselines_from_history(), indent=2, sort_keys=True))
return 0
nodes = NODES
if a.only:
want = {s.strip() for s in a.only.split(",")}
nodes = [n for n in NODES if n[0] in want]
baselines = baselines_from_history()
self_ip = local_ip()
stamp = datetime.now(timezone.utc).astimezone().strftime("%Y-%m-%d %H:%M %Z")
results = []
for name, host, root in nodes:
if not a.quiet:
print(f" benching {name} ...", file=sys.stderr, flush=True)
results.append(run_node(name, host, root, self_ip))
# Adaptive retry: only a node that looks regressed pays for extra attempts. A busy
# neighbour is far more likely than a real regression, and best-of-attempts costs nothing
# when the node is idle — so the common case stays a single pass over the fleet.
for i, r in enumerate(results):
if not usable(r) or not compare([r], baselines):
continue
name, host, root = next(n for n in nodes if n[0] == r["node"])
if not a.quiet:
print(f" {name} looks off — {a.attempts} more attempt(s) ...", file=sys.stderr, flush=True)
results[i] = merge_best(
[r] + [run_node(name, host, root, self_ip) for _ in range(a.attempts)]
)
confirmed, provisional = confirm(compare(results, baselines))
with HISTORY.open("a") as fh:
for r in results:
fh.write(json.dumps({"ts": stamp, **r}) + "\n")
# Rewrite the derived view *after* appending, so it reflects this run too.
BASELINES.write_text(json.dumps(baselines_from_history(), indent=2, sort_keys=True) + "\n")
print(report(results, confirmed, provisional, baselines, stamp))
return 1 if any(x.startswith("[!]") for x in confirmed) else 0
if __name__ == "__main__":
os.chdir(HERE)
sys.exit(main())