# HANDOFF2 — Multi-user auth + system dashboard + public VPS hosting **Audience:** the AI agent (Opus 4.8 / Claude Code) executing this build. Read `HANDOFF.md` §1–3 first for the architecture, operator contract, and working rules — they all still apply (one venv per tool, verify through the real browser UI, commit per milestone, push to origin, run `tests/smoke.sh` before each commit). This brief adds three capabilities to MODELBEAST: 1. **Accounts + login** so the owner's friends can run LOCAL models on this machine over the internet — paid/cloud operators stay owner-only by hard rule, and the owner controls how much GPU anyone can hog. 2. **A dashboard page** — btop-ish *snapshot on page refresh* (no live streaming): CPU/RAM/GPU/disk, what's running, what's queued, what ran recently. 3. **Public hosting** via the owner's spare VPS + Caddy reverse-proxy over Tailscale at `modelbeast.digalot.fyi` (and optionally `modelbeast.partyl.party`). **Non-negotiable design premise — LOCAL-ONLY FOR GUESTS, HARD RULE:** the owner's explicit intent is that this feature is 100% local generation on the M3 Ultra — guests cost him a tiny bit of electricity and NOTHING else. Guests must NEVER be able to (a) invoke ANY operator that uses a paid key (fal/OpenRouter/etc.) — this is role-based and NOT configurable per user, (b) read or write settings/keys, (c) delete other people's work, or (d) starve the GPU queue. Cloud operators aren't just "locked" in the UI for guests — they are filtered out server-side so the paid paths are structurally unreachable. Every task below serves one of those or the dashboard. Current state you inherit: 22 operators, FastAPI + SQLite + lanes runner, React/Vite UI, `mb` CLI (stdlib, `MB_HOST` env), server on :8777 via `scripts/serve.sh`, no auth (tailnet was the boundary). Key files: `server/main.py`, `server/db.py`, `server/settings.py`, `server/runner.py`, `web/src/App.jsx`, `mb`. --- ## Phase 1 — Accounts, sessions, tokens (backend) ### 1.1 Schema (additive migrations in `db.py::MIGRATIONS`, same pattern as asset_ids) ```sql CREATE TABLE IF NOT EXISTS users ( id TEXT PRIMARY KEY, -- db.new_id() username TEXT UNIQUE NOT NULL, -- lowercase, [a-z0-9_-]{2,24} pw_hash TEXT NOT NULL, -- bcrypt via passlib role TEXT NOT NULL DEFAULT 'guest', -- 'owner' | 'guest'; guests are LOCAL-ONLY (hard rule, no per-user cloud toggle) max_active_jobs INTEGER NOT NULL DEFAULT 4, created_at REAL NOT NULL ); CREATE TABLE IF NOT EXISTS api_tokens ( token_hash TEXT PRIMARY KEY, -- sha256 hex of the raw token user_id TEXT NOT NULL, name TEXT NOT NULL DEFAULT '', created_at REAL NOT NULL ); -- add columns: jobs.user_id TEXT, assets.user_id TEXT (nullable; legacy rows stay NULL = owner) ``` Deps: `uv add "passlib[bcrypt]" itsdangerous psutil` (server venv — all light). ### 1.2 Auth mechanics (`server/auth.py`) - **Sessions:** stateless signed cookie (`itsdangerous.URLSafeTimedSerializer`), payload `{user_id}`, max age 30 days. Cookie `mb_session`: HttpOnly, SameSite=Lax, `Secure` when request came via https (check `x-forwarded-proto`). Signing secret: settings key `auth_secret` — auto-generate `secrets.token_hex(32)` on first startup if unset (never expose via GET /api/settings — add it to a new `HIDDEN_KEYS` set). - **Tokens:** `Authorization: Bearer mbt_<40 hex>` for `mb` CLI / agents. Store only sha256. Token auth == that user, same permissions. - **Dependency** `current_user(request) -> user row`: cookie first, then bearer. 401 JSON on failure. Apply to EVERY /api route and the websocket (read cookie at accept; also accept `?token=` query param for WS since JS can't set WS headers). Static files stay public (the SPA must load to show the login screen). - **Login endpoints:** `POST /api/login {username,password}` → sets cookie (rate-limit: in-memory 5 fails/minute/IP → 429); `POST /api/logout`; `GET /api/me` → `{username, role, allow_cloud}`. - **No self-signup.** Owner manages users via CLI (1.4) and endpoints: `GET/POST/PATCH/DELETE /api/users` (owner-only; POST takes username+password; PATCH can reset password and set max_active_jobs). ### 1.3 Permission enforcement (mostly in `main.py`, one check in `runner.py`) | Action | Rule | |---|---| | `GET /api/operators` | for guests, **filter out every operator with non-empty `requires_env` server-side** — guests never even see the cloud operators (defense in depth, not just UI hiding). | | `GET/PUT /api/settings` | **owner only** — guests get 403. Never leak `auth_secret` or secret values to anyone (existing masking stays). | | `POST /api/jobs` | attach `user_id`. If operator has non-empty `requires_env` and role != owner → 403 `"guests run local models only"` (hard rule — enforced here regardless of what the UI showed). Count user's queued+running jobs; ≥ `max_active_jobs` → 429. | | job cancel/retry/delete | own jobs only, unless owner. | | `DELETE /api/assets/{id}` | own assets only, unless owner. (Viewing/downloading stays workspace-shared — it's a mates' box; note this choice in the UI copy.) | | `POST /api/assets` (upload) | any authed user; reject bodies > 1 GB early via Content-Length. Attach `user_id`. | | `/api/users*` | owner only. | | `/api/system` (Phase 2) | any authed user. | - `mb` CLI additions: read `MB_TOKEN` env → send bearer header on every request (including download/upload paths); `mb token create ` (owner, prints raw token ONCE); `mb user add|list|passwd` subcommands wrapping /api/users. - Update `AGENTS.md` §1/§2: agents now need `MB_TOKEN`; show how the owner mints one. ### 1.4 Bootstrap `scripts/users.py` (runs in server venv against the DB directly, for when you're locked out): `add [--owner] [--allow-cloud]` (prompts for password, no echo), `list`, `passwd `. On server startup, if users table is EMPTY: create user `monster` role=owner with a random password, print it to the server log ONCE with big banners, and require change via `mb user passwd`. (Owner = the human; never auto-create guests.) ### 1.5 Frontend (login + attribution) - `web/src/Login.jsx`: centered card, username/password, error state; App checks `GET /api/me` on load → render Login instead of the app when 401. Logout button + username chip in the header. On 401 from any API call, drop to login. - Jobs list and asset rows show the creator's username (tiny dim chip). "Mine/all" filter toggle on jobs. - Guests: cloud operators simply don't appear (the filtered /api/operators handles this automatically). Add a small dim footer note for guests: "local models only — everything you run is generated on this machine". Owner still sees everything. - WS: pass the session automatically (same-origin cookie works; add `?token=` fallback). **Acceptance (Phase 1):** unauthenticated API returns 401 but the login page renders; owner logs in, creates guest `testmate`; `testmate` in a private window sees ONLY local operators in the dropdown (no fal_*/openrouter_* anywhere), can run `flux_local`, gets 403 if a fal operator id is POSTed directly to /api/jobs (curl — prove the server-side rule, not just the UI), 403 on Settings, can't delete the owner's assets, and hits 429 at 5 active jobs; `MB_TOKEN` makes `mb ops`/`mb run` work; smoke.sh extended for: 401-without-auth, login flow, token flow, guest-cloud-403-via-direct-POST (bootstrap a test user inside the scratch-DB server). Browser-verify with screenshots. --- ## Phase 2 — Dashboard (snapshot on refresh) ### 2.1 `GET /api/system` (in `main.py`, gathering in a new `server/sysinfo.py`) Snapshot, cached 5 s (in-memory timestamp guard so refresh-spam is harmless): ```json { "cpu": {"percent": 34.2, "cores": 28, "per_core": [..28 floats..], "load1": 6.1}, "ram": {"used_gb": 181.2, "total_gb": 256.0, "percent": 70.8}, "disk": {"free_gb": 214.0, "total_gb": 926.0, "data_dir_gb": 62.3}, "gpu": {"util_percent": 71.0, "power_w": 89.2, "source": "macmon"} | {"source": "unavailable"}, "uptime_s": 123456, "server_started_s": 9876, "lanes": {"gpu": {"limit": 1, "running": 1, "queued": 3}, "cpu": {"limit": 3, "running": 0, "queued": 0}, "net": {"limit": 6, "running": 2, "queued": 0}}, "jobs_24h": {"done": 41, "error": 3, "by_operator": {"flux_local": 28, "trellis_mac": 2}}, "hf_cache_gb": 267.0 } ``` - CPU/RAM/disk/uptime: `psutil` (now in the server venv). `data_dir_gb` via a cached `du` of `data/` refreshed at most every 5 min (it's slow — never on the hot path); same for `hf_cache_gb`. - **GPU (Apple Silicon):** `powermetrics` needs sudo — do NOT use it. Use **macmon** (`brew install macmon`), which reads IOReport without sudo: run `macmon pipe -s 1` (one JSON sample), parse GPU utilization + power. **Verify the exact JSON field names empirically on this machine before coding the parser** — run it, look at the output, then write the parser. If macmon is missing or errors, return `{"source":"unavailable"}` and show "install macmon for GPU stats" in the UI — never crash the endpoint. Cache the sample with the 5 s guard (each call costs ~1 s of sampling). - Lanes: compute from DB (queued/running counts joined to each operator's `resources` via the registry) — do not reach into runner internals. ### 2.2 Dashboard UI (`web/src/Dashboard.jsx`, header tab "📊 Dashboard") btop vibe, MODELBEAST dark theme, **fetch once per mount + a ⟳ refresh button** (explicitly no polling): - Top stat cards: CPU % (with a tiny per-core bar strip), RAM used/total with bar, GPU % + watts (or the install hint), disk free (warn tint < 100 GB). - Lane strip: three cards "GPU 1/1 busy · 3 waiting" style. - "Now running / queued" table: operator, user, elapsed (running) or position (queued), cancel button when permitted. - "Recent (24 h)": last ~20 finished jobs — status icon, operator, user, duration, and a thumbnail for image outputs (reuse `assetFileURL` on the job's output asset — join via `parent_job` in a small `GET /api/jobs/recent` that embeds output asset ids; don't N+1 from the browser). - A one-line machine header: "M3 Ultra · 256 GB · up 3 d 4 h · 22 operators". **Acceptance (Phase 2):** dashboard renders real numbers while a `flux_local` job runs (GPU lane shows busy; CPU/RAM plausible); refresh button updates; endpoint answers in <1.5 s warm; works for guest accounts; screenshot-verified in the browser. --- ## Phase 3 — VPS public hosting The VPS is the only public surface; it reaches the M3 Ultra over the tailnet. Nothing about the M3's setup changes except that auth (Phase 1) is now mandatory before this phase ships. **Do not do this phase until Phase 1 acceptance passes.** ### 3.1 On the VPS (document as `docs/VPS.md` with copy-paste blocks; the owner runs these — ASK him to run them when Phase 1+2 are merged, don't assume SSH access to the VPS) ```bash # 1) join the tailnet curl -fsSL https://tailscale.com/install.sh | sh && sudo tailscale up # 2) Caddy (auto-HTTPS) sudo apt install -y caddy ``` `/etc/caddy/Caddyfile`: ``` modelbeast.digalot.fyi { encode zstd gzip reverse_proxy 100.89.131.57:8777 } # optional twin: modelbeast.partyl.party { reverse_proxy 100.89.131.57:8777 } ``` DNS: A record `modelbeast` → VPS public IP on each domain. `sudo systemctl reload caddy`. - **Subdomain, not a path.** `digalot.fyi/modelbeast` would need a Vite base-path + FastAPI root-path rework across the SPA, WS, and mb CLI for zero benefit. Use `modelbeast.`. (Say so in docs — the owner suggested the path form.) - Caddy proxies WebSockets out of the box; verify job-log streaming works through it. - Considered and rejected: Tailscale Funnel (no custom domain, bandwidth-capped, but zero-VPS — note it as the fallback if the VPS ever dies); proxy-level basic-auth (can't do per-user roles/cost-gating). ### 3.2 Hardening that lands in the app (do in Phase 1, verify here) - Trust `X-Forwarded-Proto` for the Secure-cookie flag (single known proxy). - Login rate-limit keyed on `X-Forwarded-For` when present. - Upload cap (1 GB) + request body limit in Caddy (`request_body max_size 1GB`). - The M3 keeps binding 0.0.0.0 on the tailnet — auth now protects it there too (agents/CLI use tokens; the LaunchAgent/serve.sh setup is unchanged, still 1 uvicorn worker — the in-proc queue + SQLite design requires exactly one worker; say so in a comment where serve.sh is touched, and don't "optimize" it). **Acceptance (Phase 3):** from OFF the tailnet (phone on cellular): the login page loads over https at the domain, a guest can log in, run flux_local, watch the log stream, see the dashboard; `mb` with `MB_HOST=https://modelbeast.digalot.fyi` + `MB_TOKEN` works end-to-end; unauthenticated API probes get 401; Settings 403 for guests. Record a short before/after in BENCHMARKS.md? No — this isn't a benchmark; instead append a "Public access" section to AGENTS.md and README. --- ## Build order & guardrails 1. Phase 1 backend → `tests/smoke.sh` additions → Phase 1 frontend → browser verify → commit/push. 2. Phase 2 endpoint (verify macmon output format on-machine first) → dashboard UI → verify → commit/push. 3. `docs/VPS.md` + app hardening checks → hand the VPS steps to the owner → after he runs them, do the off-tailnet acceptance with him → commit/push. Update `AGENTS.md` (tokens), `README.md` (auth + dashboard + public URL), and project memory at the end. Pitfalls from this codebase's history (respect them): - SQLite: additive `ALTER TABLE` migrations only, via the existing `MIGRATIONS` list; `check_same_thread=False` is load-bearing. - Frontend changes need `cd web && npm run build`; the server serves `web/dist`. - Restart via `./scripts/serve.sh`; manifest changes need a restart, `run.py` changes don't. - Don't break the inbox watcher, the lanes, or the operator contract; don't add Celery/Redis/Postgres — the current design is deliberate. - Secrets never in logs (the redaction layer exists — route any new logging through it). - If the user's other Claude session is mid-generation (check `./mb jobs` for running gpu jobs), wait for the lane to drain before restarting the server. **Scope guard:** no OAuth, no email, no signup flows, no billing, no live websocket dashboards, no Grafana. The whole point is: friends type a password at a nice URL, run models within the limits the owner set, and everyone can see what the box is doing on one page.