# MODELBEAST — ollama operators brief (Opus handover) **From:** the asset-campaign session (2026-07-13). **Decision (John + Fable):** MODELBEAST is THE universal queue for all local model work on the tailnet — no second queue system. Direct Ollama calls (:11434) stay fine for quick interactive hits, but BULK LLM/vision sweeps must ride the job queue so they don't fight Klein/TRELLIS for Metal. ## The job: two new operators (mirror an existing operator's manifest+run.py pattern) 1. **`ollama_vision`** [category: llm] — params: `{model: "gemma3:4b", prompt, temperature=0}`, accepts: image (the job's asset), produces: text (a .json/.txt asset with the response). Implementation = one POST to localhost:11434/api/generate with the asset base64'd into `images`. Resource lane: gpu (that's the whole point — arbitration with image/mesh gen). 2. **`ollama_llm`** [category: llm] — same but text-only (`/api/chat`), no input asset, `{model: "qwen3:8b|qwen3:32b|qwen3:235b", prompt|messages}`. Notes: - Keep operators thin — Ollama does the work; the operator is just queue admission + asset IO. - A bulk sweep is then N jobs (one per image) — the lane serialises them gracefully and `/api/jobs` gives progress/retry/cancel for free. No new batch machinery. - pgvector embeddings: NOT now (deferred by design). When un-deferred it's one more thin operator (`ollama_embed` → /api/embed); DB writes stay client-side. - Test: submit 3 ollama_vision jobs while a flux_local job runs; assert FIFO within the gpu lane and all 4 complete.