discgod/README.md
type-two cc0eabe2e0 Add launchd agent for auto-start on the Ultra
discgod_server.py now runs as a launchd LaunchAgent (RunAtLoad + KeepAlive) on
:5057 — starts on login/boot, respawns on crash. Documented in the README.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-13 20:25:54 +10:00

8.9 KiB

discgod

A lean, all-in-one Discogs seller inventory tracker. Paste a Discogs API token, pick the sellers you care about (Aussie or anyone), and discgod pulls their full inventories from the official Discogs API on a schedule into the discogs_full Postgres DB — building seller profiles and tracking price changes and sales over time. No buttons, no DOM scraping.

It's the token-driven successor to pliceclogs / marketwatcher: same database, same proven mp_* schema and upsert/diff/sold-detection engine, but the data now comes from the API (robust, no fragile selectors, no Cloudflare) instead of scraping rendered pages.

┌─────────────────────────┐    Discogs API (your token, 60/min)
│  extension (background) │ ───────────────────────────────────►  api.discogs.com
│  • holds token          │ ◄───────────────────────────────────  /users/{u}/inventory
│  • schedules + paginates│
│  • rate-limits + queue  │    raw listing pages (HTTP POST)
└───────────┬─────────────┘ ─────────────►  ┌────────────────────┐
            │                                │  discgod_server.py │ ──►  discogs_full
   popup = control panel                     │  (maps + upserts)  │      (mp_* tables)
                                             └────────────────────┘
            optional, for 24/7 even with the browser closed:
            discgod_crawl.py  ──(same engine, server-side)──►   discogs_full

Layout

discgod/
  extension/            Chrome MV3 extension (load unpacked)
    manifest.json
    background.js        engine: token, scheduler, API pagination, rate-limit, retry queue
    popup.html/.css/.js  control panel: token, watchlist, scrape-now, live stats
    content.js           "▶ Track in discgod" button on Discogs seller pages
    icons/
  server/
    discgod_server.py    STANDALONE single-file ingest server (port 5057) —
                         schema + API→mp_* mapping + upsert/diff/gone + watchlist
                         + HTTP, all in one. Coexists with the legacy :5002 server;
                         paste its handlers into pliceclogs_server.py to merge later.
    discgod_crawl.py     headless crawler for cron / 24-7 (imports the engine
                         from discgod_server; no browser needed)
    schema.sql           reference (server auto-applies it on startup)
    requirements.txt     just psycopg2
  seeds/
    aus_sellers.txt      top 100 Australian sellers, ready to seed the watchlist
  deploy.sh              backup + deploy + restart the server on the tailscale box

Quick start

1. Server (writes Postgres)

The Ultra (ultra.local = tailscale 100.91.239.7) hosts discogs_full locally, so there's nothing to deploy — just run the standalone server on the Ultra itself. It listens on :5057 and coexists with the legacy plice server on :5002 (separate process, separate port, same DB, untouched legacy endpoints):

cd server
nohup /opt/homebrew/bin/python3 discgod_server.py > discgod_server.log 2>&1 &
curl -s localhost:5057/discgod/health   # {"ok":true,"server":"discgod",...}

It creates the mp_* tables and discgod_seller_watch in discogs_full on startup (additive — touches nothing existing). deploy.sh is only needed if you ever run the server on a different box.

To keep it running across reboots it's installed as a launchd agent on the Ultra (deploy/com.monsterrobot.discgod.plist~/Library/LaunchAgents/): starts on login/boot, restarts itself if it crashes, logs to server/discgod_server.log.

cp deploy/com.monsterrobot.discgod.plist ~/Library/LaunchAgents/
launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.monsterrobot.discgod.plist
# status / stop:
launchctl print  gui/$(id -u)/com.monsterrobot.discgod
launchctl bootout gui/$(id -u)/com.monsterrobot.discgod

(LaunchAgent = runs as the johnking user, needed for local Postgres peer auth. If you want it up even with nobody logged in, convert to a LaunchDaemon in /Library/LaunchDaemons with UserName johnking — needs sudo.)

2. Extension (the engine + control panel)

  1. chrome://extensions → enable Developer mode → Load unpacked → pick discgod/extension.
  2. Open the popup, paste your Discogs API token (Settings → DevelopersGenerate token), click Test (should say "signed in as …").
  3. Set the Ingest server URL (default http://100.91.239.7:5057) and an interval, tick Auto-crawl, Save.
  4. Add sellers — type a username (or paste a /seller/NAME URL) and Track, or hit the ▶ Track in discgod button while on any Discogs seller page. Add an optional filter like format=Vinyl or ships_from=Australia.
  5. Scrape all now, or wait for the schedule. Watch the live stats fill in.

That's the whole "add token and go" loop.

3. (Optional) 24/7 headless crawling on the box

The extension only runs while Chrome is open. For always-on tracking, run the crawler on the box — it shares the same watchlist:

ssh johnking@100.91.239.7
echo YOUR_TOKEN > ~/.discgod_token && chmod 600 ~/.discgod_token
cd /Users/johnking/Documents/MONSTERPANEL/discgod

python3 discgod_crawl.py --seed /tmp/aus_sellers.txt   # seed the 100 Aussie sellers
python3 discgod_crawl.py                                # crawl everything once
python3 discgod_crawl.py Dennis-brisbane-qld           # just one seller
python3 discgod_crawl.py --loop 43200                   # crawl all, repeat every 12 h

Or cron it (daily 3am):

0 3 * * *  cd /Users/johnking/Documents/MONSTERPANEL/discgod && /opt/homebrew/bin/python3 discgod_crawl.py >> crawl.log 2>&1

What gets captured

Table Contents
mp_listing One row per listing (item_id PK, release_id joins to your XML release table). Condition, sleeve, notes, price/currency, shipping, converted total, have/want, ships-from, seller. first_seen / last_seen / times_seen / status (active/gone).
mp_listing_event Append-only history: listed, price_change (old/new), condition_change, gone (presumed sold/delisted, with last price), relisted.
mp_seller Each seller seen — id, username, rating, ships-from.
mp_seller_snapshot Inventory-size time series (source='api_inventory'), from the API's authoritative pagination total. Deduped within 6 h per (seller, filter).
mp_page_log Provenance: every captured inventory page + the item_ids it held.
discgod_seller_watch The watchlist: who to crawl, optional per-seller filter, last-crawled status.

Sold detection: every full crawl stamps each listing with a crawl_id. When the crawl finishes, listings the crawl didn't see (last_crawl_id differs) are flipped to gone with a gone event — your sold/delisted feed. Filtered crawls (e.g. format=Vinyl) skip gone-marking unless they cover the seller's whole inventory.

Useful queries

-- A seller's inventory size over time
SELECT captured_at::date, max(total_listings)
FROM mp_seller_snapshot
WHERE username = 'Dennis-brisbane-qld' AND source = 'api_inventory'
GROUP BY 1 ORDER BY 1;

-- Presumed sales last 30 days, with full release info from the XML dump
SELECT e.recorded_at::date, e.seller_username, e.old_price, r.title, ra.artist_name
FROM mp_listing_event e
JOIN release r ON r.id = e.release_id
LEFT JOIN release_artist ra ON ra.release_id = r.id AND ra.position = 1
WHERE e.event = 'gone' AND e.recorded_at > now() - interval '30 days'
ORDER BY e.recorded_at DESC;

-- Price drops
SELECT e.recorded_at, e.seller_username, e.old_price, e.new_price, r.title
FROM mp_listing_event e JOIN release r ON r.id = e.release_id
WHERE e.event = 'price_change' AND e.new_price < e.old_price
ORDER BY e.recorded_at DESC LIMIT 50;

Notes & limits

  • Rate limit: the Discogs API allows ~60 authenticated requests/min. discgod paces itself (~55/min) and backs off on 429. A 26k-listing seller is ~270 API pages ≈ 5 min — fine periodically; the headless crawler is better for big bulk sweeps.
  • Public inventories only — the API returns another user's listings that are "For Sale". That's every real seller.
  • The extension's background worker can be suspended by Chrome on very long crawls; discgod persists the per-seller queue and a 1-minute keepalive so a sweep resumes where it left off. For unattended bulk work, prefer the headless crawler.
  • Off-box server? Set PGHOST/PGUSER/PGPASSWORD (and DISCGOD_DB) before launching discgod_server.py. On the box it uses local peer auth, no creds.