Regex/stdlib passes (dedupe, alias, ANSI/timestamp/blob strip, JSON minify, caveman word-drop) measured at 8.9-12.2x on real logs with o200k counts. Optional LLMLingua-2 wrap. grunts.py mines Claude Code transcripts for retyped prompts -> slash-command stubs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|---|---|---|
| .claude | ||
| .gitignore | ||
| grunts.py | ||
| lessismore.py | ||
| README.md | ||
| test_lessismore.py | ||
lessismore
Squeeze text before it hits an LLM. Deterministic passes first (free, safe, reproducible); optional perplexity pruning via LLMLingua-2 for the last mile.
python3 lessismore.py dump.log -l 2 > small.log # file in, file out
tail -5000 app.log | python3 lessismore.py -l 2 # pipe filter
python3 lessismore.py big.txt -l 2 --ml 0.5 # + LLMLingua (pip install llmlingua)
from lessismore import compress, count_tokens
small = compress(big_log, level=2)
Token stats print to stderr, so pipes stay clean. Measured at level 2 with the
real o200k tokenizer (no ML, no deps): a 2000-line mixed error log went
86,914 → 7,895 tokens (11x); an interleaved three-service log went
55,999 → 4,587 (12.2x); an ANSI-colored CI build log went 21,198 → 2,387
(8.9x); captured pip/tqdm output with \r redraws compressed 98%.
| Level | Passes | Safe for |
|---|---|---|
| 1 | whitespace collapse, duplicate-line runs | code structure & indentation (whitespace runs inside string literals still collapse) |
| 2 | + minify whole-input JSON, keep final state of \r progress redraws, strip ANSI + ISO timestamps, squash blobs/UUIDs/venv paths, alias scattered duplicate lines, collapse same-words-different-numbers runs |
logs, dumps, tool output |
| 3 | + filler-word stripping | prose only (eats "just" everywhere) |
| 4 | + two_sticks caveman word-dropping |
gist-only prose — never instructions (38% measured on chatty prose) |
--ml RATE adds LLMLingua-2 perplexity pruning after the deterministic passes
(RATE = fraction of tokens kept). Needs pip install llmlingua; pip install tiktoken for exact token counts (falls back to chars/4).
Design notes — where the wins actually are
-
Compress the bulk, not the question. A 30-token prompt isn't worth touching; the 50k-token log dump, retrieved doc, or file dump is. That's why this is a pipe filter, not a chat-prompt rewriter.
-
Characters ≠ tokens. Measured with o200k:
[fmt:md_tbl+hdr]= 8 tokens vs"Format the output as a markdown table with headers."= 10.[out:json_strict]vs"Respond only with valid JSON."= 6 vs 6. Tokenizers already compress common English; bracket-shorthand gets shredded into fragments and adds ambiguity. Shorthand codebooks: skipped, measured, dead. -
Stop-word stripping: instructions no, gist yes. Mangled grammar hurts instruction-following, so levels 1–3 never touch structure words. But every dropped word is a whole token, so for content you only need the gist of, level 4 (
two_sticks) drops articles/copulas/auxiliaries — measured 38% on chatty prose. Negations, modals, and order words are never dropped: "do not delete" must stay "not delete", never "delete".The two-sticks recoding idea (map words to 2–3 letter codes) is dead on arrival, measured: common words are already 1 token (
database= 1) while off-distribution codes cost 2 (qx= 2), plus a ~4-token/entry codebook tax. BPE already is a 200k-entry compression codebook; you can't beat it with an 18k-entry codebook written inside its own encoding. Zipping a zip.The salvage: codes DO pay when one code replaces a repeated multi-token sequence — that's
alias_repeats, dictionary compression that composes with BPE instead of fighting it (12.2x measured on interleaved logs). -
Perplexity pruning is a dependency, not a project. Microsoft's
llmlinguaalready does the budget-model math. We wrap it in six lines behind--mlinstead of reimplementing it. Deterministic passes run first so you're not paying a classifier to delete duplicate log lines. -
Determinism matters for prompt caching. Same input → same output keeps cache prefixes stable. Corollary: never compress a stable cached system prompt — you'd bust the cache to save tokens that were already free.
-
The compressor must cost less than it saves. Regex passes are ~free at any size. The ML pass loads a model, so it only pays off on multi-KB inputs.
-
The information-loss dial from the original idea is the
-llevel +--mlrate: level 1 is lossless-ish and code-safe,--ml 0.3is maximum squeeze for prose you only need the gist of. -
Personal shorthand belongs on the keyboard, not in the model. Your typed prompts are the cheapest tokens in the context (~10 each); a skill teaching the model your abbreviations costs more per turn than it can ever save. Expand shorthand client-side instead — Claude Code slash commands are the native mechanism, and
grunts.pymines your own transcript history for what you actually retype (--emitwrites the command stubs).
Fleet setup (stupendo / m3 air)
git clone ssh://git@100.71.119.27:222/monster/lessismore.git ~/Documents/lessismore
printf '#!/bin/sh\nexec python3 ~/Documents/lessismore/lessismore.py "$@"\n' | sudo tee /opt/homebrew/bin/lm >/dev/null
sudo chmod +x /opt/homebrew/bin/lm
lm is the pipe shim: docker logs dealgod | lm -l 2. Optional: pip3 install tiktoken for exact token stats (chars/4 heuristic otherwise).
Known tradeoffs
- Levels 2+ assume line order matters but wall-clock timing doesn't. When gaps and deltas are the signal (hang hunting), stay on level 1.
- Output is for LLM consumption, not round-tripping: markdown hard breaks (trailing double-space) and diff context lines don't survive even level 1.
- At levels 3–4, dedupe counts describe the post-stripped text — five differently-phrased "retry the job" lines can legitimately merge.
- In-band markers can collide with input that already contains them;
alias_repeatsbails out if its own@Nmarkers already appear as lines.