lessismore/README.md
2026-07-07 13:52:17 +10:00

9.9 KiB
Raw Blame History

lessismore

Squeeze text before it hits an LLM. Deterministic passes first (free, safe, reproducible); optional perplexity pruning via LLMLingua-2 for the last mile.

python3 lessismore.py dump.log -l 2 > small.log     # file in, file out
tail -5000 app.log | python3 lessismore.py -l 2      # pipe filter
python3 lessismore.py big.txt -l 2 --ml 0.5          # + LLMLingua (pip install llmlingua)
from lessismore import compress, count_tokens
small = compress(big_log, level=2)

Token stats print to stderr, so pipes stay clean. Measured at level 2 with the real o200k tokenizer (no ML, no deps): a 2000-line mixed error log went 86,914 → 7,895 tokens (11x); an interleaved three-service log went 55,999 → 4,587 (12.2x); an ANSI-colored CI build log went 21,198 → 2,387 (8.9x); captured pip/tqdm output with \r redraws compressed 98%.

Level Passes Safe for
1 whitespace collapse, duplicate-line runs code structure & indentation (whitespace runs inside string literals still collapse)
2 + minify whole-input JSON, keep final state of \r progress redraws, strip ANSI + ISO timestamps, squash blobs/UUIDs/venv paths, alias scattered duplicate lines, collapse same-words-different-numbers runs logs, dumps, tool output
3 + filler-word stripping prose only (eats "just" everywhere)
4 + two_sticks caveman word-dropping gist-only prose — never instructions (38% measured on chatty prose)

--ml RATE adds LLMLingua-2 perplexity pruning after the deterministic passes (RATE = fraction of tokens kept). Needs pip install llmlingua; pip install tiktoken for exact token counts (falls back to chars/4).

Design notes — where the wins actually are

  1. Compress the bulk, not the question. A 30-token prompt isn't worth touching; the 50k-token log dump, retrieved doc, or file dump is. That's why this is a pipe filter, not a chat-prompt rewriter.

  2. Characters ≠ tokens. Measured with o200k: [fmt:md_tbl+hdr] = 8 tokens vs "Format the output as a markdown table with headers." = 10. [out:json_strict] vs "Respond only with valid JSON." = 6 vs 6. Tokenizers already compress common English; bracket-shorthand gets shredded into fragments and adds ambiguity. Shorthand codebooks: skipped, measured, dead.

  3. Stop-word stripping: instructions no, gist yes. Mangled grammar hurts instruction-following, so levels 13 never touch structure words. But every dropped word is a whole token, so for content you only need the gist of, level 4 (two_sticks) drops articles/copulas/auxiliaries — measured 38% on chatty prose. Negations, modals, and order words are never dropped: "do not delete" must stay "not delete", never "delete".

    The two-sticks recoding idea (map words to 23 letter codes) is dead on arrival, measured: common words are already 1 token ( database = 1) while off-distribution codes cost 2 ( qx = 2), plus a ~4-token/entry codebook tax. BPE already is a 200k-entry compression codebook; you can't beat it with an 18k-entry codebook written inside its own encoding. Zipping a zip.

    The salvage: codes DO pay when one code replaces a repeated multi-token sequence — that's alias_repeats, dictionary compression that composes with BPE instead of fighting it (12.2x measured on interleaved logs).

  4. Perplexity pruning is a dependency, not a project. Microsoft's llmlingua already does the budget-model math. We wrap it in six lines behind --ml instead of reimplementing it. Deterministic passes run first so you're not paying a classifier to delete duplicate log lines.

  5. Determinism matters for prompt caching. Same input → same output keeps cache prefixes stable. Corollary: never compress a stable cached system prompt — you'd bust the cache to save tokens that were already free.

  6. The compressor must cost less than it saves. Regex passes are ~free at any size. The ML pass loads a model, so it only pays off on multi-KB inputs.

  7. The information-loss dial from the original idea is the -l level + --ml rate: level 1 is lossless-ish and code-safe, --ml 0.3 is maximum squeeze for prose you only need the gist of.

  8. Personal shorthand belongs on the keyboard, not in the model. Your typed prompts are the cheapest tokens in the context (~10 each); a skill teaching the model your abbreviations costs more per turn than it can ever save. Expand shorthand client-side instead — Claude Code slash commands are the native mechanism, and grunts.py mines your own transcript history for what you actually retype (--emit writes the command stubs).

How it was tested, what won, what died

Development was measurement-first: every pass had to beat the real o200k tokenizer (tiktoken) on a realistic sample before it earned its lines, and an adversarial review then hunted for inputs that corrupt meaning. python3 bench.py reproduces the headline numbers; python3 test_lessismore.py runs the regression suite (26 asserts, stdlib only, covers every fixed bug).

Wins (real o200k counts)

Sample Tokens Why it works
Mixed error log, 2000 lines (85% one repeated ERROR, 15% INFO differing only in numbers) 86,914 → 7,895 (11.0x) timestamp strip → dedupe; similar-line collapse catches the numbered INFO lines
Interleaved 3-service log, zero consecutive repeats 55,999 → 4,587 (12.2x) alias_repeats dictionary-codes scattered duplicates consecutive dedupe can't touch (alone: 36,524 → 8,795 where plain dedupe managed ~0%)
ANSI-colored CI/docker build log 21,198 → 2,387 (8.9x) ANSI strip (colors make identical lines differ) + similar-line collapse
Captured pip/tqdm output with \r redraws 14,483 → 291 (98%) every overwritten progress frame is invisible in a terminal but real tokens in a capture
pytest failure dump 2,223 → 1,673 (25%) venv path spam: one site-packages prefix = 25 tokens → 7
Pretty-printed JSON API response 11,741 → 6,833 (42%) minify (lossless) + UUID→8-hex squash
Chatty prose, gist mode (level 4) 427 → 264 (38%) every dropped function word is a whole token; level 3 managed only 2% on the same text

Dead ends (measured so they stay dead)

  • Shorthand codebooks: [fmt:md_tbl+hdr] = 8 tokens vs the plain-English sentence = 10. BPE already compresses common English; brackets shred.
  • Word→2-3-letter-code recoding: common words are already 1 token, codes cost 2, plus ~4 tokens/entry codebook tax. Can't beat a 200k-entry codebook from inside its own encoding.
  • Known abbreviations (db, fn, env, auth): 1 → 1 tokens. Zero.
  • Personalized shorthand skill (SwiftKey-style): scanned 2,917 real typed prompts across 654 transcripts — 2,762 were unique, and the repeats were already grunts ("yes", "go", "continue"). Model-side decode skills cost ~1k tokens/turn to save ~10. grunts.py keeps the useful half: mine your history, emit client-side slash-command stubs (free).
  • Digit-masked template dedupe: 0.0% — real progress bars differ in bar glyphs, not digits; the alpha-skeleton key in dedupe_similar is what works.
  • Separator-line shortening: a 78-char ---- rule is already 1 token.
  • Common-prefix hoisting: 0.0% — one stray line kills the common prefix.
  • JSON→TSV, float truncation, hex-pointer squash: real savings, but each paid in corruption risk or lines-of-code for marginal gains over minify.

Adversarial review: 14 confirmed breaks → 9 fixed, 4 documented, 1 wontfix

Worst finds, all reproduced then fixed with regression tests: filler stripping inverted negations ("was not just the database" → "was not the database"); URLs eaten as base64 blobs (/ was in the character class); crashes on empty and non-UTF-8 stdin; two different SHA-256s squashing to the same stub and falsely merging as "repeated"; two_sticks eating "IT" and "US" as function words; dedupe markers longer than the short lines they replaced. The four survivors are documented under Known tradeoffs below.

Overall efficacy — the honest version

Savings are proportional to redundancy, not size. This tool removes repetition and machine noise; it cannot compress information.

Content Expect Verdict
Repetitive machine output (logs, CI, installs, captured progress) 512x, up to 50x on pathological repeats the sweet spot — use level 2 by default
Structured data (JSON, tracebacks) 2540% worthwhile, lossless where it fires
Varied prose ~2% (level 3) / 38% lossy (level 4) only worth it in gist mode
Clean code, unique dense text ~0% by design don't bother — that's what --ml or truncation is for

Cost side: pure regex/stdlib, measured ~6 MB/s single-core at level 2 (a 28 MB log in 4.4s) — no model, no network, no deps. The break-even input size is effectively zero; it just never wins on non-redundant text.

Fleet setup (stupendo / m3 air)

git clone ssh://git@100.71.119.27:222/monster/lessismore.git ~/Documents/lessismore
printf '#!/bin/sh\nexec python3 ~/Documents/lessismore/lessismore.py "$@"\n' | sudo tee /opt/homebrew/bin/lm >/dev/null
sudo chmod +x /opt/homebrew/bin/lm

lm is the pipe shim: docker logs dealgod | lm -l 2. Optional: pip3 install tiktoken for exact token stats (chars/4 heuristic otherwise).

Known tradeoffs

  • Levels 2+ assume line order matters but wall-clock timing doesn't. When gaps and deltas are the signal (hang hunting), stay on level 1.
  • Output is for LLM consumption, not round-tripping: markdown hard breaks (trailing double-space) and diff context lines don't survive even level 1.
  • At levels 34, dedupe counts describe the post-stripped text — five differently-phrased "retry the job" lines can legitimately merge.
  • In-band markers can collide with input that already contains them; alias_repeats bails out if its own @N markers already appear as lines.