diff --git a/README.md b/README.md index 7b0aca5..fcb2e89 100644 --- a/README.md +++ b/README.md @@ -1,12 +1,62 @@ -# lessismore +# lessismore πŸ“‰ -Squeeze text before it hits an LLM. Deterministic passes first (free, safe, -reproducible); optional perplexity pruning via LLMLingua-2 for the last mile. +**Turn 50,000 tokens of log spam into 4,500 tokens of pure signal β€” before it +ever hits your LLM.** + +lessismore is a deterministic compression filter for the bulk text that +actually fills context windows: logs, CI output, tracebacks, JSON dumps, +captured tool output. Pure stdlib Python. Zero dependencies. Every claim below +was measured against the real o200k tokenizer β€” and every idea that failed the +measurement is documented at the bottom, so you know exactly what you're +getting. + +```bash +tail -5000 app.log | python3 lessismore.py -l 2 | llm "why did this crash?" +``` + +**11x** on mixed error logs Β· **12.2x** on interleaved service logs Β· +**8.9x** on ANSI CI logs Β· **98%** on captured pip/tqdm output Β· +**~6 MB/s** single-core Β· **0** dependencies + +## Why this exists + +Three facts about LLM context, and the gap between them is this tool: + +1. **The expensive part isn't your prompt.** Your typed question is ~30 + tokens. The log dump you attach is 50,000. Compress the wall, not the + question. +2. **Machine output is mostly redundancy.** ANSI color codes, `\r` progress + redraws, timestamps, venv path spam, the same error line 400 times β€” noise + that looks small on a terminal screen but is real tokens in a capture. +3. **BPE tokenizers already compress English** (common words are 1 token β€” + you can't out-abbreviate them, we checked). What BPE *can't* see is + repetition across a file or terminal noise. That's the seam this tool + mines. + +Three payoffs when you pipe through it: + +- **Context capacity** β€” an hour of log history fits where two minutes did. + On a local model, that's the difference between full speed and crawling. +- **Prompt-cache longevity** β€” every pass is deterministic: same input, same + output, byte for byte. Follow-up questions re-hit the provider cache + instead of re-paying for the logs. An ML compressor in the loop would bust + the cache on every subtle variation; the regex passes never do. +- **Model attention** β€” LLMs lose things in the middle of walls of text. Feed + signal, not noise, and the first answer is the right one more often. + +## Quickstart + +```bash +git clone ssh://git@100.71.119.27:222/monster/lessismore.git +cd lessismore +python3 test_lessismore.py # 26 asserts, should print "ok" +pip3 install tiktoken # optional: exact token stats (else chars/4) +``` ```bash python3 lessismore.py dump.log -l 2 > small.log # file in, file out -tail -5000 app.log | python3 lessismore.py -l 2 # pipe filter -python3 lessismore.py big.txt -l 2 --ml 0.5 # + LLMLingua (pip install llmlingua) +docker logs dealgod 2>&1 | python3 lessismore.py -l 2 # pipe filter +python3 lessismore.py big.txt -l 2 --ml 0.5 # + LLMLingua last mile ``` ```python @@ -14,140 +64,93 @@ from lessismore import compress, count_tokens small = compress(big_log, level=2) ``` -Token stats print to stderr, so pipes stay clean. Measured at level 2 with the -real o200k tokenizer (no ML, no deps): a 2000-line mixed error log went -**86,914 β†’ 7,895 tokens (11x)**; an interleaved three-service log went -**55,999 β†’ 4,587 (12.2x)**; an ANSI-colored CI build log went **21,198 β†’ 2,387 -(8.9x)**; captured pip/tqdm output with `\r` redraws compressed **98%**. +Token stats print to stderr, so pipes stay clean. -| Level | Passes | Safe for | +## The dial + +Four levels, from byte-cautious to caveman. Pick by content, not by greed. + +| Level | What it eats | Point it at | Profile | +|---|---|---|---| +| **1** | whitespace runs, consecutive duplicate lines | code, scripts, anything | structure-safe: indentation and content untouched (whitespace *inside string literals* still collapses) | +| **2** | + JSON minify, `\r` redraw collapse, ANSI strip, ISO timestamps, base64/hex blobs, UUIDs, venv paths, scattered-duplicate aliasing, similar-line collapse | logs, CI dumps, traces, tool output | **the sweet spot** β€” destroys machine noise, keeps every distinct fact | +| **3** | + filler-phrase stripping ("could you please", "just") | prose, chat history | fine for text, never for strict logic | +| **4** | + `two_sticks` caveman mode: drops articles/copulas/auxiliaries, never negations or modals | gist-only prose, transcripts | lossy on style, protective of meaning β€” "do **not** delete" keeps its *not* | + +The crown jewel at level 2 is `alias_repeats`: ordinary dedupe only sees +*consecutive* repeats, so interleaved multi-service logs sail straight through +it. `alias_repeats` hunts scattered duplicates across the whole file and +dictionary-codes them (`@1 = ERROR [pool-3] psycopg2...` once in a legend, +2-token `@1` everywhere else). It's lossless β€” the legend keeps every line +verbatim β€” and it's the difference between 1.5x and 12x on real logs. + +## The receipts + +All reproducible: `python3 bench.py` (o200k counts via tiktoken). + +| Sample | Tokens | Why it wins | |---|---|---| -| 1 | whitespace collapse, duplicate-line runs | code structure & indentation (whitespace runs *inside string literals* still collapse) | -| 2 | + minify whole-input JSON, keep final state of `\r` progress redraws, strip ANSI + ISO timestamps, squash blobs/UUIDs/venv paths, alias scattered duplicate lines, collapse same-words-different-numbers runs | logs, dumps, tool output | -| 3 | + filler-word stripping | prose only (eats "just" everywhere) | -| 4 | + `two_sticks` caveman word-dropping | gist-only prose β€” never instructions (38% measured on chatty prose) | +| Mixed error log, 2000 lines | 86,914 β†’ 7,895 (**11.0x**) | timestamp strip unmasks identical lines β†’ dedupe; similar-line collapse catches the numbered stragglers | +| Interleaved 3-service log, zero consecutive repeats | 55,999 β†’ 4,587 (**12.2x**) | `alias_repeats` β€” plain dedupe managed ~0% on this input | +| ANSI-colored CI/docker build log | 21,198 β†’ 2,387 (**8.9x**) | color codes make identical lines look different; strip them and the log collapses | +| Captured pip/tqdm output | 14,483 β†’ 291 (**98%**) | every overwritten `\r` progress frame is invisible on screen but real tokens in a capture | +| pytest failure dump | 2,223 β†’ 1,673 (25%) | one site-packages path prefix = 25 tokens β†’ 7 | +| Pretty-printed JSON API response | 11,741 β†’ 6,833 (42%) | minify (round-trip verified lossless) + UUIDβ†’8-hex squash | +| Chatty prose, level 4 | 427 β†’ 264 (38%) | every dropped function word is a whole token; level 3 got 2% on the same text | -`--ml RATE` adds LLMLingua-2 perplexity pruning after the deterministic passes -(RATE = fraction of tokens kept). Needs `pip install llmlingua`; `pip install -tiktoken` for exact token counts (falls back to chars/4). +Throughput: measured **~6 MB/s single-core** at level 2 (28 MB log in 4.4s). +No model, no network β€” break-even input size is effectively zero. -## Design notes β€” where the wins actually are +## What we refused to build (measured so it stays dead) -1. **Compress the bulk, not the question.** A 30-token prompt isn't worth - touching; the 50k-token log dump, retrieved doc, or file dump is. That's why - this is a pipe filter, not a chat-prompt rewriter. +The graveyard is a feature. Each of these looks clever on paper and loses +against a real tokenizer: -2. **Characters β‰  tokens.** Measured with o200k: `[fmt:md_tbl+hdr]` = 8 tokens - vs `"Format the output as a markdown table with headers."` = 10. - `[out:json_strict]` vs `"Respond only with valid JSON."` = 6 vs 6. - Tokenizers already compress common English; bracket-shorthand gets shredded - into fragments *and* adds ambiguity. Shorthand codebooks: skipped, measured, dead. +- **Shorthand codebooks** β€” `[fmt:md_tbl+hdr]` costs **8 tokens**; "Format + the output as a markdown table with headers." costs **10**. BPE already has + English baked in; bracket syntax shreds into off-distribution fragments. +- **Wordβ†’code recoding** ("database" β†’ `qx`) β€” common words are already 1 + token, random codes cost 2, plus ~4 tokens/entry of codebook tax. You + cannot beat a 200k-entry codebook from inside its own encoding. Zipping a + zip. +- **Known abbreviations** (db, fn, env, auth) β€” measured 1 β†’ 1 tokens. Zero. +- **Personalized shorthand skills** β€” scanned 2,917 real typed prompts across + 654 transcripts: 2,762 unique, and the repeats were already grunts ("yes", + "go"). A model-side decode skill costs ~1k tokens/turn to save ~10. + `grunts.py` keeps the half that works: mine your own history, emit + *client-side* slash-command stubs β€” expansion before the model sees it is + free. +- **Digit-masked dedupe** (0.0% β€” progress bars differ in glyphs, not + digits), **separator shortening** (a 78-char `----` is already 1 token), + **prefix hoisting** (one stray line kills it), **JSONβ†’TSV / float + truncation / pointer squash** (real savings, worse trade). -3. **Stop-word stripping: instructions no, gist yes.** Mangled grammar hurts - instruction-following, so levels 1–3 never touch structure words. But every - dropped word is a whole token, so for content you only need the gist of, - level 4 (`two_sticks`) drops articles/copulas/auxiliaries β€” measured 38% on - chatty prose. Negations, modals, and order words are never dropped: "do not - delete" must stay "not delete", never "delete". +## Battle-tested - The two-sticks *recoding* idea (map words to 2–3 letter codes) is dead on - arrival, measured: common words are already 1 token (` database` = 1) while - off-distribution codes cost 2 (` qx` = 2), plus a ~4-token/entry codebook - tax. BPE already is a 200k-entry compression codebook; you can't beat it - with an 18k-entry codebook written inside its own encoding. Zipping a zip. +An adversarial review agent was told to break it and confirmed 14 real +failure modes β€” negation-inverting filler stripping ("was **not just** the +db" β†’ "was **not** the db"), URLs eaten as base64, crashes on empty and +non-UTF-8 stdin, distinct SHA-256s falsely merging as "repeated", +`two_sticks` eating "IT" and "US" as function words. Nine fixed with +regression tests, four documented below, one wontfix (adversarial in-band +marker collision). `python3 test_lessismore.py` β€” 26 asserts, no framework. - The salvage: codes DO pay when one code replaces a repeated *multi-token - sequence* β€” that's `alias_repeats`, dictionary compression that composes - with BPE instead of fighting it (12.2x measured on interleaved logs). +## Where it does nothing (on purpose) -4. **Perplexity pruning is a dependency, not a project.** Microsoft's - `llmlingua` already does the budget-model math. We wrap it in six lines - behind `--ml` instead of reimplementing it. Deterministic passes run first - so you're not paying a classifier to delete duplicate log lines. - -5. **Determinism matters for prompt caching.** Same input β†’ same output keeps - cache prefixes stable. Corollary: never compress a stable cached system - prompt β€” you'd bust the cache to save tokens that were already free. - -6. **The compressor must cost less than it saves.** Regex passes are ~free at - any size. The ML pass loads a model, so it only pays off on multi-KB inputs. - -7. **The information-loss dial** from the original idea is the `-l` level + - `--ml` rate: level 1 is lossless-ish and code-safe, `--ml 0.3` is maximum - squeeze for prose you only need the gist of. - -8. **Personal shorthand belongs on the keyboard, not in the model.** Your typed - prompts are the cheapest tokens in the context (~10 each); a skill teaching - the model your abbreviations costs more per turn than it can ever save. - Expand shorthand client-side instead β€” Claude Code slash commands are the - native mechanism, and `grunts.py` mines your own transcript history for - what you actually retype (`--emit` writes the command stubs). - -## How it was tested, what won, what died - -Development was measurement-first: every pass had to beat the real o200k -tokenizer (tiktoken) on a realistic sample before it earned its lines, and an -adversarial review then hunted for inputs that corrupt meaning. `python3 -bench.py` reproduces the headline numbers; `python3 test_lessismore.py` runs -the regression suite (26 asserts, stdlib only, covers every fixed bug). - -### Wins (real o200k counts) - -| Sample | Tokens | Why it works | -|---|---|---| -| Mixed error log, 2000 lines (85% one repeated ERROR, 15% INFO differing only in numbers) | 86,914 β†’ 7,895 (**11.0x**) | timestamp strip β†’ dedupe; similar-line collapse catches the numbered INFO lines | -| Interleaved 3-service log, zero consecutive repeats | 55,999 β†’ 4,587 (**12.2x**) | `alias_repeats` dictionary-codes scattered duplicates consecutive dedupe can't touch (alone: 36,524 β†’ 8,795 where plain dedupe managed ~0%) | -| ANSI-colored CI/docker build log | 21,198 β†’ 2,387 (**8.9x**) | ANSI strip (colors make identical lines differ) + similar-line collapse | -| Captured pip/tqdm output with `\r` redraws | 14,483 β†’ 291 (**98%**) | every overwritten progress frame is invisible in a terminal but real tokens in a capture | -| pytest failure dump | 2,223 β†’ 1,673 (25%) | venv path spam: one site-packages prefix = 25 tokens β†’ 7 | -| Pretty-printed JSON API response | 11,741 β†’ 6,833 (42%) | minify (lossless) + UUIDβ†’8-hex squash | -| Chatty prose, gist mode (level 4) | 427 β†’ 264 (38%) | every dropped function word is a whole token; level 3 managed only 2% on the same text | - -### Dead ends (measured so they stay dead) - -- **Shorthand codebooks**: `[fmt:md_tbl+hdr]` = 8 tokens vs the plain-English - sentence = 10. BPE already compresses common English; brackets shred. -- **Wordβ†’2-3-letter-code recoding**: common words are already 1 token, codes - cost 2, plus ~4 tokens/entry codebook tax. Can't beat a 200k-entry codebook - from inside its own encoding. -- **Known abbreviations** (db, fn, env, auth): 1 β†’ 1 tokens. Zero. -- **Personalized shorthand skill** (SwiftKey-style): scanned 2,917 real typed - prompts across 654 transcripts β€” 2,762 were unique, and the repeats were - already grunts ("yes", "go", "continue"). Model-side decode skills cost - ~1k tokens/turn to save ~10. `grunts.py` keeps the useful half: mine your - history, emit client-side slash-command stubs (free). -- **Digit-masked template dedupe**: 0.0% β€” real progress bars differ in bar - glyphs, not digits; the alpha-skeleton key in `dedupe_similar` is what works. -- **Separator-line shortening**: a 78-char `----` rule is already 1 token. -- **Common-prefix hoisting**: 0.0% β€” one stray line kills the common prefix. -- **JSONβ†’TSV, float truncation, hex-pointer squash**: real savings, but each - paid in corruption risk or lines-of-code for marginal gains over minify. - -### Adversarial review: 14 confirmed breaks β†’ 9 fixed, 4 documented, 1 wontfix - -Worst finds, all reproduced then fixed with regression tests: filler stripping -inverted negations ("was **not just** the database" β†’ "was **not** the -database"); URLs eaten as base64 blobs (`/` was in the character class); -crashes on empty and non-UTF-8 stdin; two different SHA-256s squashing to the -same stub and falsely merging as "repeated"; `two_sticks` eating "IT" and "US" -as function words; dedupe markers longer than the short lines they replaced. -The four survivors are documented under Known tradeoffs below. - -### Overall efficacy β€” the honest version - -Savings are proportional to **redundancy**, not size. This tool removes -repetition and machine noise; it cannot compress information. +Savings are proportional to **redundancy, not size**. This removes repetition +and machine noise; it cannot compress information, and doesn't pretend to. | Content | Expect | Verdict | |---|---|---| -| Repetitive machine output (logs, CI, installs, captured progress) | 5–12x, up to 50x on pathological repeats | the sweet spot β€” use level 2 by default | -| Structured data (JSON, tracebacks) | 25–40% | worthwhile, lossless where it fires | -| Varied prose | ~2% (level 3) / 38% lossy (level 4) | only worth it in gist mode | -| Clean code, unique dense text | ~0% by design | don't bother β€” that's what `--ml` or truncation is for | +| Repetitive machine output | 5–12x, up to 50x on pathological repeats | the reason this exists | +| Structured data (JSON, tracebacks) | 25–42% | worthwhile, lossless where it fires | +| Varied prose | ~2% (level 3) / 38% lossy (level 4) | gist mode only | +| Clean code, unique dense text | ~0% by design | that's what `--ml` or truncation is for | -Cost side: pure regex/stdlib, measured ~6 MB/s single-core at level 2 (a 28 MB -log in 4.4s) β€” no model, no network, no deps. The break-even input size is -effectively zero; it just never *wins* on non-redundant text. +`--ml RATE` bolts on Microsoft's LLMLingua-2 for perplexity pruning as a +last mile (`pip install llmlingua`) β€” runs *after* the deterministic passes so +you're not paying a classifier model to delete duplicate log lines. Only +worth it on multi-KB inputs, and it forfeits the cache-stability guarantee. ## Fleet setup (stupendo / m3 air) @@ -157,8 +160,7 @@ printf '#!/bin/sh\nexec python3 ~/Documents/lessismore/lessismore.py "$@"\n' | s sudo chmod +x /opt/homebrew/bin/lm ``` -`lm` is the pipe shim: `docker logs dealgod | lm -l 2`. Optional: `pip3 install -tiktoken` for exact token stats (chars/4 heuristic otherwise). +`lm` is the pipe shim: `docker logs dealgod | lm -l 2`. ## Known tradeoffs