# lessismore πŸ“‰ **Turn 50,000 tokens of log spam into 4,500 tokens of pure signal β€” before it ever hits your LLM.** lessismore is a deterministic compression filter for the bulk text that actually fills context windows: logs, CI output, tracebacks, JSON dumps, captured tool output. Pure stdlib Python. Zero dependencies. Every claim below was measured against the real o200k tokenizer β€” and every idea that failed the measurement is documented at the bottom, so you know exactly what you're getting. ```bash tail -5000 app.log | python3 lessismore.py -l 2 | llm "why did this crash?" ``` **11x** on mixed error logs Β· **12.2x** on interleaved service logs Β· **8.9x** on ANSI CI logs Β· **98%** on captured pip/tqdm output Β· **~6 MB/s** single-core Β· **0** dependencies ## Why this exists Three facts about LLM context, and the gap between them is this tool: 1. **The expensive part isn't your prompt.** Your typed question is ~30 tokens. The log dump you attach is 50,000. Compress the wall, not the question. 2. **Machine output is mostly redundancy.** ANSI color codes, `\r` progress redraws, timestamps, venv path spam, the same error line 400 times β€” noise that looks small on a terminal screen but is real tokens in a capture. 3. **BPE tokenizers already compress English** (common words are 1 token β€” you can't out-abbreviate them, we checked). What BPE *can't* see is repetition across a file or terminal noise. That's the seam this tool mines. Three payoffs when you pipe through it: - **Context capacity** β€” an hour of log history fits where two minutes did. On a local model, that's the difference between full speed and crawling. - **Prompt-cache longevity** β€” every pass is deterministic: same input, same output, byte for byte. Follow-up questions re-hit the provider cache instead of re-paying for the logs. An ML compressor in the loop would bust the cache on every subtle variation; the regex passes never do. - **Model attention** β€” LLMs lose things in the middle of walls of text. Feed signal, not noise, and the first answer is the right one more often. ## Quickstart ```bash git clone ssh://git@100.71.119.27:222/monster/lessismore.git cd lessismore python3 test_lessismore.py # 26 asserts, should print "ok" pip3 install tiktoken # optional: exact token stats (else chars/4) ``` ```bash python3 lessismore.py dump.log -l 2 > small.log # file in, file out docker logs dealgod 2>&1 | python3 lessismore.py -l 2 # pipe filter python3 lessismore.py big.txt -l 2 --ml 0.5 # + LLMLingua last mile ``` ```python from lessismore import compress, count_tokens small = compress(big_log, level=2) ``` Token stats print to stderr, so pipes stay clean. ## The dial Four levels, from byte-cautious to caveman. Pick by content, not by greed. | Level | What it eats | Point it at | Profile | |---|---|---|---| | **1** | whitespace runs, consecutive duplicate lines | code, scripts, anything | structure-safe: indentation and content untouched (whitespace *inside string literals* still collapses) | | **2** | + JSON minify, `\r` redraw collapse, ANSI strip, ISO timestamps, base64/hex blobs, UUIDs, venv paths, scattered-duplicate aliasing, similar-line collapse | logs, CI dumps, traces, tool output | **the sweet spot** β€” destroys machine noise, keeps every distinct fact | | **3** | + filler-phrase stripping ("could you please", "just") | prose, chat history | fine for text, never for strict logic | | **4** | + `two_sticks` caveman mode: drops articles/copulas/auxiliaries, never negations or modals | gist-only prose, transcripts | lossy on style, protective of meaning β€” "do **not** delete" keeps its *not* | The crown jewel at level 2 is `alias_repeats`: ordinary dedupe only sees *consecutive* repeats, so interleaved multi-service logs sail straight through it. `alias_repeats` hunts scattered duplicates across the whole file and dictionary-codes them (`@1 = ERROR [pool-3] psycopg2...` once in a legend, 2-token `@1` everywhere else). It's lossless β€” the legend keeps every line verbatim β€” and it's the difference between 1.5x and 12x on real logs. ## The receipts All reproducible: `python3 bench.py` (o200k counts via tiktoken). | Sample | Tokens | Why it wins | |---|---|---| | Mixed error log, 2000 lines | 86,914 β†’ 7,895 (**11.0x**) | timestamp strip unmasks identical lines β†’ dedupe; similar-line collapse catches the numbered stragglers | | Interleaved 3-service log, zero consecutive repeats | 55,999 β†’ 4,587 (**12.2x**) | `alias_repeats` β€” plain dedupe managed ~0% on this input | | ANSI-colored CI/docker build log | 21,198 β†’ 2,387 (**8.9x**) | color codes make identical lines look different; strip them and the log collapses | | Captured pip/tqdm output | 14,483 β†’ 291 (**98%**) | every overwritten `\r` progress frame is invisible on screen but real tokens in a capture | | pytest failure dump | 2,223 β†’ 1,673 (25%) | one site-packages path prefix = 25 tokens β†’ 7 | | Pretty-printed JSON API response | 11,741 β†’ 6,833 (42%) | minify (round-trip verified lossless) + UUIDβ†’8-hex squash | | Chatty prose, level 4 | 427 β†’ 264 (38%) | every dropped function word is a whole token; level 3 got 2% on the same text | Throughput: measured **~6 MB/s single-core** at level 2 (28 MB log in 4.4s). No model, no network β€” break-even input size is effectively zero. ## What we refused to build (measured so it stays dead) The graveyard is a feature. Each of these looks clever on paper and loses against a real tokenizer: - **Shorthand codebooks** β€” `[fmt:md_tbl+hdr]` costs **8 tokens**; "Format the output as a markdown table with headers." costs **10**. BPE already has English baked in; bracket syntax shreds into off-distribution fragments. - **Wordβ†’code recoding** ("database" β†’ `qx`) β€” common words are already 1 token, random codes cost 2, plus ~4 tokens/entry of codebook tax. You cannot beat a 200k-entry codebook from inside its own encoding. Zipping a zip. - **Known abbreviations** (db, fn, env, auth) β€” measured 1 β†’ 1 tokens. Zero. - **Personalized shorthand skills** β€” scanned 2,917 real typed prompts across 654 transcripts: 2,762 unique, and the repeats were already grunts ("yes", "go"). A model-side decode skill costs ~1k tokens/turn to save ~10. `grunts.py` keeps the half that works: mine your own history, emit *client-side* slash-command stubs β€” expansion before the model sees it is free. - **Digit-masked dedupe** (0.0% β€” progress bars differ in glyphs, not digits), **separator shortening** (a 78-char `----` is already 1 token), **prefix hoisting** (one stray line kills it), **JSONβ†’TSV / float truncation / pointer squash** (real savings, worse trade). ## Battle-tested An adversarial review agent was told to break it and confirmed 14 real failure modes β€” negation-inverting filler stripping ("was **not just** the db" β†’ "was **not** the db"), URLs eaten as base64, crashes on empty and non-UTF-8 stdin, distinct SHA-256s falsely merging as "repeated", `two_sticks` eating "IT" and "US" as function words. Nine fixed with regression tests, four documented below, one wontfix (adversarial in-band marker collision). `python3 test_lessismore.py` β€” 26 asserts, no framework. ## Where it does nothing (on purpose) Savings are proportional to **redundancy, not size**. This removes repetition and machine noise; it cannot compress information, and doesn't pretend to. | Content | Expect | Verdict | |---|---|---| | Repetitive machine output | 5–12x, up to 50x on pathological repeats | the reason this exists | | Structured data (JSON, tracebacks) | 25–42% | worthwhile, lossless where it fires | | Varied prose | ~2% (level 3) / 38% lossy (level 4) | gist mode only | | Clean code, unique dense text | ~0% by design | that's what `--ml` or truncation is for | `--ml RATE` bolts on Microsoft's LLMLingua-2 for perplexity pruning as a last mile (`pip install llmlingua`) β€” runs *after* the deterministic passes so you're not paying a classifier model to delete duplicate log lines. Only worth it on multi-KB inputs, and it forfeits the cache-stability guarantee. ## Fleet setup (stupendo / m3 air) ```bash git clone ssh://git@100.71.119.27:222/monster/lessismore.git ~/Documents/lessismore printf '#!/bin/sh\nexec python3 ~/Documents/lessismore/lessismore.py "$@"\n' | sudo tee /opt/homebrew/bin/lm >/dev/null sudo chmod +x /opt/homebrew/bin/lm ``` `lm` is the pipe shim: `docker logs dealgod | lm -l 2`. ## Known tradeoffs - Levels 2+ assume line *order* matters but wall-clock timing doesn't. When gaps and deltas are the signal (hang hunting), stay on level 1. - Output is for LLM consumption, not round-tripping: markdown hard breaks (trailing double-space) and diff context lines don't survive even level 1. - At levels 3–4, dedupe counts describe the post-stripped text β€” five differently-phrased "retry the job" lines can legitimately merge. - In-band markers can collide with input that already contains them; `alias_repeats` bails out if its own `@N` markers already appear as lines.