- bench.py now generates every README table row (ANSI CI, pip/tqdm, pytest, pretty JSON added) and refuses to print chars/4 estimates — reproduced numbers can no longer silently diverge from the claims - eval.py: the receipt ratios can't give — 8 seeded failure logs with planted root causes and red herrings, asked raw vs compressed, graded by deterministic keyword check. First run (haiku 4.5): raw 6/8, compressed 8/8 at 3.1x fewer tokens; both raw misses were lost-in-the-middle - README: eval section, related-work section (rtk, Headroom, Drain3, LLMLingua-2, arXiv 2604.13066), prompt-cache claim sharpened to the rebuilt-prompt scenario, agent-harness integration (--run, --hook), re-measured receipts table Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
16 KiB
lessismore 📉
Turn 50,000 tokens of log spam into 4,500 tokens of pure signal — before it ever hits your LLM.
lessismore is a deterministic compression filter for the bulk text that actually fills context windows: logs, CI output, tracebacks, JSON dumps, captured tool output. Pure stdlib Python. Zero dependencies. Every claim below was measured against the real o200k tokenizer — and every idea that failed the measurement is documented at the bottom, so you know exactly what you're getting.
tail -5000 app.log | python3 lessismore.py | llm "why did this crash?"
11x on mixed error logs · 12x on interleaved service logs · 11.8x on ANSI CI logs · 97% on captured pip/tqdm output · ~6–12 MB/s single-core · 0 dependencies
And the part a compression ratio can't prove: in the root-cause eval (eval.py), the model found the planted root cause in 8/8 compressed logs vs 6/8 raw — smaller and more often right.
Why this exists
Three facts about LLM context, and the gap between them is this tool:
- The expensive part isn't your prompt. Your typed question is ~30 tokens. The log dump you attach is 50,000. Compress the wall, not the question.
- Machine output is mostly redundancy. ANSI color codes,
\rprogress redraws, timestamps, venv path spam, the same error line 400 times — noise that looks small on a terminal screen but is real tokens in a capture. - BPE tokenizers already compress English (common words are 1 token — you can't out-abbreviate them, we checked). What BPE can't see is repetition across a file or terminal noise. That's the seam this tool mines.
Three payoffs when you pipe through it:
- Context capacity — an hour of log history fits where two minutes did. On a local model, that's the difference between full speed and crawling.
- Prompt-cache longevity — every pass is deterministic and idempotent: same input, same output, byte for byte (it's a test invariant). This matters wherever a prompt is re-generated from source each call — a CI assistant re-reading the same log, RAG context rebuilt per request, a hook re-running on every tool call. A deterministic pass produces the identical prefix every time, so the provider cache hits; an ML compressor in the loop busts it on every subtle variation. (Within a single chat, follow-ups hit the cache regardless — the transcript is append-only. The claim is about rebuilt prompts, not chat turns.)
- Model attention — LLMs lose things in the middle of walls of text. Feed signal, not noise, and the first answer is the right one more often.
Quickstart
pip install . # from a clone — installs the `lessismore` and `lm` commands
pip install '.[tokens]' # + tiktoken for exact token stats (else chars/4 estimate)
python3 test_lessismore.py # should print "ok"
lm dump.log > small.log # file in, file out (level 2 is the default)
docker logs myapp 2>&1 | lm # pipe filter
lm --run "pytest -x" # run a noisy command, print it compressed,
# keep its exit code (pipes can't)
lm big.txt --budget 4000 # compress, then hard-cap at ~4k tokens
lm big.txt --ml 0.5 # + LLMLingua last mile
lm --serve # paste-in demo page on localhost:7777
No install needed either — python3 lessismore.py works the same from a bare
clone; it's one stdlib-only file.
from lessismore import compress, count_tokens, budget
small = budget(compress(big_log, level=2), 8000)
Token stats print to stderr, so pipes stay clean.
The dial
Four levels, from byte-cautious to caveman. Pick by content, not by greed.
| Level | What it eats | Point it at | Profile |
|---|---|---|---|
| 1 | whitespace runs, consecutive duplicate lines | code, scripts, anything | structure-safe: indentation and content untouched (whitespace inside string literals still collapses) |
| 2 | + JSON minify (whole-input and embedded pretty blocks), \r redraw collapse, ANSI strip, ISO/syslog/nginx timestamps, base64/base64url/hex blobs, UUIDs, venv paths, scattered-duplicate aliasing, similar-line collapse with value summaries |
logs, CI dumps, traces, tool output | the sweet spot and the default — destroys machine noise, keeps every distinct fact |
| 3 | + filler-phrase stripping ("could you please", "just") | prose, chat history | fine for text, never for strict logic |
| 4 | + two_sticks caveman mode: drops articles/copulas/auxiliaries, never negations or modals |
gist-only prose, transcripts | lossy on style, protective of meaning — "do not delete" keeps its not |
"Keeps every distinct fact" is enforced, not hoped: when similar lines
collapse, the digits that varied are summarized in the marker —
[3 similar lines omitted; values 404/429/503] — because sometimes the
digits (status codes, ports, exit codes) are the diagnosis. An earlier
version silently ate them; the adversarial eval below is what caught it.
The crown jewel at level 2 is alias_repeats: ordinary dedupe only sees
consecutive repeats, so interleaved multi-service logs sail straight through
it. alias_repeats hunts scattered duplicates across the whole file and
dictionary-codes them (@1 = ERROR [pool-3] psycopg2... once in a legend,
2-token @1 everywhere else). It's lossless — the legend keeps every line
verbatim — and on the interleaved benchmark it's the difference between
54,999 → 34,999 (dedupe alone, 1.6x) and 54,999 → 4,585 (12x).
--budget N is the backstop, not the compressor: after the passes run, it
keeps head and tail lines and drops the middle with a [~N tokens omitted]
marker. It's the only pass with guaranteed bounded output — use it when the
context limit is a hard wall.
The receipts
Every row reproducible from one command: python3 bench.py (o200k counts via
tiktoken — it refuses to print estimate numbers, so what you measure is what
this table claims).
| Sample | Tokens | Why it wins |
|---|---|---|
| Mixed error log, 2000 lines | 86,914 → 7,895 (11.0x) | timestamp strip unmasks identical lines → dedupe; similar-line collapse catches the numbered stragglers |
| Interleaved 3-service log, zero consecutive repeats | 54,999 → 4,585 (12.0x) | alias_repeats — consecutive dedupe alone managed 1.6x on this input |
| ANSI-colored CI log | 10,440 → 888 (11.8x) | color codes make identical lines look different; strip them and the log collapses |
| Captured pip/tqdm output | 10,465 → 265 (97%) | every overwritten \r progress frame is invisible on screen but real tokens in a capture |
| pytest failure dump | 732 → 515 (30%) | one site-packages path prefix = 25 tokens → 7 |
| Pretty-printed JSON API response | 8,932 → 3,234 (64%) | minify (round-trip verified lossless) + UUID→8-hex squash |
| Chatty prose, level 4 | 427 → 264 (38%) | every dropped function word is a whole token; level 3 got 2% on the same text |
Throughput: measured 6–12 MB/s single-core at level 2 (20 MB log in 1.7s). No model, no network — break-even input size is effectively zero.
Does the model still get the answer?
A compression ratio proves the log got smaller, not that the signal survived. So this is the eval that matters, and it's in the repo: eval.py generates eight seeded failure logs, each with one planted root cause buried in machine noise plus a red herring (e.g. a cascade of connection errors caused by an OOM kill 400 lines earlier), asks a model "what's the root cause?" on the raw and the compressed version, and grades the answer with a deterministic keyword check.
Result (claude-haiku-4.5, 2026-07-07, one run; this run went through the
Claude Code agent harness with the same single-turn prompts — eval.py's
default runner is a logged-in claude -p. Rerun it yourself):
| Scenario | Raw tokens | Compressed | Raw | Compressed |
|---|---|---|---|---|
| OOM kill buried mid-log, connection-error red herring | 38,777 | 157 (247x) | ✗ blamed the red herring | ✓ found the OOM |
| 429 rate-limit burst hidden in a wall of status codes | 46,899 | 6,056 | ✗ "service overwhelmed" | ✓ named the 429s |
| Disk-full scattered through 3 interleaved services | 38,616 | 12,487 | ✓ | ✓ |
| One failed assertion in 1,400 ANSI test lines | 46,812 | 21,563 | ✓ | ✓ |
| Dependency conflict under progress-bar walls | 16,304 | 446 | ✓ | ✓ |
| Expired TLS cert in an nginx access-log wall | 56,696 | 34,474 | ✓ | ✓ |
| Missing env var in a crash-loop traceback | 7,680 | 1,610 | ✓ | ✓ |
| Java deadlock in a 40-thread dump | 8,157 | 8,041 (1.0x) | ✓ | ✓ |
| Total | 259,941 | 84,834 (3.1x) | 6/8 | 8/8 |
The two raw failures are the lost-in-the-middle effect this tool exists to
counter: the model latched onto the loud symptom cascade and never surfaced
the quiet cause. Compressed, the cause is impossible to miss. The honest
caveats: n=1 run per condition, synthetic logs authored by this project, one
model — which is exactly why the harness ships in the repo with fixed seeds.
python3 eval.py --dry costs nothing and shows the scenarios; point
LM_EVAL_CMD at any prompt-on-stdin CLI to grade your own model.
Where the tokens actually are: agent harnesses
A human remembering to pipe is the demo. The 50k-token walls of 2026 enter context windows as tool output inside agent harnesses, so lessismore ships two integrations:
lm --run — run the noisy command through the compressor and keep its
exit code (a shell pipe can't do that without pipefail games):
lm --run "pytest -x" # agent sees 500 tokens, not 20,000
lm --run "docker build ."
Tell your agent about it once (CLAUDE.md: "run noisy commands through
lm --run") and every test run gets cheaper.
lm --hook — a Claude Code PostToolUse hook that compresses Bash
output before it enters the model's context. Zero dependencies, no jq;
outputs under 2,000 chars pass through untouched. In
.claude/settings.json:
{
"hooks": {
"PostToolUse": [
{
"matcher": "Bash",
"hooks": [{ "type": "command", "command": "lm --hook" }]
}
]
}
}
Because the passes are deterministic, hook-compressed transcripts stay byte-stable across re-runs — the prompt-cache argument above is strongest exactly here.
What we refused to build (measured so it stays dead)
The graveyard is a feature. Each of these looks clever on paper and loses against a real tokenizer:
- Shorthand codebooks —
[fmt:md_tbl+hdr]costs 8 tokens; "Format the output as a markdown table with headers." costs 10. BPE already has English baked in; bracket syntax shreds into off-distribution fragments. - Word→code recoding ("database" →
qx) — common words are already 1 token, random codes cost 2, plus ~4 tokens/entry of codebook tax. You cannot beat a 200k-entry codebook from inside its own encoding. Zipping a zip. - Known abbreviations (db, fn, env, auth) — measured 1 → 1 tokens. Zero.
- Personalized shorthand skills — scanned 2,917 real typed prompts across
654 transcripts: 2,762 unique, and the repeats were already grunts ("yes",
"go"). A model-side decode skill costs ~1k tokens/turn to save ~10.
grunts.pykeeps the half that works: mine your own history, emit client-side slash-command stubs — expansion before the model sees it is free. - Digit-masked dedupe (0.0% — progress bars differ in glyphs, not
digits), separator shortening (a 78-char
----is already 1 token), prefix hoisting (one stray line kills it), JSON→TSV / float truncation / pointer squash (real savings, worse trade).
Battle-tested
An adversarial review agent was told to break it and confirmed 14 real
failure modes — negation-inverting filler stripping ("was not just the
db" → "was not the db"), URLs eaten as base64, crashes on empty and
non-UTF-8 stdin, distinct SHA-256s falsely merging as "repeated",
two_sticks eating "IT" and "US" as function words. Nine fixed with
regression tests, four documented below, one wontfix (adversarial in-band
marker collision). A second review round found the big one: similar-line
collapse silently eating distinct status codes — fixed with value summaries
and now guarded by the downstream eval. python3 test_lessismore.py —
42 asserts, no framework.
Where it does nothing (on purpose)
Savings are proportional to redundancy, not size. This removes repetition and machine noise; it cannot compress information, and doesn't pretend to.
| Content | Expect | Verdict |
|---|---|---|
| Repetitive machine output | 5–12x, up to 50x on pathological repeats | the reason this exists |
| Structured data (JSON, tracebacks) | 30–64% | worthwhile, lossless where it fires |
| Varied prose | ~2% (level 3) / 38% lossy (level 4) | gist mode only |
| Clean code, unique dense text | ~0% by design | that's what --ml or truncation is for |
--ml RATE bolts on Microsoft's LLMLingua-2 for perplexity pruning as a
last mile (pip install llmlingua) — runs after the deterministic passes so
you're not paying a classifier model to delete duplicate log lines. Only
worth it on multi-KB inputs, and it forfeits the cache-stability guarantee.
Related work (know the river you're panning)
Honest placement, so you can pick the right tool:
- rtk — Rust proxy with per-command filters (ANSI strip, dedupe, failures-only test output), hooks into agent CLIs. Command-aware and aggressive where lessismore is generic text passes and lossless-leaning: rtk decides what you need to see; lessismore removes only what is provably redundant and keeps a legend.
- Headroom — the heavyweight tool-output compressor (Python, proxy + MCP). Has ML models in the loop, so it needs a cache-alignment component to patch the non-determinism lessismore doesn't have.
- Drain3 — industrial log template
mining.
alias_repeats+dedupe_similarare a single-pass, zero-dep approximation of it; if you need streaming template state across files, use Drain3. - LLMLingua-2 — ML token
pruning; wrapped here as the optional
--mllast mile, after the deterministic passes, so you never pay a classifier to delete duplicate log lines. - Dictionary-encoding prompt compression
(2026) — independently validates the
@1 = linelegend technique academically (≥0.99 fidelity on log benchmarks); no code released. lessismore is, in effect, a reference implementation.
What none of them package together — and the reason this exists — is the combination: pure-stdlib, deterministic end to end, tokenizer-measured, with the cache-stability argument as a design constraint rather than a patch.
Try it in a browser
lm --serve runs a paste-in demo page on http://localhost:7777 — paste your
ugliest log, pick a level, watch the token count drop. Stdlib only, binds
localhost only.
License
MIT — see LICENSE.
Known tradeoffs
- Levels 2+ assume line order matters but wall-clock timing doesn't. When gaps and deltas are the signal (hang hunting), stay on level 1.
- Output is for LLM consumption, not round-tripping: markdown hard breaks (trailing double-space) and diff context lines don't survive even level 1.
- At levels 3–4, dedupe counts describe the post-stripped text — five differently-phrased "retry the job" lines can legitimately merge.
- In-band markers can collide with input that already contains them;
alias_repeatsbails out if its own@Nmarkers already appear as lines. - Blob squashing eats base64url runs (real JWTs) only when they look like blobs (digits + mixed case); a 64-char kebab-case or snake_case identifier is treated as content and survives.
- Similar-line collapse summarizes varying digits per column
(
values 404/429/503, or a min–max range past 4 distinct); if the digit runs per line differ in count, it falls back to a plain count marker.