lessismore/README.md
2026-07-07 13:52:17 +10:00

173 lines
9.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# lessismore
Squeeze text before it hits an LLM. Deterministic passes first (free, safe,
reproducible); optional perplexity pruning via LLMLingua-2 for the last mile.
```bash
python3 lessismore.py dump.log -l 2 > small.log # file in, file out
tail -5000 app.log | python3 lessismore.py -l 2 # pipe filter
python3 lessismore.py big.txt -l 2 --ml 0.5 # + LLMLingua (pip install llmlingua)
```
```python
from lessismore import compress, count_tokens
small = compress(big_log, level=2)
```
Token stats print to stderr, so pipes stay clean. Measured at level 2 with the
real o200k tokenizer (no ML, no deps): a 2000-line mixed error log went
**86,914 → 7,895 tokens (11x)**; an interleaved three-service log went
**55,999 → 4,587 (12.2x)**; an ANSI-colored CI build log went **21,198 → 2,387
(8.9x)**; captured pip/tqdm output with `\r` redraws compressed **98%**.
| Level | Passes | Safe for |
|---|---|---|
| 1 | whitespace collapse, duplicate-line runs | code structure & indentation (whitespace runs *inside string literals* still collapse) |
| 2 | + minify whole-input JSON, keep final state of `\r` progress redraws, strip ANSI + ISO timestamps, squash blobs/UUIDs/venv paths, alias scattered duplicate lines, collapse same-words-different-numbers runs | logs, dumps, tool output |
| 3 | + filler-word stripping | prose only (eats "just" everywhere) |
| 4 | + `two_sticks` caveman word-dropping | gist-only prose — never instructions (38% measured on chatty prose) |
`--ml RATE` adds LLMLingua-2 perplexity pruning after the deterministic passes
(RATE = fraction of tokens kept). Needs `pip install llmlingua`; `pip install
tiktoken` for exact token counts (falls back to chars/4).
## Design notes — where the wins actually are
1. **Compress the bulk, not the question.** A 30-token prompt isn't worth
touching; the 50k-token log dump, retrieved doc, or file dump is. That's why
this is a pipe filter, not a chat-prompt rewriter.
2. **Characters ≠ tokens.** Measured with o200k: `[fmt:md_tbl+hdr]` = 8 tokens
vs `"Format the output as a markdown table with headers."` = 10.
`[out:json_strict]` vs `"Respond only with valid JSON."` = 6 vs 6.
Tokenizers already compress common English; bracket-shorthand gets shredded
into fragments *and* adds ambiguity. Shorthand codebooks: skipped, measured, dead.
3. **Stop-word stripping: instructions no, gist yes.** Mangled grammar hurts
instruction-following, so levels 13 never touch structure words. But every
dropped word is a whole token, so for content you only need the gist of,
level 4 (`two_sticks`) drops articles/copulas/auxiliaries — measured 38% on
chatty prose. Negations, modals, and order words are never dropped: "do not
delete" must stay "not delete", never "delete".
The two-sticks *recoding* idea (map words to 23 letter codes) is dead on
arrival, measured: common words are already 1 token (` database` = 1) while
off-distribution codes cost 2 (` qx` = 2), plus a ~4-token/entry codebook
tax. BPE already is a 200k-entry compression codebook; you can't beat it
with an 18k-entry codebook written inside its own encoding. Zipping a zip.
The salvage: codes DO pay when one code replaces a repeated *multi-token
sequence* — that's `alias_repeats`, dictionary compression that composes
with BPE instead of fighting it (12.2x measured on interleaved logs).
4. **Perplexity pruning is a dependency, not a project.** Microsoft's
`llmlingua` already does the budget-model math. We wrap it in six lines
behind `--ml` instead of reimplementing it. Deterministic passes run first
so you're not paying a classifier to delete duplicate log lines.
5. **Determinism matters for prompt caching.** Same input → same output keeps
cache prefixes stable. Corollary: never compress a stable cached system
prompt — you'd bust the cache to save tokens that were already free.
6. **The compressor must cost less than it saves.** Regex passes are ~free at
any size. The ML pass loads a model, so it only pays off on multi-KB inputs.
7. **The information-loss dial** from the original idea is the `-l` level +
`--ml` rate: level 1 is lossless-ish and code-safe, `--ml 0.3` is maximum
squeeze for prose you only need the gist of.
8. **Personal shorthand belongs on the keyboard, not in the model.** Your typed
prompts are the cheapest tokens in the context (~10 each); a skill teaching
the model your abbreviations costs more per turn than it can ever save.
Expand shorthand client-side instead — Claude Code slash commands are the
native mechanism, and `grunts.py` mines your own transcript history for
what you actually retype (`--emit` writes the command stubs).
## How it was tested, what won, what died
Development was measurement-first: every pass had to beat the real o200k
tokenizer (tiktoken) on a realistic sample before it earned its lines, and an
adversarial review then hunted for inputs that corrupt meaning. `python3
bench.py` reproduces the headline numbers; `python3 test_lessismore.py` runs
the regression suite (26 asserts, stdlib only, covers every fixed bug).
### Wins (real o200k counts)
| Sample | Tokens | Why it works |
|---|---|---|
| Mixed error log, 2000 lines (85% one repeated ERROR, 15% INFO differing only in numbers) | 86,914 → 7,895 (**11.0x**) | timestamp strip → dedupe; similar-line collapse catches the numbered INFO lines |
| Interleaved 3-service log, zero consecutive repeats | 55,999 → 4,587 (**12.2x**) | `alias_repeats` dictionary-codes scattered duplicates consecutive dedupe can't touch (alone: 36,524 → 8,795 where plain dedupe managed ~0%) |
| ANSI-colored CI/docker build log | 21,198 → 2,387 (**8.9x**) | ANSI strip (colors make identical lines differ) + similar-line collapse |
| Captured pip/tqdm output with `\r` redraws | 14,483 → 291 (**98%**) | every overwritten progress frame is invisible in a terminal but real tokens in a capture |
| pytest failure dump | 2,223 → 1,673 (25%) | venv path spam: one site-packages prefix = 25 tokens → 7 |
| Pretty-printed JSON API response | 11,741 → 6,833 (42%) | minify (lossless) + UUID→8-hex squash |
| Chatty prose, gist mode (level 4) | 427 → 264 (38%) | every dropped function word is a whole token; level 3 managed only 2% on the same text |
### Dead ends (measured so they stay dead)
- **Shorthand codebooks**: `[fmt:md_tbl+hdr]` = 8 tokens vs the plain-English
sentence = 10. BPE already compresses common English; brackets shred.
- **Word→2-3-letter-code recoding**: common words are already 1 token, codes
cost 2, plus ~4 tokens/entry codebook tax. Can't beat a 200k-entry codebook
from inside its own encoding.
- **Known abbreviations** (db, fn, env, auth): 1 → 1 tokens. Zero.
- **Personalized shorthand skill** (SwiftKey-style): scanned 2,917 real typed
prompts across 654 transcripts — 2,762 were unique, and the repeats were
already grunts ("yes", "go", "continue"). Model-side decode skills cost
~1k tokens/turn to save ~10. `grunts.py` keeps the useful half: mine your
history, emit client-side slash-command stubs (free).
- **Digit-masked template dedupe**: 0.0% — real progress bars differ in bar
glyphs, not digits; the alpha-skeleton key in `dedupe_similar` is what works.
- **Separator-line shortening**: a 78-char `----` rule is already 1 token.
- **Common-prefix hoisting**: 0.0% — one stray line kills the common prefix.
- **JSON→TSV, float truncation, hex-pointer squash**: real savings, but each
paid in corruption risk or lines-of-code for marginal gains over minify.
### Adversarial review: 14 confirmed breaks → 9 fixed, 4 documented, 1 wontfix
Worst finds, all reproduced then fixed with regression tests: filler stripping
inverted negations ("was **not just** the database" → "was **not** the
database"); URLs eaten as base64 blobs (`/` was in the character class);
crashes on empty and non-UTF-8 stdin; two different SHA-256s squashing to the
same stub and falsely merging as "repeated"; `two_sticks` eating "IT" and "US"
as function words; dedupe markers longer than the short lines they replaced.
The four survivors are documented under Known tradeoffs below.
### Overall efficacy — the honest version
Savings are proportional to **redundancy**, not size. This tool removes
repetition and machine noise; it cannot compress information.
| Content | Expect | Verdict |
|---|---|---|
| Repetitive machine output (logs, CI, installs, captured progress) | 512x, up to 50x on pathological repeats | the sweet spot — use level 2 by default |
| Structured data (JSON, tracebacks) | 2540% | worthwhile, lossless where it fires |
| Varied prose | ~2% (level 3) / 38% lossy (level 4) | only worth it in gist mode |
| Clean code, unique dense text | ~0% by design | don't bother — that's what `--ml` or truncation is for |
Cost side: pure regex/stdlib, measured ~6 MB/s single-core at level 2 (a 28 MB
log in 4.4s) — no model, no network, no deps. The break-even input size is
effectively zero; it just never *wins* on non-redundant text.
## Fleet setup (stupendo / m3 air)
```bash
git clone ssh://git@100.71.119.27:222/monster/lessismore.git ~/Documents/lessismore
printf '#!/bin/sh\nexec python3 ~/Documents/lessismore/lessismore.py "$@"\n' | sudo tee /opt/homebrew/bin/lm >/dev/null
sudo chmod +x /opt/homebrew/bin/lm
```
`lm` is the pipe shim: `docker logs dealgod | lm -l 2`. Optional: `pip3 install
tiktoken` for exact token stats (chars/4 heuristic otherwise).
## Known tradeoffs
- Levels 2+ assume line *order* matters but wall-clock timing doesn't. When
gaps and deltas are the signal (hang hunting), stay on level 1.
- Output is for LLM consumption, not round-tripping: markdown hard breaks
(trailing double-space) and diff context lines don't survive even level 1.
- At levels 34, dedupe counts describe the post-stripped text — five
differently-phrased "retry the job" lines can legitimately merge.
- In-band markers can collide with input that already contains them;
`alias_repeats` bails out if its own `@N` markers already appear as lines.