alias_templates is now two tiers with exact char accounting:
- tier A masks hex/digit runs INSIDE tokens (run-level constants — dates,
IP prefixes, the 2 in ssh2 — inline into the template); groups are
re-costed under whole-token value encoding and take the cheaper
- tier B masks whole digit-bearing tokens on what tier A left, so paths/ids
varying in letters still group (compile/build logs)
- dedupe_similar skeleton is hex-aware but keeps a leading @tN verbatim:
refs from different templates can no longer merge as "similar" — an
earlier build scored 8.9x on LogHub from exactly that silent value
destruction, and v0.3.0's HDFS number had the same phantom; both gone
Measured (bench_real.py): LogHub 3.0x -> 3.9x, LogDx-CI 2.2x -> 2.3x at
unchanged 95.2% critical-signal retention, LogChunks 1.9x -> 2.1x.
Equal-budget receipt added to bench_real logdx: at a 2k cap,
compress-then-truncate retains 46.6% of critical signals vs 24.7% for
plain truncation. Synthetic eval re-run on shipped code: 8/8 at 6.4x.
Graveyard: Drain-style letter-merging measured flat (3.0x -> 3.0x, Linux
regressed) and documented dead.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Benchmarking on real corpora (LogHub) exposed that synthetic logs flatter
the compressor: real logs repeat templates, not exact lines, and level 2
managed only 1.5x. Changes, each measured:
- alias_templates: lines differing only in digit-bearing tokens become one
legend template (@t1 = Failed password for root from <*> port <*> ssh2)
plus per-line values; constant digit tokens (ssh2) inline into the
template. Deterministic, single-pass, in-band bail-out, same payoff bar
as alias_repeats. LogHub corpus: 1.8x -> 3.0x.
- four field-observed timestamp formats (HDFS yymmdd hhmmss, BGL RAS
stamps, epoch+date pairs, HealthApp) — LogHub: 1.5x -> 1.8x
- blob squash: long alpha runs (16+ letters) exempt the -/_ branch — a
camelCase test name was being eaten as base64; caught by LogDx-CI,
regression-tested
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- dedupe_similar: summarize varying digits in the marker (values 404/429/503
or min-max) — distinct status codes/ports no longer vanish silently
- minify_json: also fires on pretty JSON embedded in other text (the
captured-tool-output shape), via raw_decode at line-breaking braces
- squash_blobs: admit base64url (- and _) so real JWTs squash, guarded by
digits+mixed-case so kebab/snake identifiers survive
- strip_timestamps: syslog and nginx CLF formats join ISO-8601
- level 2 is now the default (the pitch, the examples, and the sweet spot
all said 2; level 1 stays the explicit code-safe opt-down)
- lm --run CMD: compress a command's output, keep its exit code
- lm --hook: zero-dep Claude Code PostToolUse hook (reads hook JSON on
stdin, emits updatedToolOutput; <2k-char outputs pass through)
- tests: 42 asserts incl. level-2 idempotency invariant, value-summary
regression, embedded JSON, JWT/slug boundary, --run exit passthrough
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- LICENSE (MIT) and pyproject.toml: pip install . gives lessismore + lm
commands; [tokens] and [ml] extras
- budget(): middle-out hard cap, the only guaranteed-bounded pass; --budget N
runs after compression as the backstop; char-slicing fallback for single
giant lines
- serve(): stdlib paste-in demo page, localhost only; lm --serve
- Scrubbed internal hostnames/IPs from samples and README; benchmarks
re-measured after scrub (interleaved now 54,999 -> 4,585, 12.0x; dedupe
alone 1.6x on same input)
- 30-assert suite, green on system python and tiktoken venv
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>