September 29, 2026
SuperCompress v2 is live
We built a 400M-parameter engine that looks at all of your context, figures out what’s relevant to answering the query, and cuts the rest.
Where it excels
The hard case is a coding-agent turn: a pile of tool output, logs, diffs, and stack traces, and one line that actually answers the question. Truncation often deletes that line. A summary rewrites the error, the ID, and the path.
v2 reads the whole context against the current query, keeps the original lines that matter, and drops the rest. The query itself is never compressed. Same idea for RAG chunks and long support threads, when most of the text is not about this ask.
Coding-agent benchmark
- Cut context by 64.1% on average
- Preserved the required evidence in 24/24 cases
- Took 16,647 tokens → 5,148
- Reached 96.6% reduction on an individual case at ≥99% evidence retention
Launch promo on the hosted API: 5 million tokens a month free, then $0.10 per million. Self-hosting stays free under the MIT license.
How to use it
SuperCompress v2 is open source and available now through an API, and through an agent plugin that works with 40+ harnesses.