Meaning-First Compression
Keeps required evidence in the original wording—query-aware keep/drop, not a rewrite.
SuperCompress v2 is a query-aware compression layer that removes irrelevant context before inference. 64.1% less context · 24/24 evidence passes on coding-agent dumps — so you cut API spend without losing the lines that answer the ask.
Launch promo · 5M free tokens/mo · then $0.10/1M · no credit card
Coding-agent dumps · logs · diffs · original lines kept
const result = await compress({
context: conversation,
query: "What failed in the last deploy?"
})
// → Neural Keep · original evidence lines retained
64.1% mean context cut on the coding-agent benchmark. 16,647 tokens to 5,148. Required evidence held in 24/24 cases.
Required evidence kept on every coding-agent case. One case cut 96.6% at ≥99% retention.
Tool output, logs, diffs, stack traces. Keeps the evidence. Does not rewrite it into a summary.
BUILT FOR DEVELOPERS
Keeps required evidence in the original wording—query-aware keep/drop, not a rewrite.
Use our API or SDK in minutes. Works with OpenAI, Anthropic, Mistral, and more.
Works across 40+ coding-agent harnesses. Drop in before the model call.
HOW IT WORKS
Context arrives oversized — history, docs, and traces piled into one prompt.
NEURAL V2 · CODING-AGENT BENCHMARK
We measure evidence containment — not downstream LLM completion. Full methodology on /benchmarks.
Full methodology →Neural v2 B5 (coding-agent): 64.1% mean cut, 24/24 evidence passes, 16,647 → 5,148 tokens. Evidence containment — not downstream completion. Details on /benchmarks.
curl -X POST https://api.supercompress.dev/compress \
-H "Authorization: Bearer $SC_LIVE_..." \
-H "Content-Type: application/json" \
-d '{
"context": "$(cat conversation.txt)",
"query": "What failed in the last deploy?"
}'
// Response · Neural Keep (hosted)
{
"compressed": "...",
"original_tokens": 16647,
"compressed_tokens": 5148,
"tokens_saved_pct": 64.1,
"engine": "neural-keep-v2"
}
SuperCompress
Codex
Gemini CLI
Windsurf
VS Code
OpenCodeEvery coding agent. One compression layer.
CODING AGENT PLUGIN · MCP-FIRST
Auto-detect Cursor, Claude Code, Codex, OpenCode, Windsurf, and more. The MCP plugin compresses huge dumps before they burn tokens — keep your normal login.
Or the long name: npm install -g supercompress-proxy then supercompress setup.
npx -y supercompress-ai setupUSE CASES
From agentic coding to document-heavy work, meaning-first compression helps teams do more with less—without losing what matters.
Compress task history, repo context, tool traces, and diffs so agents reason farther inside the same context window—and spend less per turn.
Keep long chats useful and coherent. SuperCompress retains the decisions, constraints, and facts that drive better answers—not every filler turn.
Retrieved chunks often drown the query. Compress retrieved context so the model sees the evidence that matters for the current ask.
Ticket history, macros, and knowledge-base hits add up fast. Compress before generation to keep replies accurate without burning tokens.
Specs, tickets, research notes, and PRDs are dense. Compress for reviews and synthesis while retaining requirements and decisions.
Query-aware keep before every model call — shrink context, keep the evidence lines.
Live polling…
—
tokens processed
across SuperCompress · updates as people compress
Launch promo · 5M free tokens/mo · then $0.10/1M · no credit card.