SuperCompress
Benchmarks Agents Blog Changelog Docs Get API key Log in Playground GitHub
SuperCompress on Product Hunt

Cut your API Costs by 65%.

SuperCompress v2 is a query-aware compression layer that removes irrelevant context before inference. 64.1% less context · 24/24 evidence passes on coding-agent dumps — so you cut API spend without losing the lines that answer the ask.

Launch promo · 5M free tokens/mo · then $0.10/1M · no credit card

SuperCompress launch video preview

Coding-agent dumps · logs · diffs · original lines kept

Consumer AI apps send chat history, retrieved docs, tool traces, and user context on every request. That context drives LLM API cost, latency, and context-window pressure before the answer is generated. SuperCompress runs before inference, selecting the context most relevant to the current user request. Use it for chatbots, AI search, support agents, copilots, RAG, and any feature where context grows with users.

EXAMPLE: LONG CONVERSATION20,814 TOKENS
const result = await compress({
  context: conversation,
  query: "What failed in the last deploy?"
})
// → Neural Keep · original evidence lines retained
NEURAL V2 · HOSTED 0%+

Mean context cut

64.1% mean context cut on the coding-agent benchmark. 16,647 tokens to 5,148. Required evidence held in 24/24 cases.

GATE 24/24

Evidence held

Required evidence kept on every coding-agent case. One case cut 96.6% at ≥99% retention.

BEST AT agents

Coding dumps

Tool output, logs, diffs, stack traces. Keeps the evidence. Does not rewrite it into a summary.

BUILT FOR DEVELOPERS

Control AI Feature Cost
Before the Model Call

☼

Meaning-First Compression

Keeps required evidence in the original wording—query-aware keep/drop, not a rewrite.

⌗

Drop-In Simple

Use our API or SDK in minutes. Works with OpenAI, Anthropic, Mistral, and more.

♢

Built for agents

Works across 40+ coding-agent harnesses. Drop in before the model call.

HOW IT WORKS

Compress before
the model call.

  1. 01 Ingest Chat, RAG hits, tool traces
  2. 02 Score Relevance vs current ask
  3. 03 Keep Meaning-first retention
  4. 04 Ship Fewer tokens to inference

Context arrives oversized — history, docs, and traces piled into one prompt.

NEURAL V2 · CODING-AGENT BENCHMARK

Cut hard.
Keep required evidence.

We measure evidence containment — not downstream LLM completion. Full methodology on /benchmarks.

Full methodology →
METHODMEAN TOKEN CUT
(HIGHER IS BETTER)
EVIDENCE PASS
(B5 · 24 CASES)
HOW IT WORKS
LLMLingua-2 52.7% 12/24 Token pruning · can drop evidence
Truncation 64.3% 11/24 Head/tail only · blind to the ask
Headroom 0.36.4 14.8% 18/24 Type / tool-output heuristics

Neural v2 B5 (coding-agent): 64.1% mean cut, 24/24 evidence passes, 16,647 → 5,148 tokens. Evidence containment — not downstream completion. Details on /benchmarks.

curl -X POST https://api.supercompress.dev/compress \
  -H "Authorization: Bearer $SC_LIVE_..." \
  -H "Content-Type: application/json" \
  -d '{
    "context": "$(cat conversation.txt)",
    "query": "What failed in the last deploy?"
  }'
// Response · Neural Keep (hosted)
{
  "compressed": "...",
  "original_tokens": 16647,
  "compressed_tokens": 5148,
  "tokens_saved_pct": 64.1,
  "engine": "neural-keep-v2"
}
SuperCompress SuperCompress

Every coding agent. One compression layer.

CODING AGENT PLUGIN · MCP-FIRST

One Install. Every Agent.
Fewer Tokens.

Auto-detect Cursor, Claude Code, Codex, OpenCode, Windsurf, and more. The MCP plugin compresses huge dumps before they burn tokens — keep your normal login.

Or the long name: npm install -g supercompress-proxy then supercompress setup.

$npx -y supercompress-ai setup
  • ✓ Works across 40+ coding-agent harnesses
  • ✓ MCP plugin installed
  • ✓ Works with login — no API-key mode

USE CASES

Use Cases That
Compound

From agentic coding to document-heavy work, meaning-first compression helps teams do more with less—without losing what matters.

Coding Agents

Compress task history, repo context, tool traces, and diffs so agents reason farther inside the same context window—and spend less per turn.

  • Shrink multi-file diffs before the model call
  • Keep plans + decisions, drop stale tool noise
  • Works as an MCP plugin across Cursor, Claude Code, Codex

Long Conversations

Keep long chats useful and coherent. SuperCompress retains the decisions, constraints, and facts that drive better answers—not every filler turn.

  • Lower cost on multi-hour support or sales threads
  • Preserve user preferences and prior commitments
  • Reduce latency as history grows

RAG & Search

Retrieved chunks often drown the query. Compress retrieved context so the model sees the evidence that matters for the current ask.

  • Cut redundant passages across top-k results
  • Keep citations and key claims intact
  • Fit more evidence into the same budget

Support Copilots

Ticket history, macros, and knowledge-base hits add up fast. Compress before generation to keep replies accurate without burning tokens.

  • Prioritize current issue + account state
  • Drop repeated boilerplate from prior tickets
  • Ship faster replies at lower per-ticket cost

Document Workflows

Specs, tickets, research notes, and PRDs are dense. Compress for reviews and synthesis while retaining requirements and decisions.

  • Faster design and compliance reviews
  • Keep must-have constraints and open questions
  • Pair with human-in-the-loop editing

Production Inference

Query-aware keep before every model call — shrink context, keep the evidence lines.

  • Compress context before the provider call
  • 64.1% mean B5 cut · 24/24 evidence containment
  • Drop-in before OpenAI, Anthropic, and more

Live polling…

—

tokens processed

across SuperCompress · updates as people compress

SuperCompress

Ship AI features with lower per-request cost.

Launch promo · 5M free tokens/mo · then $0.10/1M · no credit card.