v2 launch · Sep 2026 · 390 public cases · evidence containment

Cut your API costs by 65%.
Prove what stayed.

SuperCompress v2 is a ~150M-parameter Neural Keep engine on the hosted API (GPU). We measure whether required evidence survives compression — not whether a downstream chat model finishes the task. Downstream LLM eval has not been run yet.

Public cases 390 Launch suite across needles, LongBench slices, multihop, coding agents, scaling.
B5 mean cut 64.1% Coding-agent dumps: 16,647 → 5,148 tokens. Max 96.6% at ≥99% retention.
Evidence held 24/24 Required evidence lines kept on every coding-agent case.
B5 latency p50 ~5.8s Hosted GPU Neural Keep path — not the millisecond compiler path.

Two engines — do not mix numbers

Recommended · hosted Neural v2

~150M query-aware cross-encoder. Hosted GPU inference. Highest-quality context selection. These launch numbers.

Fast / local Compiler

Lightweight local path. CPU. Millisecond-class preprocessing. Legacy results live in the labeled section below — separate product.

Metric honesty Evidence, not answers

We score containment of required lines. Prefer “preserves required evidence” over “same answers” until downstream eval ships.

Coding-agent suite (B5) — headline

Held-out coding-agent dumps. Same scorer for every method. Tokenizer: tiktoken cl100k_base.

Method Mean cut Evidence pass Tokens in → out p50 latency
SuperCompress v2 (Neural Keep) 64.1% 24 / 24 16,647 → 5,148 ~5.8 s
Headroom 0.36.4 14.8% 18 / 24 16,647 → 14,009 ~0.3 s
LLMLingua-2 52.7% 12 / 24 16,647 → 8,030 ~0.7 s
Truncation (keep last ~35%) 64.3% 11 / 24 16,647 → 5,851 <1 ms
SuperCompress v1 (compiler) 60.7% 21 / 24 16,647 → 5,918 ~10 ms

Raw JSON: launch-benchmark.json · writeup: v2 launch.

Aggregate honesty (all 390 cases)

Mean cut across the full 390-case mix is ~3.6%. That is intentional, not a failure. On needle / LongBench / multihop slices the engine often keeps nearly everything when every line could matter — it knows when not to compress. The 64.1% figure is the coding-agent (B5) suite, where dumps are noisy and query-aware keep shines.

Max reduction at ≥99% evidence retention: 96.6%. Aggregate mean quality retention (containment scorer): ~85%. Downstream LLM completion eval: not run.

How we measured

  • Quality metric: answer / needle / evidence containment in the compressed text (original wording).
  • Not measured yet: downstream LLM task success after compression.
  • v2 inference: query-conditioned cross-encoder (product: distill-v2 ~150M on Fly).
  • Comparators: Headroom 0.36.4, LLMLingua-2, truncation, and SuperCompress v1 compiler on B5.

Limitations

  • Latency: Neural v2 B5 p50 is ~5.8s on the hosted GPU path — not millisecond-class.
  • Downstream eval: evidence containment ≠ proven answer quality from a chat model.
  • Workload fit: strongest on coding-agent / tool-dump noise; conservative on dense needle/QA mixes.
Legacy · compiler engine benchmarks

Compiler / v1 results (not Neural v2)

These numbers are from the lightweight compiler path (CPU, millisecond-class). They are not SuperCompress v2 Neural Keep results. Kept here for history and for teams still on the local compiler.

Legacy metric Result Notes
Held-out gold containment 99.4% (180/181) Compiler suites real / fresh4 / fresh5 · Aug 2026
Mean cut (compiler) ~58–66% Maximize drop with ≥98% answer-keep gate
Fixed 35% budget oracle recall 100% Public 8-seed suite vs ~25% truncation
Typical latency ~10–47 ms CPU Local compiler — not the hosted Neural Keep path

Older JSON snapshots remain under /assets/data/*-benchmark-latest.json. Prefer launch-benchmark.json for v2.

FAQ

Why isn’t aggregate mean cut ~64%?

64.1% is the coding-agent (B5) suite. Across all 390 cases many slices are already dense; v2 refuses to over-cut when evidence is everywhere, so the aggregate mean sits near ~3.6%.

Does “24/24” mean same answers from GPT/Claude?

No. It means required evidence lines were present after compression. Downstream completion quality is a separate eval we have not published yet.

GPU or CPU?

Hosted Neural v2 → GPU, seconds-scale on long dumps. Compiler → CPU, milliseconds. Pick the path that matches the product you ship.

Ship with less context

5M free tokens/mo · then $0.10/1M · MIT

Get API key   v2 launch →   Arena →