~150M query-aware cross-encoder. Hosted GPU inference. Highest-quality context selection. These launch numbers.
Cut your API costs by 65%.
Prove what stayed.
SuperCompress v2 is a ~150M-parameter Neural Keep engine on the hosted API (GPU). We measure whether required evidence survives compression — not whether a downstream chat model finishes the task. Downstream LLM eval has not been run yet.
Two engines — do not mix numbers
Lightweight local path. CPU. Millisecond-class preprocessing. Legacy results live in the labeled section below — separate product.
We score containment of required lines. Prefer “preserves required evidence” over “same answers” until downstream eval ships.
Coding-agent suite (B5) — headline
Held-out coding-agent dumps. Same scorer for every method. Tokenizer: tiktoken cl100k_base.
| Method | Mean cut | Evidence pass | Tokens in → out | p50 latency |
|---|---|---|---|---|
| SuperCompress v2 (Neural Keep) | 64.1% | 24 / 24 | 16,647 → 5,148 | ~5.8 s |
| Headroom 0.36.4 | 14.8% | 18 / 24 | 16,647 → 14,009 | ~0.3 s |
| LLMLingua-2 | 52.7% | 12 / 24 | 16,647 → 8,030 | ~0.7 s |
| Truncation (keep last ~35%) | 64.3% | 11 / 24 | 16,647 → 5,851 | <1 ms |
| SuperCompress v1 (compiler) | 60.7% | 21 / 24 | 16,647 → 5,918 | ~10 ms |
Aggregate honesty (all 390 cases)
Mean cut across the full 390-case mix is ~3.6%. That is intentional, not a failure. On needle / LongBench / multihop slices the engine often keeps nearly everything when every line could matter — it knows when not to compress. The 64.1% figure is the coding-agent (B5) suite, where dumps are noisy and query-aware keep shines.
Max reduction at ≥99% evidence retention: 96.6%. Aggregate mean quality retention (containment scorer): ~85%. Downstream LLM completion eval: not run.
How we measured
- Quality metric: answer / needle / evidence containment in the compressed text (original wording).
- Not measured yet: downstream LLM task success after compression.
- v2 inference: query-conditioned cross-encoder (product: distill-v2 ~150M on Fly).
- Comparators: Headroom 0.36.4, LLMLingua-2, truncation, and SuperCompress v1 compiler on B5.
Limitations
- Latency: Neural v2 B5 p50 is ~5.8s on the hosted GPU path — not millisecond-class.
- Downstream eval: evidence containment ≠ proven answer quality from a chat model.
- Workload fit: strongest on coding-agent / tool-dump noise; conservative on dense needle/QA mixes.
Compiler / v1 results (not Neural v2)
These numbers are from the lightweight compiler path (CPU, millisecond-class). They are not SuperCompress v2 Neural Keep results. Kept here for history and for teams still on the local compiler.
| Legacy metric | Result | Notes |
|---|---|---|
| Held-out gold containment | 99.4% (180/181) | Compiler suites real / fresh4 / fresh5 · Aug 2026 |
| Mean cut (compiler) | ~58–66% | Maximize drop with ≥98% answer-keep gate |
| Fixed 35% budget oracle recall | 100% | Public 8-seed suite vs ~25% truncation |
| Typical latency | ~10–47 ms CPU | Local compiler — not the hosted Neural Keep path |
FAQ
Why isn’t aggregate mean cut ~64%?
64.1% is the coding-agent (B5) suite. Across all 390 cases many slices are already dense; v2 refuses to over-cut when evidence is everywhere, so the aggregate mean sits near ~3.6%.
Does “24/24” mean same answers from GPT/Claude?
No. It means required evidence lines were present after compression. Downstream completion quality is a separate eval we have not published yet.
GPU or CPU?
Hosted Neural v2 → GPU, seconds-scale on long dumps. Compiler → CPU, milliseconds. Pick the path that matches the product you ship.