Benchmarks
These first-party results are reproducible from published evaluation manifests. We report bands rather than best runs and exclude degraded-validity runs from anchors. Scorecard date: 2026-08-14.
Headline
| Metric | Value | Scope |
|---|---|---|
| Context reduction | 99.507% | Terminal-Bench v2: 160,509 → 791 tokens per query |
| Payload gold recall | 99.7% | LME v2-Small, n = 630 evidence recall |
| Reader-token savings | 8.4× | vs full-context replay, 20.2k reader tokens/query (LoCoMo arm) |
| LoCoMo accuracy | 67.9–69.2% | full 10 conversations, 1,489 questions |
| LongMemEval v2-Small | 48.1% | 451 questions, fixed reader + judge |
| Agentic SWE-Atlas | 9/9 | 10-checkpoint probe; N ≤ 9, indicative only |
How to read these
- First-party evidence standard. These results use our harness and runs, the same evidence class as vendor self-reports. Depending on protocol, academic LoCoMo reproductions place Mem0 at 63–66.9 and Zep at 63.8–71.2; vendor pages self-report higher. The published harness keeps the comparison falsifiable.
- Bands, not bests. The LoCoMo result is a band across three full runs. We discarded a higher single run (74.1%) under our drift discipline.
- Small N is annotated. The agentic probe is 9/9 on both arms with validity caveats: a directional signal, not a superiority claim.
- Cost honesty. The 99.507% context reduction pays back in ~4 queries under favorable conditions and 31–53 queries under adverse cache behavior. The report includes both figures.
Method
Benchmarks use frozen-anchor paired evaluations. Each checkpoint has one anchor configuration, and both arms (with-memory and baseline) receive the same canonical evidence stream. Runs missing forced compaction or transport symmetry are retained but marked degraded. Reader and judge models are pinned per anchor; readers cannot change mid-comparison.