Benchmarks
This page reports only currently valid evidence. Results measured with a different recall stack, embedding service, or incomplete runner are not presented as current accuracy. These are observations from stated cases, not guarantees for every request.
Measured in production
| Observation | Result | Scope |
|---|---|---|
| Long-history request (extreme case) | One eligible request compressed 66,265 input tokens to 365 forwarded (279 recalled) — a ~99.5% reduction on that single request only | One real, eligible hosted Provider Proxy request; an upper bound, not a representative rate |
| Agentic tool-output removal | 43k–53k fewer input tokens per request (about 85% of the removable tool-output bulk); upstream response time fell from ~2 minutes to ~35 seconds | Sampled production agentic sessions, 2026-08 |
| Starter limiter burst | 30 successful requests + 15 HTTP 429 responses from a 45-request burst | One measured Starter 30 req/min window |
| New memory readiness | ~8.2 seconds from save to recallable | One production remember/status round trip |
| Saved-text integrity | Stored and re-read text matched byte-for-byte | Production remember → status → recall → read |
| Warm recall | ~5 seconds | One warm production observation; not an SLO |
Network distance, corpus size, provider latency, protected request content, and request shape all affect results. Do not extrapolate a single observation to every request.
Recall depth: cost and quality trade-offs
Measured on a frozen probe set (English and Chinese arms, reproducible from this repository) with the production recall stack, 2026-08-25:
| Depth | Unanswerable questions declined | Answerable questions found | Provider tokens spent | Typical latency |
|---|---|---|---|---|
fast |
23/23 | 15/15 | 0 | ~3 ms |
auto (default) |
all | all | spent only on low-confidence queries (~55% of probes), at 52–66% of deep's cost | ~3.5 ms |
deep |
all | all | highest; multi-round gathering | ~1.3 s |
Reading: the free fast path already declines every unanswerable question in the probe set. auto keeps that quality and only pays for deeper gathering when it is genuinely needed. deep exists for the hardest multi-part questions.
Note: the latencies above come from the internal probe harness and exclude network and provider time. End-to-end production calls are slower: a 2026-09-11 single-conversation production probe measured fast at roughly 1-3 seconds and deep at roughly 5-6 seconds per question, including the multi-round selector.
When token reduction appears
Reduction needs both: a history long enough to matter, and older exchanges whose content is already stored as memory. Short, new, or fully protected conversations correctly show no reduction.
A 0% reduction is a common and correct result for real agentic and chat traffic: only long, eligible, already-memorized histories reduce at all. The 66,265-token example above is one extreme eligible request, not a user-facing average saving rate, and must not be read as one.
Recall accuracy (LoCoMo-Refined)
Measured on the public LoCoMo-Refined long-conversation memory benchmark: 1,310 scored questions, one frozen production-near run.
| Category | Questions | Accuracy |
|---|---|---|
| single-hop | 204 | 48.0% |
| temporal | 276 | 63.0% |
| multi-hop | 68 | 67.6% |
| open-domain | 762 | 79.0% |
| All scored | 1,310 | 70.2% |
SPM returns evidence or abstains: in this run it returned evidence for about 95% of scored questions and declined the rest rather than guessing.
What these numbers mean
- A single frozen run on the public LoCoMo-Refined question set with the current recall stack (privately operated embedding service, no hosted reranker).
- Scored with a general open-ended answer model over the evidence SPM returned; a binary judge checked answers against gold strings.
- Not a guarantee for every request, and not an Agent Memory Leaderboard figure: that leaderboard applies its own answer model, prompt, and scoring, while SPM contributes the evidence it returns.
Methodology rules
- Publish the exact commit, configuration, model/embedding identity, corpus, question set, and launch command.
- Keep degraded or interrupted runs, but label and exclude them from headline anchors.
- Separate retrieval/evidence recall from downstream answer phrasing.
- Report medians and tail latency where sample size permits.
- Never present a single request as an SLO or a general savings guarantee.