SPM — StellarPath Memory Operating System Docs

Benchmarks

Published Last reviewed Applies to SPM-Polaris

This page reports only currently valid evidence. Results measured with a different recall stack, embedding service, or incomplete runner are not presented as current accuracy. These are observations from stated cases, not guarantees for every request.

Measured in production

Observation Result Scope
Long-history request (extreme case) One eligible request compressed 66,265 input tokens to 365 forwarded (279 recalled) — a ~99.5% reduction on that single request only One real, eligible hosted Provider Proxy request; an upper bound, not a representative rate
Agentic tool-output removal 43k–53k fewer input tokens per request (about 85% of the removable tool-output bulk); upstream response time fell from ~2 minutes to ~35 seconds Sampled production agentic sessions, 2026-08
Starter limiter burst 30 successful requests + 15 HTTP 429 responses from a 45-request burst One measured Starter 30 req/min window
New memory readiness ~8.2 seconds from save to recallable One production remember/status round trip
Saved-text integrity Stored and re-read text matched byte-for-byte Production remember → status → recall → read
Warm recall ~5 seconds One warm production observation; not an SLO

Network distance, corpus size, provider latency, protected request content, and request shape all affect results. Do not extrapolate a single observation to every request.

Recall depth: cost and quality trade-offs

Measured on a frozen probe set (English and Chinese arms, reproducible from this repository) with the production recall stack, 2026-08-25:

Depth Unanswerable questions declined Answerable questions found Provider tokens spent Typical latency
fast 23/23 15/15 0 ~3 ms
auto (default) all all spent only on low-confidence queries (~55% of probes), at 52–66% of deep's cost ~3.5 ms
deep all all highest; multi-round gathering ~1.3 s

Reading: the free fast path already declines every unanswerable question in the probe set. auto keeps that quality and only pays for deeper gathering when it is genuinely needed. deep exists for the hardest multi-part questions.

Note: the latencies above come from the internal probe harness and exclude network and provider time. End-to-end production calls are slower: a 2026-09-11 single-conversation production probe measured fast at roughly 1-3 seconds and deep at roughly 5-6 seconds per question, including the multi-round selector.

When token reduction appears

Reduction needs both: a history long enough to matter, and older exchanges whose content is already stored as memory. Short, new, or fully protected conversations correctly show no reduction.

A 0% reduction is a common and correct result for real agentic and chat traffic: only long, eligible, already-memorized histories reduce at all. The 66,265-token example above is one extreme eligible request, not a user-facing average saving rate, and must not be read as one.

Recall accuracy (LoCoMo-Refined)

Measured on the public LoCoMo-Refined long-conversation memory benchmark: 1,310 scored questions, one frozen production-near run.

Category Questions Accuracy
single-hop 204 48.0%
temporal 276 63.0%
multi-hop 68 67.6%
open-domain 762 79.0%
All scored 1,310 70.2%

SPM returns evidence or abstains: in this run it returned evidence for about 95% of scored questions and declined the rest rather than guessing.

What these numbers mean

Methodology rules