Ranked by what ran.
Task success on executable checkers, paired by task against the
claude_md baseline. Numbers are generated from published run artifacts.
This benchmark tests failure as well as recall. The corpus is empty, stale, contradictory, or out of scope in 65% of the scored cells. A clean governing fact is present in 35%.
The table measures useful memory under realistic liability, not retrieval on a clean prompt.
| arm | success | cost / cell or task | mean session | vs baseline | read |
|---|
Products and controls
ranked when a run is official| rank | arm | integration | task success | Δ vs baseline | 95% CI | discarded cells | total tokens | cost / task |
|---|
Small deltas are labeled as noise. Discarded cells failed the admission gate and are reported rather than scored. Tokens and cost are end-to-end, including ingestion.
Products, condition by condition
where a memory layer earns and where it costsThe condition breakdown shows where each product helps and where memory
becomes a liability. present holds a clean governing fact; the other four test
abstention, recency, conflict, and scope.
Values are solved cells over admitted cells. A held product publishes neither a headline nor a breakdown.
Reference tracks
diagnostics, never rankedprotocol carries the shared memory instruction with no memory
behind it. Its delta is the cost of asking the agent to use memory. recall_prefetch
retrieves with the exact task prompt, removing query formulation from the path.
| track | what it isolates | task success | Δ vs baseline |
|---|
Reference tracks diagnose the memory path. They run in the grid but are never ranked as products. See the method page.
What makes a result official
or it does not appear hereA run appears on this page only when all of the following hold, in this order:
| gate | requirement |
|---|---|
| 1 · preregistered | protocol committed under preregistration/ before the first session; the harness refuses to start while that directory is dirty |
| 2 · announced | the run is announced before it happens, not after it succeeds |
| 3 · gated | every scored cell passed the admission gate; discard counts published per arm |
| 4 · published in full | per-session logs, streams, admission verdicts and costs land in results/<run_id>/, wins and losses alike |
The results/ tree holds the evidence behind this page.
The dated project status and verification commands are in
docs/STATUS.md.