agent-memory-benchAMB / AMB
Leaderboard

Ranked by what ran.

Task success on executable checkers, paired by task against the claude_md baseline. Numbers are generated from published run artifacts.

Adversarial by design

This benchmark tests failure as well as recall. The corpus is empty, stale, contradictory, or out of scope in 65% of the scored cells. A clean governing fact is present in 35%.

The table measures useful memory under realistic liability, not retrieval on a clean prompt.

01

Products and controls

ranked when a run is official
rank arm integration task success Δ vs baseline 95% CI discarded cells total tokens cost / task

Small deltas are labeled as noise. Discarded cells failed the admission gate and are reported rather than scored. Tokens and cost are end-to-end, including ingestion.

02

Products, condition by condition

where a memory layer earns and where it costs

The condition breakdown shows where each product helps and where memory becomes a liability. present holds a clean governing fact; the other four test abstention, recency, conflict, and scope.

Values are solved cells over admitted cells. A held product publishes neither a headline nor a breakdown.

03

Reference tracks

diagnostics, never ranked

protocol carries the shared memory instruction with no memory behind it. Its delta is the cost of asking the agent to use memory. recall_prefetch retrieves with the exact task prompt, removing query formulation from the path.

track what it isolates task success Δ vs baseline

Reference tracks diagnose the memory path. They run in the grid but are never ranked as products. See the method page.

04

What makes a result official

or it does not appear here

A run appears on this page only when all of the following hold, in this order:

gaterequirement
1 · preregistered protocol committed under preregistration/ before the first session; the harness refuses to start while that directory is dirty
2 · announced the run is announced before it happens, not after it succeeds
3 · gated every scored cell passed the admission gate; discard counts published per arm
4 · published in full per-session logs, streams, admission verdicts and costs land in results/<run_id>/, wins and losses alike

The results/ tree holds the evidence behind this page. The dated project status and verification commands are in docs/STATUS.md.