Architecture
How RE-call works
RE-call builds an immutable, calibrated generation of your memory once, then answers every search against that one pinned generation. Before anything reaches the caller, a trust gate gives each hit a verdict. When nothing qualifies, it abstains with a reason instead of returning the nearest match.
Both diagrams on this page are interactive. Click any box to see the source lines it describes, on GitHub, pinned to the commit the diagram was drawn from. Open a diagram full screen to trace a path between two boxes, switch to the light theme, or export it as an image.
1 · Build once
A generation is built, checked, and only then served.
A build starts from a source manifest: a fixed list of the objects that go into memory. The generation manager fetches each object, parses its frontmatter and body, splits it into chunks, and embeds the chunks. Embeddings go through a content addressed cache, and a source whose content and pipeline have not changed is carried forward instead of embedded again.
The result is a generation: an immutable set of chunks, vectors and full text entries in your PostgreSQL database with pgvector. It is validated, and in production it can only become the active generation once its calibration is certified. Calibration fits the abstention threshold for that generation against a labelled set of queries, and publishes it only when the threshold certifies.
2 · Serve against a pin
Every search reads one generation from start to finish.
A search first pins the active generation, and every read in that search uses it, so a promotion that lands halfway through cannot mix two corpora into one answer. The search also loads that generation's certified threshold.
The hybrid retriever embeds the query and runs two searches against the pinned generation: a dense cosine search in pgvector and a PostgreSQL full text search. Reciprocal rank fusion merges the two lists. A reranker can reorder the result, but only when you turn one on.
3 · The trust gate
Each hit gets a verdict, and an empty verdict means no answer.
Every hit is given a verdict, such as ok, superseded, expired, not_yet_valid or low_confidence. A hit below the certified threshold is low confidence. Valid hits are ranked first, so a newer memory outranks the stale one it replaced.
If no hit earns ok, the search abstains and says why. It does not hand back the nearest match as if it were an answer.
Step by step
One recall_search call, from question to verdict.
The same path, followed through a single call from an agent to the MCP server. Read it from top to bottom. The red arrows are the trust checks, including the refusal that happens before anything is retrieved.
Three outcomes
Refuse, abstain, or answer.
| Outcome | When | What the caller gets |
|---|---|---|
| Refuse | Under strict trust, the generation has no certified calibration, or the embedder is not the model the generation was built with. | An error naming the reason. Nothing was retrieved. |
| Abstain | Retrieval ran, but no hit earned the verdict ok. |
abstained: true and the reason. |
| Answer | At least one hit earned ok. |
The hits, each with its verdict, confidence and provenance, ok hits first. |
Beyond the default path
Optional paths go through the same gate.
Explicit reasoning queries and one hop graph expansion both re-enter trusted search, so their candidates face the same trust gate before a cited answer can use them. Applying a reviewed fact goes through the provenance controller into an append only ledger. In generation mode, ingesting files through the MCP server builds, validates and promotes a generation through the same generation manager.
For the full account, read the architecture writeup, the guides to generations and calibration, or how the diagrams are made. To set it up yourself, start with the install guide.
