THE MEMORY SYSTEM
408 facts, one per file — persistent agent memory with integrity monitoring on what matters most
The Problem
Every agent session starts cold. Context windows close; state evaporates. An agent that spent three sessions debugging a subtle infrastructure issue wakes the next morning with no memory of what it learned — and starts the same debugging loop again.
The obvious fix is a shared knowledge file. But shared files create a different problem: two agents updating the same document simultaneously corrupt it. One agent overwrites the other's additions. The file that was supposed to accumulate knowledge instead loses it. And once a shared index grows past a few hundred entries, the entire blob loads into every session — burning token budget on things that are irrelevant 95% of the time.
The fleet needed memory that persisted across restarts, scaled without blowing context budgets, survived concurrent agents without corruption, and could be trusted not to quietly rot.
The Design
The core principle: one fact, one file. Each memory is a standalone Markdown file with frontmatter. It contains exactly one recalled observation — user preference, project state, learned pattern, reference — and nothing else. 408 files total.
A curated index loads hot every session: 139 lines, roughly 15 KB after compaction, one line per memory. The line is a hook — a recall trigger — not a summary. Bodies load cold, on demand, only when the hook matches something in the current task. Most sessions load two or three memory bodies, not four hundred.
Write-rights are namespaced by agent role. Each agent owns the files in its own prefixed lane. Shared files — user profile, feedback patterns, cross-cutting reference — belong to the curator alone. No two agents can fight over the same file because no two agents share write-rights to the same file. Siblings may append to the index; only the curator may rewrite it.
Four memory types cover the full surface: user (who the principal is, what they value, how they work), feedback (what to do differently), project (what exists and where), reference (operational facts that don't change). Three overflow indexes handle the long tail without bloating the main index.
The index loads hot. The bodies load cold. You only pay for what the task actually needs.
Craft Details
The compaction protocol. When the index exceeds its size budget, a compaction pass rewrites hooks to 150 characters or fewer — concise enough to scan in one glance. The protocol enforces three invariants: backup before any edit, additive-only changes to target files (a compaction pass may shorten hooks but never drop facts), and a fail-open trap that restores from backup mid-run if any step errors. A compaction that partially completes is worse than one that didn't start — the protocol treats them the same way.
Before compaction the index was 32.6 KB. After: 15 KB. Same 408 facts.
Staleness discipline. Memory files are point-in-time observations, not live state. An agent that reads a memory file from six months ago and acts on it as if it's current is worse than an agent that had no memory — it acts with false confidence. Any memory recalled past an age threshold carries an automatic warning: verify against current state before asserting as fact.
The phrase that governs this: recalled, not obeyed. A memory is a prior, not an instruction. The agent checks; it doesn't blindly inherit.
Integrity monitoring. Sensitive memory sits under integrity monitoring. If protected content changes unexpectedly, something touched the memory system that shouldn't have: prompt injection, config drift, an agent acting outside its write-rights, or outright tampering. An unexplained change is not noise to be reconciled later — it is evidence, and it is treated that way.
The Honest Failure
The memory corpus has a semantic search layer — a vector index that lets agents retrieve relevant memories by meaning, not just keyword hooks. This worked correctly at launch, and then silently stopped working correctly. Nobody noticed.
The ingest job that fed the vector index had no refresh schedule. It ran once and was never rerun. As new memory files were added over two weeks — the system's most productive period — the index drifted. When the gap was finally measured against a gold-standard set of 57 known memories, the result was unambiguous: 35 of 57 were missing from search. Recall@5 was 22%. The semantic index was confidently returning results from 61% of the corpus it was supposed to cover.
The fix was a full re-index. After: recall@5 at 62%, 35 recovered entries, all 57 now reachable. The fix took less than an hour. The drift had been accumulating for two weeks.
This incident produced a permanent lesson in the fleet's engineering archive: "a derived index with no refresh schedule rots silently — 'the query works' is not the same as 'the corpus is complete.'" The lesson is now enforced by a scheduled re-index job that did not exist before.
The query worked. The index was missing 61% of what it was supposed to know. Those are not the same thing.
Stack
Result
The fleet's memory system currently holds 408 facts across 4 memory types, per-agent namespace lanes, and 3 overflow indexes. The hot-load index is 139 lines — small enough to fit in any session's context budget, specific enough that matching hooks point to the right body on first read. Index size: 15 KB after compaction (down from 32.6 KB, same fact count).
UNMEASURED: end-to-end recall benefit on task outcomes (no controlled baseline — the fleet with and without memory were never run in parallel); optimal index size (15 KB is the post-compaction observation, not a calibrated target).
The integrity monitor has never caught tampering. That is the most important number in this system, and it has no receipt — only the absence of one.
Recalled, not obeyed. A memory is a prior. The agent checks; it doesn't blindly inherit.