Persistent Memory for LLM Agents

View on GitHub → Updated 2026-07 · v4.0.0

Persistent memory for LLM agents, built on keyword search and typed knowledge graphs with no embeddings anywhere. Benchmarked in public and re-reported lower wherever the re-runs came back lower. At v4.0.0 the staged retrieval lanes are live on the production path -- temporal spine, intentional clustering, structural HRR, junk demotion.

Jump to the latest update ↓ -- What changed since the original write-up.


Abstract

LLM agents have no memory across conversations. Corrections are lost when the context window closes, and the agent repeats the same mistakes. Existing memory systems treat this as a retrieval problem, but the harder unsolved problems are write correctness and governance: what gets stored, how conflicts are resolved, and whether wrong memories can be corrected. I built aelfrice, a persistent memory system running on FTS5 keyword search, typed knowledge graphs, and entity-index retrieval, with no embeddings. The system detects user corrections at 92% accuracy without LLM calls, recovers 99.5% of the 31% of stored directives unreachable by keyword search, and reduces injected tokens by 55% with zero retrieval loss. Across five benchmarks, aelfrice achieves 66.1% F1 on LoCoMo (+14.5pp over GPT-4o), 90% on MemoryAgentBench single-hop (+45pp), 60% on multi-hop (8.6x the published 7% ceiling), 100% on StructMemEval state tracking, and 59.0% on LongMemEval (-1.6pp vs GPT-4o pipeline, different judge). All code, benchmarks, and experiment data are open source under MIT license. Throughout, pp means percentage points, the plain arithmetic difference between two percentages rather than a relative change. (aelfrice was developed and benchmarked under the working name "agentmemory"; the two names refer to the same system. The benchmark figures above are from the v1 substrate, and a v3.0.1 re-run did not reproduce the MemoryAgentBench pair. The 90% single-hop and 60% multi-hop scores, and the +45pp and 8.6x framings built on them, were lab-side projections of a SUPERSEDES mechanism keyed on (subject, predicate) collisions, and that mechanism is not in the public retrieval substrate. That substrate scores 57% single-hop and 6% multi-hop. See §12, which also carries a corrected account of what did and did not drive the LoCoMo change.)


1. Introduction

LLM agents have no memory across conversations. Every session starts from zero. When a user corrects the agent, that correction is lost the moment the context window closes. The next session, the agent makes the same mistake. The user corrects again. To make matters worse, the agent will often ignore corrections in the exact same context window session.

No Implementation comic - CS-002/CS-006

Memory failures are one of the two largest categories of LLM behavioral failures documented in this project's failure taxonomy: 7 of 38 cataloged patterns, just behind code-quality failures at 8. The problem compounds across sessions. The MemoryAgentBench benchmark (ICLR 2026) tested multi-hop conflict resolution and found a ceiling of 7% accuracy across the methods in the 2025 preprint this project benchmarked against; the camera-ready has since added a reasoning model that reaches 28%.

Existing approaches overwhelmingly treat memory as a retrieval problem. Across the architectures cataloged by Zhang et al. (2024), Hu et al. (2025), and Leonard Lin's independent architectural analysis of 35+ papers and 14 community memory systems, the dominant pattern is: store text, embed it, retrieve by similarity. StructMemEval found that retrieval-only systems built on this pattern score near zero on state tracking at scale. They can't tell you what's currently true vs. what was superseded.

Embedding is the middle step of that pattern, and it is the one this system never takes. An embedding turns a piece of text into a list of numbers arranged so that text with a similar meaning lands near it, and retrieval then compares the numbers rather than the words. Keyword search compares the words. When a query and a stored directive share none, keyword search cannot reach that directive at all, and no amount of re-ranking fixes it: the directive never becomes a candidate.

Lin concluded:

"The biggest differentiator is not 'vector DB vs SQLite' — it's write correctness and governance: provenance/audit trail, write gates / confirmation, conflict handling, reversibility (inspect/edit/delete)."

On every turn, the memory system retrieves stored content, the LLM reads it and generates a response, and the turn ends. Whether the retrieved content helped is never written back, and a user correction and an LLM inference enter the store as the same kind of thing. Feedback paths do exist, but they are narrow. The 47-author survey by Hu et al. catalogs a reinforcement-learning-assisted class (RMM, Mem-alpha, Memory-R1, Memory-as-Action) that learns memory decisions from task reward. Lin's 14+ system analysis finds retrieval-utility feedback proposed in two community systems, and one complete citation→usage→retention loop shipped in Codex. What none of them carry is a signal derived from the user's own reactions: reinforcing a directive the user found helpful, weakening one the user has overridden.

Letta showed that a gpt-4o-mini agent with ordinary file tools reaches 74% on LoCoMo (Maharana et al., ACL 2024). That's the bar.

Now for my Fosbury Flop.

Agentmemory system architecture: ingestion, retrieval, and feedback pipelines

I built this system because I got sick and tired of asking Claude for the latest on my test runs, which were burning CPU time on cloud compute, only for Claude to tell me "huh? what test dispatches? oh those. yeah they've been hanging for 2 hours because I didn't follow the runbook you told me to follow."

Scale Before Validate comic - CS-011

Before building, I surveyed prior work: the survey papers (Zhang et al. 2024, Hu et al. 2025, Yang et al. 2026, "Memory in the LLM Era" 2026), the benchmarks, the community systems, and Leonard Lin's independent architectural analysis of 35+ papers and 14 community memory systems, which compares mechanisms against the effectiveness their authors claim.

Four findings shaped the project direction:

  1. Human memory is the wrong target. Zhang et al. (2024) and Hu et al. (2025) show that the dominant design paradigm maps psychological memory models onto LLM architectures. Human memory is notoriously unreliable: Ebbinghaus (1885; replicated by Murre & Dros, 2015) showed ~56% of learned material is forgotten within one hour. Eyewitness misidentification contributed to roughly 69% of DNA-based exonerations, 252 of 367 cases (Innocence Project). Computer memory is perfect at storage; the hard problem is retrieval. Where human memory is useful: retrieving gists. Brainerd & Reyna (2005) showed that gist traces are far more durable than verbatim traces. We're trying to replicate that associative retrieval: connecting things that share no surface-level vocabulary.

  2. An off-the-shelf agent with file tools reaches 74% on the most-used benchmark. Letta reports 74.0% on LoCoMo (Maharana et al., ACL 2024) from a gpt-4o-mini agent with no specialized memory tools. The transcript is uploaded as a file, which Letta parses and embeds, and the agent searches it with grep and an embedding-backed search_files. Any memory system has to beat that.

  3. Multi-hop conflict resolution topped out at 7%. MemoryAgentBench (ICLR 2026) tested the ability to follow a chain of related decisions across sessions and determine which is currently in effect. That 7% envelope is the one this project benchmarked against.

  4. Lin's independent analysis pointed to an underexplored axis. The majority of systems focus their novelty on retrieval: better embeddings, better similarity metrics, better ranking. The harder unsolved questions are about what gets stored, how conflicts are resolved, and whether wrong memories can be corrected. Lin's concept of "write correctness and governance" breaks into four components:

    • Provenance: Where did this memory come from? A user prompt, an LLM inference, or derived content? Every stored entry should carry its lineage.
    • Write gates: Not everything should be stored. Expanding a 586-node graph to 16,463 nodes without quality filtering dropped grep coverage from 92% to 85% and FTS5 coverage from 85% to 69% (Exp 48).
    • Conflict handling: When two memories contradict, which wins? StructMemEval found that vector stores score near zero on state tracking at scale.
    • Reversibility: Lin scores reversibility (inspect/edit/delete) as one of four governance axes, and most of the systems he surveyed do not have it. The ones that come closest keep supersedes or version chains rather than an undo.

Full benchmark comparison tables are in Appendix B.


3. Approach

The unit this system stores is a belief: one record holding a short statement, where it came from, a confidence number, and typed links to other beliefs. Four questions drove the design:

  1. How do you detect user corrections, stated preferences, and behavioral rules without extra LLM inference?
  2. How do you retrieve relevant content when the query shares zero vocabulary with the stored content?
  3. How do you track whether a retrieved memory was actually useful?
  4. How do you distinguish "the LLM should consider this" from "the LLM must obey this"?

Correction detection: 92% accuracy without any LLM calls (Exp 1 V2, 38 labelled corrections from a single project, no holdout set). When LLM classification is enabled (~$0.005/session), accuracy reaches 99%. The zero-LLM pipeline runs on every conversation turn with zero marginal cost.

Vocabulary gap recovery: 31% of directives extracted from five codebases are unreachable by keyword search (Exp 53, 3,321 directives examined). The query and the stored directive share zero vocabulary. For example: a user says "never mock the database in tests." Later, the agent is about to write a test with unittest.mock.patch('db.connect'). Zero overlapping words, but a human immediately sees the connection. 99.5% of those gapped directives are structurally bridgeable: co-located with, or sharing content words with, a directive keyword search does reach. The system walks those bridges with a graph traversal that needs no embeddings and no LLM inference.

Life of one belief, in three panels. WRITE, session 3: the user types 'never mock the database in tests'; a correction detector made of regexes and keyword heuristics catches it with no LLM call, at 92% accuracy on 38 labelled corrections (Exp 1 V2), where paying an LLM classifier about $0.005 a session would buy 99%. STORE: the row the database actually holds is the sentence, type correction, origin user_corrected, lock_level none, and a Beta-Bernoulli prior of alpha 9.0 and beta 0.5. Two typed edges sit beside it in the edges table, both RELATES_TO, both pointing at neighbouring beliefs that keyword search does reach. That prior means a correction the user made starts 94.7% confident before any evidence, while the same sentence merely inferred by the agent starts deflated at 1.8 / 0.5, or 78.3%, and has to earn the rest. That row is a belief, and the belief counts in this article count these. READ, session 7: the agent is about to write unittest.mock.patch('db.connect'), which shares no word with the stored sentence. The query enters four lanes. The locked-beliefs lane is empty here, because this belief is not locked. FTS5 keyword search misses, because the query and the stored sentence share zero words. That is the gap that hides 31% of stored directives, 1,030 of 3,321 across five codebases. The entity-index lane chains structured triples, and it took MemoryAgentBench multi-hop from 6% to 35% chain-valid. The typed-edge graph walk finds it in two hops, from a belief keyword search does reach along a typed edge to this one, recovering 99.5% of that 31% with no embeddings and no LLM call, though section 12 reports that walk ships off by default. The merge stage packs the result at 55% fewer injected tokens with zero retrieval loss, and the belief arrives above the next user turn before the mock.patch call is written. The argument of the figure: the write side is the part the literature skips, and it is where this system's arguments live.Vocabulary gap bridge: graph traversal connects mock.patch to never-mock directive
Vocabulary Gap Prevalence (Exp 53)
5 codebases, 3,321 directives examined
MetricValueRate
Total directives examined3,321
Directives with vocab gap1,03031%
Recovered by graph layer1,02599.5%
Gap CategoryRateExample
Uncategorized34%—
Emphatic prohibitions29%"NEVER do X"
Domain jargon13%Tool names
Tool bans12%"don't use Y"
Implicit rules8%Context-dependent
99.5% of gaps (1,025 of 1,030) are structurally bridgeable.
Graph traversal hub: research_HRR_FINDINGS node with edges radiating to connected beliefs
A high-connectivity hub node (research_HRR_FINDINGS) in the knowledge graph. Edges radiate outward to beliefs across multiple topics. When keyword search misses a directive due to vocabulary mismatch, graph traversal follows these edges to recover it -- the 31% of directives FTS5 cannot reach.
Task 41 comic - CS-020

Entity-index retrieval: To address multi-hop conflict resolution, the system extracts structured triples (entity, property, value, serial_number) from ingested text using regex relation-family patterns, then chains through entity relationships at query time. This layer (L2.5, between FTS5 and holographic reduced representations, or HRR) took the MemoryAgentBench multi-hop score from 6% to 35% chain-valid, 5x the 7% envelope this project benchmarked against (see Section 5.3). HRR is a fixed arithmetic construction rather than a learned one, so the same input produces the same vectors on every run. Chain-valid counts only the answers the system reached by following the entity chain. A raw score also counts the ones that matched without it.

Confidence tracking: The system tracks retrieval outcomes and updates confidence accordingly (Exp 66: +22% MRR gain over 10 feedback rounds; Bayesian calibration ECE 0.066, target < 0.10). MRR is mean reciprocal rank: 1 over the position of the first correct hit, averaged across queries, so an answer that lands second every time scores 0.5. ECE is expected calibration error: the average gap between how confident the system says it is and how often it turns out to be right. A 0.066 ECE is a gap of 6.6 percentage points. A retrieved memory moves in whichever direction the outcome went. A memory that was never retrieved does not move at all, so confidence can only be lost by being used and being wrong.

Correction enforcement: Storing a correction and enforcing it are different problems (see CS-006, earlier). The system distinguishes between content the LLM should consider and constraints the LLM must obey (Exp 84: 10/10 locked directives retrieved and enforced across 5 sessions).

Correction hub node with edges fanning outward to connected beliefs across topics
A correction hub in the knowledge graph. When a user corrects the agent, the correction belief becomes a high-degree node with SUPERSEDES and CONTRADICTS edges linking it to the beliefs it replaces. Long-range edges (upper-left, lower-right) connect to beliefs in distant topic clusters, so the correction propagates across the full graph during retrieval.
Retrieval pipeline layers: L0 locked through L3 graph traversal

The architecture uses keyword search as the primary retrieval layer and wraps it with structural gap recovery, confidence tracking, and constraint injection. The combined system handles both the content keyword search reaches and the 31% it misses. Grep is the plain text-search command that ships with Unix, and it is the baseline any of this has to beat. Grep won on the keyword retrieval benchmark (Exp 47, 92% coverage vs. 85% for the prototype), and I accepted that: grep is fast, precise, and costs no LLM calls. But grep cannot bridge the vocabulary gap. The locked-directive MRR improved from 0.589 to 0.867 after retrieval tuning (Exp 60, LOCK_BOOST_TYPED re-ranking at K=30).

Retrieval Coverage (Exp 47, 586-node graph)
MethodCoverageTokensPrecision
grep (decision)92%lowhigh
grep (sentence)92%highmoderate
Prototype A85%lowmoderate
Prototype B85%lowmoderate
Null hypothesis: the architecture beats filesystem+grep. Result: NOT rejected. Grep achieved 92% vs 85% for both prototypes.

Token reduction: Type-aware token reduction achieves 55% savings with zero measured retrieval loss (Exp 42). Constraints survive verbatim; rationale compresses to 0.4x; metadata compresses to 0.3x. At 19K+ stored nodes, naive injection would consume the entire context window before the agent reads the user's message.

Token Reduction Results (Exp 42)
MetricValue
Before35,741 tokens
After15,926 tokens
Savings55%
Retrieval coverage100% (all 6 topics, 18 queries)
Content TypeTarget Reduction Factor
Constraints1.0x (never reduce)
Rationale0.4x
Context0.3x
Measured context compression came in harder than the target, at 0.23x.

Scale effects: When the graph expanded from 586 to 16,463 nodes without filtering, retrieval coverage dropped from 92% to 85% for grep and from 85% to 69% for prototypes (Exp 48). Decision-level directives were 3.6% of the expanded graph. The fix was to filter at ingestion time. Only decision-level directives pass the write gate.

Scale vs. Coverage (Exp 48)
Graph SizegrepProto AProto B
586 nodes92%85%85%
16,463 nodes85%69%69%
Decision-level directives: 3.6% of expanded graph. The other 96.4% is noise that dilutes signal.
aelfrice knowledge graph visualization showing node clusters and edge density across a 19K-node production database
Production knowledge graph (19K nodes, Obsidian graph view). The large central cluster is the primary project's belief network. Two smaller satellite clusters (upper-left, lower-left) are isolated project databases. Scattered peripheral nodes are low-edge beliefs awaiting graph integration. Dense interior regions correspond to heavily cross-referenced topics where multiple beliefs reinforce each other through SUPPORTS, RELATES_TO, and SUPERSEDES edges.

4. Evaluation

Layer 1 (programmatic checkers): Deterministic checks like "does the output contain the required section headings?" They catch structural violations and cannot evaluate semantic correctness.

Layer 2 (structural validators): Added after CS-007b, where an LLM satisfied every individual constraint and violated the relationships between them: correct sections in the wrong order, or a reference to a decision that contradicts an earlier one.

Layer 3 (LLM-as-judge with anti-contamination): Standard LLM-as-judge approaches (Zheng et al., 2023) hand the evaluating LLM both the input prompt and the response. When the judge can see the reasoning that produced a violation, it finds that reasoning plausible and rationalizes the violation. So the judge here is isolated. It receives only the constraint and the output, never the conversation that produced them.

Layer 4 (adversarial follow-up): Added after CS-024 (sycophantic collapse). An LLM can pass all three previous layers and still fail when a user pushes back.

Sycophantic Collapse comic - CS-024

For agent memory retrieval, precision matters more than recall. A false negative (relevant directive missed) is invisible. A false positive (irrelevant directive injected) makes the LLM act on wrong context, and the user has to notice it, diagnose it, and correct it. The scale experiment (Exp 48) showed this directly: expanding the graph without filtering retrieved more content, but the wrong content.

Retrieval Error Cost Matrix
Error TypeUser Impact
True positiveCorrect directive retrieved and followed. No user intervention.
True negativeIrrelevant directive correctly excluded. No user intervention.
False negative (recall failure)Relevant directive missed. Nothing in the output flags the omission -- the failure is invisible.
False positive (precision failure)Irrelevant directive injected. LLM acts on wrong context. User must notice, diagnose, and correct. Active harm.
Big Numbers comic - CS-008

5. Results

Cross-benchmark comparison: aelfrice vs published best across 5 benchmarks
Key Results Summary
MetricResultNotes
Benchmarks
LoCoMo F1 (Opus 4.6)66.1%+14.5pp vs GPT-4-turbo 128K (51.6%)
MAB SH 262K57% (lab prototype scored 90%; see §12)+12pp vs GPT-4o-mini (45%)
MAB MH 262K60% Opus (lab prototype; 35% chain-valid, and 6% on the shipped substrate — see §12)8.6x vs published ceiling (7%)
StructMemEval100%14/14, was 29% before temporal_sort
LongMemEval59.0%-1.6pp vs GPT-4o pipeline (60.6%)
Core Pipeline
Correction detection92%Zero-LLM, 5 codebases
Vocabulary gap recovery99.5%31% of directives gapped
LLM classification99%~$0.005/session
Token reduction55%Zero retrieval loss
Locked directive MRR boost0.589->0.867After retrieval tuning
Bayesian calibration (ECE)0.066Target < 0.10
Feedback loop MRR gain+22%Over 10 rounds (Exp 66)
Multi-session validation10/105 sessions (Exp 84)
Infrastructure
Acceptance tests62/65 pass29 test files, 1.65s
Test suite362 pass (as recorded in the v1 project summary; no run artifact survives)Unit, integration, behavioral
Retrieval latency0.7s avg19K-node production DB
Onboarding speed (scan)~2.5s16,690 nodes from 35 commits + 163 docs
Onboarding speed (full pipeline)6.5sScan + ingest + edge storage + vault sync

5.1 Core pipeline metrics

Correction detection reaches 92% accuracy without LLM calls (Exp 1 V2, 38 labelled corrections from a single project, no holdout set). Adding LLM classification takes it to 99% at ~$0.005/session. Keyword search misses 31% of directives, and the vocabulary gap recovery layer walks the graph to recover 99.5% of those (Exp 53). Token reduction saves 55% with zero retrieval loss (Exp 42). The feedback loop improves MRR by +22% over 10 feedback rounds (Exp 66), with Bayesian calibration at ECE 0.066.

5.2 Single-hop results

LoCoMo

The LoCoMo benchmark (Maharana et al., ACL 2024) tests whether a system can answer questions about past conversations across five categories. The run ingested the 10-conversation locomo10.json release (5,882 turns, 272 sessions, 1,986 QA pairs — counted from the distributed file; the project page reports only per-conversation averages) through the standard onboarding pipeline. Retrieval used FTS5+HRR+BFS with a 2,000-token budget. Scoring followed LoCoMo's F1 methodology, with a forced-choice heuristic for the adversarial category and 'and' stripped alongside the articles. F1 scores the word overlap between the answer and the reference answer, not whether the answer is right, so a correct answer worded differently still loses points. That is part of why the human ceiling on this benchmark is 87.9% rather than 100%.

Answer Key comic - Benchmark Contamination

The initial run was contaminated: the agents had access to ground truth (Appendix A carries the full narrative and the protocol that replaced it). The protocol-correct results:

LoCoMo Per-Category F1 (Protocol-Correct, Opus 4.6)
CategoryF1n
Single-hop69.4%841
Temporal45.4%321
Multi-hop42.2%282
Open-ended30.5%96
Adversarial97.5%446
Overall66.1%1986
LoCoMo Leaderboard Context
SystemF1Notes
Human87.9%Ceiling
GPT-4-turbo (128K full context)51.6%Best long-context in paper
RAG (DRAGON + gpt-3.5, top-5 obs)43.3%Best RAG in paper
Claude-3-Sonnet (200K)42.8%Long-context
gpt-3.5-turbo (16K)35.9%Long-context
aelfrice + Opus 4.666.1%FTS5+HRR, no embeddings

Single-hop is strongest at 69.4% and adversarial is near-perfect at 97.5%. Multi-hop and temporal sit at 42-45%; both need cross-session reasoning and date arithmetic. Open-ended is weakest at 30.5%, and it needs a kind of synthesis the retrieval pipeline does not directly support. Ingest time for all 10 conversations was ~25s. Average query latency was ~16ms.

MAB single-hop

MemoryAgentBench (Hu et al., ICLR 2026) tests conflict resolution: when facts change over time, can the system track which version is current? Single-hop asks direct questions: "What is X's current Y?" The scores are SEM, substring exact match: an answer counts as correct when the reference answer appears inside it.

MAB Single-Hop 262K
ReaderSEMPaper GPT-4o-miniPaper GPT-4o
Opus 4.690% (lab substrate)45%60%
Haiku 4.562%45%88%

MAB single-hop rose from 60% at v1.0 to 90% at v1.1, across a release whose main ingestion change was triple extraction with automatic SUPERSEDES edges; no triple-extraction-off ablation was run at v1.1, so the attribution is a reading of the version delta, not an isolated measurement. An ablation turns one piece off, changes nothing else, and re-runs the benchmark. Haiku still beats GPT-4o-mini (62% vs 45%) with the same retrieval, so retrieval is carrying part of the gain — but the reader matters here too: on identical retrieval Opus scores 90% to Haiku's 62%.

LongMemEval

LongMemEval (Wu et al., ICLR 2025) is a 500-question benchmark spanning six categories. GPT-4o reading the full LongMemEval_S history scores 60.6% without Chain-of-Note, and 64.0% with it; that is a long-context baseline, not a retrieval pipeline, and the paper's stronger configurations score higher still.

LongMemEval Per-Category Accuracy (Opus Judge)
CategoryAccuracyn
single-session-user91.4%70
single-session-preference80.0%30
single-session-assistant73.2%56
knowledge-update70.5%78
temporal-reasoning59.4%133
multi-session24.1%133
Overall59.0%500

The strengths are single-session recall at 91.4% and knowledge updates at 70.5%. The weakness is multi-session, at 24.1%. Failure analysis of the 101 incorrect multi-session answers split them 67% retrieval misses and 33% reasoning failures, and 84% of the retrieval misses were counting or aggregation questions. Budget and top_k sweeps, each varied alone, did not help. The later wide-retrieval work (§12) showed why: the two knobs only move together, and the bottleneck for this category is candidate recall, not BM25 ranking. Scoring uses Opus as judge rather than GPT-4o, so the comparison carries an asterisk.

5.3 Multi-hop and state tracking

MAB multi-hop

Multi-hop chains entity relationships. The rows that follow hold Alice's spouse as Bob in session 42 and as Carol in session 78. A multi-hop question asks about a property of Alice's current spouse, so the system has to land on Carol rather than Bob before the second hop can be answered at all. The entity-index (described in Section 3) extracts structured triples and chains through them at query time.

Entity-Index Triple Extraction
FieldExtracted TripleUpdated Triple
Input: "In session 42, Alice's spouse is Bob."
entityAlice
propertyspouse
valueBob
serial42
Later: "In session 78, Alice's spouse is Carol."
entityAlice
propertyspouse
valueCarol
serial78 (supersedes serial 42)
MAB Multi-Hop 262K
ReaderRaw SEMChain-ValidBenchmarked Envelope
Opus 4.647%35%<=7%
Haiku 4.546%35%<=7%

Opus and Haiku score identically at 35% chain-valid. That is the strongest evidence that the entity-index retrieval, not the LLM reader, drives the improvement. When retrieval provides the right entity chain, even Haiku can follow it.

Multi-hop experiment progression: 6% baseline to 60% across 6 experiments
Multi-Hop Progression (Experiments 1-6)
ExpMethodMH SEMKey Finding
--v1.0 Baseline (FTS5 chunks)6%Single FTS5 query
1Per-hop failure analysis--58% chaining, 17% world knowledge, 11% retrieval miss
2SUPERSEDES edges4%Helps SH, not MH — filtering shrinks context and hurts multi-hop
3Triple decomposition10%Granular helps
4Entity-index 2-hop35%Core breakthrough
5Extended regex (+7 patterns)55%+8pp over the Exp 5 control (47% raw SEM)
LLM entity extraction43%-4pp vs the control, -12pp vs extended regex
6Temporal coherence (resolve_all + branching)60%96% GT-reachable; reader bottleneck

Experiment 6 settled the retrieval question: branching through all historical values at each hop made 96 of 100 ground truth answers reachable. The remaining 36pp gap is a reader chain resolution problem.

Reader Quality Analysis
MetricOpusHaikuGapInterpretation
SH 262K90%62%28ppReader matters
MH chain-valid35%35%0ppRetrieval does all the work
MH raw SEM47%46%1ppRetrieval does all the work

StructMemEval

StructMemEval (Shutova et al., 2026) tests state tracking: given location updates across sessions, can the system answer "where is X now?"

StructMemEval Results
VersionAccuracyFix
v1.04/14 (29%)--
v1.114/14 (100%) retrieval recall, single runtemporal_sort + narrative timestamps (a later 4-attempt variance probe returned 0%, 7%, 100%, 100% — the headline is not robust)

The fix: assign narrative timestamps (30 days apart per session) and enable temporal_sort=True so the reader sees the most recent session content first. Whether this generalises beyond StructMemEval has not been measured — no other benchmark exercises the lane.

5.4 Scale and onboarding

Onboarding Performance
Metricaelfricealpha-seek-memtest
Git commits35619
Git date range2 days16 days
Documents1631,726
Nodes extracted16,69090,793
Edges extracted32,538302,268
Final belief count (after dedup + supersession)31,86360,641
Scan time~2.5s~5.8s
Full pipeline time----
Scale factor1x5.4x (nodes)
Time factor1x2.3x

The scan phase scales sublinearly: 5.4x more nodes in 2.3x the time. AST parsing is the bottleneck, at 32-38% of scan time. A full-pipeline run on a smaller 10,872-node repo was logged at 6.5s after the v1.2.1 performance fixes, which batched edge inserts and deferred per-belief FTS5 checks during bulk ingestion; the two repos measured in the table above were scan-phase only. Temporal decay was validated on the larger codebase: 2-day-old beliefs score 0.92, 18-day-old beliefs score 0.43, 14-month-old beliefs score ~0.

There is no section 6. It was a discussion section, withdrawn because its A/B comparison rested on artifacts that were never retained, and the numbering stays as it was so existing links still resolve.


7. Failure taxonomy

Failure taxonomy: Memory, Calibration, Behavioral, and Operational failure families

I documented 35 behavioral failures across Claude and Codex and classified them into recurring patterns. Most case studies include a verbatim exchange, root cause analysis, pattern classification, what the memory system should do to prevent it, and a concrete acceptance test with pass/fail criteria. All 35 have an acceptance test. Thirty state it inline; the five briefest entries (CS-001 to CS-004, and CS-020) state theirs in the separate test registry.

Failure Pattern Families
FamilyPattern
Memory Failures
P4Repeated procedural instructions
P1Repeated decisions
Context drift within and across sessions
Calibration Failures
P5Provenance-free status reporting
P7Output volume presented as validation
Result inflation in reporting
Behavioral Failures
P6Correction stored but not enforced
P9Sycophantic collapse under pressure
P10Point-fix without generalization
P11Intent completion gated by permission
Operational Failures
P7Namespace collision across parallel sessions
P8Multi-hop query collapse
Scale-before-validate bias
Extensive Research comic - CS-005

The acceptance suite runs the real SQLite store and retrieval pipeline against a fresh temporary database per test: 62/65 tests pass. The 3 skips cover capabilities requiring behavioral hooks not yet implemented (CS-012: PostEdit hook, CS-024: sycophantic collapse detection, CS-026: permission-gated intent completion).

Case Study to Acceptance Test Mapping (selected)
CSFailureWhat the Test Validates
002Premature implementation push (3 corrections ignored)Locked correction created on first user correction; persists indefinitely
006Correction stored but not enforced (implementation ban violated in new session)Locked prohibition retrieved AND enforced across session boundaries; output gating blocks violations
009Correction lost across session reset ("use B not A")SUPERSEDES edges preserve latest correction; holds across resets
022Multi-hop query collapse (4 agents, wrong machine)All entities identified via graph traversal; correct state aggregated
025Correction not generalized (fix one instance, miss others)Correction applies to pattern class, not just the specific instance

The component counts in the following table were hand-assigned per case study; no machine-readable mapping is retained.

Component Dependency (by case study count)
ComponentCase StudiesPriority
Locked beliefs / L0 behavioral11Critical
COMMIT_BELIEF (git-derived)6High
FTS5 retrieval6High
Triggered beliefs (TB-01-15)6High
Source priors / provenance5High
SUPERSEDES edges4Medium
IMPLEMENTS / CALLS / CO_CHANGED5Medium
Output gating (enforcement)2Critical*
HRR typed traversal3Medium
TESTS / coverage edges1Low
* Output gating covers only 2 case studies but both are severity-critical: CS-006 and CS-016 are multi-session correction violations, the most painful failure class.
Validating the Validation comic - CS-007b

8. What was abandoned (and why)

SimHash clustering was the first attempt at deduplication. It works well on near-duplicate text, which stored directives are not: they are short and semantically dense, and "always use mocks" and "never use mocks" differ by the one word that inverts the meaning.

Mutual information re-ranking was supposed to improve retrieval by scoring candidates on their statistical relationship to the query. In practice it only reordered: PMI cut MRR 12.4% and NMI 4.4% against BM25, with precision and recall unchanged, and BM25 was already near its ceiling at 0.843 MRR (Exp 54).

Global holographic superposition was theoretically promising and it failed. At 775 edges, the representation exceeded its information-theoretic capacity by 7.6x and produced pure noise (Exp 30).

Pre-prompt compilation attempted to pre-compute relevant directives for common query patterns. It performed worse than random selection (23.1% vs. 33.1%, Exp 64) and added no coverage that on-demand FTS5 (69.2%) did not already reach.

Abandoned Approaches
ApproachWhy it failed
SimHash clusteringNot viable for deduplication in this domain
Mutual information re-rankingHurts more than helps in retrieval
Rate-distortion optimizationUnnecessary complexity for marginal gains
Pre-prompt compilationWorse than random selection (23% vs 33%)
Global holographic superpositionCapacity exceeded 7.6x at 775 edges; pure noise
Multi-layer graph expansionSignal diluted to 3.6% of graph at 16K nodes
Autonomous edge discoveryPrecision 0.001, recall 0.005
Zero-LLM classification as sufficient4% precision on corrections (805 found, 32 correct)

9. Conclusion

aelfrice demonstrates that persistent memory for LLM agents does not require embeddings, vector databases, or expensive inference. A pipeline built on FTS5 keyword search, typed knowledge graphs, and entity-index retrieval achieves competitive results across five benchmarks on the v1 substrate while running at roughly 0.7s per warm query on a 19K-node production database (the first call after a cold start is ~10s).

Correction detection runs at 92% accuracy without LLM calls. Vocabulary gap recovery reaches the 31% of directives keyword search cannot. Entity-index retrieval, on the v1 lab substrate, reached 35% chain-valid multi-hop where the envelope this project benchmarked against was 7%, though the shipped public substrate measures 6% on that split (see §12). The confidence tracking loop improves retrieval quality over time.

Two limitations remain. LongMemEval multi-session accuracy is 24.1%, bottlenecked by FTS5's inability to aggregate scattered mentions. Contradiction detection during retrieval does not work yet.

The failure taxonomy and its acceptance tests are the part I would hand to someone else first: a catalog of the specific ways LLM agents fail at memory, each with a reproducible test that blocks recurrence. All code and benchmark adapters are available at github.com/robotrocketscience/aelfrice under MIT license. Research frozen 2026-04-16 at the pre-rename v1.2.1; the underlying experiment data is not part of the public repository.


10. Research breadth

The project drew on multiple fields:

  • Information theory: The information bottleneck (Tishby et al., 1999) framed the compression work, but the 55% token savings came from a type-ratio heuristic; full IB optimization was tested and rejected as worth 0-5% more (Exp 42). Mutual information was tested for retrieval re-ranking and abandoned. Rate-distortion theory was explored for optimal token budget allocation but proved unnecessary.

  • Bayesian inference: Beta-Bernoulli conjugate pairs for confidence tracking. Thompson sampling for the exploration/exploitation tradeoff in retrieval. Calibration measured at ECE 0.066 (target < 0.10).

  • Cognitive architectures: SOAR's impasse-driven substates resemble retrieval failure escalation, CLARION's meta-cognitive subsystem parallels confidence tracking, and ACT-R's declarative/procedural distinction parallels the system's separation of factual content from behavioral constraints. The design borrows structure without inheriting human-like decay and distortion.

  • Bio-inspired optimization: Slime mold network dynamics (Tero-Kobayashi equations) and evolutionary algorithms were both surveyed as candidate mechanisms for dynamic edge creation and pruning. Algorithm sketches were written; neither was prototyped or adopted.

  • Graph theory: Typed knowledge graphs with weighted edges are the structural backbone, and multi-hop traversal over them is what recovers the vocabulary gap.


11. Technical details

  • Source: github.com/robotrocketscience/aelfrice, MIT license.
  • Language: Python with strict typing enforced by pyright in strict mode.
  • Storage: SQLite with WAL mode. The entire memory store is a single file.
  • Dependencies: Minimal by design. No PyTorch, no TensorFlow, no embedding models. LLM classification (~$0.005/session) brings accuracy to 99% and is the recommended configuration.
  • Deployment: MCP server with 19 tools, integrating with Claude Code, Cursor, Windsurf, and other MCP-compatible tools. Also ships as a CLI with 23 commands.
  • Modules: 18 production modules, plus benchmark adapters and scoring scripts.
  • Scale tested: 600 to 90,000+ nodes across five codebases. Largest production deployment: 0.7s average retrieval latency on 19K-node graph.
  • Benchmarks: Four benchmarks (five scored cuts) tested with a contamination-proof protocol. Three contamination modes identified and two incidents caught and documented during development.
  • Test suite: 362 passing tests (self-reported; no run artifact survives) plus 62 acceptance tests (29 files, 1.65s).
  • Experiments: 69 numbered experiments enumerated in the v1 research ledger (IDs run to 111; the ledger's index stopped being maintained at 65), plus 6 benchmark-phase experiments, each with hypotheses and decision criteria written up alongside the result. The hypothesis documents were committed after the runs, so the pre-registration is not independently timestamped. Negative findings documented with the same rigor as positive findings.
  • Case studies: 35 documented LLM behavioral failures, each with verbatim transcripts, root cause analysis, and derived acceptance tests.
  • Version: 1.2.1 (research frozen 2026-04-16)

12. Update: v1.2 to v4.0 (2026)

Everything above documents the v1 research. The system kept evolving and was renamed aelfrice. This section covers what changed through v4.0.0 (June–July 2026).

What the v1 write-up got right

The core architecture held. FTS5 keyword search, typed knowledge graphs, Bayesian confidence tracking, correction detection, and the feedback loop all survived into v4.0 without fundamental redesign. The benchmarks were re-run at v3.0.1, and the scores changed (see Benchmarks re-run at v3.0.1).

What changed

The shipped retrieval pipeline stayed at four lanes. In v3.5.0 they are locked beliefs (L0), an entity-index lane (L2.5), FTS5 keyword search (L1, now BM25F with anchor-augmented scoring, default-on since v1.7), and an optional typed-edge graph walk (BFS, off by default), plus a structural HRR lane (Plate-FFT bind/probe, default-on since v2.1) that answers structural queries such as CONTRADICTS:<belief-id>. Two ideas from the research line landed here in their shipping form:

  • Intention clustering ships as a packing stage. Default-on in retrieve_v2 since v3.0, and live on the production retrieval path since v4.0.0. It replaces the L1 pack loop with a diversity-aware greedy fill that biases the returned set toward distinct graph-connected clusters, so the token budget is not spent on near-duplicates. Locked and L2.5 results are pre-included unchanged.
  • HRR ships as a structural-query lane. The deterministic Plate-FFT codec binds role-filler structure rather than learned embeddings, and surfaces structurally-related beliefs with no LLM call and no vector index.

Confidence stayed Beta-Bernoulli. Each belief carries a Beta-Bernoulli (α, β) posterior, and L1 scoring blends BM25 with the posterior mean. A regime classifier (SUPERSEDE / IGNORE / MIXED / INSUFFICIENT_DATA) ships as a read-only health audit surfaced through aelf health / aelf regime.

Beliefs can cross project boundaries, read-only. v3 added cross-project federation. A project lists peer stores in knowledge_deps.json; peer databases are opened read-only (mode=ro&immutable=1), and any attempt to mutate a foreign belief id is rejected at the API surface (ForeignBeliefError). Shared scopes live under ~/.aelfrice/shared/{scope}/memory.db, and a scope column (project / global / shared:<name>) tracks each belief's visibility.

Wonder and reason turned research into conversation. Two new capabilities let the agent investigate open questions using the memory graph as context. The slash-command syntax (/aelf:wonder, /aelf:reason) makes them look like separate formal operations. In practice they run mid-discussion: the user types something like "please /aelf:wonder about X", and the system launches the research pipeline with the full conversational context already loaded. A wonder query spawns parallel research agents; a reason query builds evidence chains. Both save their findings as beliefs that persist across sessions.

A case study (in the project's internal docs, which didn't make the public-repo cut) documents a real session where wonder + reason produced a marketing strategy for the project itself: parallel research agents, findings synthesized, README rewritten, all grounded in beliefs accumulated over prior weeks.

Updated numbers

v1.2 vs v3.8
Metricv1.2v3.8
Tests362 + 62 acceptance5,632 collected (4,757 test functions)
MCP tools1915
Production modules1891
Retrieval lanes44 (+ BM25F, HRR, entity-index, clustering on by default)
Version1.2.13.8.0

The movement worth reading there is that the tool surface shrank while the internals grew. Four MCP tools went away, the module count went up roughly five-fold, and retrieval did not gain a lane -- BM25F, HRR, entity-index and clustering were all already built, and what changed at v3.8 is that they run by default.

Benchmarks re-run at v3.0.1

The suite was re-run at v3.0.1 (captured 2026-05-13, single-run, with MAB and StructMemEval back-fills added 2026-05-21 to 05-22). The scores moved in both directions. Neither version used embeddings, so that is not the variable. An earlier revision of this section attributed the LoCoMo movement to a retrieval-lane change: the write-up ran a deterministic HRR FFT vocabulary-bridge expansion lane that aelfrice gates off under the #605 ratification. A 2026-07 forensic audit of the archived v1 code and the later controlled ablations revised that attribution; see the LoCoMo row and the paragraph following the table.

v1 write-up vs v3.0.1 re-run
Benchmarkv1 write-upv3.0.1Note
LoCoMo F166.1%40.88%The v1 run used a deterministic, embeddings-free HRR FFT vocabulary-bridge expansion lane; in aelfrice that lane is off (#605). An earlier revision of this row called the lane's removal the proximate cause of the drop. That attribution did not survive a forensic audit: no with/without-HRR isolation was ever captured; the v1 66.1% was itself unstable (the archived v2.2.2 re-run scored 49.1% with the lane intact, attributed in the archive to reader variance); part of the remaining gap is an adversarial-category scoring-heuristic artifact; and a later controlled ablation with the lane restored and fully fed measured ≈0 recall effect. The drop is real; its decomposition is mostly reader variance plus scoring artifact, and the lane's isolated contribution was never measured and is bounded small.
LongMemEval59.0%74.0% (76.8% strict re-judge)Up. Multi-session rose from 24.1% to 56.4%.
MAB single-hop (262K)90% (projection)57%The 90% was a lab projection of a (subject, predicate)-collision SUPERSEDES mechanism not present in the public substrate. 57% sits above the GPT-4o-mini 45% baseline.
MAB multi-hop (262K)60% (projection)6%Also a projection; 6% sits inside the 7% envelope this project benchmarked against.
StructMemEval100%100% raw-retrieval / 57.1, 47.4, 7.0, 0.0% answer-correctnessThe 100% is retrieval recall over the 14-row aggregate; under the canonical answer-correctness judge the per-task scores are location 57.1%, recommendations 47.4%, tree 7.0%, accounting 0.0%.

An earlier revision of this section got the lane story wrong. The v1 write-up hit 66.1% with a deterministic, embeddings-free HRR vocabulary-bridge expansion lane: circular-convolution FFT binding over fixed-seed vectors, seeded from the top keyword hits, traversing typed edges written at ingest (the semantics live in those edges; HRR is the deterministic index over them). aelfrice gates that lane off under the #605 ratification, which classifies HRR-class associative recall as the same semantic-relatedness signal aelfrice delegates to the consuming agent. So this is not embeddings versus no-embeddings (HRR is deterministic FFT, not learned vectors), and it is not a determinism-forced loss (the lane is deterministic); it is a settled scope decision about what retrieval is for.

The earlier text said how much LoCoMo the lane would recover "needs a controlled ablation, not an assumption." Two ablations and a forensic audit later, the answer is: almost nothing. v3.7.0 wired the lane back in behind a flag (#981, use_hrr_expand, backed by an hrr_expand_neighbors cache) and ablated it over LoCoMo: +0.13pp — 2 of 1,540 queries — so it ships off by default. A 2026-07 forensic audit went further: with the lane restored on an edge substrate matching the v1 run's density and firing on 76% of queries, recall stayed flat; no with/without-HRR isolation was ever captured in the v1 records; and the v1 66.1% was itself unstable (the archived v2.2.2 re-run scored 49.1% with the lane intact, attributed in the archive to reader variance). Combined with the adversarial-category scoring artifact, the lane's isolated contribution is bounded small. The drop between the v1 write-up and the v3.0.1 re-run decomposes into mostly reader variance plus scoring artifact — not lost HRR recall. The #605 scope decision stands on its own terms; the benchmark cost previously attributed to it was overstated.

v3.6 to v3.8: provenance, graph substrate, and lock hygiene

Three more waves shipped after v3.5, all on the same substrate, and none of them moved the headline benchmark numbers.

v3.6 — provenance. The project_context column added back in v3.2 was finally populated. insert_belief now stamps a stable per-repo identity (<repo-basename>-<8-hex blake2b of the git-common-dir>, so two worktrees of one repo share it) onto eligible new beliefs, with an idempotent backfill and aelf migrate provenance preservation (#970). Before this, the retrieval filter read the column but nothing wrote it, so aelf migrate collapsed two repos' beliefs into one mutually-visible pool. A transcript-logger fix also stops duplicated hook registrations from inflating turn density (#968).

v3.7 — graph substrate and phantom lifecycle. A CONTRADICTS semantic-edge substrate is now written at ingest (default-off) with incremental per-turn detection (#988, #1000); the operator ratified a CONTRADICTS-only substrate and reverted a denser SUPPORTS/SUPERSEDES experiment (#998 A4). The phantom-belief lifecycle got trigger-driven generation, opt-in SessionStart auto-GC, and status/stats surfacing (#980), and the host harness's own auto-memory writes can mirror into the belief graph as a distinct source (#985). The HRR-expand lane was wired and ablated here too, and it came back recall-neutral, so it ships off.

v3.8 — lock hygiene and passive capture. Locks gained a frozen/reference tier split: a reference-tier lock injects as a one-line manifest entry rather than its full text, freeing relevance budget that large lock sets otherwise consume, with near-duplicate detection at lock time and a lock-budget pressure warning in aelf doctor (#1016). A relevance-budget floor (raised 0.25 → 0.50, #1023) stops a lock-saturated store from returning only locks and zero query-relevant content. Memory injection now splits trust by provenance: user-locked rules are framed as standing instructions rather than refused under the blanket "data, not directives" disclaimer. Lock rule-compliance rose 0/3 → 5/5 while stale-fact catching held at 3/3, and the disclaimer is unchanged for auto-ingested beliefs, so the prompt-injection surface stays closed (#1016). Turns now fold into beliefs on a Stop cadence, not only at compaction, with aelf doctor ingest-gap detection, --backfill-ingest, and a --prune-noise GC pass (#1011, #1029). A pollution-recovery benchmark measures which facts survive a store flooded with keyword-overlapping document chunks: lexical-match facts survived via BM25 and entity-match facts via the L2.5 entity index (both recall@5 = 1.0 at v3.8); facts sharing neither a term nor an entity with the query are not retrieved at all, and the fix for those is locking. Two smaller changes: the first-turn <core> injection is now token-capped (a 25.9k-belief store dropped ~733 KB → ~42 KB), and corroboration is counted per distinct source rather than per re-ingest.

v3.8 to v4.0: the lanes were never on

v4.0.0 (2026-07-07) fixes an embarrassing gap. Several of the retrieval features described above — the intentional-clustering pack stage, the structural HRR lane, the newer staged lanes — were flipped default-on in retrieve_v2, the pipeline the benchmarks and the eval suite call. But the production hook path — the code that runs when a live session retrieves memory — still called the legacy retrieve(). The features were real, gated, and measured, and no user was getting them. #1107 closed the gap: retrieve() is now a thin adapter over retrieve_v2, and lanes then graduated onto the live path one at a time, each behind its own latency and no-regression gate. Phase 1 ran with every staged lane forced off; the equivalence test that accompanied it turned out to assert an identity that held by construction (#1160).

Four lanes made it to production:

  • Temporal spine (#1064). The largest retrieval-coverage gain measured on this codebase, deterministic and embeddings-free: per-session TEMPORAL_NEXT chains written at ingest, plus a lane that walks them from the top keyword seeds and appends chronological neighbours. It reaches gold that shares zero salient terms with the question through chronological adjacency, which is the hole the v3.8 pollution-recovery benchmark left open. Winner's curse is what you get when you pick the best of a batch of candidates and have partly picked the luckiest one, so the winner shrinks on fresh data. Confirmed on a fresh LoCoMo sample: +14.6pp gold-set coverage (0.460 → 0.606; temporal questions +17.2pp, multi-hop +10.4pp), 10× a seeded shuffled control, and the out-of-sample gain exceeded dev (+12.7pp on LongMemEval) — the opposite of winner's curse.
  • Junk demotion (#1096). A mild log-additive rerank penalty for beliefs whose only referential grounding is transient coordination tokens (bare PR numbers, version tags) rather than durable entities (file paths, error codes, symbols). On 118 hand-labeled coordination beliefs, the grounding score separates durable content from ephemeral status at mean 0.56 vs 0.06, lifting the durable-above-ephemeral ranking AUC from 0.48 (posterior-only) to 0.87. Pure demotion, never a promotion: entity-free beliefs are never touched, and well-grounded ones are penalised only marginally.
  • Intentional clustering (#436). The multi-fact pack stage, now actually reaching users: cluster coverage 0.500 → 1.000 on the public ablation with recall non-degrading, p99 latency 0.328 ms against a 5 ms budget.
  • Structural HRR (#152). Marker queries (CONTRADICTS:<belief-id>) now route through the graph lane on the production path. The public ablation is constructed so the answers share no vocabulary with the query: recall 0.000 → 1.000, because keyword search cannot reach a structural answer by definition.

And two lanes deliberately did not. Origin-priority reranking was refuted on LoCoMo (#1013). The failure there was BM25 recall, which reranking cannot fix: the trusted fact never became a candidate. So it survives only as a within-tier tie-break, default-off. And HRR-expand stays off, on the +0.13pp ablation.

The same distrust of recurrence drove the release's biggest scoring fix: retrieval exposure no longer updates the posterior (#1086). Being retrieved had counted as positive evidence, so whatever kept getting retrieved kept rising. On a real 24,883-belief store, junk (session scaffolding, fragments) accumulated ~3× the exposure of clean beliefs and outscored it (mean μ 0.554 vs 0.446). Exposure is now audit-only: the event still lands in feedback history, but the posterior moves only on real signal. The controlled check: a belief retrieved 50 times no longer outranks a belief seen once. Under the old behaviour it inflated to μ 0.857.

The rest of the release, with numbers where they exist: curation grew honest tooling (aelf introspect surfaces per-session posterior, recurrence, grounding, and noise signals; aelf retire / restore make soft-delete reversible end-to-end, #1081); the host harness's own memory files now reconcile into the belief graph by default (#1089); dispatched subagents inherit memory context instead of running blind to the store (#1068); valence propagation finally turned on (#1058 — shipped in v0.1.0, never invoked by any production path until now; mean fan-out 1.7 recipients per explicit event, walk p99 < 1 ms); and Codex joined Claude as a first-class host (#1052–#1055). One more result from the wide-retrieval knobs (#1045): raising the BM25 candidate cap from 50 to 200 with an 8,000-token budget takes LongMemEval-S from 58.8% to 68.6%, past the GPT-4o paper baseline of 60.6%. The multi-hop misses were a recall-cap problem, not a ranking problem, because candidates cap before the pack trim, which leaves budget alone inert.

One research thread closed negative this cycle. A posterior-informed reranking term (ζ) showed a real directional signal on dev data: a three-sample sign test at p ≈ 5×10⁻⁴, though the pooled effect was only +0.013 MRR. The best of ten candidate formulations reached +0.023 MRR on dev; its pre-registered confirmatory run on fresh samples came back −0.001 MRR. The direction was real; the effect size was winner's curse. ζ stays off, the PR is parked, and the thread reopens only on a new signal.

What's still not solved

A full canonical benchmark re-run has not been done since v3.0.1. The v4.0 cycle ran targeted fresh-data measurements instead — the temporal-spine coverage gate on a fresh LoCoMo sample, the wide-retrieval knob on LongMemEval-S, the fully-fed HRR ablation — but suite-wide numbers on the current substrate remain outstanding. A v2 reproducibility harness completes 6 of 11 canonical invocations; the other 5 fail on absent benchmark data rather than on tolerance.

Known limitations that remain from Section 7:

  • Contradiction detection during retrieval still isn't wired into scoring. A CONTRADICTS semantic-edge substrate is written at ingest (v3.7, default-off, with incremental per-turn detection), and as of v4.0 an explicit structural query (CONTRADICTS:<belief-id>) reaches those edges on the production path — but free-text retrieval scoring still doesn't consume them for internal consistency.
  • Correction rates were not re-measured on the v3/v4 substrate.
  • LongMemEval multi-session was 24.1% in v1 and 56.4% at the v3.0.1 re-run; the wide-retrieval knobs later took LongMemEval-S overall to 68.6% (opt-in configuration, not the hook-path default). What the default configuration scores on the current substrate is unmeasured.

Cross-project scoping was solved (v3 shared scopes with content-hash dedup).


Appendix A: benchmark methodology

Any deviation from these protocols invalidates the results. They were written after two contamination incidents during development, and the protocol itself, the contamination verification script, and every benchmark adapter live in the public repository.

Contamination protocol

Three contamination modes were identified during development:

  1. Ground truth in retrieval output. The retrieval JSON carried answer fields, so the LLM reader saw correct answers while generating predictions. That produced the invalid 87.8% LoCoMo score. Nothing about it looked wrong at first: the opening Opus score of 61.6% F1 was plausible. It surfaced only when slow-finishing agents overwrote the merged predictions file and the score came back at 87.8% F1, against a human ceiling of 87.9%. Exact-match analysis at the time showed most batches reproducing ground-truth strings far above what extraction from noisy retrieved context would produce; the per-batch output was not retained. Four further isolation failures turned up: a renamed _ground_truth field, pre-computed prediction and f1 fields, category_name labels leaking evaluation strategy, and no separation between question-context and scoring metadata. All results from this run were retracted. Prevention: Adapter code writes two separate files. A mandatory contamination check (verify_clean.py) scans for 23 banned keys before any reader touches the data.

  2. LLM self-judging with answer visible. Prevention: Generation and judging are strictly separate passes.

  3. World knowledge override. The LLM reader answered from real-world knowledge instead of retrieved context, most often on counterfactual benchmarks. Mitigation: Reader prompts include explicit instructions to use only the provided context. Documented as a known limitation (~17% of MAB failures).

General protocol

Every benchmark run follows these steps:

Step 1: Data acquisition. Download from published source. Verify row counts and field names.

Step 2: Retrieval. Run the adapter in --retrieve-only mode. It writes a retrieval file and a ground truth file, and each test case uses a fresh SQLite database.

uv run python benchmarks/<adapter>.py \
  --retrieve-only /tmp/benchmark_<name>.json

Step 3: Contamination check. Mandatory before any reader touches the data:

uv run python benchmarks/verify_clean.py /tmp/benchmark_<name>.json

Step 4: Answer generation. LLM reader receives only the retrieval file. Never sees ground truth.

Step 5: Scoring. The scorer reads predictions and ground truth. Metrics follow exact published formulas.

Step 6: Reporting. The run report carries exact commands, contamination check output, adapter commit hash, dataset version, reader model, scoring metric, published baselines, and known limitations.

Per-benchmark specifics

LoCoMo (Maharana et al., ACL 2024)

  • Dataset: the 10-conversation locomo10.json release: 5,882 turns and 1,986 QA pairs across 5 categories, counted from the distributed file. The project page reports only per-conversation averages.
  • Ingestion: All 10 conversations through standard onboarding pipeline. Session boundaries preserved.
  • Retrieval: FTS5 + HRR + BFS, 2,000-token budget, batch size 1.
  • Reader model: Claude Opus 4.6.
  • Prompts: Exact LoCoMo protocol prompts. Categories 1/3/4: "Based on the above context, write an answer in the form of a short phrase..." Category 2 appends: "Use DATE of CONVERSATION to answer with an approximate date." Category 5: forced-choice "(a) Not mentioned (b) [adversarial_answer]" with randomized option order (seed=42).
  • Scoring: Token-level F1 with Porter stemming and article removal.
  • Score: 66.1% F1.

MemoryAgentBench FactConsolidation (Hu et al., ICLR 2026)

  • Dataset: HuggingFace ai-hyz/MemoryAgentBench, Conflict_Resolution split.
  • Ingestion: Context chunked at 4,096 tokens using NLTK sent_tokenize and tiktoken gpt-4o encoding.
  • Retrieval (single-hop): FTS5 with triple extraction. SUPERSEDES edges created automatically.
  • Retrieval (multi-hop): Entity-index adapter. 42 regex patterns. 4-hop chaining with breadth cap of 30.
  • Reader models: Claude Opus 4.6 and Claude Haiku 4.5.
  • Scoring: substring_exact_match per the paper. Chain validation for multi-hop.
  • Scores: SH: 90% Opus, 62% Haiku. MH: 60% Opus raw SEM at v1.2.1 after temporal branching; the conservative chain-validated number, 35% for both Opus and Haiku, was measured at the earlier entity-index stage and not recomputed.

StructMemEval (Evaluating Memory Structure in LLM Agents, arXiv:2602.11243; code at yandex-research/StructMemEval)

  • Dataset: benchmark/data/state_machine_location/small_bench, 14 cases.
  • Ingestion: Narrative timestamps (30 days apart per session). Standard pipeline.
  • Retrieval: FTS5 with temporal_sort=True.
  • Disclosure: temporal_sort was developed after seeing the initial 29% result.
  • Score: 14/14 (100%).

LongMemEval (Wu et al., ICLR 2025)

  • Dataset: HuggingFace xiaowu0162/longmemeval-cleaned, longmemeval_s_cleaned split — the 500-question LongMemEval_S set defined in Wu et al., with noisy history sessions removed.
  • Retrieval: FTS5 + HRR + BFS, 2,000-token budget, top_k=50.
  • Judge: Claude Opus 4.6 binary judge (non-standard; paper specifies GPT-4o).
  • Disclosure: Using Opus as judge instead of GPT-4o means the comparison is not apples-to-apples.
  • Score: 59.0% (295/500).

Reproducibility

All benchmark adapters, scoring scripts, and the contamination verification script are in the benchmarks/ directory of the public repository. To reproduce any result:

# Clone and install
git clone https://github.com/robotrocketscience/aelfrice
cd aelfrice
uv sync

# Run retrieval (example: MAB single-hop)
uv run python benchmarks/mab_adapter.py \
  --split Conflict_Resolution \
  --source factconsolidation_sh_262k \
  --retrieve-only /tmp/mab_sh.json

# Verify clean
uv run python benchmarks/verify_clean.py /tmp/mab_sh.json

Scoring for MAB runs through the public adapter (benchmarks/mab_adapter.py, score_multi_answer); exp6_score.py is a lab-only script and is not in the public repository.

Complete per-benchmark commands and adapter documentation are in docs/concepts/BENCHMARKS.md.


Appendix B: literature tables

Prior Art: LoCoMo Benchmark Results
SystemLoCoMoNotes
EverMemOS92.3%Cloud LLM backend; source released under Apache-2.0
Hindsight89.6%Cloud LLM
SuperLocalMemory C87.7%*Single-conversation subset (Conv-30, 81 questions); text-embedding-3-large + GPT-4.1-mini
Zep/Graphiti75.1%Temporal knowledge graph; Zep's own corrected figure
SuperLocalMemory A74.8%Local retrieval, cloud synthesis (60.4% fully local)
Letta (filesystem)74.0%gpt-4o-mini; no specialized memory tools, but files are auto-embedded for vector search
Mem0 (self-reported)~66%Hybrid store
aelfrice66.1%FTS5+HRR+BFS, no embeddings
* Three rows from earlier drafts (Letta/MemGPT ~83%, Supermemory ~70%, an "independent Mem0" ~58%) were removed after a citation audit found no primary source for those scores; the ~58% figure turned out to be Mem0's rerun of Zep, not an evaluation of Mem0. Zep's row uses the 75.14% ± 0.17 figure Zep published after correcting how it had calculated its own LoCoMo score; Zep has since published a higher 80% result separately.
* Scores measured under different conditions and LLM backends. Not directly comparable.
* aelfrice score from protocol-correct run with full input isolation. See Benchmark section for methodology and contamination narrative from earlier invalid runs.
Cross-Benchmark Summary
BenchmarkaelfricePaper BestDelta
LoCoMo (ACL '24)66.1% F1 (v1 substrate; 40.88% at v3.0.1)51.6% GPT-4-turbo+14.5pp
MAB SH 262K (ICLR '26)90% Opus45% GPT-4o-mini+45pp
MAB MH 262K (ICLR '26)60% Opus<=7% (the envelope this project benchmarked against)8.6x
StructMemEval ('26)100% (14/14)*vector stores fail (retrieval-only baseline: 0.00 on state-tracking)-- (*single-run retrieval recall on location/small_bench; a later variance probe returned [0, 7, 100, 100], and under the canonical answer-correctness judge the per-task scores are 57.1 / 47.4 / 7.0 / 0.0%)
LongMemEval (ICLR '25)59.0%60.6% GPT-4o-1.6pp
* MAB = MemoryAgentBench FactConsolidation
* SH = single-hop, MH = multi-hop
* LongMemEval uses Opus as judge (paper uses GPT-4o); comparison carries an asterisk until same judge is used
* MAB MH "chain-valid" score (reader-independent): 35% for both Opus and Haiku, measured against the 47% raw SEM of the 4-hop entity-index run — the remaining 12pp were incidental matches. No chain validation was run on the 60% Exp 6 temporal-branching configuration.
Benchmarks Studied
BenchmarkKey FindingOurs
LoCoMo (ACL '24)gpt-4o-mini with grep and auto-embedded file search = 74% baseline66.1% F1
MemoryAgentBench (ICLR '26)Single-hop: 45% GPT-4o-mini; Multi-hop: 7% ceilingSH: 90%; MH: 60%
LongMemEval (ICLR '25)500 questions, scales to 1.5M tokens59.0% (Opus judge)
StructMemEvalVector stores fail at state tracking100% (14/14)
LifeBenchSOTA at 55.2%Not yet tested
AMA-BenchGPT 5.2 achieves 72.26%Not yet tested