Agent Memory Architecture Beyond Vector Stores
Agent memory needs five lifecycle decisions, not just a retrieval store.

Most teams building production agents are solving a retrieval problem when they should be solving a lifecycle problem, and that single mismatch explains why so many agents degrade in predictable ways as conversations grow. The word "memory" is partly to blame. It names a store, a place to put facts and get them back, and a system built around a store only optimizes two moments: the write and the read. A production agent, as Gaurav Dadhich argues in a paper on agentic context management (arXiv:2607.21503), has to make five distinct decisions at five distinct moments in every single turn, and a store answers none of them on its own.
Those five decisions are the ones a plain store leaves unaddressed. An agent has to decide which of the things just said are worth retaining. It has to decide in what structure to retain them. At each new turn, it has to pick out the small fraction of everything retained that actually belongs in this turn's context. It has to anticipate what the next turn is likely to need before that turn arrives. And when relevant context grows past what the model can usefully hold, it has to decide what gets compressed, summarized, or dropped.
Skip those decisions and specific, recognizable failure modes follow: agents that forget something a user said twenty turns earlier, multi-agent handoffs that contradict each other because two agents hold two different versions of the same fact, hallucinations that come from a context window stuffed past the point of coherence, and a cost curve that climbs with every additional turn regardless of whether anything useful is happening. Dadhich's answer is to name the discipline properly. Agentic Context Management, or ACM, treats this as a lifecycle: deciding what to remember, extracting and structuring it, choosing the right storage for the right kind of data, consolidating it consciously, forgetting what's gone stale while keeping a record of where it came from, and compacting context to fit a budget without losing what matters.
What the memory lifecycle contains
The ACM lifecycle breaks into five primitives, and mixing up any two of them tends to produce its own distinct kind of production failure. Dadhich's paper (arXiv:2607.21503) lays these out as architecting, ingesting, scoping, anticipating, and compacting and consolidation, and a reference implementation built on this decomposition reports strong results on both LongMemEval and LoCoMo under the configuration the paper describes.
Architecting is the decision, made before any data flows, about which storage substrate will handle which memory type. It is not a setup step to get through once; it's a structural commitment that shapes every trade-off the system makes afterward. Ingesting happens at write time: deciding what from a raw conversation is actually worth keeping, and in what form, since a raw chat log is noise rather than knowledge until something extracts and structures it. Scoping happens at every turn: pulling the small slice of everything stored that belongs in this particular context, which is the part of the problem most developers mistake for the whole problem. Anticipating means predicting what the next turn will probably need and loading it ahead of time, a prefetching step that most architectures skip. Compacting and consolidation is what happens when relevant context exceeds the model's usable budget: deciding what gets compressed, summarized, or evicted without triggering a cliff in accuracy.
The cost argument for taking all five seriously is concrete. Dadhich shows that naive context accumulation, where nothing is ever consolidated, grows token cost quadratically as a conversation lengthens. Crude summarization trades that for linear cost, but it buys that improvement at the price of an accuracy cliff once summarization throws away something that mattered. Skipping the lifecycle isn't just a correctness risk: it's a bill that grows on its own.
Memory type versus memory architecture
Most teams that fail at agent memory don't fail because they picked the wrong database. They fail because they never separated two questions that feel like one question: what kind of memory does this agent need, and how should that memory be stored and retrieved. These are independent decisions, and each one fails in its own way when handled badly.
Memory type describes content. Episodic memory covers events: what happened, when, and in what order. Semantic memory covers facts and concepts: what's true about the world or the domain the agent operates in. Procedural memory covers how things should be done, distinct from both facts and events, and it's now showing up in production systems as a third recognized category alongside the other two. Working memory is simpler: it's whatever the agent is actively holding in its reasoning context at this moment.
Architecture, by contrast, is a three-part decision that sits on top of type. First, which storage substrate handles a given type, whether that's the context window itself, a vector database, a relational database, a knowledge graph, or a governed metadata catalog. Second, which retrieval mechanism applies, whether that's injecting the full context, running a top-k semantic search, traversing a graph, or calling a function. Third, whether the agent passively receives whatever context gets injected into it or actively manages its own memory.
Two axes determine which architecture pattern actually fits a given job: accuracy against latency, and governance against freshness. These axes move independently of each other. Making the context window bigger shifts the accuracy-latency trade-off, but it does nothing at all for governance or freshness, because a bigger window doesn't know when a fact became true or when it stopped being true. Confusing memory type with memory architecture is how teams end up trying to solve a governance problem by buying a bigger context window, and wondering why nothing improves.
What vector stores do well
Vector stores earn their place in a production memory system. The mistake isn't using one, it's treating it as the whole system rather than one component among several. The core operation behind a vector store is approximate nearest-neighbor search: embed a query, find the k stored embeddings closest to it, and return whatever content those embeddings point to. That operation is fast, it parallelizes well, and it's well matched to unstructured text. That combination is why it became the default starting point for agent memory.
Vector stores do their best work in a few specific settings: single-session agents working over a fixed corpus, surface-level personalization, and retrieval-augmented generation, where relevant chunks get pulled into a prompt at inference time. None of that is in question. What changes is what happens once the corpus starts growing and the conversation history stops being fixed.
As the corpus grows, ingestion rates climb and query latency degrades with them. Pushing recall higher requires deeper traversal and larger candidate sets, and both of those push latency up further, so the system is fighting itself. A second gap appears when one agent updates a record while another agent is mid-decision: that situation calls for transactional semantics, which similarity search was never built to provide. The deepest problem is staleness, and it isn't something a better embedding model or a smarter index fixes. Two embeddings stored months apart look equally "similar" to a query if their content matches on the surface, because vector search has no staleness flag, no concept of contradiction, and no sense of time passing. Simply appending every interaction to the store eventually produces retrieval noise, dilutes the context with irrelevant matches, and causes latency spikes, and the fix for that, consolidation, is not something vector search does on its own.
Mem0's 2026 benchmark report puts a number on where the ceiling sits. Mem0's selective memory algorithm beats full-context in-process memory on the LoCoMo benchmark on both accuracy and token efficiency. This trade-off between stuffing everything into context and selecting what matters is measurable.
How the five architecture patterns compose across the lifecycle
No single pattern covers the whole memory lifecycle on its own. Production systems combine several, with each pattern assigned to whichever lifecycle stage it's actually good at, rather than competing to be the one system that does everything.
External retrieval over a vector store, handles the ingesting and scoping primitives: it decides what to keep at write time and pulls the relevant slice back out at each turn. The ceiling described above applies directly here, so this pattern alone leaves consolidation and anticipation unaddressed. A second pattern combines episodic, semantic, and state memory with temporal reasoning and fact versioning, adding exactly the consolidation and anticipating capabilities that a vector-only setup lacks, since none of those are things vector search does by itself. Graph-enhanced memory captures relational structure and temporal edges between facts, which closes the consistency gap between concurrent agents and resolves stale retrieval, at the cost of added latency. The MAGMA architecture (arXiv:2601.03236) takes this further by proposing separate, orthogonal graphs for semantic, temporal, causal, and entity relationships rather than one monolithic graph, on the reasoning that entangling those dimensions together limits how interpretable the system is and degrades the accuracy of its reasoning.
The clearest working model of how these layers compose comes from Letta, formerly known as MemGPT, which borrows directly from operating system design. Core memory stays always present in the model's context window, functioning like RAM. Recall memory is searchable conversation history, functioning like a disk cache. Archival memory is long-term storage the agent queries only when it needs to, functioning like cold storage. Each tier maps onto a different lifecycle primitive: developers already understand why an operating system doesn't keep everything in RAM, and the same logic applies to what an agent keeps in its active context versus what it leaves on disk.
None of these patterns ranks above the others in some general sense. The question for a given agent is which lifecycle primitives its failure modes actually require, not which architecture scores highest on a leaderboard.
Graph memory's specific role in the lifecycle (and its real cost)
Graph memory is the right substrate for two specific primitives, consolidation and provenance, and the benchmark evidence shows the cost of using it is real rather than hypothetical.
What a graph adds that a vector store structurally cannot provide includes multi-hop relational reasoning, temporal edges that track when a fact became valid and when it was superseded, entity resolution across separate sessions, and contradiction detection. Zep, built on Graphiti, is one of the more mature production implementations of this idea, offering a bi-temporal knowledge graph for agent memory (Rasmussen et al., 2025) that tracks both when an event happened and when the system learned about it. Mem0's ECAI 2025 paper (arXiv:2504.19413) measured this directly on the LOCOMO benchmark: graph-enhanced memory outperforms flat vector retrieval on relational tasks, at a modest cost in latency, confirming that the trade-off shows up in practice and not just on paper.
MAGMA's case for orthogonal graphs (arXiv:2601.03236) rests on the same logic raised earlier: representing each memory item across separate semantic, temporal, causal, and entity graphs, rather than folding everything into one store, keeps the system interpretable and avoids the reasoning degradation that comes from tangling those dimensions together.
The strongest argument for graph memory in agents meant to run for a long time comes from a provenance result. Quentin Spencer's paper (arXiv:2607.21962) tested a provenance-typed graph, one that tracks where each assertion came from, against planted injection probes designed to smuggle false information into the agent's memory. The provenance-typed graph produced zero unsupported assertions across every planted, non-adaptive probe. Flattened assertional memory stores, by comparison, failed on a subset of the same probes. The boundary that provenance draws around each fact, marking where it came from and what supports it, is what made the graph resistant to injection in a way a flat store wasn't. That resistance comes at the cost of the added latency and engineering complexity graphs require. Graph memory belongs in the architecture for the specific failure modes it solves rather than as a wholesale replacement for simpler stores.
How history length exposes architecture rankings
Short-horizon evaluations tend to rank memory architectures incorrectly, because the architecture that performs best over a few weeks of conversation is often not the one that holds up over months of accumulated history. Most published benchmarks don't run long enough to catch this.
Spencer's paper (arXiv:2607.21962) addresses the problem by building an evaluation instrument from the ground up rather than extracting it after the fact. Standard benchmarks generate conversations first and pull answer keys out of them afterward, a pipeline with documented problems around label errors and contamination. Spencer's method instead emits facts first, each with a validity interval, a volatility class, and a source-channel record of provenance, before any text exists. Five memory architectures were then benchmarked against a no-memory control across two different history horizons using this ground-truth-first method.
The rankings inverted. At the short horizon, a budgeted, curated-map style of memory came out ahead. As history grew longer, that same architecture lost its hold on early content that had been evicted to stay within budget, and its recall fell behind architectures that had made different trade-offs about what to keep. A benchmark run only at the short horizon would have crowned the wrong winner. Any evaluation of a memory architecture needs to say how long the conversation ran before it says which architecture won.

Sources
- Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings
- Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems
- MAGMA: A Multi-Graph based Agentic Memory Architecture for AI Agents


