Persistent Disk Layout for Long-Running AI Agents
Agent state lives across four incompatible memory types that flat storage layouts silently corrupt.

Whether a prototype agent becomes production-grade has little to do with which model sits behind it and comes down to whether the agent can hold onto authoritative, versioned state across sessions without stale overwrites, privilege escalation, or lost continuity. Most agents deployed today start every session from scratch: stateless responders that reprocess context on every invocation, unable to build continuity from one interaction to the next. It is an infrastructure limitation, and it starts with how the disk beneath the agent is organized.
State itself is not a single thing to be stored. Red Hat's architecture survey distinguishes session memory, long-term episodic memory, long-term semantic memory, and procedural memory, and each carries different durability requirements, different retrieval patterns, and different ownership rules. A flat disk layout, the kind where everything lands in one directory tree with no differentiation, collapses all four into an undifferentiated pile. Once that happens, nothing enforces which process gets to write where, and an update from a model, a tool, or a background worker can silently overwrite state that should have been immutable.
Practitioners already reach for an operating-system analogy when they think about this problem, and it is a useful one. Letta's design treats the LLM's context window like RAM and treats persistent storage like a hard drive, letting the agent decide what to page in and out of active context. That analogy only works if the hard drive has a structure worth paging into. An operating system without a filesystem hierarchy, without permissions, without a clear separation between kernel space and user space, is not an operating system anyone would trust with real workloads. The same standard applies to agent disk layout. It is a control-plane decision, on par with deciding who can escalate privileges and who can read a secret, not a detail to leave to whichever directory structure a framework happens to default to. The sections that follow build that control plane zone by zone.
What an agent accumulates on disk
A long-running agent does not produce one kind of artifact. It produces at least four categories of state, each with incompatible durability, ownership, and access requirements, and conflating them is the root cause of the layout failures described above.
Episodic memory is the record of what happened: conversation turns, task outcomes, errors encountered across prior sessions. This category is append-only by nature. Once a session concludes, its record needs to survive sleep cycles and redeployments without being altered after the fact.
Semantic or factual memory is different in kind. It holds extracted entities, user preferences, and domain facts, and unlike episodic records, these are mutable and need versioning, because facts get superseded rather than simply added to. MinnsDB's design illustrates the requirement concretely: every fact is stored as a graph edge carrying a valid_from and valid_until timestamp, so that when a user states "I moved to New York," the system invalidates "lives in London" and cascades that invalidation to dependent facts, rather than letting the old and new facts sit side by side as contradictory duplicates. A layout that only supports appending, with no mechanism for marking a fact superseded, cannot host this kind of memory correctly no matter what database sits on top of it.
Working memory artifacts make up a third category: intermediate files, tool outputs, and scratch data generated in the middle of a task. These are ephemeral within a single session, but many agent workflows pause and resume, and a working-memory artifact that gets discarded between a pause and its resume undoes whatever continuity the rest of the architecture was built to preserve.
Credentials and authorization tokens form the fourth category, and they carry the highest sensitivity of anything on the disk. OAuth tokens, API keys, and other secrets need a dedicated zone with access restricted to a credential manager, never exposed to arbitrary tool calls. Alongside credentials sit logs and audit trails: immutable records of which tool the agent called and what it returned. These need to be append-only and agent-write, never agent-rewritable, because a log an agent can rewrite becomes closer to a diary the agent edits after the fact.
Four categories, four different sets of rules. That mismatch is what a flat directory structure erases, and a deliberate zone design has to restore it.
Structuring the disk into deliberate zones with defined ownership
The right disk layout maps each of those artifact categories onto a zone with an explicit owner, a defined access mode, and a stated durability contract. The layout functions as an access-control manifest, not a filing convention, and that shapes whether the design holds up once an agent is doing real work under load.
A workable zone sketch looks like this. An episodic zone, something like /memory/episodic/, is agent-write and append-only: no tool call ever overwrites an entry there, and the zone's contents survive a snapshot-restore cycle intact. A semantic zone, /memory/semantic/, is agent-write but with versioned supersession built into the structure itself, so that old facts are archived rather than deleted and a write cannot silently clobber the prior value. A workspace zone, /workspace/, is both agent-write and tool-write, ephemeral within a single task but persisted across a pause and resume, and it is this zone that gets mounted as a block volume in multi-tenant deployments. A credentials zone, /credentials/, is writable only by the credential manager and readable by the agent solely through a controlled interface, never directly accessible to tool calls or subprocesses, because the most common privilege-escalation path in agent systems is an agent that can write to its own credential store. And a logs zone, /logs/, is agent-write, append-only, and immutable once written, forming the audit trail that lets a reviewer reconstruct agent behavior without depending on the agent's own account of what it did.
Ownership semantics carry more weight than the path names themselves. Two agents sharing a host need zone ownership enforced at the filesystem level, not by convention, and at that point the isolation layer beneath the layout becomes load-bearing rather than incidental. A convention such as a note telling an agent not to write to a directory is not an access control. A convention is not a boundary, and an agent that generates its own code has no particular reason to respect either.
The versioned-supersession requirement inside /memory/semantic/ is not something a database bolted on top can supply by itself. It has to be reflected in the directory structure: a write to a fact produces a new timestamped file and archives the version it replaces, rather than overwriting in place, so that rolling back a bad update never requires a full database restore. That structural discipline is what turns a directory into a zone with a durability contract, rather than just a folder with a name that sounds like one.
The isolation layer beneath the disk and the zone design
A zone design enforced only by path convention survives right up until the moment an agent executes code, installs a package, or spawns a subprocess. At that point, isolation shifts from an architectural preference to a security boundary, often with a named CVE sitting on the wrong side of it.
The distinction between containers and microVMs is not a matter of taste. Containers share a host kernel that runs tens of millions of lines of code, so a single kernel vulnerability compromises every container on the platform at once. AI agents make this exposure worse than it would be for a conventional application, because unlike software whose behavior is predictable and auditable in advance, agents autonomously generate code, recursively fork processes, and experiment with memory-mapped I/O in ways that can trigger kernel bugs no security team has ever had to think about before.
A zone design enforced only by path convention fails the moment an agent executes code, installs a package, or spawns a subprocess, at which point the isolation question is no longer architectural preference but a security boundary question with a named CVE on the wrong side of it. Whatever zone design sat inside that container, however carefully the credential store was separated from the workspace, none of it mattered once the boundary around the container itself failed. A zone map is only as trustworthy as the wall around it.
The rule of thumb that emerged from the 2026 practitioner community draws the line cleanly: agents that only process text can run safely inside containers, but any agent that executes code, installs packages, or reaches the network needs a microVM. Firecracker exists to serve exactly that requirement. Infrastructure research describes it as purpose-built for AI agents, developer platforms, and multi-tenant workloads, booting in well under a second with minimal memory overhead, providing hardware-level isolation, and allowing operators to pack thousands of microVMs onto a single machine. Inside that model, each agent's disk zones are enforced by the boundary of the virtual machine itself, not by filesystem permissions an agent could potentially find a way around.
AWS Lambda MicroVMs, launched in June 2026, brought this pattern into mainstream production use. Each session runs inside its own dedicated microVM with no shared kernel and no shared resources between users, with up to 32 GB of disk available per microVM as a ceiling rather than a default allocation. The zone design described above maps directly onto that per-session disk: episodic, semantic, workspace, credentials, and logs all fit within the allocation for a single isolated session, with the isolation boundary guaranteeing that one session's zones cannot bleed into another's.
GPU passthrough into microVMs remains slow and fragile across driver versions, so the pattern that actually works in practice is running CPU-bound agent code inside the sandbox and calling out to a separate GPU-backed inference service over an authenticated network allowlist. That has a direct layout consequence: the credential for that outbound call belongs in /credentials/ and must stay unreachable from the agent's own tool execution environment.
Snapshot-restore economics and the layout contract
Snapshot-restore is the mechanism that makes persistent, per-user agents economically workable. A disk layout that was not designed with snapshot semantics in mind will either lose working-memory state on every pause or accumulate storage costs no one is watching.
The speed of restore is what makes the pattern practical rather than theoretical. On self-hosted Firecracker fleets, restoring from a baked snapshot happens fast enough that an agent can be torn down the instant it goes idle and rebuilt before the next user interaction arrives, with no perceptible delay. Research prototypes are pushing that floor lower still: Zeroboot claims sub-millisecond restores using copy-on-write Firecracker snapshots, an indication that restore latency has not finished falling.
The layout has to be built to be snapshotted correctly for that speed to pay off. The /workspace/ zone has to be explicitly included in the snapshot, or mounted as a persistent block volume, because if it is left sitting on ephemeral tmpfs, a pause-resume cycle silently erases mid-task working state, which is precisely the continuity failure the entire zone design was meant to prevent. The multi-tenant pattern that solves this mounts a per-workspace block volume into the microVM filesystem at /workspace/, and block volumes are generally preferred over network file systems for agents doing heavy file I/O, since they carry lower overhead. The zone design from earlier in this piece maps onto that mount point without modification.
Forking deserves to be treated as a first-class operation rather than an edge case. Same-host forks complete quickly by mapping guest memory as MAP_PRIVATE and reflinking the root filesystem on XFS, so branching an already-warm agent's state costs only metadata rather than a fresh package install. That means the zone layout itself can be forked to run parallel task branches without duplicating the entire disk for each one, which matters enormously for any workload that explores several approaches to a task before committing to one.
The economics carry a trap that is easy to miss. Pausing an agent typically saves both filesystem and memory state and stops the compute clock, which is the entire point of the mechanism, but the saved snapshot keeps consuming storage until it is explicitly deleted. A fleet that pauses thousands of sandboxes and never reaps the resulting snapshots accumulates storage cost that stays invisible on a compute bill anyone is watching closely. The layout has to include a reaping policy for the /workspace/ zone, and that policy is as much a part of the design as the mount point itself, not an operational afterthought bolted on once storage bills start climbing.
Where the memory framework layer sits
Memory frameworks operate above the disk layout, not instead of it. They implement the retrieval and versioning logic that runs against /memory/episodic/ and /memory/semantic/, but they inherit whatever ownership and access semantics the underlying layout actually provides. A framework layered on top of a flat, unzoned disk can implement versioning logic in its own code while providing no real enforcement, since a different process can write directly to the same files and bypass the framework.
Letta, the framework that grew out of the MemGPT research, treats core memory blocks as always-in-prompt, the RAM layer, and archival memory as agent-callable storage, the disk layer. That tiered structure maps directly onto the episodic and semantic zone split described earlier, and it presupposes a durable, structured backing store that the framework itself does not provision. Letta can enforce its own rules about what belongs in active context, but it depends entirely on the disk beneath it to actually keep archival memory durable and properly separated.
Mem0 takes a different approach to the same problem. It extracts and stores facts across multiple memory scopes, user-level, session-level, and agent-level, using an accumulation model that only adds new facts, with no built-in update or delete and no version control for existing memories. Its hybrid retrieval, combining vector search with metadata filtering, operates entirely on whatever the disk layout exposes to it. If the layout beneath Mem0 has no mechanism for marking old facts superseded, that gap in the framework's own model is a gap in the deployed system, not just a documented limitation.
Cognee runs a Remember, Recall, Improve, Forget pipeline over a graph-vector hybrid store, and its research shows graph-enhanced memory retrieval achieves substantially higher accuracy on contextual queries than plain retrieval-augmented generation. The Forget stage places real demands on the layout: deleting a memory has to happen without corrupting the audit trail sitting in /logs/, so the logging zone needs to record that a deletion occurred even as the memory itself disappears from /memory/.
Laying these pieces side by side reveals the mismatch to avoid. A framework that implements temporal supersession, the pattern Zep and MinnsDB both follow, sitting on top of a /memory/ zone with no append-only enforcement is a framework whose own invariants can be violated by a direct filesystem write that never goes through the framework at all. The layout and the framework have to enforce the same contract, or the framework's guarantees are only as good as the discipline of whichever process happens to be touching the disk at a given moment.
Credential zones, log zones, and the two most common agent failure modes
When two agents share a host, zone ownership must be enforced at the filesystem level rather than by convention. Most agent security failures and most audit failures trace back to the same root cause, which is that these zones were never given structural enforcement in the first place.
A credentials zone that restricts write access to a dedicated credential manager, and restricts agent access to a controlled read interface rather than direct file access, closes off the most common privilege-escalation path in agent systems: an agent that can write its own credential store can mint itself new permissions, and no amount of prompt-level instruction prevents that once the filesystem allows the write. The outbound credential for a network call to an inference service has to live in /credentials/ and stay unreachable from the tool execution environment directly, because a credential reachable by tool calls is no longer meaningfully separated from the rest of the agent's operating surface.
A logs zone that is agent-write, append-only, and immutable once written closes the parallel gap on the audit side. An agent's own account of what it did is not a reliable record if the agent can rewrite that account after the fact. The audit trail only holds value if it survives independently of the agent's memory, its incentives, and any errors it might later want to paper over. Treating /credentials/ and /logs/ as structurally enforced zones, rather than directories an agent happens to be told not to touch, is what separates a disk layout that merely looks secure from one that actually is.


