Est.

Credential and Secret Management Inside Agent VMs

AI agents break six core assumptions that traditional credential management relies on.

Contributing Editor · · 11 min read
Cover illustration for “Credential and Secret Management Inside Agent VMs”
Persistent Agents · October 6, 2026 · 11 min read · 2,408 words

A traditional application receives its credentials once, at startup, from a short and predictable list: a database password here, an API key there, all fixed before a single request arrives. An AI agent breaks that model because it decides at inference time, not at deploy time, which APIs to call, which databases to query, and which external services to reach, so no static review of its code can tell you in advance what credentials it will need or when. This matters because the agent reasons inside the same context window that holds its working memory, so tool descriptions, API error messages, and documents pulled in from outside all sit next to whatever the agent already knows. If a failed authentication attempt echoes a key back in its error message, an attacker watching the agent's inputs and outputs can now reach that key. One agent runtime commonly holds OAuth tokens for a cloud provider, a handful of SaaS tools, a production database, and a messaging platform all at once, so a single compromise can hand an attacker access to everything that agent can touch. Chains of delegated agents make the accounting worse still: if a downstream agent's MCP server is compromised, it may be holding credentials passed down through an entire chain of agents above it, and no conventional audit log was built to trace that path back to its source. Taken together, these are six separate ways agents break the assumptions traditional credential management was built on: credentials acquired dynamically at runtime, credentials sitting inside the reasoning chain itself, non-deterministic behavior from one run to the next, sessions that run far longer than a typical request, delegation across multiple agents, and the aggregation of many credentials into one runtime.

Diagram: Six Ways Agents Break Traditional Credential Management. Visualizes: Visualize the six distinct failure modes that AI agents introduce to credential management, as enumerated in the article: (1) credentials acquired dynamically at runtime…

Tooling layers can't contain a prompt-injected agent without hard process boundaries

Guardrails built at the application layer, whether they take the form of environment variable checks, access-control libraries, or callback hooks, all depend on one assumption: that untrusted code never runs with a live credential sitting next to it. Once that assumption fails, none of those guardrails matter, because there is no real process or kernel boundary, so nothing stops that code from reading whatever the rest of the process can see. Google Cloud's default configuration shows what this looks like at the infrastructure level: Compute Engine assigns a default service account to new VMs, and any code running on one of those machines can request a token from the metadata server at 169.254.169.254 with no API key and no credential file required. This failure appears at the framework level too. A LangChain vulnerability, tracked as CVE-2025-68664, let an attacker's injected prompt trigger LangChain's serialization system and pull out an environment variable named in the attack payload, a path that only existed because secrets_from_env was enabled by default before the patch and because the agent had direct, unmediated access to its own runtime environment. LangChain does offer a way to intercept tool calls for governance, through its AgentMiddleware system and the wrap_tool_call hook, but that mechanism has to be turned on by the developer to run. Security that depends on someone remembering to opt in is security that gets skipped the first time a deadline gets tight. If the compute layer itself leaks, no amount of application-layer logic secures anything. A sandbox has to start with no ambient credentials at all, no IAM role, no API key baked into an image, and no reachable metadata endpoint, because the instant untrusted code runs alongside a live credential, that credential belongs to whoever controls the code. Isolation, understood this way, is the condition that has to hold before least-privilege injection, short-lived leases, audit logs, or killswitches mean anything more than good intentions.

What VM-level isolation provides over containers and serverless

A hard process boundary means something specific: separation enforced by the CPU's own virtualization features, with each workload running its own guest kernel, rather than separation enforced by namespaces and control groups inside a kernel shared across tenants. Micro-VM isolation, as implemented in projects built around this approach to virtualization, therefore sits on a fundamentally different footing from container-based sandboxing. Firecracker itself is a building block, not a finished security posture. How safe an agent sandbox built on Firecracker is comes down to how the surrounding platform configures the guest image, the filesystem mounts, the network path, secret injection, logging, and the lifecycle controls that sit around the microVM, all choices made above the virtualization layer. Five specific risks have to be closed off inside any such sandbox: filesystem access reaching outside the directories an agent should touch, network egress to arbitrary destinations, exposure of the host kernel's syscall surface to code the agent runs, data leakage between tenants sharing the same host, and secrets leaking out through environment variables or the /proc filesystem. Of these five, egress control carries the most weight on its own. A default-deny network policy neutralizes the most common exfiltration path outright: if a sandbox can reach the open internet, any injected code that manages to read a secret can simply send it to a server the attacker controls, but routing all outbound traffic through a proxy with an explicit allowlist of permitted hosts turns that same attack into a blocked connection that shows up in a log. Without that one control in place, every other protection built on top of it is closer to a formality than an actual defense. When you get the platform configuration right here, you also make the credential-handling patterns described next something you can rely on, not just something that looks good in documentation.

Secret flow into an isolated agent VM

Diagram: The Credential Exposure Spectrum: From Raw Secrets to Workload Identity. Visualizes: Show a linear spectrum of five secret-injection patterns ordered from highest to lowest raw credential exposure, using the exact stages the article names…

Once a sandbox has no ambient credentials and no open egress path, the question of how secrets get into it becomes a design choice rather than an afterthought, and a handful of patterns are available depending on how much raw exposure a team is willing to accept. You should never bake API keys into a sandbox template at all; that is the simplest starting point. Instead, fetch the needed keys from a secrets store at the moment a sandbox is created, inject them as environment variables, and include only the specific keys that session's task requires, putting the least-privilege decision at the point of injection. A stronger pattern inserts a proxy between the agent's reasoning engine and the outside world: when the agent needs to call an external API, it goes through the proxy, which fetches the real credential from a vault, makes the call, and hands back only the result, so the credential itself never passes through the agent's own code or context. The most complete version of this idea takes the credential out of the agent's reach. The AgentSecrets approach, built around the agentsecrets project, has no get() method. There is no function call that returns a raw secret value into application code; credentials are instead injected directly into HTTP headers or connection strings at the transport layer, so that neither the agent's own logic nor a prompt injection attack targeting that logic has any way to read a secret value, because the value is never present anywhere the agent's code executes. Dynamic secrets address the problem from a different angle: rather than storing a long-lived database password, a platform like HashiCorp Vault generates a temporary credential with a short time-to-live the moment an agent requests access, and automatically revokes it once that window closes, so a successfully stolen credential is only a brief window. The furthest point on this spectrum does away with the idea of a stored secret. SPIFFE and the broader workload-identity model issue a workload a cryptographically verifiable, short-lived identity document instead of a secret to protect, and that identity document is what the workload uses to obtain credentials just in time. This model's operating principle: use identity wherever you can, and fall back to secrets only where no alternative exists. None of these patterns hold up without the isolation layer underneath them: a proxy or a transport-layer injection scheme is only as strong as the boundary preventing code from reading the process environment directly, which is what VM-level isolation is built to guarantee.

Session-scoped leases, audit trails, and killswitches as the operational layer above injection

Getting secrets into a sandbox safely is not enough; what happens to them while the agent is running has to be governed too, through time-bounded controls that compound on top of isolation. A session-scoped lease ties a credential to a short time-to-live: the agent requests access to a named secret, uses it for the duration of that session, and the lease is automatically revoked when the session ends or the time-to-live expires, so a leaked lease is a narrow window. One open-source secrets project implements exactly this pattern, storing secrets with Age encryption so they are never written to disk in plaintext, issuing TTL-based session leases that expire on their own, supporting auto-rotation hooks, and including a killswitch that, when triggered, revokes every active lease, rotates every credential, and wipes the store in one motion. The OWASP Secrets Management Cheat Sheet explains why this kind of automation has to be built in rather than handled manually: manual rotation is a place human error creeps in, restarting an application does nothing on its own to revoke a credential that's already been stolen (it stays usable until the backing service itself expires or revokes it), and rotating an encryption key can trigger a full or partial re-encryption of stored data, all of which argue for automated lifecycle management rather than a rotation schedule someone has to remember to run. Granular revocation depends on granular issuance in the first place: if credentials are issued per task rather than per agent, revoking one task's grant shuts down exactly that task's access without touching any other agent running in the same fleet. But none of this logging means anything if the boundary underneath it leaks. An audit trail built around a secrets manager only records what passes through that secrets manager, so if an agent can exfiltrate a credential through an unrestricted network path, the theft never touches the system meant to be watching it, and the log stays empty. Closing that gap is what VM-level egress control is for: it forces every outbound connection, including an attempted exfiltration, through a path the audit system can actually see.

MCP servers as a new credential concentration point

The Model Context Protocol has become one of the fastest-growing ways agents connect to outside tools, and it has also become the single densest concentration of credentials in many agent architectures. A single MCP server commonly holds tokens for GitHub, Slack, Google Workspace, a production database, and a cloud provider all at once, which makes it an unusually attractive target relative to its size. Research by Trend Micro found hundreds of publicly exposed MCP servers with no client authentication at all, each one exposing many tools to anyone who found it, so this exposure is not a hypothetical risk analysts are preparing for but a condition that already exists at scale. Most of these servers still rely on static, long-lived API keys rather than anything short-lived or scoped, so moving them to OAuth 2.0 or to dynamic, auto-expiring secrets is the single most important fix teams running them can make today. The protocol itself moved to address part of this problem in its June 2025 revision (2025-06-18), which tackled the specific class of attacks built around token relay: tokens are now bound to the specific server they were issued for, so a token obtained for one server cannot be replayed against a different one, and the spec now explicitly forbids an MCP server from forwarding a token it received as though it were its own. For a fleet of agents serving many users, this means an MCP gateway has to track token state per user rather than per deployment: each person authorizes access to their own resources, the gateway stores that person's refresh token, keeps an encrypted record mapping their identity to an access token along with its expiry and scope, and rotates that access token on its own schedule. Keeping those access-token lifetimes as short as OAuth 2.1's security guidance allows, paired with transparent refresh and the ability to revoke immediately, applies the same lease principle already seen at the credential layer, just moved up to the protocol level. A documented pattern of attack against MCP, acknowledged by CrewAI among others, has malicious servers hide instructions inside their own tool descriptions, turning the very text an agent is supposed to read in order to use a tool into the mechanism of the attack. That pattern is a direct argument for the isolation boundary discussed earlier: if the MCP server itself runs inside an isolated VM with default-deny egress, a successful injection still cannot get harvested credentials out, because the network path needed to send them anywhere is simply not there.

Snapshot-restore scheduling and the economics of isolated agent VMs at scale

The strongest objection to running every agent in its own isolated VM is cost: a dedicated micro-VM per agent sounds expensive next to a shared container pool, and at real fleet sizes that difference adds up fast. Agents spend most of their actual lifetime idle, waiting on a model response, a tool call, or a human to come back with the next instruction. Idle time, not active compute, drives the cost of running a large fleet of isolated agents. The fix that makes per-agent isolation affordable is to stop paying for compute during that idle time: capture the full state of a VM in a snapshot, tear the VM down to free up the host it was running on, and restore that exact state when the agent is needed again. So a given amount of host capacity can support a much larger number of agents than keeping every VM running continuously ever could, because idle agents pay for stored state on disk rather than reserved compute on a host. It also introduces an obligation that has to be handled deliberately rather than assumed away: a snapshot taken while an agent's VM holds live credentials is itself a copy of those credentials sitting somewhere, and the same discipline applied to the running VM, encryption at rest, short-lived leases that expire on their own schedule, and access controls around who can trigger a restore, has to extend to every snapshot the scheduling system creates and stores.

Sources

  1. Secrets Management - OWASP Cheat Sheet Series

More in Persistent Agents