Classic painting used as the article cover
← Back to blog

PLAYBOOK

Provisioning AI Agents at Scale

A repeatable six-stage path to production for agent fleets, with the identity, role and quota ceilings on AWS, Google Cloud and Entra that decide what scale actually means.

•Jul 6, 2026•Updated Sep 29, 2026•15 min
ProvisioningLifecycle ManagementEngineering

TL;DR

Provisioning breaks at scale not because any one step is hard, but because teams skip, reorder, or improvise the same six steps differently every time.

There is also a hard ceiling nobody checks until they hit it. AWS gives an account 1,000 IAM roles by default and 10,000 as the maximum, with 20 managed policies per role. Microsoft Entra caps an Agent Blueprint managed by a non-Microsoft platform at 250 Agent Identities. Google Cloud allows 1,500 principals in a single resource's allow policy. Those numbers, not your review board, define what "at scale" means for your fleet.

The instinct is to fix inconsistency with more review. That's the wrong lever. What breaks is repeatability: the same six stages, in the same order, every time.

We are not arguing for one uniform gate on every agent. A read-only summarizer and a payments agent should not queue behind the same committee. The argument is that the stages are fixed and the scrutiny is variable, rather than the reverse.

Start by counting. How many agent identities exist, which are shared, and how close is any account to a quota you can't raise? That number is the honest baseline for everything below.

Overview

Provisioning one AI agent is easy. Wire it up, hand it a credential, point it at a system, watch it work. Provisioning hundreds is a different problem, and most of the difficulty has nothing to do with the model.

At scale the bottleneck is process, and underneath it is arithmetic. Each agent needs an identity, an owner of record, narrow permissions, managed credentials, monitoring, and a review gate sized to its risk. Done once that's a checklist. Done a thousand times it becomes a backlog that throttles adoption or a free-for-all nobody can account for. And every identity you mint consumes a quota with a documented, sometimes unraisable limit.

How a credential is scoped, stored and rotated is its own discipline, covered elsewhere in this series. This piece is the pipeline that gets an agent from nothing to a governed system, plus the platform limits that constrain it.

NIST's AI Risk Management Framework puts the first stage in GOVERN almost word for word, asking that "mechanisms are in place to inventory AI systems", with a companion subcategory on decommissioning. Provisioning is where that inventory gets written or doesn't.

Provisioning is the step that decides whether the rest of the lifecycle is even possible. Skip it, or do it inconsistently, and offboarding, auditing, and incident response all inherit the gap.

Why It Breaks at Scale

Ad-hoc provisioning doesn't fail because a step is hard. It fails because the steps get skipped, reordered or improvised differently by each team, and at low volume that's invisible: five agents, five shortcuts, everyone still remembers which is which. The tenth team doesn't inherit the first nine teams' shortcuts. It invents its own.

That's how an enterprise ends up with no single answer to "how many agents do we have." The two failure modes are predictable:

  • Chaos. Agents launch with no record, no scoped identity, no monitoring. Adoption is fast and every agent is a blind spot.
  • Bureaucracy. Every agent needs a bespoke review thread, a manual credential request and a custom permission set. Governance exists and is slow enough that teams route around it.

OWASP's Top 10 for Agentic Applications, published in December 2025, names the security shape of the first one. ASI03 is Identity and Privilege Abuse, on the grounds that for agents "leaked credentials let them operate far beyond their intended scope". ASI10, Rogue Agents, is the chaos case once an unregistered agent outlives everyone who knew about it. A good process resolves both by giving teams one well-lit path to production.

1. Identity

2. Registration

3. Authorization

4. Credentials

5. Monitoring

6. Review tier

Active in production

Figure 1 — The six-stage provisioning pipeline. Nothing here is optional; the point is that it runs the same way for agent one and agent one thousand.

The Platform Ceilings

Before designing the pipeline, find out what your identity platform will let you have. These are the limits we could confirm in vendor documentation:

PlatformLimit that bites firstValueRaisable
AWS IAMRoles per account1,000 default, 10,000 maximumYes, auto-approved
AWS IAMManaged policies attachable per role20 default, 25 maximumYes, to 25 only
AWS IAMSize of one customer managed policy6,144 charactersNo
AWS STSAssumeRole and friends, per account per Region600 requests per secondBy support ticket
Google Cloud IAMPrincipals in one resource's allow policy1,500No
Google Cloud IAMKeys per service account10No
Microsoft EntraAgent Identities per non-Microsoft Agent Blueprint250No
Microsoft EntraDirectory objects per tenant50,000, or 300,000 with a verified domainBy contacting support

Table 1 — Identity and authorization ceilings, from each vendor's own quota page. The unraisable ones are the interesting column.

Two deserve comment. The 20-managed-policies-per-role cap punishes fine-grained design: give every tool its own policy and a capable agent runs out of attachment slots long before it runs out of things to do. Consolidate into fewer, larger policies and you hit the 6,144-character limit on one. Past a point, per-agent roles stop scaling and you need a role per archetype with conditions doing the narrowing.

Entra's numbers are newer. A Blueprint managed by a platform Microsoft doesn't own is capped at 250 Agent Identities, a restriction that "does not apply to Blueprints managed by Microsoft-owned platforms, such as Foundry and Copilot Studio". A new tenant is also held to 600 directory objects for two days, long enough to break a proof of concept that bulk-provisions on day one.

Read your own quota pages before you design the identity model, not after. A design that needs 2,000 roles in one AWS account is a support ticket; one that needs 400 managed policies on a role is a rewrite.

1. Identity First

Provisioning starts with identity, because everything downstream depends on it. Every agent needs a distinct, first-class identity: not a shared credential, not a service account borrowed from a human team, not a token passed around between projects.

This is not a house opinion. NIST SP 800-53's IA-9 requires organizations to "uniquely identify and authenticate" system services and applications "before establishing communications with devices, users, or other services or applications". AC-2 treats shared accounts as something to justify rather than assume. Google Cloud says it with less ceremony: "Create dedicated service accounts for each application, and avoid using default service accounts."

The consequence is evidence. If three agents share a service account you can't scope them differently, and a log line says only that "the account" acted. Give each its own identity and you can attach permissions, trace actions, and retire it without disturbing anything else.

A workable agent identity record is small but precise:

agent_id:      agt_fpa_variance_07
display_name:  FP&A Variance Explainer
owner:         finance-platform@company
created_by:    a.okafor@company
identity_type: managed_agent_identity
status:        provisioning

The identity is the agent's own, not the same as the user who created it. The agent acts on behalf of users but is accountable as itself, which is what lets you answer "what did this agent do" without untangling it from the humans nearby. NIST's zero-trust document assumes that shape: "client identity can include the user account (or service identity)".

2. Registration

Once an agent has an identity it should be registered: captured as a record the enterprise can govern. That's the difference between an agent that exists and an agent the organization knows about. At minimum:

  • Owner: a named person or team accountable for the agent's behavior. AC-2 calls this assigning account managers; it's the field most often left blank.
  • Purpose: what the agent is for, in one plain sentence.
  • Connected tools: the systems, APIs and data sources it will reach.
  • Requested permissions: the specific actions it intends to take.
  • Lifecycle state: proposed, provisioning, active, suspended or retired.
  • Risk level: the classification that drives how much scrutiny it gets.

Take a hypothetical HR onboarding agent that drafts welcome packets and creates calendar invites. Its record names the People Operations platform team as owner, lists the HRIS and calendar APIs, requests read on new-hire records and write only on drafts, and is moderate risk: personal data, no irreversible action. That one record tells a reviewer almost everything.

Registration has an expiry dimension people forget. AC-2 requires disabling accounts that have expired or are "no longer associated with a user or individual", and asks that account management align with personnel termination. An agent whose named owner left in March is already out of compliance.

3. Scoped Authorization

After registration comes authorization, where most of the safety actually lives. The principle is old and stated plainly in AC-6: least privilege, "allowing only authorized accesses for users (or processes acting on behalf of users) that are necessary to accomplish assigned organizational tasks". The parenthesis is the part that matters. An agent is a process acting on behalf of a user, so the control already covers it.

The failure pattern is granting by category instead of by action. "Access to the data warehouse" is a category. "Run read-only queries against the sales schema" is an action. The first quietly includes dropping tables.

  • An analytics agent summarizing weekly metrics reads from the warehouse and never modifies source data or pipelines.
  • A procurement agent creates draft purchase requests and never releases payment.
  • An IT operations agent reads system status and assigns tickets, and never restarts production services unattended.
  • A legal-intake agent tags and routes contracts, and never signs or sends them.

Enforce this at provisioning time rather than hoping the agent behaves. Authorization is a structural guarantee: if the permission was never granted, no prompt, jailbreak or model mistake produces the action. SP 800-207's third tenet pushes further, granting access "on a per-session basis" with trust evaluated each time rather than inherited.

Default deny, then add. Start every agent with no permissions and grant each action explicitly against its registered purpose. It's far easier to widen a scope on request than to discover, after an incident, that an agent quietly had more reach than anyone intended.

4. Credentials

Authorization decides what an agent is allowed to do. Credentials let it reach the systems where it does it, and they belong inside provisioning rather than bolted on afterward. Four rules apply regardless of domain:

  • Stored securely: in a managed secret store, never in prompts, config files or repositories.
  • Scoped narrowly: the minimum reach the authorized actions require, and no more.
  • Short-lived by default: prefer a session over a key. An assumed AWS role runs from 15 minutes to 12 hours, and one hour is the default if you say nothing.
  • Cut off in one action: when the agent is suspended, changes purpose or retires, its credentials stop working everywhere they were handed out.

Google Cloud is blunt on the first rule: "We recommend that you avoid using service account keys whenever possible," with workload identity federation or impersonation as the alternatives. A service account also tops out at 10 keys, which is a low ceiling for anything on a rotation schedule.

Tying credentials to the agent's own identity is what makes cutoff clean. Decommission a sales-forecasting agent and you invalidate its token without touching the CRM access of the dozen other agents reading the same system. Shared credentials make that impossible: pulling one breaks everything downstream, so nobody pulls it.

One more number to plan around. If agents assume roles per request instead of holding sessions, AWS STS allows 600 requests per second per account per Region, raisable only by support ticket.

5. Monitoring by Default

An agent shouldn't enter production unless the enterprise can see what it's doing and investigate later. Monitoring isn't a maturity upgrade once an agent proves valuable. It's a precondition.

The AI RMF states it as a measurement requirement: MEASURE 2.4 asks that "the functionality and behavior of the AI system and its components" be "monitored when in production", and MANAGE 4.1 expects post-deployment plans covering "appeal and override, decommissioning, incident response, recovery, and change management". Provisioning is when those get switched on, because nothing later has a natural moment for it.

By the time an agent acts, three things should be true. Every action is logged against its identity, with the tool called and the data touched: AC-2 puts "monitor the use of accounts" in the control text itself. Its behavior is visible on a dashboard alongside other agents. And an investigator can reconstruct what happened without asking the owning team to dig through scattered logs.

The test is uncomfortable and clarifying: if a hypothetical data-quality agent began silently rewriting records it was only supposed to read, how long until someone noticed, and could you prove what it changed? If the answer is "we're not sure," it was provisioned without the monitoring it needed.

6. Review Tiers

Not every agent deserves the same scrutiny, and pretending otherwise is how governance becomes the bottleneck. The shape we'd defend:

TierWhat qualifiesGate before launchCredential patternRe-review
LowRead-only, no personal data: meeting notes, doc searchSelf-service, owner named, automatic registry entryShort-lived read scopeOn model or scope change
ModerateWrites internally or touches personal data: the HR drafterOwner sign-off plus a security read of scopesScoped token, scheduled rotationAnnually
HighMoves money, changes production, acts externallyNamed reviewers, scope justification, human confirmation on the riskiest actionsShort sessions, no long-lived keysQuarterly and after any incident

Table 2 — Review sized to consequence. The stages never vary; only this column does.

AC-2 supports the middle column more literally than people expect, requiring "approvals by [organization-defined personnel or roles] for requests to create accounts". It does not say every account needs the same approver, which is the latitude the tiering uses. A read-only summarizer shouldn't wait behind the committee that clears a payments agent, and a payments agent should never slip through on a summarizer's light touch.

Templates & Workflows

At enterprise scale, running six stages by hand for every agent is the bottleneck. Two mechanisms make the safe path the fast one.

Reusable permission templates

Most agents fall into a handful of recognizable shapes. A template encodes the identity type, typical scopes, credential pattern and review tier for that shape, so teams configure rather than invent.

template: read_only_analytics_agent
risk_tier: moderate
scopes:
  - warehouse.query.read       # sales, marketing schemas
  - dashboard.read
denied_by_default:
  - warehouse.write
  - pipeline.modify
credentials:
  type: scoped_service_token
  rotation_days: 30
review:
  required: owner + security_scope_check

A new analytics agent starts here, inherits sane defaults, and only the deviations need discussion. The reviewer reads a diff against a known-good baseline instead of a blank slate. Templates also keep you inside the ceilings above, because a role per archetype consumes far fewer attachment slots than a role per agent.

Standardized provisioning workflow

The stages should run as one automated flow rather than six disconnected tickets, and it should refuse to mark an agent active until every stage has completed.

ObservabilitySecret storeAgent registryIdentity providerProvisioning workflowTeamObservabilitySecret storeAgent registryIdentity providerProvisioning workflowTeamsubmit template + deviationscreate agent identityidentity refwrite record (owner, purpose, tier)apply template scopes, default denyissue scoped credentialenable logging and dashboardset status = activeendpoint + audit link

Figure 2 — One flow, six stages, one gate. The registry status flips to active last, which is what makes "provisioned" a checkable claim rather than an assertion.

That is what turns provisioning into infrastructure.

How This Goes Wrong

Four failures show up repeatedly, each with a recognizable symptom.

The quota wall. Provisioning fails for new agents while existing ones work fine. The cause is almost always a ceiling from the table above, and the give-away is that the failure is account-wide rather than agent-specific. Check roles per account, managed policies per role and directory object count first.

The shared identity that grew. An audit asks which agent performed an action and the answer is a service account used by four of them. The AC-2 shared-account case, unfixable retroactively. The log is already written.

Scope creep by ticket. Each widening was reasonable; the total is an agent with warehouse write access nobody would grant today. The symptom is a permission set that no longer matches the registered purpose, which is why the registry needs a requested-permissions field and a periodic diff against reality.

The orphan. The owner changed teams, the agent kept running, nothing pointed at it until something broke. ASI10 in slow motion. The defence is an owner field that fails validation once the named person is no longer an employee.

Where Fabriq Fits

We build Agentic Fabriq around the middle stages above. An agent is registered, scoped, then activated with a one-time secret and an in-console test, and its effective tool list is the intersection of the acting user's scopes and the agent's own, enforced in the backend rather than in a prompt. Access is default-deny across org, team and member. Credentials sit in a vault and are attached server-side at the moment of the call, so the agent never handles the secret. Every call carries two identities, the agent and the acting user, into a BigQuery-backed audit record.

One limit worth stating plainly: clearing our vault copy of a token stops Fabriq using it and does not withdraw the underlying grant at Google or Microsoft. That happens upstream, and your offboarding runbook should say so in writing.

Frequently Asked Questions

How many AI agents can one AWS account hold? If each gets its own IAM role, 1,000 by default and 10,000 after an auto-approved increase. The tighter constraint is usually 20 managed policies per role, raisable only to 25, which is what pushes teams toward a role per archetype.

Should each agent get its own service account? Yes, and Google Cloud says so in its own words: "Create dedicated service accounts for each application, and avoid using default service accounts." The reason is evidence, not security theatre. A shared account makes audit logs unattributable and permissions unscopable.

What's the difference between an agent identity and a service account? Mechanically, often nothing; an agent identity is frequently implemented as one. The difference is the record around it: named owner, stated purpose, risk tier, lifecycle state. Microsoft now models the distinction directly, with Agent Identities created under a Blueprint and quotas of their own.

How do I provision hundreds of agents without a bottleneck? Fix the stages and vary the scrutiny. Template the archetypes so a launch is a diff rather than a design review, and reserve named reviewers for the tier that moves money or changes production.

Does provisioning need to handle offboarding too? It needs to make offboarding possible, which is a smaller claim. Distinct identity, recorded owner, and a credential you can invalidate in one action are the preconditions. The AI RMF keeps decommissioning in GOVERN 1.7 for the same reason: a commitment made at the start, not a task discovered at the end.

Conclusion

Good provisioning makes adoption faster, not slower, because teams know exactly how to reach production. Bad provisioning produces either chaos, agents everywhere with no accountability, or enough friction that teams route around governance entirely.

We don't think those are the only two options. The repeatable path avoids both: identity, registration, scoped authorization, managed credentials, monitoring by default, review proportionate to risk, standardized so the safe path is the quick one. And before any of it, read the quota pages. The number of agents your platform will hold is a fact you can look up, better learned from documentation than from a failed provisioning run.

Provisioning is where speed and control stop being a trade-off. Get it right once, and every agent after the first inherits a path to production that is both fast and accountable.

Sources

All URLs read 2026-09-29.