An agent is a workload, not a user, and a long-lived static API key baked into a tool definition or an environment variable is the wrong authorization primitive for it: it grants standing privilege forever, cannot be revoked per task, and turns a single exfiltration into the entire blast radius. Credential scoping fixes that by minting a short-lived, audience-bound, least-privilege token at the task boundary through workload identity federation or token exchange — never a shared secret. What follows is the primitives, the 2026 platform support, the exploit chain that punishes the old model, and a production pattern you can deploy this week. If your current governance story is one key per tool plus a calendar reminder to rotate, read this before the next deploy.
The agent is a workload, not a user
The agent is a workload, not a user, because it holds no human identity of its own and acts only on delegated authority — a distinction NIST SP 800-207 draws when it demands per-request authorization instead of persistent trust, and one that OWASP’s ASI03 entry treats as a central form of agentic identity abuse.
A user principal can consent, be phished, be MFA-challenged, and be reasoned about by a help desk. A workload principal can only prove what it is: which platform it runs on, which namespace, which service account, which signed identity document it holds. That is a stronger claim than any password, and it is the claim a federation endpoint actually wants to verify. The OWASP Non-Human Identities Top 10 makes the same point from the other direction: machine identities routinely inherit privileges that nobody deliberately chose, and that inheritance accumulates into standing privilege. Blast radius is what it costs you the moment one credential walks out the door.
Why API keys break agent permission models
API keys break agent permission models because a static bearer secret carries no task context: OWASP’s ASI03 identity and privilege abuse names exploited delegated trust and inherited credentials as its core mechanism, and a key in an env var cannot be revoked for one task without breaking every task that shares it.
Three failure modes do the damage. First, no per-task revocation: the key is valid until someone rotates it, so a runaway agent loop and a compromised agent look identical to your authorization layer. Second, env-var leakage: tool definitions, container images, crash dumps, and .env files copied into debug bundles all carry the secret, and the agent runtime almost never needs to see it. Third, no audience restriction: a token minted for your ticketing tool is silently accepted by your database proxy if both read the same key. This is why a boring MCP token waste audit usually finds more live standing credentials than anyone expected — most of them minted once, never scoped, never expired.
Credential scoping defined: short-lived, audience-bound, least-privilege
Credential scoping means satisfying three properties simultaneously — short lifetime, audience binding, and least privilege — which is what Anthropic’s workload identity federation does when it trades an OIDC JWT for a token with token_lifetime_seconds between 60 and 86400, and what RFC 8693 token exchange standardizes.
Each property closes a different hole. Short lifetime bounds the value of theft: a stolen token that dies in ten minutes is a nuisance, not an incident. Audience binding means the token is useless anywhere except the one resource it was minted for, which is exactly what RFC 8707 resource indicators exist to express. Least privilege means the scope list reflects the tool call, not the integration. Anthropic’s federation rule sets the default scope to workspace:developer and issues tokens bound to a service account, so the credential is meaningful only inside one workspace and only for the duration the rule allows.
The OAuth and MCP authorization base layer
The OAuth and MCP authorization base layer requires an MCP server on HTTP transports to behave as an OAuth 2.1 resource server, publish Protected Resource Metadata per RFC 9728, and advertise scopes in its WWW-Authenticate challenge, per the MCP authorization specification.
Two changes matter for anyone shipping in 2026. Dynamic Client Registration (RFC 7591) is deprecated in MCP authorization in favor of OAuth Client ID Metadata Documents, which removes the open-registration footgun that let any client obtain credentials from an authorization server. And scopes are no longer a static list in a config file — they are advertised per challenge, enabling a step-up authorization flow where the client requests more scope only when a tool actually needs it. The 2026-07-28 MCP revision shipped this alongside the stateless core (see our MCP stateless 2026 spec deep dive), and the ext-auth repository tracks the extension surface. Bearer token semantics still come from RFC 6750, the authorization code flow from the OAuth 2.1 draft, and PKCE from RFC 7636.
A resource server publishes its metadata so clients never guess:
{"resource":"https://mcp.example.com/tools","authorization_servers":["https://auth.example.com"],"scopes_supported":["mcp:read","mcp:write","tool:sql"],"bearer_methods_supported":["header"],"code_challenge_methods_supported":["S256"]}
And it tells a caller what it is missing, in the response, not in documentation:
HTTP/1.1 401 Unauthorized
WWW-Authenticate: Bearer scope="mcp:read mcp:write"
Audience binding to the MCP server is enforced by the resource parameter per RFC 8707.
Workload identity federation across major agent platforms
Workload identity federation now ships across the major agent platforms: Anthropic exchanges an IdP JWT for a service-account-bound token using the RFC 7523 jwt-bearer grant, OpenAI accepts an OIDC JWT, a SPIFFE JWT-SVID, or an X.509 certificate, and Microsoft Entra Agent ID issues identities that hold no credentials at all.
| Option/Standard | Token type | Audience binding | Default/max lifetime | Revocation model | Exchange protocol | MCP/OAuth alignment | Production state |
|---|---|---|---|---|---|---|---|
| Anthropic WIF | OIDC JWT → short-lived access token bound to a service account | Federation rule plus resource parameter | 60–86400 s; default 3600 s, wizard default 600 s | Expiry, plus single-use jti — re-presenting an exchanged JWT fails with jti_reused |
RFC 7523 jwt-bearer grant at /v1/oauth/token |
OAuth 2.1 token endpoint; no refresh token | Shipping; documented in Claude platform docs |
| OpenAI WIF | OIDC JWT, SPIFFE JWT-SVID, or X.509 → short-lived access token | Service-account mapping, scoped per resource | Short-lived; renew by repeating the exchange | Expiry only — no refresh token issued | Token exchange at the platform endpoint | OAuth-style exchange | Beta on the Codex path |
| Microsoft Entra Agent ID | Federated token for a credential-less service principal | Agent identity blueprint plus FIC scope | Per-token; nothing stored | Blueprint or FIC change, or identity deletion | Federated identity credentials and agent OAuth protocols | Entra OAuth protocols | Documented; single-tenant only |
| Okta XAA | Identity-based token scoped to the target app | Per-app, enterprise-managed | Short-lived per task or session | Central policy revocation in Okta | Cross App Access OAuth extension | Official MCP Enterprise-Managed Authorization extension | Anthropic production beta; Okta Workforce OIN access since August 2026 |
| SPIFFE/SPIRE | X.509-SVID and JWT-SVID | SPIFFE ID per workload, not a shared secret | Minutes to hours, auto-rotated | Attestation failure stops renewal | JWT-SVID presented to a token endpoint | Supplies the subject token for OAuth exchange | Open standard with a maintained reference implementation |
| AWS/GCP cloud STS | Temporary role credentials | Role or resource policy | 15 min–12 h on AWS; configurable, 1 h default on GCP | Policy change or session revocation | AssumeRoleWithWebIdentity / STS token exchange |
Supplies subject tokens | Generally available |
The Anthropic path is the one most teams start with, because the exchange is a single small HTTP call:
import httpx
# idp_jwt is the short-lived OIDC token your identity provider issued to the workload
response = httpx.post(
"https://api.anthropic.com/v1/oauth/token",
data={
"grant_type": "urn:ietf:params:oauth:grant-type:jwt-bearer",
"assertion": idp_jwt,
"scope": "workspace:developer",
"token_lifetime_seconds": 600,
},
)
response.raise_for_status()
access_token = response.json()["access_token"]
Three things are absent on purpose: no key material, no refresh token to store, and no long-lived artifact to rotate. Renewal is a repeat of the same call, which is a feature — every renewal re-proves workload identity, so a captured token cannot be renewed once it has expired.
SPIFFE/SPIRE and Vault: platform-agnostic workload identity
SPIFFE is the open standard for workload identity, and SPIRE is the implementation that attests a workload against platform signals — Kubernetes service accounts, process selectors, cloud instance metadata — and then issues rotating identity documents such as the JWT-SVID, so the agent proves what it is without ever sharing a secret, as described in the SPIFFE overview.
This is the layer that makes federation portable. If you run agents on Kubernetes, on bare EC2, and on a colleague’s laptop under Docker, you do not want three identity stories. SPIRE gives you one SPIFFE ID per workload and a rotating SVID that any federation-capable token endpoint will accept as a subject token. The downstream half matters just as much: a workload that authenticated with a JWT-SVID can use Vault’s SPIFFE secrets engine to obtain short-lived credentials, and Vault’s database secrets engine to mint a PostgreSQL user that lives for the length of one task. The agent never holds a database password, because there is no database password to hold.
Okta Cross App Access and MCP enterprise-managed authorization
Okta Cross App Access is an open OAuth extension that removes standing privilege by issuing identity-based tokens scoped to what the agent needs, and it was formally incorporated as an official MCP authorization extension for Enterprise-Managed Authorization, per Okta’s partner announcement.
The ecosystem list is the interesting part: 25+ partners including Anthropic, Cloudflare, Cursor, Docker, Figma, Slack, Supabase, Datadog, VS Code, Zoom, and WorkOS. That breadth is what turns XAA from a vendor feature into a plausible enterprise default — the IT admin governs agent access from the same console that governs human SSO, rather than from a spreadsheet of tool keys. Okta made OIN access available to Workforce customers in August 2026, and Anthropic runs a production beta, which means the pattern is testable now rather than theoretical. For MCP specifically, this is the piece that closes the gap between “the server supports OAuth” and “the enterprise can actually govern it.”
Cloud workload identity: AWS, GCP, and Azure
Cloud workload identity is where most teams already have federation available, since AWS IAM OIDC identity providers let a workload exchange an external JWT for temporary role credentials, and GCP offers the same shape through Workload Identity Federation.
On EKS the ergonomic version is IAM roles for service accounts, which projects a signed token into the pod and lets the SDK assume a role without any static credential in the image. GCP’s federation works the same way with workload identity pools and providers: the agent presents an OIDC token, the pool validates the issuer and subject, and the caller receives short-lived credentials to impersonate a service account. The composition that matters for agents is chaining — cloud STS to prove the workload, then an OAuth token exchange to prove the task, with the cloud credential never leaving the runtime.
Failure modes and the OWASP ASI03 exploit chain
Failure modes around agent credentials all ladder into ASI03 identity and privilege abuse, which the OWASP Top 10 for Agentic Applications describes as exploiting delegated trust, inherited credentials, and role chains, with mitigations expressed as per-tool least-privilege profiles attached to each tool.
| Failure mode | Attack path (OWASP ASI03) | Detection signal | Mitigation |
|---|---|---|---|
| Long-lived API key in env var or tool definition | Secret harvested from the runtime and reused with full standing privilege | Key material appearing in env dumps, logs, or tool JSON | Exchange at the task boundary; no key enters the runtime |
| Missing audience restriction | Token minted for service A replayed against service B | A resource accepts a token it never issued for | RFC 8707 resource indicators; validate aud at every resource server |
| Bare bearer token without DPoP | Stolen token replayed from an attacker host | Same token seen from two client key sets or geographies | RFC 9449 proof-of-possession binding |
| Over-scoped token | Agent performs actions beyond the task because the token permits them | Requested scopes exceed the scopes the task declared | Per-tool least-privilege scopes and egress allowlists |
| Static service account treated as a user principal | Inherited credentials enable role-chain escalation | Service account authenticating interactively or from a new host | Federated identity credentials; credential-less agent identities |
Missing jti replay check |
An exchanged identity token is re-presented to mint more tokens | Repeated assertion carrying the same jti |
Enforce single-use identity tokens and reject repeats |
The pattern worth internalizing is that the detection signals are cheap and the mitigations are policy, not code. You can read more open-source approaches in this write-up on agent tool governance open-source patterns, which maps the same control set onto tools you can actually install.
Production implementation pattern
The production implementation pattern is to bootstrap the workload from an IdP JWT, exchange it at the task boundary per RFC 8693, bind audience and scope per tool, and make the resulting token proof-of-possession with RFC 9449 DPoP.
The exchange itself is a single form post:
POST /token HTTP/1.1
Content-Type: application/x-www-form-urlencoded
grant_type=urn:ietf:params:oauth:grant-type:token-exchange&subject_token=<IDP_JWT>&subject_token_type=urn:ietf:params:oauth:token-type:jwt&audience=https://mcp.example.com/tools&resource=https://mcp.example.com/tools&scope=mcp%3Aread%20mcp%3Awrite&requested_token_type=urn:ietf:params:oauth:token-type:access_token
The five-step migration that works in practice:
- Bootstrap an IdP JWT. Have the platform issue a signed identity token to the workload — cloud STS, SPIRE, or the platform’s own federation endpoint. This is the only credential with a lifetime longer than a task, and it lives in the platform, not the repo. RFC 7523 defines the assertion format.
- Exchange at the task boundary. Mint the downstream token when a task starts, not when the process starts. If a task runs for an hour and the token lives ten minutes, mint two.
- Bind audience per tool. Every tool definition names the resource it will call. The token for the SQL tool must be rejected by the ticketing API.
- Scope per tool. Copy scopes from the tool’s declared capability, not from the integration’s maximum. A read tool gets a read scope.
- Enable DPoP and enforce
jti. Bind the token to a client key pair so a copied token is inert, and reject any re-presented identity assertion by itsjti.
FAQ
Four questions come up most often when teams start replacing API keys with credential-scoped agents.
Is a long-lived API key ever acceptable in 2026? A long-lived API key is acceptable only for local development and initial bootstrap, never for a deployed agent. Once the workload can reach an identity provider, replace it with a federated token: Anthropic’s WIF documentation frames short-lived OIDC tokens as the direct replacement for long-lived provider keys in production.
Do I need to replace all my MCP tool keys today? No — prioritize by blast radius. Start with tools that touch production data or spend money, migrate those to audience-bound tokens per the MCP authorization spec, and leave read-only internal tools on keys until their next major change. The metadata surface every MCP server already needs makes this incremental.
Does WIF work for self-hosted agents? Yes, when the host can present a verifiable platform identity. SPIRE attests a workload against platform signals and issues a rotating JWT-SVID that any federation-capable token endpoint accepts, and Vault’s SPIFFE secrets engine then mints downstream credentials from it. A VM with no attestable identity needs a bootstrap credential first.
What is the operational overhead of short-lived tokens? Mostly caching and retry logic. Tokens expire, so the agent must re-exchange near expiry and treat a 401 as “refresh and retry” rather than “fail”; OpenAI’s WIF flow returns no refresh token at all, so renewal means repeating the exchange. In return, revocation becomes automatic at expiry.
The Bottom Line
The bottom line is that agents should hold short-lived, audience-bound, least-privilege tokens instead of API keys, because OWASP’s ASI03 identifies delegated trust and inherited credentials as a core mechanism of agentic identity abuse, while NIST SP 800-207 requires per-request authorization rather than persistent trust.
The migration is deliberately unglamorous. You are not adopting a new security product; you are moving where the credential is minted. Do the highest-blast-radius tool first, prove the exchange works, then repeat. The failure mode of doing nothing is not a dramatic breach — it is a twelve-month-old key in a container image that nobody can attribute, revoke, or safely rotate.
How This Guide Was Built
This guide was built from desk research only; I did not run any credential system or test any token exchange in a live environment. It draws on the published specifications and vendor documentation cited inline rather than on any live deployment, so the honest scope is narrow: it describes how these systems are documented to work, not how a specific estate behaves under load. Every source URL cited here was fetched on 2026-10-03 and returned HTTP 200. Where the specification language is permissive — “MUST implement” versus “SHOULD support” — the distinction has been preserved rather than smoothed over, and any platform behavior described as beta reflects the vendor’s own labeling at that date.
← Back to all posts


