AXONN Vantis logo
AXONN VantisAgentic eXperience, Open Neural Network,Complete Governance
EN
← Blog

Delegation Without Handing Over the User's Account — Workload Identity and Scoped Tokens for Agents

Picture a travel-booking agent that receives the user's OAuth access token and calls a booking API with it. The API's audit log records bookings the user made and bookings the agent made under the same principal, and if a prompt injection leads the agent to make an unintended booking, there is no way to revoke the agent's access alone. Revoking the token kills the user's other sessions too; leaving it in place lets the agent keep exercising everything the user can do. Once an agent takes on actions like booking, ordering, or querying data, the authorization question of "who can do what" has to be answered not for a person but for a pair: the person and the software acting on their behalf.

This post is about giving agents their own non-human identity and narrowly scoped delegated credentials instead of a person's account. It starts with how the attack surface is moving from model output to the identity an agent holds and the authority delegated to it, and why human and agent identities need to be separated. It then walks through how SPIFFE and SPIRE issue workload identity, and how OAuth Token Exchange (RFC 8693) and RAR (RFC 9396) turn a user's token into a restricted token for an agent. It closes with the open problems around dynamic tool selection and revoking and auditing delegation chains.

The attack surface shifts from model answers to agent identity and authority

The attack surface of agent security is shifting from what a model says to whose authority an agent acts under and what it can do with it. To send mail, request a payment, or query internal data, an agent has to authenticate to the target system and be authorized, and the credential it uses sets the upper bound on what a compromised agent can do. Individual breach cases make poor evidence for this shift: each depends on its own product configuration and on how much was disclosed, so their causes are hard to generalize and the facts hard to verify. This post therefore traces the shift through risk taxonomies and standards documents, which turn the structure common to many cases into named items. The OWASP Top 10 for Agentic Applications 2026 lists identity and privilege abuse (ASI03 Identity and Privilege Abuse) as a risk of its own, separate from goal hijacking and tool misuse. Identity Management for Agentic AI, a report published in October 2025 by the OpenID Foundation's AI Identity Management Community Group, then observes that in today's deployments agents often impersonate users in ways external services can't detect, so an agent's API calls end up logged indistinguishably from actions the user took directly. Whether a token is stolen or a tool is misused, the damage ultimately flows through that credential, so the focus of defense moves from the model's output to the principal and scope of the credential.

After risk classification and diagnosis, standards work is moving the same way. The architecture draft from the IETF WIMSE working group (Workload Identity in a Multi System Environment Architecture) treats AI and ML intermediaries as a special case of delegated workloads in a section of their own. It says they should propagate upstream context, distinguish autonomous actions from delegated ones, and explicitly re-scope at every hop of a multi-agent delegation chain. The OAuth working group's draft on identity and authorization chaining across domains is making its way through the RFC publication process. Agent identity and delegation, in other words, are being worked out as extensions of existing standards for workload identity and OAuth rather than as one-off implementation conventions. How tool calls are judged just before execution was covered in "From Guardrails to the Execution Boundary — Policy Enforcement for Agent Tool Calls"; this post is about how the principal and authority that feed that decision are created.

Separating human and agent identity

The design principle that follows is to separate human identity from agent identity and to express the delegation between them as a credential of its own. RFC 8693 draws this line as two semantics: impersonation and delegation. Under impersonation, the actor takes on the subject's rights and is indistinguishable from the subject to the recipient. Under delegation, the actor keeps its own identity while acting for the subject, and the issued token carries information about both parties. An agent using a person's token or API key as-is is impersonation. The audit log records only the person as the actor, and revoking one misbehaving agent means revoking the person's credential, cutting off the person's other access along with it. And because the agent can exercise the person's full scope rather than what the task needs, least privilege doesn't hold either.

So an agent needs two things issued separately. One is an identity for the agent as a workload, which proves cryptographically which agent made a call. The other is a delegated credential stating that it may perform a particular action on a particular resource on behalf of a particular user, with scope, audience, and lifetime limited to the task. With that separation, the audit log can record "agent A acting for user U," and revocation can target agent A's delegation instead of the user as a whole. Authenticated Delegation and Authorized AI Agents (2025) proposes extending OAuth 2.0 and OpenID Connect with agent-specific credentials and metadata, and Identity Management for Agentic AI recommends explicit on-behalf-of delegation; both come from the same principle. Implementing it does not require a new protocol. SPIFFE already standardizes workload identity, and OAuth Token Exchange and RAR already standardize scoped delegation.

Issuing and verifying workload identity: SPIFFE and SPIRE

SPIFFE (Secure Production Identity Framework for Everyone) is a specification for giving workloads a service identity, and SPIRE is an implementation of it. An identity is a SPIFFE ID, a URI made of the spiffe:// scheme, a trust domain, and a path. For example, spiffe://example.org/agent/travel-booking names the travel-booking agent in the example.org trust domain. The document that carries this ID with a signature and is presented to a peer is the SVID (SPIFFE Verifiable Identity Document), and an SVID is valid only if signed by an authority in that trust domain. An X.509-SVID puts the SPIFFE ID in the certificate's URI SAN; there must be exactly one URI SAN, and the cA flag must be false. A verifier can therefore read one identity unambiguously from one certificate and confirm that the certificate can't sign others. A JWT-SVID is a short-lived token issued for a specified audience, and the recipient verifies its signature against the JWK Set in the trust bundle.

SPIRE issues an SVID to an agent process in this order. First, an operator creates a registration entry on the SPIRE Server, pairing a SPIFFE ID with the selectors a workload must have, such as a Kubernetes namespace and service account or a Unix user ID. Next, the SPIRE Agent on each node goes through node attestation, presenting evidence such as a cloud instance identity document, receives its own SVID, and uses it to fetch the registration entries it is authorized for. The agent process then calls the Workload API, exposed on a Unix domain socket, and presents no credential at all: the Workload API specification deliberately omits direct client authentication and relies on out-of-band checks instead. The SPIRE Agent identifies the caller's PID, its workload attestors query the kernel or the kubelet for selectors, it matches those against registration entries to decide the identity, and it returns an X.509-SVID, the private key, and the trust bundle. The agent's code and prompts hold no secret from the start; its identity is derived from properties of the environment it runs in.

What happens after issuance determines how revocation works. The SPIRE Server's default TTL is one hour for X.509-SVIDs and five minutes for JWT-SVIDs. The Workload API is a stream that resends the full state whenever an SVID rotates, and clients must drop anything that disappears from a response. Deleting a registration entry therefore stops that identity from being issued or renewed, and a banned node cannot re-attest. The X.509-SVID specification, however, defines no certificate revocation mechanism such as CRLs or OCSP, so an SVID already issued can still pass validation while its signature is valid. Revocation in SPIFFE is thus implemented as short TTLs plus stopping renewal, and the delay before it takes effect is bounded by the TTL. Note, too, that SPIFFE proves only which workload made a call. Which user that workload acts for, and with what scope, is not in the SVID; that is the job of the delegation token.

Scoped delegation: OAuth Token Exchange and RAR

RFC 8693, OAuth 2.0 Token Exchange, defines how to present a token you already hold at an authorization server's token endpoint and receive a different one. The request's grant_type is urn:ietf:params:oauth:grant-type:token-exchange, and it carries a subject_token for the party being represented, an actor_token for the party actually acting, a resource or audience naming where the token will be used, and the requested scope. In the running example, the travel-booking agent presents the user's access token as the subject_token and the JWT-SVID it got from the Workload API as the actor_token, and names the booking API as the audience. The authorization server can use the may_act claim in the user's token to decide whether this agent may act for this user, and if so, it issues a new token with the user in sub, the agent's SPIFFE ID in the act claim, and a short expires_in. When the agent hands work to a sub-agent, the exchange happens again and the act claims nest: the outermost act is the current actor, and the most deeply nested one is the earliest actor.

A scope string can express "create a booking," but struggles with "create a booking for this itinerary, up to this amount." RFC 9396, Rich Authorization Requests, handles this with a JSON array called authorization_details. Each object in the array has a required type and common fields — locations, actions, datatypes, identifier, privileges — and can add API-specific fields such as an amount or a payee. Within one object, the requested rights are the product of the listed values, so finer restrictions mean splitting into more objects. Clients can include the parameter in token requests too; the authorization server checks that the underlying grant allows it and returns the authorization_details it actually granted in the token response. Resource servers read the same structure from JWT access tokens or token introspection responses. An agent's token can therefore carry rights scoped to resources and actions for a single task.

The catch is that narrowing is authorization-server policy, not a property the protocol guarantees. RFC 8693 places no requirements on which exchanges are allowed and does not require the issued token to be narrower than the subject_token; its security considerations only suggest scope and limited token lifetime as mitigations against abuse. It is the cross-domain chaining draft that makes narrowing normative, requiring the authorization server to verify that the requested scopes are not more privileged than those of the presented subject_token. RAR is harder still. RFC 9396 defines no general way to compare two authorization_details and warns against relying on simple object equality, because relationships between fields — write implying read, for instance — differ by API. Monotonic narrowing at each hop of a delegation chain therefore holds only if the authorization server implements, for each type, logic that decides whether a request is a subset of the original grant. The standards supply the vocabulary for narrowing; enforcing it is the responsibility of the deployment environment.

RFC 9635, GNAP (Grant Negotiation and Authorization Protocol), redesigns the same problem without extending OAuth. In GNAP a client instance is identified by its own key rather than by pre-registration and signs its requests with it, and rights are requested through an access array with the same type, actions, locations, datatypes structure as RAR. Issued tokens are bound to the client's key unless the bearer flag is set, and the client can rotate or revoke them through a token management URI. Giving each agent its own key and preventing reuse of stolen tokens fits the principle of this post, but the specification itself says it is not directly compatible with OAuth 2.0, which makes it hard to adopt where existing authorization servers stay in place.

Limits of dynamic tool selection and delegation-chain auditing

Applying this separation of identity and delegation to real agents still runs into problems without settled answers.

First, dynamic tool selection versus least privilege. Scoped delegation assumes you know in advance what a task needs, but an LLM agent discovers and connects to new tools and services at runtime based on what the user wants. Pre-authorizing an agent that can't know its tools before it runs means granting a scope wide enough to cover every tool it might use, which runs directly against least privilege. Grant broadly and least privilege collapses; grant narrowly and the task stalls. The alternative is a token exchange each time a tool is chosen, yielding a token for that call — at the price of more round trips to the authorization server and, where delegation involves human approval, more frequent prompts and the consent fatigue that comes with them, as users start approving without reading.

Second, revoking delegation chains. A token issued through token exchange is a new token, separate from the original subject_token, and RFC 8693 defines no procedure linking the revocation of the two. A resource server that validates a JWT by signature and expiry alone has no way to learn that the upstream delegation was revoked. If a user revokes the top-level agent's delegation, then, a sub-agent's token obtained by exchanging it may stay valid until it expires, and there is still no standard mechanism to propagate revocation down the chain immediately. Bounding impact with execution-count limits, or supplementing with shared security-event signals, are proposed directions but only partial ones. As with SPIFFE's short TTLs, the practical revocation tool is still a short lifetime; making revocation take effect sooner means keeping downstream token lifetimes short or having resource servers check each token's active state through token introspection.

Third, auditing delegation chains. RFC 8693 makes prior actors in nested act claims informational only and says access control decisions should use only the top-level claims and the current actor. Resource servers therefore can't decide based on the full chain, and the chain is useful for audit only if every authorization server passed it along faithfully. The WIMSE architecture draft's requirement to re-scope and re-bind context at every hop aims at this gap, but the requirement itself is still a draft.

Fourth, identity across trust domains. A SPIFFE ID is meaningful only to those who know and control the infrastructure of its trust domain, so the model doesn't carry across organizational boundaries as-is. The cross-domain chaining draft defines a flow in which a token exchange at domain A's authorization server yields a JWT authorization grant that is presented to domain B's authorization server as a JWT assertion grant, but it does not define the format of transcribed claims, so both sides must agree on their meaning separately. W3C's Decentralized Identifiers (DIDs) v1.0, which defines identifiers verifiable without a central registry, and the Verifiable Credentials Data Model v2.0, which defines issuer, holder, and verifier roles and validity periods, are discussed as candidates for portable agent identity and proof of delegation, but no profile for expressing agent identity and delegation interoperably across organizations has been standardized yet.

Summary

As agents take on real actions, the attack surface has moved from model answers to the identity an agent holds and the authority delegated to it. Impersonation through a person's credentials keeps audit logs from telling the person and the agent apart and makes it impossible to revoke the agent alone, so agents need a workload identity and a delegated credential issued separately. SPIFFE and SPIRE attest properties of the runtime environment to issue short-lived SVIDs; RFC 8693 Token Exchange produces a token carrying both user and agent as sub and act; and RFC 9396 RAR expresses that token's rights at the level of resources and actions. For synchronous agents within a single trust domain, this combination is enough to build on. Least privilege for agents that choose tools at runtime, propagating revocation through delegation chains and auditing them, and identity across domains are still areas the standards have not answered.

In the end, non-human identity for agents isn't a matter of picking a new authentication technology. It is a delegation design problem: stating in every token who is acting for whom, to do what, until when, and making the authorization server ensure that scope never widens at any step of the delegation.

References

← Blog
© 2026 AXONN Vantis Inc. All rights reserved.