AXONN Vantis logo
AXONN VantisAgentic eXperience, Open Neural Network,Complete Governance
EN
← Blog

From Guardrails to the Execution Boundary — Policy Enforcement for Agent Tool Calls

A prompt injection against a chatbot produces one bad answer. The same injection against an agent that can send mail, run code, and request payments produces side effects that can't be undone. Yet most of what gets called a "guardrail" inspects the model's input and output. What an agent actually does is decided not by the tokens the model emits but by the runtime that turns those tokens into tool calls and executes them, and in many systems the controls still sit upstream of that runtime.

This post is about the control point moving from the model's input and output to the execution boundary, the point where a tool call becomes a real action. It covers how the unit of risk changed and where standards and research are converging, the design principles that follow, how a policy enforcement point works and the main ways of writing policy, data-flow control for the cases policy can't reach, and the problems that remain when you deploy it.

The unit of security: from responses to actions

The unit of agent security has shifted from the sentences in a response to the side effects of tool calls, and the cause is architectural. When an agent reads an email, a web page, or a tool's return value, that content lands in the same context, as the same token sequence, as the user's instructions, and the model has no structural way to tell which tokens are instructions and which are data. The OWASP Top 10 for Agentic Applications 2026 puts this first, as ASI01 Agent Goal Hijack. Prompt injection, in other words, is less a defect of a particular model than a property of an architecture in which data and instructions share one channel.

The evidence that detection-based defenses can't get past this structure has piled up. The Attacker Moves Second (USENIX Security 2026), a joint study by researchers from several institutions, attacked twelve defenses across four families — prompting, training on known attacks, detection models, and secret-knowledge schemes — using gradient descent, reinforcement learning, random search, and human-guided exploration. Most of these defenses had originally reported near-zero attack success, yet adaptive attacks tuned against each defense's design got through most of them more than 90% of the time, and a human red-teaming competition with over 500 participants beat all 40 challenges. The authors point out that detectors are neural networks too, and can be fooled along with the model they protect. A June 2026 study of automated attacks found that task-universal injections transfer to unseen tasks and domains, but that their effectiveness depends heavily on the target model. A defense whose strength varies with the model and the attacker's budget is not something you can treat as a design guarantee.

The control point moves: from the model to the execution boundary

Standards and research are converging on moving control out of the model and to where actions happen. The OWASP agentic list keeps tool misuse (ASI02 Tool Misuse and Exploitation) and identity and privilege abuse (ASI03 Identity and Privilege Abuse) as items separate from goal hijack: even if the goal is hijacked, damage has to flow through tools and privileges, so that path can be controlled on its own. Design Patterns for Securing LLM Agents against Prompt Injections (2025), by researchers from ETH Zürich, EPFL, and other institutions, states it as a single principle: once an agent has ingested untrusted input, it must be constrained so that the input cannot trigger consequential actions. Regulation looks at the same point. Article 14 of the EU AI Act requires that overseers of high-risk AI systems be able to disregard or override the output and to halt the system in a safe state through a stop button or a similar procedure.

There are three places a control point can sit. The table compares what each can see, what kind of decision it makes, and how it can be bypassed.

Location What it can see Nature of the decision Bypass paths
Model (alignment, I/O classifiers) A token sequence mixing instructions and data Probabilistic Adaptive attacks, classifier errors
Agent loop (framework hooks) Plans, actions, observations, conversation state Deterministic for rules, probabilistic for LLM checks Built-in tools and direct calls that skip the hook
Execution boundary (tool runtime, gateway) Caller, tool name, arguments, resource, environment context Deterministic Execution paths left outside the boundary, abuse within authorized call forms

The execution boundary understands the least about meaning, but it has the narrowest bypass surface. The point is not to discard model-level or loop-level controls; it is to push final responsibility for the decision down to a place that can't be routed around.

Three principles for execution-boundary design

Three things follow for anyone designing these systems. First, put the policy enforcement point where every side effect has to pass. If even one of shell execution, file writes, or network egress skips it, a single detour neutralizes everything else. This is the old principle of complete mediation applied to agent tool calls. Second, make the decision outside the model, deterministically, and record why. The same call under the same policy must produce the same result for policies to be testable, and the audit log is only useful for post-incident analysis if it records which rule determined each decision.

Third, design on the assumption that decisions will sometimes be wrong. Policies are written by people, people don't anticipate every case, and wrong calls will get through. That is why isolation and least privilege to bound the blast radius, human approval for high-impact calls, and a stop path outside the agent's control are needed as the next lines of defense. If guardrails are probabilistic filters, the execution boundary is a deterministic mediator, and approval, isolation, and the kill switch are the layer that catches what the mediator misses.

Policy enforcement at the execution boundary

The policy enforcement point (PEP) and policy decision point (PDP) are terms defined in RFC 2753 as a framework for policy-based admission control. The PEP is where decisions are actually enforced; the PDP is where they are made. RFC 2753 requires the PEP to consult the PDP even when a request carries no policy information, specifically so that requests can't skip the policy check. Mapped onto agents, the PEP is the tool runtime or gateway that receives the model's tool call just before execution, and the PDP is a policy engine that evaluates policy separately from the model.

A small example shows the sequence. An agent asked to summarize an inbox reads an instruction hidden in an email and emits a send_email call that sends an attachment to an external address. First, the PEP intercepts the call before it runs. Next, it sends the PDP a request containing the principal (the agent session acting for this user), the action (send_email), the resource (the recipient address), and the context (the current task and the attachment's classification). The PDP evaluates policies such as "allow sending within the organization's domain" and "forbid sending confidential attachments externally" and returns deny. The PEP then blocks execution and, according to a predefined fallback action, either returns the reason to the agent or asks the user to confirm. Finally, it writes the decision and the identifier of the determining policy to the audit log.

How the policy itself is written differs by approach. AgentSpec (ICSE 2026), from researchers at Singapore Management University, is a domain-specific language that attaches rules to events in the agent loop. A rule has a trigger, a check, and an enforce clause; events include just-before-action (before_action), state change (state_change), and completion (agent_finish). Enforcement picks from user confirmation (user_inspection), LLM self-examination (llm_self_examine), running a predefined action (invoke_action), and stopping (stop). On code agents it prevented over 90% of unsafe executions, and predicate evaluation took a few milliseconds against agent runs lasting tens of seconds. The authors acknowledge that enforcement happens only at discrete checkpoints, so it can't reason about the safety of a long trajectory as a whole.

Progent (2025), from researchers at UC Berkeley and other institutions, narrows the unit to a single tool call. A policy is an ordered list of rules, each with an effect (allow or forbid), a target tool, conditions on its arguments (membership tests, regular-expression matching, and so on), and a fallback action. Forbid rules are checked before allow rules, a call that matches no rule is blocked, and the fallback is one of terminating execution, asking the user, or returning an error message to the agent. Policies are written in JSON Schema so an LLM can generate them for the task at hand, and runtime policy updates go through an SMT solver: changes that narrow what is allowed are applied automatically, while changes that widen it require approval. The authors call this monotonic confinement, and report that with human-written rules, attack success on one benchmark fell to 0%. They place attacks that stay within a task's least privilege, and built-in tools that bypass the MCP interface, outside the protection scope.

For a general-purpose authorization language there is the open-source Cedar. A Cedar request has four parts — principal, action, resource, and context. If any forbid policy is satisfied the result is deny; otherwise, if any permit is satisfied, it is allow; otherwise, deny. Default deny and forbid-overrides-permit are fixed in the evaluation algorithm, and the response lists the policies that determined the decision, which can serve directly as the basis of an audit record. Where agent-specific languages understand the loop and the meaning of tools, a general authorization language offers a precisely defined decision rule and lets agent policy be managed in the same form as existing authorization.

Where policy stops, and data-flow control

Policy judges the form of a call; it cannot tell whether the call matches what the user intended. If mail to a colleague is an allowed call form, a confidential summary that a hidden instruction slipped into the body passes in the same form. Progent's exclusion of manipulation within least privilege, and the design-patterns paper's note that even a quarantined LLM can be steered to change a résumé ranking or an email's parameters, point to the same limit. Stopping abuse inside authorized call forms requires feeding the decision with where an argument's value came from — its provenance.

CaMeL (Defeating Prompt Injections by Design, 2025), from researchers at ETH Zürich and other institutions, tackles this by concretely implementing the dual LLM pattern. A privileged LLM (P-LLM) sees only the trusted user query, writes the plan as Python-like code, and never sees tool outputs. A quarantined LLM (Q-LLM), with no tool access, only parses untrusted data into a given schema. A custom interpreter runs the code, tracks a data-flow graph, and attaches a capability to every value recording its provenance and permitted readers; before each tool call it checks security policies written as Python functions, blocking violations and asking the user for approval. As a result, untrusted data cannot change control flow, and the path for exfiltrating information over unauthorized data flows is closed. On AgentDojo it solved 77% of tasks with provable security, close to the 84% of an undefended system, at roughly 2.82x the input tokens and 2.73x the output tokens for the median task. The authors list as limitations the burden on users of writing and maintaining policies, side channels through exceptions and timing, the inability to solve tasks whose required actions depend on untrusted data, and the possibility of attacks that chain together allowed control-flow blocks. Data-flow control widens what the PEP sees from the form of the call to the provenance of its arguments, and pays for it in expressiveness and cost.

Design decisions and open problems

Bringing execution-boundary enforcement into a real system leaves several questions without settled answers.

First, writing and verifying policy. Having an LLM generate policy reduces the authoring burden, but the AgentSpec authors report that generated rules overfit to examples and produce both false negatives and false positives. Policy needs to be tested and reviewed like code, and using an SMT solver to decide mechanically whether a change narrows or widens scope, as Progent does, is one direction.

Second, the operational cost of default deny. When unanticipated legitimate requests are blocked and approval prompts pile up, users start approving without reading. Evaluation errors deserve attention too: Cedar drops a policy that errors during evaluation from the decision, so an erroring forbid policy cannot produce a deny, and the result falls to whatever other policies say.

Third, privilege propagation across multiple agents. When an orchestrator delegates work to sub-agents, there is still no established way to pass privileges down so that the delegate never exceeds the delegator. That delegation chain is part of why identity and privilege abuse is a risk item of its own.

Fourth, no standard for policy formats or decision records. AgentSpec uses its own DSL, Progent uses JSON Schema, CaMeL uses Python functions, and Cedar is a general authorization language. With no agreed format for the request a tool call sends to the PDP or for the decision record, every implementation ends up reshaping its audit logs.

Fifth, what makes a kill switch real. For a kill switch to actually work, issued credentials and delegated privileges must be revocable immediately, in-flight tool calls and child processes must be stoppable, and state at the moment of the stop must be preserved so it can be recovered or rolled back. Above all, the stop path must sit outside the agent's control. A stop implemented as a setting the agent can modify or invoke fails under exactly the attack the policy missed. Credential isolation for bounding the blast radius was covered in "The MCP Gateway, A Deep Dive", and the layers of isolation technology in "The Software Development Harness Is Turning Into an Agent".

Summary

Agent risk has moved from the sentences in a response to the side effects of tool calls, and in an architecture where data and instructions arrive as one token sequence, model-level detection doesn't hold up against adaptive attacks. Standards and research are therefore moving the control point to the execution boundary that every side effect passes through. There, a PEP intercepts the call, a PDP outside the model makes a deterministic decision over principal, action, resource, and context, and AgentSpec, Progent, and Cedar show different ways of writing that policy. Abuse that the form of a call can't reveal is covered by data-flow control such as CaMeL, which tracks the provenance of values as capabilities, at a cost in expressiveness and tokens.

In the end, agent execution control isn't a matter of picking a smarter guardrail. It is a trust-boundary design problem: where to put a decision point that can't be bypassed, how far to bound the damage when that decision is wrong, and who can stop the agent, through which path.

References

← Blog
© 2026 AXONN Vantis Inc. All rights reserved.