CAGE Certifies Tool-Using AI Agent Actions Under Return Uncertainty

Blaise Delattre, Cong Wang, Yang Cao· August 3, 2026 View original

Key takeaways

  • Tool-using AI agents face safety risks from small errors or uncertainties in tool return data.
  • CAGE provides certified authorization, ensuring agent actions remain safe even with plausible data perturbations.
  • Joint certification of categorical and numerical channels is crucial, as separate certification is insufficient.
  • CAGE reduces false allowances by authorization gates while preserving agent autonomy.

Who benefits

CybersecurityAutonomous SystemsFinanceIndustrial AutomationRegulatory Compliance

Summary

CAGE is a new framework that provides certified authorization for tool-using LLM agents, ensuring actions remain authorized even with small errors in tool returns. It directly certifies joint categorical and numerical perturbations, preventing unsafe actions that pointwise gates might miss.

Large Language Model (LLM) agents often interact with external tools, translating their outputs into real-world actions. These tools return structured data, including categorical fields and numerical values. A critical safety concern arises when runtime permission gates, which authorize actions based on observed tool returns, fail to account for minor errors or uncertainties in how these returns are processed or bound to their source. This research introduces CAGE (Certified Authorization under Typed-Return Uncertainty for Tool-Using Agents), a framework designed to address this vulnerability. CAGE asks whether a proposed action remains authorized within a defined "neighborhood" of plausible, correctly bound tool returns, considering both a single admissible binding fault and bounded numerical drift. The core innovation of CAGE is its ability to certify this joint neighborhood directly. It demonstrates that certifying categorical and numerical channels separately is insufficient, as perturbations safe in isolation can become unsafe when combined. CAGE enumerates discrete branches exactly and certifies continuous perturbations within each branch. Across various settings, CAGE effectively eliminates "in-budget false allows" that accurate but pointwise gates might permit, while still maintaining a useful level of autonomous decision-making.

Why it matters

For professionals deploying AI agents in critical systems, ensuring robust safety and authorization is paramount. CAGE provides a method to certify agent actions against potential tool return errors, significantly reducing the risk of unintended or unsafe real-world consequences.

How to implement this in your domain

  1. 1Evaluate CAGE for enhancing the safety and reliability of tool-using AI agents in your organization's critical applications.
  2. 2Implement certified authorization mechanisms that account for both categorical and numerical uncertainties in tool returns.
  3. 3Develop robust testing protocols that simulate plausible binding faults and numerical drift in tool outputs to validate agent safety.
  4. 4Integrate CAGE-Exact for policies that are executable or CAGE-Lip/CAGE-RS for learned gates under an explicit, measured fidelity assumption.

Original post by Blaise Delattre, Cong Wang, Yang Cao

"arXiv:2607.29190v1 Announce Type: new Abstract: Tool-using LLM agents act on typed tool returns, records pairing provenance and categorical fields with numerical values. Runtime permission gates generally authorize the observed return and action, leaving the decision unprotected…"

View on X

Originally posted by Blaise Delattre, Cong Wang, Yang Cao on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses