Irreversibility Budget Proposed for Fleet-Level AI Agent Risk Control

Bardia Mohammadi, Laurent Bindschaedler· September 2, 2026 View original

Key takeaways

  • Individual agent controls are insufficient for managing cumulative risks from LLM agent fleets.
  • The "irreversibility budget" tracks cumulative value-at-risk across agents and workflows.
  • It prevents overdrawing a principal's risk limit by denying marginal irreversible effects.
  • Accurate, dependency-aware pricing of effects is the main challenge for implementation.

Who benefits

FinanceCybersecurityIT OperationsManufacturingLegal

Summary

A new concept, the "irreversibility budget," is proposed to manage fleet-level risks from LLM agents that perform irreversible actions like moving money or deleting data. This budget acts as a cumulative account of residual value-at-risk, denying marginal effects when the aggregate risk exceeds a principal's limit, addressing the inadequacy of current per-effect controls.

As large language model (LLM) agents increasingly perform actions with irreversible consequences—such as financial transactions, code deployment, data deletion, or information disclosure—current security controls are proving insufficient. These controls typically check individual effects, allowing a fleet of agents, each individually authorized, to collectively exceed a principal's risk tolerance under a shared trigger, even if each local gate functions correctly. To address this, researchers propose the "irreversibility budget," a novel framework for fleet-level risk accounting and admission control. This budget functions as a cumulative account of residual value-at-risk, maintained by a trusted runtime for each principal across all agents, workflows, and tenants. Under this system, irreversibility is treated as a first-class resource. The runtime charges each agent action its residual loss, denying any marginal effect that would cause the aggregate risk to overdraw the budget. A controlled study demonstrated that while per-effect gates allowed fleet-level overdraws up to 48 times the tenant's risk limit, the irreversibility budget successfully kept all correctly charged runs within the specified limit. The primary challenge for deployment remains establishing accurate, dependency-aware pricing for heterogeneous and potentially adversarially declared effects.

Why it matters

This concept is critical for organizations deploying autonomous AI agents, providing a framework to prevent catastrophic cumulative risks that individual agent controls cannot address, ensuring responsible and safe AI operation.

How to implement this in your domain

  1. 1Assess the irreversible actions your LLM agents can perform (e.g., financial, data modification, code deployment).
  2. 2Quantify the potential value-at-risk associated with each irreversible action.
  3. 3Explore developing a centralized "irreversibility budget" mechanism to track cumulative risk across agent fleets.
  4. 4Integrate admission control logic into your agent operating system that checks against this budget before executing irreversible actions.
  5. 5Research and develop robust, dependency-aware pricing models for heterogeneous agent effects to accurately charge against the budget.

Original post by Bardia Mohammadi, Laurent Bindschaedler

"arXiv:2609.00275v1 Announce Type: new Abstract: Fleets of LLM agents now externalize effects that cannot be fully undone: they move money, deploy code, delete data, and disclose information. Current controls check one effect at a time, so a fleet of individually authorized agents…"

View on X

Originally posted by Bardia Mohammadi, Laurent Bindschaedler on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses