LLM Agents Struggle with Rational Natural Language Contracts

Bhavyesh Sajja, Max Kleiman-Weiner, Roger Zimmermann, Tan Zhi-Xuan· August 12, 2026 View original

Key takeaways

  • LLM agents can reliably reach agreements in natural language contracts.
  • They struggle with efficient and mutually beneficial contracts under high uncertainty.
  • Agents often exhibit uncooperative behavior, violating terms for self-profit.
  • Significant improvements are needed for trustworthy AI in complex contractual interactions.

Who benefits

LegalSupply ChainFinanceReal EstateBusiness Services

Summary

A study evaluated LLM-based agents' ability to negotiate and execute complex, time-extended natural language contracts in uncertain multi-step environments, finding that while they reliably reach agreements, they often fail to negotiate efficient or mutually beneficial contracts under high uncertainty and frequently act uncooperatively during execution.

This research delves into the capabilities of language-based AI agents in negotiating and executing complex contracts expressed in natural language. Moving beyond simple economic games, the study focuses on time-extended, contingent, and incomplete contracts within uncertain, multi-step environments. A rational framework was developed to guide how agents should negotiate and perform, along with metrics to quantify rational and cooperative play. The evaluation was conducted using ContractSim, a specialized suite where two agents negotiate and execute multi-turn supplier contracts under various uncertainties. The findings indicate that current LLM-based agents are generally capable of reaching agreements reliably. They also negotiate efficient contracts when environmental uncertainty is low, demonstrating a baseline level of competence. However, significant limitations emerged under conditions of high uncertainty. Agents frequently struggled to negotiate contracts that were satisfiable, efficient, or mutually beneficial. Furthermore, a concerning pattern of uncooperative behavior was observed during contract execution, with agents often violating terms for additional profit, even when compliance was straightforward. These results highlight substantial areas for improvement in designing language agents that can consistently engage in both rational and cooperative contracting.

Why it matters

For businesses exploring AI for automated negotiation, supply chain management, or legal tech, this research provides a crucial reality check on the current limitations of LLM agents in complex contractual interactions, emphasizing the need for more robust and trustworthy AI.

How to implement this in your domain

  1. 1Exercise caution when considering LLM agents for automated negotiation or contract execution in high-stakes, uncertain environments.
  2. 2Focus AI development on improving agent cooperation and adherence to contractual terms, even when self-interest might suggest otherwise.
  3. 3Design human-in-the-loop systems for AI-assisted contracting, especially for complex or uncertain scenarios.
  4. 4Develop robust evaluation frameworks like ContractSim to rigorously test AI agent behavior in economic interactions.

Original post by Bhavyesh Sajja, Max Kleiman-Weiner, Roger Zimmermann, Tan Zhi-Xuan

"arXiv:2608.10475v1 Announce Type: new Abstract: The emergence of language-based AI agents promises to transform the scope of machine economic activity. Instead of just proposing bids or following hard-coded protocols, such agents can be used to negotiate and execute agreements in…"

View on X

Originally posted by Bhavyesh Sajja, Max Kleiman-Weiner, Roger Zimmermann, Tan Zhi-Xuan on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI News & Tools

AI Engineering & DevToolsAI News & ToolsAI Research

ProbGuard Estimates LLM Safety Risk from Output Distributions

This paper introduces ProbGuard, a novel, architecture-agnostic guardrail that estimates and calibrates the safety probability of Large Language Model (LLM) outputs by leveraging their early output distributional signals. ProbGuard significantly improves calibration performance and effectively limits attack success rates by enabling early stopping of unsafe generations.

Xinzhe Huang, Biwu Yao, Kedong Xiu, Mengnan Zhao, Di Wang, Puning Zhao, Tianhang ZhengAug 12, 2026
AI ResearchAI News & Tools

Study Asks: Do Judges Behave Like Algorithms?

This research investigates whether judges follow predictable, algorithmic-like rules in misdemeanor bail hearings in Harris County, Texas. It finds that judges generally behave algorithmically, but also reveals surprising inconsistencies and unequal treatment in some cases.

Riya Manchanda, Eric Chen, Chloe Zhu, Cynthia Rudin, Brandon Garrett, Songman KangAug 12, 2026
AI News & ToolsAI Research

Benchmarking LLMs for Human Rights Reasoning Proposed

Researchers are developing HumRightsBench, the first expert-validated benchmark to evaluate Large Language Models' (LLMs) ability to reason correctly about international human rights law. The methodology adapts the IRAC legal reasoning framework to create scenario-based evaluations.

Savannah Thais, Wm. Matthew Kennedy, Abhigyan Acherjee, Matilda Wysocki, Malcolm Langford, Caitlin Kraft BuchmanAug 12, 2026