FM-Bench Evaluates AI Agents in Long-Horizon Management Tasks

Tianyou Wang, Chongyang Gao, Kezhen Chen, Chen Dong, Yinghao He, Donghan Li, Wangcheng Xu, Hongjiu Zhang, Chi Li· August 20, 2026 View original

Key takeaways

  • AI agents can sustain decision-making over long horizons in complex simulations.
  • Managerial behavior, not model size or cost, drives long-term performance.
  • Top models still struggle with learning hidden market dynamics and effective memory management.
  • Benchmarks like FM-Bench are crucial for evaluating strategic AI capabilities.

Who benefits

GamingStrategic ConsultingFinancial ServicesOperations ManagementAI Development

Summary

FM-Bench is a new benchmark that tests AI agents' sustained decision-making over 20 in-game years in a football club management simulation, revealing that managerial behavior, not model scale or cost, dictates performance. All 15 frontier models completed the horizon, with claude-fable-5 topping the board, but none learned hidden market prices or managed memory effectively.

While language model agents are becoming proficient at bounded tasks, their ability to make effective, sustained decisions over long periods, where actions have cumulative effects and the environment adapts, remains largely unmeasured. To address this, FM-Bench (Football Management Benchmark) has been introduced. FM-Bench simulates a football club manager's role over 20 in-game years, requiring an LLM agent to utilize 26 tools and make hundreds of decisions, including drafting, trading, contract negotiation, investment, and lineup setting. A deterministic engine scores performance without human or LLM judgment. The benchmark includes both solo play against a frozen world and an "Arena" where models compete head-to-head. All 15 frontier models tested successfully completed the 20-year horizon, unlike scripted baselines that often failed. Claude-fable-5 emerged as the top performer in both solo and Arena modes, though titles rotated among models in the Arena. Interestingly, model scale, price, or vendor did not predict performance. Instead, managerial behaviors like strategic investment timing and proactive contract renewals were key differentiators. Models struggled with learning hidden market prices and exhibited issues with self-managed memory, either growing archives indefinitely or rewriting plans every season.

Why it matters

For professionals developing or deploying AI for complex, long-term strategic decision-making, FM-Bench highlights that sustained performance depends more on nuanced behavioral strategies than raw model size, and reveals persistent challenges in memory management and learning from experience.

How to implement this in your domain

  1. 1Adopt long-horizon, multi-agent simulation benchmarks for evaluating AI systems intended for strategic decision-making.
  2. 2Focus AI agent development on improving strategic behavioral capabilities rather than solely increasing model scale.
  3. 3Investigate advanced memory management techniques for AI agents to prevent issues like unbounded growth or constant plan rewriting.
  4. 4Design AI agents to learn from complex, dynamic environments, including implicit market dynamics.

Original post by Tianyou Wang, Chongyang Gao, Kezhen Chen, Chen Dong, Yinghao He, Donghan Li, Wangcheng Xu, Hongjiu Zhang, Chi Li

"arXiv:2608.18423v1 Announce Type: new Abstract: Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over long horizons, where actions have cumulative consequences and the environment responds to their choices, remains large…"

View on X

Originally posted by Tianyou Wang, Chongyang Gao, Kezhen Chen, Chen Dong, Yinghao He, Donghan Li, Wangcheng Xu, Hongjiu Zhang, Chi Li on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses