CivBench Benchmarks LLM Agents in Long-Horizon, Tool-Mediated Environments

Austin Tudor David Andrews, Liam Wilkinson, Jamie Heagerty, Harry Coppock, Jakob Nicolaus Foerster, Rui Ponte Costa· September 3, 2026 View original

Key takeaways

  • CivBench evaluates LLM agents in complex, long-horizon, tool-mediated environments like Civilization VI.
  • It introduces metrics (PMR, RAG@10) to assess proactive monitoring and commitment execution.
  • Agents consistently under-monitor strategic state and fail to execute near-term plans despite guidance.
  • The benchmark highlights challenges in instruction following and sustained planning for AI agents.

Who benefits

AI DevelopmentGamingRoboticsAutonomous SystemsSoftware Engineering

Summary

CivBench is an open-source benchmark for evaluating language model agents in the complex, long-horizon game Civilization VI, requiring sustained planning and tool use under partial observability. It introduces metrics like Proactive Monitoring Rate (PMR) and RAG@10 to assess agent behavior, revealing consistent patterns of under-monitoring and failure to execute near-term commitments despite explicit guidance.

This paper introduces CivBench, an open-source benchmark designed to evaluate language model agents within long-horizon, tool-mediated environments, specifically using the game Civilization VI. The benchmark leverages the Model Context Protocol (MCP), with each episode spanning over 300 turns and involving thousands of tool calls across a vast action space. This setup demands continuous planning, state monitoring, and execution from agents, all under conditions of partial observability. The environment provides 76 MCP tools and a narration layer that translates visual game states into structured text. The researchers utilized CivBench to characterize agent behavior across four different model families in 23 admissible runs, noting that this sample is a pilot study rather than a definitive model ranking. They introduced two interface-level metrics: Proactive Monitoring Rate (PMR), which measures how actively agents query latent strategic state, and RAG@10, which assesses whether commitments stated in structured planning reflections are executed within ten subsequent turns. Across the runs, two consistent behavioral patterns emerged despite a shared playbook protocol. Agents frequently under-monitored strategically relevant state, querying victory progress far less often than guided. Additionally, agents often failed to execute near-term commitments from their own planning reflections, with RAG@10 scores ranging from 48.2% to 65.8%. These deviations occurred despite tool access and explicit instructions, suggesting issues with instruction following rather than capability absence. The environment, scenarios, logs, metrics, and analysis pipeline are openly released.

Why it matters

For professionals developing or deploying AI agents in complex, dynamic environments, CivBench provides a robust framework to identify and address critical weaknesses in long-term planning, state monitoring, and tool-use execution, leading to more reliable and autonomous agents.

How to implement this in your domain

  1. 1Explore CivBench as a testing ground for evaluating the long-horizon planning and tool-use capabilities of your AI agents.
  2. 2Adopt metrics like Proactive Monitoring Rate (PMR) and RAG@10 to assess agent performance in complex, multi-step tasks.
  3. 3Analyze agent logs from CivBench to diagnose failures in state monitoring and commitment execution.
  4. 4Develop strategies to improve agent adherence to explicit instructions and proactive information gathering.
  5. 5Contribute to or leverage the open-source CivBench environment to benchmark and improve your own agent architectures.

Original post by Austin Tudor David Andrews, Liam Wilkinson, Jamie Heagerty, Harry Coppock, Jakob Nicolaus Foerster, Rui Ponte Costa

"arXiv:2609.02459v1 Announce Type: new Abstract: We present CivBench, an open-source benchmark for evaluating language model agents in long-horizon, tool-mediated environments through the Model Context Protocol (MCP). A single episode spans 300+ turns and produces thousands of too…"

View on X

Originally posted by Austin Tudor David Andrews, Liam Wilkinson, Jamie Heagerty, Harry Coppock, Jakob Nicolaus Foerster, Rui Ponte Costa on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses