GameXpert-Bench Evaluates AI Coding Agents for Game Development.

Kun Chen, Haorong Hong, Peizhong Gao, Jianfeng Lin, Tongxu Luo, Yuxuan Xie, Chenxu Liu, Jieling He, Zhongyuan Liu, Zeno Zeng· August 25, 2026 View original

Key takeaways

  • GameXpert-Bench is a new benchmark for evaluating AI coding agents in game development across three stages: generation, fixing, and optimization.
  • It assesses agents on program logic, visual/audio content, interfaces, interaction, and playability.
  • Current agents excel at initial generation and explicit requirements but struggle with defect discovery and preserving functionality.
  • The benchmark uses live interaction, behavioral tests, and product criteria for evaluation.

Who benefits

Software DevelopmentGamingEdTechAI-EngineeringCreative Industries

Summary

Researchers introduce GameXpert-Bench, a new benchmark suite designed to comprehensively evaluate AI coding agents across the entire game development lifecycle, including initial generation, bug fixing, and multi-turn optimization. The benchmark reveals current agents are better at foundational tasks than at defect discovery or maintaining functionality across changes.

Large language models (LLMs) are increasingly capable of acting as coding agents, generating complete games from natural language descriptions. Game development is a particularly demanding domain, requiring the seamless integration of program logic, visual and audio assets, user interfaces, and playability. Existing benchmarks often fall short by evaluating only the final game product or isolated stages of development. This new research emphasizes the need to assess the full development process. The authors analyzed human-agent game development trajectories and identified three critical stages: initial game generation, bug diagnosis and repair, and iterative optimization. To address this, they introduce GameXpert-Bench, a comprehensive benchmark suite with three tracks corresponding to these stages. GameGen assesses the creation of a complete game from a single request. GameFix evaluates an agent's ability to diagnose and repair reported or self-discovered defects. GameOpt measures cumulative optimization through chains of requests, mimicking real-world iterative development. Each track is evaluated using a combination of live game interaction, deterministic behavioral tests, and final product criteria with regression checks. The suite includes 97 generation tasks, 100 repair tasks with injected bugs, and 17 optimization chains. The findings indicate that current AI agents are more proficient at producing playable foundations and implementing explicit requirements than they are at discovering defects, verifying runtime behavior, or preserving functionality during changes.

Why it matters

This benchmark provides a standardized, comprehensive way to assess and drive improvements in AI coding agents for complex software development, particularly in creative domains like game development.

How to implement this in your domain

  1. 1Utilize GameXpert-Bench to evaluate the capabilities of internal or third-party AI coding agents for software development.
  2. 2Focus AI agent development efforts on improving bug diagnosis, repair, and maintaining functionality across iterative changes.
  3. 3Integrate multi-turn optimization and defect discovery into the training objectives for coding agents.
  4. 4Adopt a lifecycle-oriented approach to AI agent evaluation, mirroring the stages of human development.

Original post by Kun Chen, Haorong Hong, Peizhong Gao, Jianfeng Lin, Tongxu Luo, Yuxuan Xie, Chenxu Liu, Jieling He, Zhongyuan Liu, Zeno Zeng

"arXiv:2608.21833v1 Announce Type: new Abstract: Recent large language models (LLMs) can operate as coding agents that build complete games from natural language requests. Game development is especially demanding because program logic, visual and audio content, interfaces, interac…"

View on X

Originally posted by Kun Chen, Haorong Hong, Peizhong Gao, Jianfeng Lin, Tongxu Luo, Yuxuan Xie, Chenxu Liu, Jieling He, Zhongyuan Liu, Zeno Zeng on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses