Survey Maps AI Agents in Command-Line Environments

Yi Bin, Xiaoyang Yuan, Haoxi Zeng, Wencheng Ye, Wenqi Shao, Chen Qian, Wei Ye, Yujuan Ding, Zheng Wang, Pengpeng Zeng, Jingkuan Song, Heng Tao Shen· August 24, 2026 View original

Key takeaways

  • Terminal agents are AI systems primarily interacting via command-line execution and textual feedback.
  • Their behavior is a complex function of the model, interface, harness, runtime, and environment.
  • Current evaluations often overlook process quality and recovery, focusing mainly on outcomes.
  • Better reporting of system conditions and replayable traces are crucial for robust evaluation.

Who benefits

Software DevelopmentCybersecurityIT OperationsDevOpsAI Development

Summary

This survey provides a comprehensive overview of AI agents operating within command-line environments, establishing a unified framework to understand their architecture, competence acquisition, and evaluation. It highlights that agent behavior is influenced by the model, interface, harness, runtime, and environment, and calls for better reporting of system conditions and replayable traces for evaluation.

This paper presents a comprehensive survey of AI agents that primarily interact with systems through command-line environments. It aims to consolidate disparate research across software engineering, tool use, and computer-use by defining "terminal agents" as systems whose core action-observation loop is mediated by terminal command execution and textual feedback. The survey introduces a seven-dimensional terminal competence profile, linking system architecture, how agents acquire skills, and how they are evaluated. It emphasizes that an agent's actual behavior is a complex interplay of the underlying language model, its interface, the testing harness, the runtime environment, and the target system itself. A key finding is that current evaluations often focus on final outcomes, neglecting the process quality, recovery mechanisms, and governance aspects of agent behavior. The authors advocate for explicit reporting of system and runtime conditions, along with replayable traces and process-level evidence, to enable more thorough and comparable assessments of terminal agents.

Why it matters

Professionals developing or deploying AI agents can use this framework to better design, evaluate, and understand the capabilities and limitations of agents operating in command-line interfaces, improving reliability and transparency.

How to implement this in your domain

  1. 1Adopt the proposed seven-dimensional terminal competence profile for evaluating new AI agent projects.
  2. 2Implement explicit logging of system and runtime conditions for all terminal agent deployments.
  3. 3Develop tools to generate replayable traces of agent interactions for debugging and auditing.
  4. 4Focus evaluations on process quality, error recovery, and governance, not just final outcomes.

Original post by Yi Bin, Xiaoyang Yuan, Haoxi Zeng, Wencheng Ye, Wenqi Shao, Chen Qian, Wei Ye, Yujuan Ding, Zheng Wang, Pengpeng Zeng, Jingkuan Song, Heng Tao Shen

"arXiv:2608.20485v1 Announce Type: new Abstract: Large language model agents increasingly act through terminals, yet existing surveys disperse terminal-mediated behavior across software engineering, tool use, and computer-use research. We regard terminal agents as systems whose do…"

View on X

Originally posted by Yi Bin, Xiaoyang Yuan, Haoxi Zeng, Wencheng Ye, Wenqi Shao, Chen Qian, Wei Ye, Yujuan Ding, Zheng Wang, Pengpeng Zeng, Jingkuan Song, Heng Tao Shen on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI Engineering & DevTools