EvoHarness-RL Enables LLM Agents to Learn Self-Evolving Runtime Harnesses.

Xuying Ning, Dongqi Fu, Tianxin Wei, Hanqing Zeng, Yuanchen Bei, Bingxuan Li, Zihao Li, Qifan Wang, Xiang Shen, Yifan Wu, Jiayi Liu, Hong Li, Yinglong Xia, Xiangjun Fan, Hanghang Tong, Jingrui He· August 7, 2026 View original

Key takeaways

  • LLM agents can learn to manage their own external state and tool use through trainable harness policies.
  • EvoHarness-RL introduces "Belief, Progress, and Experience" as key policy-facing harness states.
  • The framework enables "harness annealing" (internalizing patterns) and "harness evolution" (refining state).
  • Trainable harness policies significantly improve long-horizon task success for LLM agents.

Who benefits

Software DevelopmentRoboticsCustomer ServiceGamingEducation

Summary

EvoHarness-RL is a new framework that allows long-horizon LLM agents to learn and deploy self-evolving runtime harness policies, managing external state for tasks like tool invocation and progress tracking. This approach improves task success by internalizing recurring harness-use patterns and refining external state over time.

Long-horizon LLM agents often struggle with managing external state, tracking progress, and effectively using tools across complex interactions. Current solutions typically rely on manual engineering, prompts, or heuristics for these "harness" functions, leading to inefficiencies and limited adaptability. EvoHarness-RL introduces a novel approach where agents learn and deploy their own harness policies. It exposes Belief, Progress, and Experience (BPE) as policy-facing harness states, enabling agents to construct and update this external state dynamically. The framework uses supervised fine-tuning for basic harness actions and cost-aware GRPO for selective state management. Empirical results on ALFWorld show EvoHarness-RL achieving high success rates. It demonstrates "harness annealing," where agents internalize common patterns and shift towards selective external state access, and "harness evolution," where the harness refines into a compact, task-adaptive state. This suggests that trainable policies for external workspaces significantly benefit long-horizon agents beyond just stronger tools or larger memories.

Why it matters

This research offers a path to more autonomous and capable LLM agents by enabling them to intelligently manage their own external context and tools, reducing manual engineering effort and improving performance on complex, multi-step tasks.

How to implement this in your domain

  1. 1Investigate EvoHarness-RL's principles for designing more robust LLM agents for multi-step tasks.
  2. 2Experiment with dynamic external state management for agents, moving beyond static prompting or hardcoded tool use.
  3. 3Consider implementing "harness annealing" concepts to optimize agent interaction with external resources over time.
  4. 4Explore how to expose "Belief, Progress, and Experience" as structured states for agent policy learning.

Original post by Xuying Ning, Dongqi Fu, Tianxin Wei, Hanqing Zeng, Yuanchen Bei, Bingxuan Li, Zihao Li, Qifan Wang, Xiang Shen, Yifan Wu, Jiayi Liu, Hong Li, Yinglong Xia, Xiangjun Fan, Hanghang Tong, Jingrui He

"arXiv:2608.05446v1 Announce Type: new Abstract: Long-horizon LLM agents increasingly rely on external execution support to maintain state, track progress, invoke tools, verify outcomes, and reuse experience across interactions. However, effective harness use raises two coupled ch…"

View on X

Originally posted by Xuying Ning, Dongqi Fu, Tianxin Wei, Hanqing Zeng, Yuanchen Bei, Bingxuan Li, Zihao Li, Qifan Wang, Xiang Shen, Yifan Wu, Jiayi Liu, Hong Li, Yinglong Xia, Xiangjun Fan, Hanghang Tong, Jingrui He on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses