Recursive Synthesis Creates Thousands of Long-Horizon AI Tasks.

Zhongzhi Li, Yucheng Shi, Zongxia Li, Ruhan Wang, Anhao Li, Zixun Huang, Junyao Yang, Lei Ke, Ninghao Liu, Haitao Mi, Leowei Liang· August 7, 2026 View original

Key takeaways

  • Recursive synthesis can dramatically reduce the cost and increase the scale of long-horizon training data generation.
  • The RST framework ensures consistency between instructions, environments, solutions, and verifiers.
  • Synthesized data can significantly improve the performance of LLM agents on complex tasks.
  • The recursive process allows for the generation of increasingly difficult and diverse tasks.

Who benefits

AI DevelopmentRoboticsSoftware DevelopmentGamingAutomotive

Summary

This paper introduces Recursive Synthetic Terminal Tasks (RST), a framework that recursively synthesizes high-quality, long-horizon training data for terminal agents at low cost. RST generates tens of thousands of increasingly difficult tasks, significantly improving LLM agent performance on complex benchmarks through supervised fine-tuning and PPO.

Generating high-quality, long-horizon training data for terminal agents is notoriously expensive and time-consuming, often costing hundreds to thousands of dollars per task due to the need for consistent instructions, environments, solutions, and verifiers. Human authoring struggles to scale, and direct generation by large language models (LLMs) frequently fails to maintain these critical dependencies.This research presents Recursive Synthetic Terminal Tasks (RST), a novel framework designed for the scalable and verified synthesis of long-horizon terminal-agent tasks. RST begins with verified seed tasks and then recursively extends their reference solutions, realigns the verifier and instructions, and validates the new workflow in a sandbox. Accepted tasks are then reused as seeds for subsequent rounds.Over fifteen recursive rounds, RST successfully produced 37,484 synthesized tasks at approximately $0.05 per task. The difficulty of these tasks increased substantially with each round, with reference solutions growing significantly in length and complexity. The synthesized data proved highly effective for training: supervised fine-tuning and agentic PPO on these trajectories led to significant performance improvements (up to 10 points and 41.2% relative gains) for Qwen3.5 models on challenging terminal benchmarks. The process showed no signs of ceiling, indicating its potential for even greater scale.

Why it matters

AI engineers and product developers can leverage this method to generate vast amounts of complex, high-quality training data for autonomous agents at a fraction of the traditional cost, accelerating the development of advanced AI systems.

How to implement this in your domain

  1. 1Explore recursive synthesis techniques for generating complex, long-horizon training data.
  2. 2Design robust verification and validation steps within your data generation pipelines.
  3. 3Implement self-correction or realignment mechanisms for instructions and solutions in synthetic data.
  4. 4Utilize synthesized data for supervised fine-tuning and reinforcement learning of autonomous agents.
  5. 5Investigate cost-effective methods for scaling data generation for AI training.

Original post by Zhongzhi Li, Yucheng Shi, Zongxia Li, Ruhan Wang, Anhao Li, Zixun Huang, Junyao Yang, Lei Ke, Ninghao Liu, Haitao Mi, Leowei Liang

"arXiv:2608.05466v1 Announce Type: new Abstract: High-quality long-horizon training data for terminal agents is expensive to produce, often costing hundreds to thousands of dollars per task, because each task must keep the instruction, environment, reference solution, and verifier…"

View on X

Originally posted by Zhongzhi Li, Yucheng Shi, Zongxia Li, Ruhan Wang, Anhao Li, Zixun Huang, Junyao Yang, Lei Ke, Ninghao Liu, Haitao Mi, Leowei Liang on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses