Janus Framework Enhances Long-Horizon AI Agent Safety

Yuan Xiong, Linji Hao, Shizhu He, Yequan Wang, Lijun Li· July 23, 2026 View original

Summary

Researchers propose Janus, a foresight-oriented framework that trains AI guards to anticipate delayed risks from partial trajectories in tool-using agents. Its Vanguard model improves protection against unsafe actions while maintaining task completion rates.

This research introduces Janus, a novel framework aimed at improving the safety of AI agents, particularly those that use tools and operate over long horizons. Unlike traditional content moderation, Janus focuses on preventing operational failures before an agent takes action. The framework trains "guards" to foresee potential delayed risks by analyzing partial agent trajectories. Janus employs multi-agent simulation to synthesize diverse trajectories and then learns a shared policy with two interconnected tasks: anticipating safety-relevant future states and adjudicating safety based on both observed and anticipated future actions. This joint optimization, using CoAA-RL, rewards accurate forecasts that aid in safety judgments. The resulting guard model, named Vanguard, effectively blocks unsafe actions proactively. Evaluations across multiple agent-safety benchmarks show Vanguard significantly enhances protection while minimally impacting benign task completion.

Why it matters

As AI agents become more autonomous and capable of long-term planning, ensuring their safety and preventing unintended consequences is paramount for responsible deployment in critical applications.

How to implement this in your domain

  1. 1Investigate integrating foresight-oriented safety frameworks like Janus into your AI agent development.
  2. 2Develop internal simulations to generate diverse agent trajectories for safety training and evaluation.
  3. 3Prioritize the development of "guard" models that can anticipate and adjudicate risks before agent actions.
  4. 4Establish clear metrics for balancing safety protection with task completion rates in autonomous systems.

Who benefits

Autonomous VehiclesRoboticsAI/ML PlatformsHealthcareDefense

Key takeaways

  • Janus is a framework for proactive, long-horizon AI agent safety.
  • It trains guards to anticipate delayed risks from partial agent trajectories.
  • The Vanguard model improves safety protection while preserving task completion.
  • Multi-agent simulation is used to synthesize diverse training data for safety.

Original post by Yuan Xiong, Linji Hao, Shizhu He, Yequan Wang, Lijun Li

"arXiv:2607.19913v1 Announce Type: new Abstract: Agent safety is moving from content moderation toward preventing operational failures before tool-using agents act. We propose Janus, a foresight-oriented framework for long-horizon agent safety that trains guards to anticipate dela…"

View on X

Originally posted by Yuan Xiong, Linji Hao, Shizhu He, Yequan Wang, Lijun Li on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses