Recast Forecasts LLM Safety Risks in Multi-Turn Interactions

Shi Lin, Peng Qian, Dinghao Liu, Renjie Sun, Sifan Wu, Dezhang Kong, Chenpei Wang, Xun Wang· July 30, 2026 View original

Summary

Recast is a new framework designed to predict trajectory-level safety risks in multi-turn interactions with large language models, moving beyond reactive, turn-level violation detection. It identifies emerging risks before safety failures occur by analyzing dialogue progression and historical context.

As large language models transition from simple assistants to autonomous agents, ensuring their safety requires a more proactive approach than merely detecting violations after they occur. Current safeguards are largely reactive, missing the ability to predict how risks might evolve over extended interactions. Malicious intent can be subtly introduced across multiple turns, gradually building up to a safety failure. To address this, researchers propose Recast, a framework that forecasts safety risks at the trajectory level. Recast first gathers relevant evidence from both the immediate dialogue and long-term historical context using a dual-scale view. It then models how compositional risks evolve by capturing the current risk configuration and its temporal dynamics. Finally, a causal temporal encoder learns these latent risk evolution patterns to predict when future risk emergence turns might occur. Extensive experiments across seven risk categories show Recast can predict 88.3% of future safety failures with an average lead time of 2.41 turns, while maintaining a low false alarm rate of 12.3%. This demonstrates its effectiveness in identifying emerging risks preemptively.

Why it matters

For professionals developing or deploying LLMs, this research provides a crucial tool for enhancing safety and reliability by enabling proactive risk mitigation, preventing potential misuse or harmful outputs before they manifest. It shifts the paradigm from reactive fixes to predictive prevention.

How to implement this in your domain

  1. 1Integrate trajectory-level risk forecasting into LLM safety pipelines.
  2. 2Develop monitoring systems to track compositional risk evolution in real-time LLM interactions.
  3. 3Utilize dual-scale trajectory views to gather comprehensive risk-relevant evidence.
  4. 4Train LLM safety teams on identifying and responding to forecasted risks.
  5. 5Implement preemptive intervention mechanisms based on risk predictions.

Who benefits

Software DevelopmentCybersecurityCustomer ServiceGamingEducation

Key takeaways

  • Recast predicts LLM safety risks at the trajectory level, not just turn-by-turn.
  • It uses dual-scale context to understand risk evolution over time.
  • The framework can forecast future safety failures with significant lead time.
  • Proactive risk mitigation is possible before violations fully manifest.

Original post by Shi Lin, Peng Qian, Dinghao Liu, Renjie Sun, Sifan Wu, Dezhang Kong, Chenpei Wang, Xun Wang

"arXiv:2607.26820v1 Announce Type: new Abstract: As large language models (LLMs) evolve from standalone assistants into autonomous agents, ensuring their safety requires shifting beyond pointwise risk assessment to understand how risks emerge and unfold over long-horizon trajector…"

View on X

Originally posted by Shi Lin, Peng Qian, Dinghao Liu, Renjie Sun, Sifan Wu, Dezhang Kong, Chenpei Wang, Xun Wang on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses