Agent-Safety Benchmarks Lack Consistency, Often Confuse Safety with Capability

Youting Wang, Xiao Han, Dingyan Shang, Yuan Tang, Bowen Liu· August 3, 2026 View original

Key takeaways

  • Current agent-safety benchmarks are inconsistent and often conflate safety with capability.
  • Metrics like F1 can be misleading, allowing non-discriminating policies to score high.
  • Model capability can negatively correlate with misalignment safety.
  • Clear definitions and rigorous validation are essential for meaningful safety claims.

Who benefits

AI DevelopmentCybersecurityRegulatory AffairsEthics & GovernanceSoftware Engineering

Summary

A validity audit of four agent-safety benchmarks reveals significant inconsistencies in how they measure safety, often conflating it with model capability. The study finds that different benchmarks rank models differently, and capability can negatively correlate with misalignment safety, highlighting issues with metrics and panel artifacts.

A recent validity audit critically examines four prominent agent-safety benchmarks (R-Judge, InjecAgent, AgentHarm, AgentDojo), revealing that their scores are often quoted interchangeably despite measuring distinct behaviors. The study highlights fundamental problems with the metrics used, noting that simple "always positive" policies can achieve deceptively high F1 scores on some benchmarks, outperforming models that actually discriminate. The audit found that the three broad-coverage benchmarks rank the same models inconsistently, and the observed disagreements are often artifacts of small panel sizes. Crucially, the research indicates a negative correlation between model capability and misalignment safety on certain panels, suggesting that more capable models might not necessarily be safer in terms of alignment. While AgentHarm showed the strongest association with jailbreak safety, the authors caution that this indicates convergent validity for harmful compliance rather than general safety. The study concludes that any safety claim requires explicit naming of the benchmark, metric, target behavior, and model panel to be meaningful, emphasizing that current benchmarks often confuse safety with mere capability.

Why it matters

For professionals developing and deploying AI agents, understanding the limitations and inconsistencies of safety benchmarks is crucial for accurately assessing and ensuring the ethical and reliable behavior of AI systems.

How to implement this in your domain

  1. 1Critically evaluate the metrics and methodologies of any AI safety benchmark before relying on its results.
  2. 2Develop internal safety testing protocols that go beyond standard benchmarks, focusing on specific failure modes.
  3. 3Ensure clear definitions of "safety" and "capability" when designing or evaluating AI agents.
  4. 4Advocate for standardized, robust, and transparent safety evaluation practices across the AI industry.

Original post by Youting Wang, Xiao Han, Dingyan Shang, Yuan Tang, Bowen Liu

"arXiv:2607.28685v1 Announce Type: new Abstract: Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, InjecAgent, AgentHarm, AgentDojo) as measurements to be validated, running each u…"

View on X

Originally posted by Youting Wang, Xiao Han, Dingyan Shang, Yuan Tang, Bowen Liu on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses