Item Response Theory Improves AI Safety Benchmark Evaluation

Joshua Fonseca Rivera (Independent), Neil Shah (Independent), David Demitri Africa (UK AI Security Institute), Konstantinos Voudouris (UK AI Security Institute)· August 6, 2026 View original

Key takeaways

  • Item Response Theory (IRT) offers a robust statistical framework for evaluating LLM safety.
  • Three key factors (refusal strictness, truthfulness, contextual harm) explain most model safety variance.
  • IRT can significantly reduce evaluation costs by identifying efficient, psychometrically sound test items.
  • The method helps audit models, detecting sandbagging or changes in API behavior.

Who benefits

AI DevelopmentRegulatory ComplianceCybersecurityResearch & DevelopmentEthics & Governance

Summary

This paper applies Item Response Theory (IRT) to analyze eight AI safety benchmarks across 192 language models, revealing three key factors: refusal strictness, truthfulness, and contextual harm. IRT enables more efficient and reliable safety evaluations by identifying psychometrically sound items and detecting model sandbagging.

Evaluating the safety of language models is crucial, but current benchmark scores are often difficult to trust due to duplication, high correlation, and models potentially "sandbagging" or detecting evaluation. This makes it hard to accurately interpret differences in model safety. Researchers have now applied Item Response Theory (IRT), a statistical method used in psychometrics, to address these issues. They conducted the largest psychometric analysis of LLM safety evaluations to date, fitting IRT models to eight safety benchmarks across 192 language models. The analysis identified three primary, interpretable factors explaining most of the variance between models: refusal strictness, truthfulness, and contextual harm. Furthermore, IRT allows for the selection of psychometrically robust items, significantly reducing evaluation costs by requiring fewer items to recover full benchmark scores. It also provides a tool for auditing individual models, capable of detecting sandbagging or changes in API-backed models.

Why it matters

Professionals involved in AI development, deployment, and governance can use IRT to create more reliable, efficient, and interpretable safety evaluations for large language models, ensuring safer AI systems.

How to implement this in your domain

  1. 1Adopt Item Response Theory (IRT) as a methodology for designing and analyzing internal AI safety benchmarks.
  2. 2Analyze existing safety evaluation datasets using IRT to identify redundant or less informative items.
  3. 3Develop adaptive testing strategies based on IRT to reduce the cost and time required for model safety evaluations.
  4. 4Implement IRT-based auditing tools to monitor model behavior for consistency and detect potential sandbagging in production.
  5. 5Communicate safety evaluation results using IRT-derived factors for clearer, more actionable insights.

Original post by Joshua Fonseca Rivera (Independent), Neil Shah (Independent), David Demitri Africa (UK AI Security Institute), Konstantinos Voudouris (UK AI Security Institute)

"arXiv:2608.05086v1 Announce Type: new Abstract: Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores are hard to trust and interpret, because benchmarks duplicate one another, correlate heavily,…"

View on X

Originally posted by Joshua Fonseca Rivera (Independent), Neil Shah (Independent), David Demitri Africa (UK AI Security Institute), Konstantinos Voudouris (UK AI Security Institute) on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses