Autointerpretability Scores Unreliable for Cross-Model Comparison

Sinie van der Ben, Neele Roch, Anna Hedstr\"om, Mennatallah El-Assady· July 23, 2026 View original

Summary

A study reveals that autointerpretability scores, used to compare sparse autoencoder (SAE) interpretability, are highly unstable and influenced more by evaluation pipeline choices than by the underlying feature properties. This instability undermines meaningful cross-paper comparisons in interpretability research.

When comparing the interpretability of sparse autoencoders (SAEs) across different research papers, a common practice involves using autointerpretability scores. This process typically entails a language model explaining each feature, followed by another language model scoring the explanation. However, new research indicates that these scores are not stable indicators of feature properties. Systematic experiments across various metrics, models, and methodological variations demonstrate that the evaluation pipeline choices introduce more variance than the architectural differences of the SAEs themselves. Specific metrics exhibit distinct instability profiles, with detection being the most stable and fuzzing proving unreliable. Crucially, feature rankings do not remain consistent across different conditions, masking per-feature instability. These findings suggest that current cross-paper comparisons based on autointerpretability scores may be misleading, reflecting evaluation setup differences rather than true architectural merits. The paper proposes a variance decomposition approach, a Stability Check, and a Minimum Reporting Checklist to improve evaluation reliability.

Why it matters

Professionals relying on interpretability scores to select or compare AI models need to be aware of their inherent instability, which can lead to flawed conclusions and hinder progress in developing truly understandable AI systems.

How to implement this in your domain

  1. 1Adopt the proposed Stability Check and Minimum Reporting Checklist when evaluating AI interpretability methods.
  2. 2Prioritize interpretability metrics that have demonstrated higher stability, such as detection, over less reliable ones like fuzzing.
  3. 3Conduct sensitivity analyses on interpretability evaluation pipelines to understand the impact of methodological choices.
  4. 4Advocate for standardized and robust evaluation protocols within AI research and development teams.

Who benefits

AI/ML DevelopmentResearch & AcademiaRegulatory ComplianceSoftware Engineering

Key takeaways

  • Autointerpretability scores for SAEs are highly sensitive to evaluation pipeline choices, not just model architecture.
  • Methodological variations often outweigh architectural differences in influencing interpretability scores.
  • Different interpretability metrics have varying stability profiles, with some being unreliable.
  • Reliable evaluation tools and standardized reporting are crucial for advancing AI interpretability research.

Original post by Sinie van der Ben, Neele Roch, Anna Hedstr\"om, Mennatallah El-Assady

"arXiv:2607.19386v1 Announce Type: new Abstract: Cross-paper comparison of sparse autoencoder (SAE) interpretability often relies on autointerpretability scores. In this evaluation pipeline, a language model (LM) explains each feature, and another LM scores the explanation. For th…"

View on X

Originally posted by Sinie van der Ben, Neele Roch, Anna Hedstr\"om, Mennatallah El-Assady on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses