New Method Measures AI Explainer Stability and Reliability.

Eddie Conti, \'Alvaro Parafita, Axel Brando· August 5, 2026 View original

Key takeaways

  • Attribution methods (AMs) can produce variable explanations due to stochasticity.
  • A new framework measures AM stability via attribution separability.
  • It identifies the largest index for which feature rankings are reliable.
  • The framework helps compare AM robustness across datasets, improving trustworthiness.

Who benefits

HealthcareBFSIAutomotiveLegalCompliance

Summary

This paper proposes a distribution-based framework to measure the stability of attribution methods (AMs), which explain black-box AI models. The approach captures the degree of separability in ranked attribution vectors, identifying the largest index for reliable feature ranking and comparing AM robustness across datasets.

Attribution methods (AMs) are widely used to explain the decisions of black-box AI models by assigning importance scores to features. However, many AMs suffer from variability in their attribution scores due to stochastic components, making their explanations less reliable. This research introduces a novel distribution-based framework designed to quantify the stability of these attribution scores. The proposed method focuses on understanding the "separability" within ranked attribution vectors. It can identify the highest index at which a feature ranking remains dependable, providing a concrete measure of an explainer's reliability. Furthermore, the framework can be extended to compare different attribution methods based on how robust their rankings are across various datasets. Experiments demonstrate the practical application of this approach in evaluating the stability of AMs, offering a complementary criterion for assessing their trustworthiness.

Why it matters

Professionals relying on AI explanations for critical decisions, such as in healthcare or finance, can use this framework to select more stable and trustworthy attribution methods, improving confidence in model interpretability and accountability.

How to implement this in your domain

  1. 1Evaluate the stability of your current attribution methods using the proposed distribution-based framework.
  2. 2Identify the "largest reliable index" for feature rankings to understand the trustworthiness of top features.
  3. 3Compare different attribution methods based on their ranking robustness across your datasets.
  4. 4Integrate stability metrics into your model interpretability evaluation pipeline.
  5. 5Use these insights to select more reliable explainers for high-stakes AI applications.

Original post by Eddie Conti, \'Alvaro Parafita, Axel Brando

"arXiv:2608.02697v1 Announce Type: new Abstract: Attribution methods (AMs) assign an importance score to each feature and are widely adopted to explain black-box models. However, most methods can produce variable attribution scores due to stochastic components in their definition.…"

View on X

Originally posted by Eddie Conti, \'Alvaro Parafita, Axel Brando on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses