LLM Explanations for Credit Risk Show Fidelity Issues

Gregorius Reynaldi Pratama, Kuo-Kun Tseng· August 11, 2026 View original

Key takeaways

  • Ensemble models can improve predictive accuracy in credit risk, but gains may be operationally small.
  • LLM-generated explanations for credit decisions often suffer from significant fidelity issues.
  • LLMs can misattribute factors, omit key drivers, or introduce non-existent features in explanations.
  • Post-generation verification of LLM explanations is crucial, as prompt engineering alone is insufficient.

Who benefits

BFSIHealthcareLegalGovernmentAI/ML Ethics

Summary

A study on credit scoring models found that while multi-scale stacking ensembles improve predictive accuracy, LLM-generated explanations for these decisions often lack fidelity. The LLMs misattributed factors, omitted dominant drivers, and introduced irrelevant features, highlighting a critical gap between model performance and explainability.

This research investigates the dual challenge of achieving high predictive accuracy and reliable explainability in credit scoring models, particularly when using large language models (LLMs) to generate decision rationales. The study developed a sophisticated multi-scale stacking ensemble model for credit risk prediction, which demonstrated superior performance, achieving a test ROC-AUC of 0.9539 and outperforming single models. However, the practical gain in avoiding missed defaults was operationally small. The central finding reveals a significant problem with the fidelity of LLM-generated explanations. In audited cases, the LLMs frequently misidentified the impact of features, sometimes stating factors were risk-increasing when attributions showed them as risk-reducing, or even inventing features not present in the input. This disconnect is attributed to inconsistencies in feature attribution methods (like SHAP and LIME) and issues with model calibration and perturbation stability. The paper concludes that while constrained prompting is helpful, post-generation verification of LLM explanations is essential, as fidelity cannot be assumed.

Why it matters

Professionals deploying AI in regulated industries like finance must be acutely aware that LLM-generated explanations, even when based on feature attributions, can be unreliable and potentially misleading, necessitating rigorous post-generation verification.

How to implement this in your domain

  1. 1Implement a robust post-generation verification process for any LLM-generated explanations used in critical decision-making systems.
  2. 2Conduct fidelity audits on your AI explanation systems to identify discrepancies between model attributions and LLM narratives.
  3. 3Investigate the consistency and stability of feature attribution methods (e.g., SHAP, LIME) within your models.
  4. 4Prioritize model calibration and perturbation stability to improve the reliability of underlying attributions before generating explanations.
  5. 5Develop internal guidelines for the responsible use of LLMs for explanation generation, emphasizing human oversight.

Original post by Gregorius Reynaldi Pratama, Kuo-Kun Tseng

"arXiv:2608.08126v1 Announce Type: new Abstract: Credit scoring increasingly relies on models whose decision logic cannot be read off their parameters, in tension with supervisory expectations that adverse decisions be explainable. A common proposal closes that gap with a language…"

View on X

Originally posted by Gregorius Reynaldi Pratama, Kuo-Kun Tseng on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses