Comparison of explainability methods for hallucination analysis in LLMs

Ioannis Papagiannopoulos, Hercules Koutalidis, Panagiota Rempi, Christos Ntanos, Dimitris Askounis · Open Research Europe · 2025

The primary obstacle to the safe application of artificial intelligence (AI) in sensitive fields such as healthcare and law is the phenomenon of Large Language Models (LLM) hallucinations, which produce fluent yet factually incorrect or illogical outputs. This paper investigates hallucinations as reasoning failures, distinguishing between factual inaccuracies and flawed logical derivations. We examine the limitations of current interpretability tools including Retrieval-Augmented Generation (RAG), attention mechanisms, SHapley Additive exPlanations (SHAP), and Local Interpretable Model-agnostic Explanations (LIME), which often fail to identify the root causes of these errors. To address this gap, we analyze reasoning in LLMs through two complementary paradigms: deterministic rule-based and stochastic data-driven approaches. We argue that hybrid reasoning techniques can improve the accountability and transparency of model behavior. A real-world use case in decision support systems illustrates the risks associated with transparent reasoning in high-stakes applications. We also present a comparative framework to evaluate the ability of common explainability techniques to detect and account for hallucinated outputs. Finally, we suggest useful techniques to improve trustworthiness, such as training data attribution, confidence calibration, and prompt-based abstention mechanisms. Our results highlight how crucial reasoning-aware explainability is as a starting point for ensuring that LLMs behavior complies with safety, dependability, and responsible AI deployment standards.

Read the paper · More papers on PaperTik