Inference to the Best Explanation in Large Language Models

Dhairya Dalal, Marco Valentino, Andre Freitas, Paul Buitelaar · 2024

While Large Language Models (LLMs) have found success in real-world applications, their underlying explanatory process is still poorly understood.This paper proposes IBE-Eval, a framework inspired by philosophical accounts on Inference to the Best Explanation (IBE) to advance the interpretation and evaluation of LLM explanations.IBE-Eval estimates the plausibility of natural language explanations through a combination of explicit logical and linguistic features including: consistency, parsimony, coherence, and uncertainty.Extensive experiments are conducted on Causal Question Answering (CQA), where IBE-Eval is tasked to select the most plausible causal explanation amongst competing ones generated by the LLM (e.g.GPT 3.5 or LLaMA 2).The experiments reveal that IBE-Eval can successfully identify the best explanation with up to 77% accuracy (≈ 27% above random), improving upon a GPT 3.5-as-a-judge baseline (≈ +17%) while being intrinsically more efficient and interpretable.Additional analysis suggests that, despite LLM-specific variances, generated explanations tend to conform to IBE criteria and that IBE-Eval is significantly correlated with human judgment, opening up opportunities for future development of automated explanation verification tools. Inference to the Best Explanation (IBE)

Read the paper · More papers on PaperTik