LLM Interpretability: Tracing How LLMs Answer Factual Queries and Math Questions

Shweta Agrawal, He Nan Tony Li, Lucas Lu · 2025

In this paper, we study how LLMs like GPT store and retrieve facts to answer factual queries. Using a pretrained GPT-2 model, we evaluate factual accuracy on knowledge datasets, math questions, and hand-crafted prompts, employing metrics such as weighted first token accuracy, F1 score for token overlap, perplexity, and BLEU. We leverage interpretability techniques like attention visualization, logit-lens analysis, and causal tracing to identify layers responsible for knowledge retrieval. To enhance the model's factual and mathematical capabilities, we implement prompt engineering, retrieval-augmented generation (RAG), and fine-tuning, comparing their impact on accuracy and fact retrieval mechanisms.

Read the paper · More papers on PaperTik