INTERPRETING RANKING MODELS: FROM FEATURE ATTRIBUTIONS TO MECHANISTIC COMPREHENSION
Tanya Chowdhury · ScholarWorks@UMassAmherst (University of Massachusetts Amherst) · 2026
The rapid adoption of Neural Ranking Models (NRMs) in Information Retrieval (IR) underscores the urgent need for robust interpretability methods to ensure their reliability and transparency. Despite their widespread use, how NRMs model query-document relevance remains poorly understood, limiting their application in high-stakes domains. This dissertation addresses this gap by developing generalizable, scalable, and theoretically grounded interpretability methods for ranking systems, spanning both post-hoc feature attribution and mechanistic analysis of internal model structure. First, we extend the Local Interpretable Model-agnostic Explanations (LIME) framework to NRMs, proposing RankLIME. This framework is able to generate local feature attributions for pointwise, pairwise, as well as listwise ranking models. Using RankLIME, practitioners gain a scalable tool to interpret the individual contributions of features to the ranking decisions of complex neural models. Next, recognizing the inconsistencies and contradictions in existing empirical ranking feature attribution methods, we adopt an axiomatic approach, establishing a set of fundamental principles that ranking-based attribution methods should satisfy. To this end, we introduce the RankSHAP framework, using the Shapley value from cooperative game theory. RankSHAP ensures that feature attributions are generalizable, consistent, and aligned with human intuition. Extensive experiments on benchmark datasets and diverse ranking models, coupled with a user study, validate the effectiveness of RankSHAP. An axiomatic analysis further clarifies the compliance and deviations of other attribution methods relative to the proposed axioms, enhancing our understanding of their reliability. Beyond feature attributions, this dissertation delves into the internal mechanisms of transformer-based ranking LLMs to understand the abstract features these models internally encode to model relevance in ranking decisions. Through a layer-wise analysis of LLM neuron activations, we investigate whether known statistical and human-engineered features—such as term frequency, inverse document frequency, and covered query term ratio—are embedded within network representations across different LLM configurations. These findings provide insights into the internal workings of ranking LLMs, laying the groundwork for designing more transparent and effective ranking systems. This dissertation further advances mechanistic interpretability by introducing a coalitional framework for studying how neurons cooperate within fine-tuned ranking LLMs. While prior analyses often treat neurons independently, empirical evidence suggests that task-specific features in LoRA-tuned models emerge through interactions among groups of neurons, particularly in mid-layer MLP blocks. To capture this structure, the dissertation models neurons as agents in a hedonic game, where preferences are defined by their synergistic contributions to layer-local computations. Using top-responsive utilities and the PAC-Top-Cover algorithm, it identifies stable coalitions of neurons whose joint ablations produce non-additive effects, and tracks how these coalitions persist, split, merge, or disappear across layers. Applied to LLaMA, Mistral, and Pythia rerankers fine-tuned for scalar output tasks, this framework discovers coalitions with consistently higher synergy than clustering-based baselines. By revealing higher-order functional units beyond individual neurons, hedonic coalitions provide a new lens into how fine-tuned language models encode task-relevant features and computations. Taken together, these contributions connect feature attribution, axiomatic analysis, probing, and coalition-based mechanistic interpretability into a broader framework for understanding neural ranking systems. More broadly, they point toward an emerging research agenda for the community: building a systematic science of reverse-engineering large language models, with particular attention to the emergent features and statistical regularities encoded in their weights and internal computations, especially within MLP submodules. Advancing this agenda will require weight-based and mechanism-based interpretability methods capable of recovering the latent priors, abstractions, and invariances that models internalize during training. Such progress would not only deepen scientific understanding of how learning systems compress, organize, and represent knowledge, but also help extract actionable insights about the real-world phenomena these models capture. Ultimately, this dissertation contributes toward bridging interpretability and performance in neural ranking models, and toward a more principled foundation for transparent, accountable, and scientifically informative machine learning systems.