Poster: Exploring Explainability Techniques for Large Language Model Classification Predictions
Sriya Ayachitula · 2024
Large Language Models (LLMs), such as OpenAl's GPT series, are fundamental tools in machine learning that excel in various tasks, including in-context learning. However, these models often operate as “black boxes,” offering limited insight into their decision-making processes. This paper addresses the challenge of explainability in LLMs through a novel approach, using simpler linear models like the Passive-Aggressive Classifier (PAC) as an initial jump-start mechanism to teach the LLM various rules using prompt engineering. We extract and analyze pertinent positives (PPs) as explainable rules and formulate Disjunctive Normal Form (DNF) rules. When taught to the LLM, these rules provide a clear and logical explanation of the decision-making process, enhancing the transparency of LLM predictions. Our methodology improves the interpretability of LLM outputs. The findings suggest that even complex LLM decisions can be distilled into understandable logic, facilitating better user comprehension and trust in classification predictions.