Evaluating the Explainability of Large Language Models for Ethical Decision Making

Gabriela Sanchez San Miguel, Henry Griffith, Jacob Silva, Heena Rathore · 2025

The innate ethical decision-making capabilities of large language models (LLMs) has been previously assessed in the literature using a classification-based workflow. An analysis performed for morally unambiguous scenarios indicated strong alignment between LLM and human ethical judgement. This paper proposes an enhanced explainability-based workflow for exploring the ethical decision-making capabilities of LLMs. In addition to producing binary labels indicating moral acceptability, models are prompted to produce justifications for their decisions. We propose and verify an approach for assessing similarity in the produced justifications across prompting style using latent semantic analysis. We demonstrate that justifications produced for identical moral scenarios are considerably more similar than those produced for arbitrary scenario combinations. We also show that variability in the similarity of justifications produced across prompting styles is negligible as expected for the morally unambiguous scenarios considered.

Read the paper · More papers on PaperTik