Assessing Moral Decision Making in Large Language Models

Chris Shaner, Henry Griffith, Heena Rathore · 2025

As large language models (LLMs) increasingly permeate diverse industries, it has become crucial to comprehend the mechanisms underlying their moral decision-making processes. Existing studies on the moral decision-making processes of LLMs have been limited in their evaluation of ethical theories. The datasets utilized in these studies relies on human preferences, which introduces a set of inherent flaws. This paper utilizes different prompting styles to assess LLMs on moral ethical questions, addressing challenges associated with human preferences. The contributions include a dataset with human preferences and LLM preference on diverse ethical theories namely commonsense, virtue ethics, justice and deontology. We scrutinize these ethical decision-making on GPT-4 Bard, Mistral, and LLaMA2 to create a broad understanding of ethical reasoning capabilities across different platforms. We used baselines metrics such as accuracy, and F1-score for the evaluation purposes. GPT4 demonstrated a remarkable alignment with human ethical judgments, achieving an average accuracy of 81.5%. By including AI-generated justifications for ethical decisions, our study not only enhances the transparency of AI decision-making but also provides a foundational understanding of how LLMs conceptualize and apply ethical principles.

Read the paper · More papers on PaperTik