Values in science and AI alignment research

Leonard Dung · Inquiry · 2026

Roughly, empirical AI alignment research (AIA) is an area of AI research which investigates empirically how to design AI systems in line with human values. It involves investigations of techniques such as reinforcement learning from human feedback and adversarial robustness. This paper examines the role of non-epistemic values in AIA. It argues for several claims. First, AIA is highly value-laden along several dimensions, more than many other sciences, since it centrally involves partly evaluative concepts and high inductive risk. Second, this influence of values is currently not managed appropriately and thus threatens the epistemic integrity and ethical beneficence of AIA. Third, in response AIA should strive to achieve value transparency, critical scrutiny from inside and outside the discipline – involving the public –, and to empower actors without strong commercial interests to assume a more important role in AIA. An overarching lesson is that AIA would benefit from the scrutiny of philosophy of science.

Read the paper · More papers on PaperTik