Red-Teaming in the Public Interest
AI Risk and Vulnerability Alliance (ARVA), Ranjit Singh, Borhane Blili-Hamelin, Carol Anderson, Emnet Tafesse, Briana Vecchione, Beth Duckles, Jacob Metcalf · 2025
The increasing power and availability of generative AI (genAI) systems has led regulators, technologists, and members of the public to call for new safety practices to anticipate harms and protect the public interest. One early and promising approach, drawing from cybersecurity and military practices, is red-teaming, in which designated teams use adversarial methods to identify vulnerabilities in systems. Drawing on 26 semi-structured interviews and participant observation at three public red-teaming events - and based on a collaborative research project between Data & Society and AI Risk and Vulnerability Alliance (ARVA) - Red-Teaming in the Public Interest examines how red-teaming methods are being adapted to evaluate genAI.