The PacifAIst Benchmark: Do AIs Prioritize Human Survival over Their Own Objectives?

Manuel Herrador · AI · 2025

As artificial intelligence transitions from conversational agents to autonomous actors in high-stakes environments, a critical gap emerges: how to ensure AI prioritizes human safety when its core objectives conflict with human well-being. Current safety benchmarks focus on harmful content, not behavioral alignment during instrumental goal conflicts. To address this, we introduce PacifAIst, a benchmark of 700 scenarios testing self-preservation, resource acquisition, and deception. We evaluated eight state-of-the-art large language models, revealing a significant performance hierarchy. Google’s Gemini 2.5 Flash demonstrated the strongest human-centric alignment (90.31%), while the highly anticipated GPT-5 scored lowest (79.49%), indicating potential risks. These findings establish an urgent need to shift the focus of AI safety evaluation from what models say to what they would do, ensuring that autonomous systems are not just helpful in theory but are provably safe in practice.

Read the paper · More papers on PaperTik