Against Purposeful Artificial Intelligence Failures

Roman V. Yampolskiy · SuperIntelligence - Robotics - Safety & Alignment · 2024

Thousands of researchers are currently of opinion that advanced artificial intelligence could cause significant damage if developed without appropriate safety measures, but such measures are not currently deployed or even developed. A fringe theory suggests that a severe AI accident could serve as a fire alarm for humanity to take existential dangers of AI seriously and so it is desirable to create such a failure on purpose ASAP to prevent greater harm in the future. In this paper we rely on analogy to inoculation theory to argue against creating purposeful AI failures.

Read the paper · More papers on PaperTik