The Maverick Nanny with a Dopamine Drip: Debunking Fallacies in the Theory of AI Motivation
Richard P. W. Loosemore · National Conference on Artificial Intelligence · 2014
We examine the validity of various widely publicized scenarios that predict dire and almost unavoidable negative behavior from future artificial general intelligences, even if they are programmed to be friendly to humans. This entire class of doomsday scenarios is found to be logically incoherent at such a fundamental level that they can be dismissed as extremely implausible. In addition, we find that the most likely outcome of attempts to build such unstable AGI systems would be that the system itself would immediately detect the offending logical contradiction in its design, and spontaneously self-modify to make itself safe.