Evaluating the performances of artificial systems
Antonio Lieto · 2021
This chapter introduces the main proposals that have been developed in order to evaluate the performance of artificial systems (cognitively inspired or not) and to justify the ascription of faculties from the “cognitive” vocabulary (like “intelligence”) to such systems. After introducing the Turing Test, its problematic aspects, and some of the main modifications proposed (e.g., the Super Turing Test and other variations), we will analyze other frameworks like the Newell Test for a theory of cognition and other tasks and challenges that have been used – with different purposes – as a testbed for the evaluation of artificial systems. These tasks range from the RoboCup World Soccer to the DARPA Challenges for autonomous vehicles to the recently proposed Winograd Schema Challenge and the RoboCup@Home. We will analyze these proposals both in light of their eventual explanatory role in the context of a computationally driven science of the mind and with respect to their actual capacity for evaluating the “intelligence” of artificial systems.