Benchmarks, test beds, controlled experimentation, and the design of agent architectures
Steve Hanks, Martha E. Pollack, Paul R. Cohen · 1993
The methodological underpinnings of AI are slowly changing. Benchmarks, testbeds, and controlled experimentation are becoming more common. While we are optimistic that this change can solidify the science of AI, we also recognize a set of difficult issues concerning the appropriate use of this methodology. We discuss these issues as they relate to research on agent design. We survey existing testbeds for agents, and argue for appropriate caution in their use. We end with a debate on the proper role of experimental methodology in the design and validation of planning agents. Hanks was supported in part by NSF grants IRI-9008670 and IRI-9206733. y Pollack was supported by the Air Force Office of Scientific Research, Contracts F49620-91-C-0005 and F49620-92-J-0422, by the Rome Laboratory (RL) and the Defense Advanced Research Projects Agency, Contract F30602-93-C-0038, and by an NSF Young Investigator's Award, IRI-9258392 z Cohen was supported in part by the Defense Advanced Researc...