On Evaluating Agent Performance in a Fixed Period of Time
José Hernández‐Orallo · 2010
The evaluation of several agents over a given task in a finite period of time is a very common problem in experimental design, statistics, computer science, economics and, in general, any experimental science.It is also crucial for intelligence evaluation.In reinforcement learning, the task is formalised as an interactive environment with observations, actions and rewards.Typically, the decision that has to be made by the agent is a choice among a set of actions, cycle after cycle.However, in real evaluation scenarios, the time can be intentionally modulated by the agent.Consequently, agents not only choose an action but they also choose the time when they want to perform an action.This is natural in biological systems but it is also an issue in control.In this paper we revisit the classical reward aggregating functions which are commonly used in reinforcement learning and related areas, we analyse their problems, and we propose a modification of the average reward to get a consistent measurement for continuous time.