A learning algorithm for the finite-time two-armed bandit problem
Motoyuki Sato, Ken Abe, Hiroshi Takeda · IEEE Transactions on Systems Man and Cybernetics · 1984
A simple algorithm for the finite-time two-armed bandit problem is proposed. In this algorithm, the whole process is divided into the first estimating process and the next controlling process. Efficient methods involving the use of approximation for computing the optimal length of the estimating process are provided.