A learning algorithm for the finite-time two-armed bandit problem

Motoyuki Sato, Ken Abe, Hiroshi Takeda · IEEE Transactions on Systems Man and Cybernetics · 1984

A simple algorithm for the finite-time two-armed bandit problem is proposed. In this algorithm, the whole process is divided into the first estimating process and the next controlling process. Efficient methods involving the use of approximation for computing the optimal length of the estimating process are provided.

Read the paper · More papers on PaperTik