Estimating Value of Information Arm Allocation Indices in Contextual Ranking and Selection Problems
Andres Alban, Stephen E. Chick, Spyros I. Zoumpoulis · 2024
Contextual ranking & selection is attracting increasing attention in simulation and other fields. A successful approach to addressing related challenges uses arm allocation indices that compute the Bayesian expected value of information of one-step look-ahead policies. We recall recent work on such indices for linear contextual bandits that take advantage of structural information about the nature of the covariates that describe contexts. Such indices can be computed exactly with a finite number of contexts and no delay in observing outcomes, but may require Monte Carlo simulation otherwise. Our contribution is to describe and quantify the benefits of two variance reduction techniques (conditional Monte Carlo and common random numbers) to estimate such allocation indices for contextual ranking & selection problems when some covariates are continuous or outcomes are observed with delay. We find that both techniques significantly improve estimates and the speed of inference, but conditioning is particularly useful.