OPTIMAL STOPPING PROBLEM WITH UNCERTAIN RECALL
Seizo Ikuta · Journal of the Operations Research Society of Japan · 1988
Consider a discrete-time optimal stopping problem with a finite planning horizon in which an offer passed up j ≧ 0 periods ago becomes unavailable at the next time with a known probability p_j, provided that it remains available at present. The objective is to maximize the expected discounted gain where the term gain means the value of the offer accepted less the total search cost paid up to the termination of the process with its acceptance. The main results obtained are the next four. (1) Let a and b be, respectively, the lower bound and the upper bound in the distribution of offer w. Then the optimal stopping rule has the following property. For at least one set consisting of past offers available at present, there exists the following two critical numbers ξ and ξ' such as a a where β, μ, and c are, respectively, a discount factor, an expectation of offer w, and a cost per search. The property is called a double reservation value property or DRV-property for short. (2) The property gradually disappears as a planning horizon tends to infinity, with totally vanishing in its limit. (3) In the limit of a planning horizon, it suffices to memorize only the present offer with neglecting all past offers; in other words, the problem is eventually reduced to an infinite horizon optimal stopping problem with no recall. (4) Under the optimal stopping rule, each of the maximum expected discounted gain attained, the expected number of searches made, and the expectation of the offer accepted is less than or equal to one in an optimal stopping problem with recall, and both become the same in the limit of a planning horizon.