Learning to plan probabilistically from neural networks
Ruiji Sun, Chad Sessions · 2002
This paper discusses the learning of probabilistic planning without a priori domain-specific knowledge. Different from existing reinforcement learning algorithms that generate only reactive policies and existing probabilistic planning algorithms that requires a substantial amount of a priori knowledge in order to plan, we devise a two-stage bottom-up learning-to-plan process, in which the reinforcement learning/dynamic programming is first applied, without the use of a priori domain-specific knowledge, to acquire a reactive policy and then explicit plans are extracted from the learned reactive policy. Plan extraction is based on a beam search algorithm that performs temporal projection in a restricted fashion guided by the value functions resulting from the reinforcement learning/dynamic programming. The experiments and theoretical analysis are presented.