Sampling Based Approximate Algorithm for POMDP
Xiaoping Chen · Jisuanji fangzhen · 2006
Partially observable Markov decision procedure is a kind of problem model which describes the continuous decision making for robot within dynamic uncertain environment.This paper introduces a fast approximate algorithm for special POMDP models which have sparse state transmit matrix.First,this algorithm makes use of the policy from QMDP approximate algorithm for sampling.Then it can use these samples with point based iteration algorithm to create the value function for POMDP.Finally,the optimal policy for action choosing will be generated from the value function.In the same experiment model,the policy generated by this algorithm will make the reward as much as other algorithms.But this algoritm can run faster than others,and can generate a smaller vector set to represent the policy.So,it is more suitable for solving large POMDPs with sparse state transmit matrix than other approximate algorithms.