Improved QMDPPolicy for Partially Observable Markov Decision Processes in Large Domains: Embedding Exploration Dynamics
Giorgos Apostolikas, Spyros G. Tzafestaş · Intelligent Automation & Soft Computing · 2004
Abstract Artificial Intelligence techniques were primarily focused on domains in which at each time the state of the world is known to the system. Such domains can be modeled as a Markov Decision Process (MDP). Action and planning policies for MDPs have been studied extensively and several efficient methods exist. However, in real world problems pieces of information useful for the process of action selection are often missing. The theory of Partially Observable Mazkov Decision Processes (POMDP’s) covers the problem domain in which the full state of the environment is not directly perceivable by the agent. Current algorithms for the exact solution of POMDP’s are only applicable to domains with a small number of states. To cope with more extended state spaces, a number of methods that achieve sub-optimal solutions exist and among these the QI,IDP approach seems to be the best. We introduce a novel technique, called Explorative [Qtilde]P (EQI-IDP) which constitutes an important enhancement of the [Qtilde]P ...