DAQL-Enabled Autonomous Vehicle Navigation in Dynamically Changing Environment
Chi Kit, N.H.C. Yung · InTech eBooks · 2011
Advances in Reinforcement Learning 386this idea, we propose in this chapter a new approach, which incorporates two major features that are not found in solutions for static environments: (1) actions performed by obstacles are taken into account when the agent determines its own action; and (2) reinforcement learning is adopted by the agent to handle destination seeking (DS) and obstacle actions.Reinforcement Learning (RL) (Sutton & Barto, 1998) aims to find an appropriate mapping from situations to actions in which a certain reward is maximized.It can be defined as a class of problem solving approaches in which the learner (agent) learns through a series of trial-anderror searches and delayed rewards (Sutton & Barto, 1998;Kaelbling, 1993;Kaelbling et al., 1996;Sutton, 1992).The purpose is to maximize not just the immediate reward, but also the cumulative reward in the long run, such that the agent can learn to approximate an optimal behavioral strategy by continuously interacting with the environment.This allows the agent to work in a previously unknown environment by learning about it gradually.In fact, RL has been applied in various CA related problems (Er & Deng, 2005;Huang et al., 2005;Yoon & Sim, 2005) in static environments.For RL to work in a DE containing multiple agents, the consideration of actions of other agents/obstacles in the environment becomes necessary (Littman, 2001).For example, Team Q-learning (QL) (Littman, 2001;Boutilier, 1996) considered the actions of all the agents in a team and focused on the fully cooperative game in which all agents try to maximize a single reward function together.For agents that do not share the same reward function, Claus and Boutilier (Claus & Boutilier, 1998) proposed the used of JAL.Their results showed that by taking into account the actions of another agent, JAL performs somewhat better than the traditional QL.However, JAL depends crucially on the strategy adopted by the other agents and it assumes that other agents maintain the same strategy throughout the game.While this assumption may not be valid, Hu and Wellman proposed Nash Q-learning (Hu & Wellman, 2004) which focuses on a general sum game that the agents are not necessarily working cooperatively.Nash equilibrium is used for the agent to adopt a strategy which is the best response to the other's strategy.This approach requires the agent to learn others Q-value by assuming that the agent can observe other's rewards.In this chapter, we propose an improved QL method called Double Action Q-Learning (DAQL) (Ngai & Yung, 2005a;Ngai & Yung, 2005b) that similarly considers the agent's own action and other agents' actions simultaneously.Instead of assuming that the rewards of other agents can be observed, we use a probabilistic approach to predict their actions, so that they may work cooperatively, competitively or independently.Based on this, we further develop it into a solution for the two goal navigation problem in a dynamically changing environment, and generalize it for solving multiple goal problems.The solution uses DAQL when it is required to consider the responses of other moving agents/obstacles.If agent action would not cause the destination to move, then QL (Watkins & Dayan, 1992) would suffice for DS.Given two actions from two goals, a proportional goal fusion function is employed to maintain a balance in the final action decision.Extensive simulations of the proposed method in environments with single constant speed obstacle to multiple obstacles at variable speed and directions indicate that the proposed method is able to (1) deal with single obstacle at any speed and directions; (2) deal with two obstacles approaching from different directions; (3) cope with large sensor noise; (4) navigate in high obstacle density and high relative velocity environments.Detailed comparison with the Artificial Potential Field method (Ratering & Gini, 1995) reveals that the proposed method improves path time and the number of collisionfree episodes by 20.6% and 23.6% on average, and 27.8% and 115.6% at best, respectively.The rest of this chapter is organized as follows: Section 2 introduces the concept of the proposed DAQL-enabled reinforcement learning framework.Section 3 describes the implementation method of the proposed framework in solving the autonomous vehicle www.intechopen.com