Characterization of Motion Forms of Mobile Robots Generated in Q-Learning Process
Masayuki Hara, Jian Huang, Testuro Yabut · InTech eBooks · 2011
to acquire the robotic motions, e.g., advancement motions of a caterpillar-shaped robot and a starfish-shaped robot (Yamashina et al., 2006; Motoyama et al., 2006), gymnast-like giant-swing motion of a humanoid robot (Hara et al., 2009), etc.However, most of the conventional studies have discussed the mathematical aspect such as the learning speed, the convergence of learning, etc. Very few studies have focused on the robotic evolution in the learning process or physical factor underlying the learned motions.The authors believe that to examine these factors is also challenging to reveal how the robots evolve their motions in the learning process.This article discusses how the mobile robots can acquire optimal primitive motions through Q-learning (Hara et al., 2006;Jung et al., 2006).First, Q-learning is performed to acquire an advancement motion by using a caterpillar-shaped robot.Based on the learning results, motion forms consisting of a few actions, which appeared or disappeared in the learning process, are discussed in order to find the key factor (effective action) for performing the advancement motion.In addition to this, the environmental effect on the learning results is examined so as to reveal how the robot acquires the optimal motion form when the environment is changed.As the second step, the acquisition of a two-dimensional motion by Q-learning is attempted with a starfish-shaped robot.In the planar motion, not only translational motions in X and Y directions but also yawing motion should be included in the reward; in this case, the yawing angle have to be measured by some external sensor.However, this article proposes Q-learning with a simple reward manipulation, in which the yawing angle is included as a factor of translational motions.Through this challenge, the authors demonstrate the advantage of the proposed method and explore the possibility of simple reward manipulation to produce attractive planer motions.