Q-learning with Condition Reduced Fuzzy Rules and Its Applications
Hideki YAMAGISHI, Hiroshi Kawakami, Tadashi Horiuchi, Osamu Katai · Transactions of the Society of Instrument and Control Engineers · 2001
A new Q-learning method for the cases where the states (conditions) and actions of systems are assumed to be continuous is proposed. The components of Q-tables are interpolated by fuzzy inference. The initial set of fuzzy rules consists of all the combinations of conditions and actions relevant to the problem. Each rule is then associated with a value by which the Q-value of a condition/action pair is estimated. The values are revised by the Q-learning algorithm so as to make the fuzzy rule system effective. Although this framework may require a huge number of the initial fuzzy rules, we will show that considerable reduction can be done by adopting what we call “Condition Reduced Fuzzy Rules (CRFRs)”. The antecedent parts of CRFRs consist of all the actions and appropriately selected conditions, and their consequents are set to be their Q-values. Finally, experimental results show that controllers with CRFRs perform equally well compared to the system with the most detailed fuzzy control rules, while the total number of parameters that have to be revised through the whole learning process is considerably reduced with increasing the number of revised parameters at each learning step.