Online connectionist Q-learning produces unreliable performance with a synonym finding task
Ileana C. Johnson, Mark D. Plumbley · 2000
Neural networks (NNs) trained with reinforcement learning (RL) have the ability to produce complex, and robust behaviour which may be beneficial to language processing tasks. A method is proposed using RL to train NNs so that they might find synonyms that exist within a regular language. The learning algorithm and exploration strategy produces agents which yield consistently sub-optimal policies for expressions containing one operator, and unreliable performance over all expressions. This is surprising since previous work with lookup tables produced synonyms using a larger set of expressions for a wide range of learning rates and very little exploration.