Optimising the Hystereses of a Two Context Layer RNN for Text Classification
Garen Arevian, Christo Panchev · IEEE International Conference on Neural Networks/IEEE ... International Conference on Neural Networks · 2007
Established techniques from information retrieval (IR) and machine learning (ML) have shown varying degrees of success in the automatic classification of real-world text. The capabilities of an extended version of the Simple recurrent network (SRN) for classifying news titles from the Reuters-21578 Corpus are explored. The architecture is composed of two hidden layers where each layer has an associated context layer that takes copies of previous activation states and integrates them with current activations. This results in improved performance, stability and generalisation by the adjustment of the percentage of previous activation strengths kept "in memory" by what is defined as the hysteresis parameter. The study demonstrates that this partial feedback of activations must be carefully fine-tuned to maintain optimal performance. Correctly adjusting the hysteresis values for very long and noisy text sequences is critical as classification performance degrades catastrophic ally when values are not optimally set.