Behaviors of Reservoir Computing Models for Textual Documents Classification
Nils Schaetti · 2019
Reservoir Computing is a paradigm of recurrent neural network (RNN) models, attractive because of its ease of training and new neuromorphic optoelectronic implementations. Applied with success to time series prediction and speech recognition, few works have so far studied the behavior of these networks on natural language processing (NLP) tasks. Therefore, we decided to explore the ability of Echo State Network-based Reservoir Computing (ESN) models with additional embedding layers to classify text documents of the Reuters C50 data set based on authorship. We explored various learned representations such as word and character embedding and deep feature extractors. Our experiments demonstrate that ESN models can achieve state-of-the-art results on this task and are competitive with common models such as Support Vector Machines (SVM). Moreover, we show that these models compute documents as data streams and could then be able to handle other tasks such as event detection and text segmentation. The best performance is obtained by an ESN with a large reservoir of 1,500 neurons based on word vectors. We think that these results demonstrate the possibility of processing massive quantities of textual data in the future using Reservoir Computing-based systems.