Text mining and sentiment extraction in central bank documents
Giuseppe Bruno · 2016
The deep transformation induced by the World Wide Web (WWW) revolution has thoroughly impacted a relevant part of the social interactions in our present global society. The huge amount of unstructured information available on blogs, forum and public institution web sites puts forward different challenges and opportunities. Starting from these considerations, in this paper we pursue a two-fold goal. Firstly we review some of the main methodologies employed in text mining and for the extraction of sentiment and emotions from textual sources. Secondly we provide an empirical application by considering the latest 20 issues of the Bank of Italy Governor's concluding remarks from 1996 to 2015. By taking advantage of the open source software package R, we show the following: 1) checking the word frequency distribution features of the documents; 2) extracting the evolution of the sentiment and the polarity orientation in the texts; 3) evaluating the evolution of an index for the readability and the formality level of the texts; 4) attempting to measure the popularity gained from the documents in the web. The results of the empirical analysis show the feasibility in extracting the main topics from the considered corpus. Moreover it is shown how to check for positive and negative terms in order to gauge the polarity of statements and whole documents. The evaluation of these synthetic indexes is quite relevant for increasing the transparency of the central banks' communications.