Text as Data: Text Mining and Sentiment Analysis

Johannes Ledolter · 2013

Text is a vast source of data for business. Text data is extremely high dimensional. The analysis of phrase counts from text documents is the current state of the art. Information retrieval and the appropriate “tokenization” of the information are very important. This chapter discusses inverse multinomial logistic regression in detail. The analysis conducted by Gentzkow and Shapiro is actually slightly different. They consider an ordinary regression (not a multinomial logistic regression) of the relative frequencies of trigrams on the vote-shares that they then use, in a second step, to determine the estimate of the political “slant” of each representative. From the frequency distribution of the speech's trigrams, one wants to infer the political sentiment of the speech. This is can be achieved with the inverse multinomial logistic regression model. The chapter considers two examples namely restaurant reviews, and political sentiment.

Read the paper · More papers on PaperTik