Discovering word associations in news media via feature selection and sparse classification

Brian Gawalt, Jinzhu Jia, Luke W Miratrix, Laurent El Ghaoui, Bin Yu, Sophie M Clavier · 2010

We analyze the "image" of a given query word in a given corpus of text news by producing a short list of other words with which this query is strongly associated. We use a number of feature selection schemes for text classification to help in this task. We apply these classification techniques using indicators of the query word's appearance in each document used as the document "labels" and the indicators for all other words as document predictors/features. The features selected by any scheme is then considered the list of words comprising the query word's "image".

Read the paper · More papers on PaperTik