Topic Detection Using Independent Component Analysis

Scott Grant, David B. Skillicorn, James R. Cordy · 2007

If the documents of a large text corpus can be modeled as the rows of a matrix, it can be shown that existing mathematical methods can be used to extract previously unseen information about their relationships. In particular, it can be shown that Independent Component Analysis offers a way of identifying threads of related conversations in a large data set such as VAST. By treating each document as a vector, with word frequencies representing the components, we can extract two interesting pieces of information from the set: a list of the topics used in each document, and a list of the documents that best fit each of these topics. 1

Read the paper · More papers on PaperTik