Chunks of thought: Finding salient semantic structures in texts

Mei Mei, Aashay Vanarase, Ali A. Minai · 2014

As the availability of large, digital text corpora increases, so does the need for automatic methods to analyze them and to extract significant information from them. A number of algorithms have been developed for these applications, with topic modeling-based algorithms such as latent Dirichlet allocation (LDA) enjoying much recent popularity. In this paper, we focus on a specific but important problem in text analysis: Identifying coherent lexical combinations that represent "chunks of thought" within the larger discourse. We term these salient semantic chunks (SSCs), and present two complimentary approaches for their extraction. Both these approaches derive from a cognitive rather than purely statistical perspective on the generation of texts. We apply the two algorithms to a corpus of abstracts from IJCNN 2009, and show that both algorithms find meaningful chunks that elucidate the semantic structure of the corpus in complementary ways.

Read the paper · More papers on PaperTik