Measuring Popularity of Machine-Generated Sentences Using Term Count, Document Frequency, and Dependency Language Model

Jong Myoung Kim, Han Cheol Park, Young Seob Jeong, Ho Jin Choi, Gah Gene Gweon, Jeong Hur · Institutional Repositories DataBase (IRDB) · 2015

We investigated the notion of “popularity” for machine-generated sentences. We defined a popular sentence as one that contains words that are frequently used, appear in many docu-ments, and contain frequent dependencies. We measured the popularity of sentences based on three components: content morpheme count, document frequency, and dependency relation-ships. To consider the characteristics of agglu-tinative language, we used content morpheme frequency instead of term frequency. The key component in our method is that we use the product of content morpheme count and doc-ument frequency to measure word popular-ity, and apply language models based on de-pendency relationships to consider popularity from the context of words. We verify that our method accurately reflects popularity by us-ing Pearson correlations. Human evaluation shows that our method has a high correlation with human judgments. 1

Read the paper · More papers on PaperTik