Auto-Clustering of Conversation Corpus Based on Syntactic, Semantic and Pragmatic Features

Baojian Chen, Minghu Jiang · 2013

To understand natural language accurately, we not only need to do natural language morphology and syntactic analysis, but also need to combine semantic knowledge and pragmatic information with a specific context. Due to short knowledge and lack in background information of conversation corpus which related to the pragmatic, there is a long way to go for computer fully understand natural language. In this paper, the pragmatic features were added to the text vector space model of language spoken conversation, and hierarchical clustering is executed. Our experimental results show that the clustering effect with pragmatic features outperforms than non-pragmatic features, and precision, recall rate and F values of the former were increased by 6.67%, 6.34% and 6.6%, respectively. It indicates that pragmatic information has played an important role in enhancing the effect of the text clustering.

Read the paper · More papers on PaperTik