Optimized Clustering based on Semantic Similarity of Components for Short Text

Wensong Liu, Feng Lin, Zhuqing Hu, Jinhui Zhang · 2019 IEEE International Conference on Signal, Information and Data Processing (ICSIDP) · 2019

Short text is usually made up of several words and one sentence at most. Considering sparse features and complicated expressions, the similarity measurement and the clustering of the texts may not work well. The semantic clustering for short text is studied. Firstly, according to the analysis of the dependency syntax structure, the event component and modified components are extracted. Then, the semantic similarity of the texts is analyzed with the component as the basic unit. The strategy is that the text would be similar when not only the event component but also the modified components are similar. Further, considering the proposed semantic similarity may lead to increased topics, which means various cluster shape, indefinite cluster number, and increased noise point for the cluster, the density peak clustering is selected, and a regression parameter is designed to improve the cluster number and the noise. Based on the public data set, the proposed semantic clustering is tested: purity$P$is 96% and$F$measure is 71.97%. The proposed method has been used in the electrical power industry and is worth promoting.

Read the paper · More papers on PaperTik