Similarity Algorithm Based on Common Chunks Between English Short Texts
Huang Xian-yin · 2015
It is short text similarity computation that has been the focus of the natural language processing. Only the words are considered in the traditional text similarity algorithm based on the terms,with words order ignored. A new method based on common chunks was presented to calculate the short text similarity,which considers the number and the sequence of the same words. The similarity of the short texts was gotten through making automatic coefficient between the similarity based on the same words and the similarity based on the order of the same words. The simulation results show that,compared with conventional similarity algorithms,the presented algorithm has a better performance in the stability and the harmonic-mean towards the precision and the recall.