Text similarity computing based on Data Compression
Lian Xiong-jie · Journal of Yanbian University · 2004
In the process of information retrieval, the traditional method is to compute similarity between texts. The coefficient similitude figures the degree of compatibility. There are two main methods: Correlation coefficient and Cosine. We base on the theory of Data Compression and use the compression ratio to (express) the similarity between texts. It has some advantages over the others.