Text similarity computing based on Data Compression

Lian Xiong-jie · Journal of Yanbian University · 2004

In the process of information retrieval, the traditional method is to compute similarity between texts. The coefficient similitude figures the degree of compatibility. There are two main methods: Correlation coefficient and Cosine. We base on the theory of Data Compression and use the compression ratio to (express) the similarity between texts. It has some advantages over the others.

Read the paper · More papers on PaperTik