TEXT SIMILARITY COMPUTING BASED ON ATTRIBUTE THEORY

Qian Pan · Chinese Journal of Computers · 1999

Generally, in the process of information retrieval(IR), the users first put forward their key words to the system that they want to search. Then the key words are analyzed to the special format, they are matched with the document database whose results are considered as the results that are related to the users' interests. There are several IR models, such as reverse document model, vector space model, generalized vector space model and latent semantic model and so on. According to attribute theory, this paper analyses the relationship between textual attributes and the attribute barycenter coordinate model, and establishes the text attribute barycenter coordinate model. Within the coordinate, a text vector and a query vector can be represented. After deciding the criterion and computing the distance between the vectors, a formula that computes the similarity between the texts and the queries is shown.

Read the paper · More papers on PaperTik