CASIT: Content Based Identification of Textual Information in a Large Database

Larbi Guezouli, Hassane Essafi · 2010

This paper describes CASIT model (CAlculation of SImilarity of Text). Starting from a coarse confrontation of text documents, based on the Latent Semantic Indexing model (LSI), CASIT method calculates in a finer way, the rate of similarity between model documents of text and others which are confronted to them. Our approach takes into account the neighbourhood of the words, which makes it possible to balance the words in the calculation of the score.

Read the paper · More papers on PaperTik