Using concept structures for efficient document comparison and location

Andrew N. Edmonds · 2007

A method is discussed for comparing and locating similar documents in a computationally efficient manner by making use of inferred concept statistics, rather than word frequencies. This novel technique uses natural language structures to create a short 'concept signature' vector, which locates a document in 'concept space'. Similar documents can be located in large corpora in O(log(n)) time by making use of this space for indexing. Results from trials with reference and real world data sets are presented, along with a comparison of the method's document similarity characteristics and the cosine metric

Read the paper · More papers on PaperTik