Using concept structures for efficient document comparison and location
Andrew N. Edmonds · 2007
A method is discussed for comparing and locating similar documents in a computationally efficient manner by making use of inferred concept statistics, rather than word frequencies. This novel technique uses natural language structures to create a short 'concept signature' vector, which locates a document in 'concept space'. Similar documents can be located in large corpora in O(log(n)) time by making use of this space for indexing. Results from trials with reference and real world data sets are presented, along with a comparison of the method's document similarity characteristics and the cosine metric