Document representation and multilevel measures of document similarity
Irina Vladimirovna Matveeva · 2006
We present our work on combining large-scale statistical approaches with local linguistic analysis and graph-based machine learning techniques to compute a combined measure of semantic similarity between terms and documents for application in information extraction, question answering, and summarisation.