High Dimensional Feature Based Word Pair Similarity Measuring For Web Database Using Skip-Pattern Clustering Algorithm

Erode Arts · 2015

Measuring the semantic similarity between words is an important component in various tasks on the web such as relation extraction, community mining, document clustering, and automatic metadata extraction. Accurately measuring the semantic similarity between words is an important problem in web mining, information retrieval, and natural language processing. In information retrieval, one of the main problems is to retrieve a set of documents that is semantically related to a given user query. Text processing plays an important role in information retrieval, data mining, and web search. In text processing, the bag-of-words model is commonly used. In this paper a new scheme proposes an empirical method to estimate semantic similarity using page counts and text snippets retrieved from a web database for two words. Specifically, it defines various word co-occurrence measures using page counts and integrates those with lexical patterns extracted from text snippets.

Read the paper · More papers on PaperTik