Automatic Acquisition of Synonyms Using the Web as a Corpus

Svetlin Nakov · 2008

We present an original algorithm for automatic acquisition of synonyms from text. The algorithm measures the semantic similarity between pairs of words by comparing their local contexts extracted from the Web by series of queries against the Google search engine. The results show 11pt average precision of 63.16%. In the present paper, we set the objective to design an algorithm for automatic extraction of pairs of synonyms from a text corpus. The results can be used to create linguistic resources, such as general and domain-specific thesauri and lexicons. We use the Web as a large corpus which can be efficiently searched. Our approach is based on performing series of queries against a Web search engine and analyzing the returned excerpts of texts (snippets) in order to extract contextual semantic information which we use to measure the semantic similarity between pairs of words and thus to approximate synonymy. It is considered that the local context of a given word (few words before and after the target captured word) contains words that are semantically related to it (Hearst, 1991). Given a pair of words, we extract their local contexts from the snippets returned by the search engine and we measure the semantic similarity between these words by calculating the similarity between their local contexts. Finally, the measured similarity is used to determine whether the words are likely to be synonyms or not.

Read the paper · More papers on PaperTik