Unsupervised Domain Relevance Estimation for Word Sense Disambiguation.

Alfio Gliozzo, Bernardo Magnini, Carlo Strapparava · 2004

This paper presents Domain Relevance Estima-tion (DRE), a fully unsupervised text categorization technique based on the statistical estimation of the relevance of a text with respect to a certain cate-gory. We use a pre-defined set of categories (we call them domains) which have been previously as-sociated to WORDNET word senses. Given a cer-tain domain, DRE distinguishes between relevant and non-relevant texts by means of a Gaussian Mix-ture model that describes the frequency distribution of domain words inside a large-scale corpus. Then, an Expectation Maximization algorithm computes the parameters that maximize the likelihood of the model on the empirical data. The correct identification of the domain of the text is a crucial point for Domain Driven Dis-ambiguation, an unsupervised Word Sense Disam-biguation (WSD) methodology that makes use of only domain information. Therefore, DRE has been exploited and evaluated in the context of a WSD task. Results are comparable to those of state-of-the-art unsupervised WSD systems and show that DRE provides an important contribution. 1

Read the paper · More papers on PaperTik