Domain Specific Word Extraction from Hierarchical Web Documents: A First Step Toward Building Lexicon Trees from Web Corpora
Jing-Shin Chang · 2005
Domain specific words and ontological information among words are important resources for general natural language applications. This paper proposes a statistical model for finding domain specific words (DSW s) in particular domains, and thus building the association among them. When applying this model to the hierarchical structure of the web directories node-by-node, the document tree can potentially be converted into a large semantically annotated lexicon tree. Some preliminary results show that the current approach is better than a conventional TF-IDF approach for measuring domain specificity. An average precision of 65.4% and an average recall of 36.3% are observed if the top-10% candidates are extracted as domain-specific words. 1 Domain Specific Words and Lexicon Trees as Important NLP Resources Domain specific words (DSW s) are important anchoring words for natural language processing applications that involve word sense disambiguation (WSD). It is appreciated that multi-sense words appearing in the same document tend to be tagged with the same word sense if they belong to the same common domain in the semantic hierarchy (Yarowsky, 1995). The existence of some DSW s in a document will therefore be a strong evidence of a specific sense for words within the document. For instance, the existence of basketball in a document would strongly suggest the sport sense of the word ( Pistons ), rather than its mechanics sense. It is also a personal belief that DSW-based sense disambiguation, document classification and many similar applications would be easier than sense-based models since sense-tagged documents are rare while domain-aware training documents are abundant on the Web. DSW identification is therefore an important issue. On the other hand, the semantics hierarchy among words (especially among sets of domain specific words) as well as the membership of domain specific words are also important resources for general natural language processing applications, since the hierarchy will provide semantic links and ontological information (such as is-A and part-of relationships) for words, and, domain specific words belonging to the same domain may have the synonym or antonym relationships. A hierarchical lexicon tree (or a network, in general) (Fellbaum, 1998; Jurafsky and Martin, 2000), indicative of sets of highly associated domain specific words and their hierarchy, is therefore invaluable for NLP applications. Manually constructing such a lexicon hierarchy and acquiring the associated words for each node in the hierarchy, however, is most likely unaffordable both in terms of time and cost. In addition, new words (or new usages of words) are dynamically produced day by day. For instance, the Chinese word (pistons) is more frequently used as the sport or basketball sense (referring to the Detroit