Train1 vs. Train2: Tagging Word Senses in Corpus

Uri Zernik · Psychology Press eBooks · 2021

Homonymy and polysemy have hampered information-retrieval accuracy. The use of a word with multiple meanings as a retrieval key might yield diluted results. So far efforts to alleviate this problem were hampered, first, by the practical problem of tagging words in coprus, and second, by the theoretical problem of substantially enhancing the information contents of a query using word-sense tags. To alleviate thess problems we have implemented a scheme for tagging word senses in a corpus. Accordingly, users’s queries are made specific to particular senses of words. In this chapter we show how word-sense tagging is carried out in the I.M.Toolset system. First, the text is pre-processed as to morphology and local syntax; second, signatures are generated for each manifestation of a word, based on the local context of the word; third, signatures are clustered and word senses are extracted. We discuss the advantages of this scheme, i.e., improved accuracy, and its limitations, i.e., word senses are specific to a narrow corpus over which training took place. 1

Read the paper · More papers on PaperTik