Train-O-Matic: Large-Scale Supervised Word Sense Disambiguation in Multiple Languages without Manual Training Data
Tommaso Pasini, Roberto Navigli · 2017
Annotating large numbers of sentences with senses is the heaviest requirement of current Word Sense Disambiguation.We present Train-O-Matic, a languageindependent method for generating millions of sense-annotated training instances for virtually all meanings of words in a language's vocabulary.The approach is fully automatic: no human intervention is required and the only type of human knowledge used is a WordNet-like resource.Train-O-Matic achieves consistently state-of-the-art performance across gold standard datasets and languages, while at the same time removing the burden of manual annotation.All the training data is available for research purposes at http://trainomatic.org.