Challenges in Multilingual Domain-Specific Sense-marking
Jaya Saraswati, Rajita Shukla, Sonal Pathade, Tina Solanki, Pushpak Bhattacharyya · 2009
Annotation plays a key role in today’s NLP scenario and this paper discusses challenges involved in one of the toughest annotation tasks- sense marking. In an effort to train the machine to understand the written language and thus to ensure speedy and high-quality translation, a huge amount of data needs to be sense-marked accurately by humans using an authentic and standard lexicon. In the work reported here, the corpus is taken from tourism domain and the Princeton wordnet (Version 2.1) is used as the sense inventory for English text while the Hindi and Marathi wordnets have been used for Hindi and Marathi texts respectively. A word may have a number of senses and in identifying which particular sense has been used in the given context, word sense disambiguation becomes a critical necessity. The corpus was independently tagged by different sense-markers and it was found that the inter annotator agreement on word sense disambiguation was about 80 % across the three languages, i.e., English, Hindi and Marathi. Though the sense distinctions in the wordnets are quite fine-grained, there have been cases when the senses provided there have been inadequate and the human sense-markers have faced problems. The study records such challenges and their handling.