Building The Sense-Tagged Multilingual Parallel Corpus
Shan Wang, Francis T. Bond · 2014
Sense-annotated parallel corpora play a crucial role in natural language processing.This paper introduces our progress in creating such a corpus for Asian languages using English as a pivot, which is the first such corpus for these languages (Chinese, Japanese and Indonesian).Two sets of tools have been developed for sequential and targeted tagging, which are also easy to be set up for any new languages.This paper also briefly presents the general guidelines for doing this project.The current results of the monolingual sensetagging and multilingual linking are illustrated, which indicate the differences among genres and language pairs.All the tools, guidelines and the manually annotated corpus will be freely available at http://compling.ntu.edu.sg/ntumc.