Construction of a Word Sense Tagged Corpus for SENSEVAL-2 Japanese Dictionary Task

Kiyoaki Shirai · 2002

This paper reports the details of a Japanese word sense tagged corpus developed as an evaluation data for SENSEVAL-2 Japanese dictionary task.The corpus made up of 2,130 newspaper articles.Not all but only 10,000 words in the articles were manually annotated with sense IDs, which was used as a gold standard data.Word senses were deÞned according to the Iwanami Kokugo Jiten, a Japanese dictionary published by Iwanami Shoten.Two annotators chose a sense ID for each instance separately.If they did not agree, the third annotator chose the correct sense ID between them.Inter-tagger agreement and Cohen's κ was 86.3% and 0.677, respectively.

Read the paper · More papers on PaperTik