Japanese SemCor: A Sense-tagged Corpus of Japanese

Francis T. Bond, Timothy J. Baldwin, Richard Fothergill, Kiyotaka Uchimoto · 2012

In this paper we describe the creation of the Japanese SemCor (JSEMCOR) sensetagged corpus of Japanese. The corpus is a translation of the English SEMCOR, with senses projected across from English. The final corpus consists of 14,169 sentences with 150,555 content words of which 58,265 are sense tagged. The corpus is one of the corpora used to provide sense frequency data for the Japanese Wordnet. 1

Read the paper · More papers on PaperTik