A Japanese preprocessor for syntactic and semantic parsing

Tsuyoshi Kitani, Teruko Mitamura · 2002

The authors describe a Japanese preprocessor which includes a morphological analyzer called MAJESTY, and a proper noun identification program. The original morphological analyzer was modified to disambiguate its output when multiple possibilities for segmentations and parts of speech are found. Ambiguous segments are packed locally in the output enabling a syntactic and semantic parser to perform efficiently. Then the proper noun identification program groups several segments constructing a proper noun to present a meaningful set of segments to the parser. Tested on financial news articles, the preprocessor successfully segmented text and tagged parts of speech with greater than 98% accuracy. Over 80% of company names and 90% of personal and place names have been identified.>

Read the paper · More papers on PaperTik