Extraction Japanese Slang from Weblog Data based on Script Type and Stroke Count

Kazuyuki Matsumoto, Kyosuke Akita, Xielifuguli Keranmu, Minoru Yoshida, Kenji Kita · Procedia Computer Science · 2014

Young people commonly use slang in the texts for weblogs or Social Networking Sites. How to treat such slang words properly is one of the problems in the field of text mining. In this paper, we examined several methods to extract Japanese slang called “Wakamono Kotoba,” which is particularly used by young people, by focusing on its script type and stroke count. In the evaluation experiment, a high precision was obtained when we adopted script type for extraction.

Read the paper · More papers on PaperTik