Pronunciation Ambiguities in Japanese Kanji

Wen Zhang · 2023

Japanese writing is a complex system, and a large part of the complexity resides in the use of kanji.A single kanji character in modern Japanese may have multiple pronunciations, either as native vocabulary or as words borrowed from Chinese.This causes a problem for text-tospeech synthesis (TTS) because the system has to predict which pronunciation of each kanji character is appropriate in the context.The problem is called homograph disambiguation.To solve the problem, this research provides a new annotated Japanese single kanji character pronunciation data set and describes an experiment using the logistic regression (LR) classifier.A baseline is computed to compare with the LR classifier accuracy.This experiment provides the first experimental research in Japanese single kanji homograph disambiguation.The annotated Japanese data is freely released to the public to support further work. IntroductionJapanese uses a mixed writing system with three distinct scripts and one romanization.Kanji 漢字 is the writing script that borrows directly from Chinese characters which were introduced in Japan from China through Korea from the third century CE.There are 2,136 commonly used kanji characters termed Joyo kanji in present-day Japanese. 1 A single kanji character in modern Japanese may have multiple pronunciations derived from the linguistic history of the kanji characters as either native vocabulary words or as terms borrowed from Chinese.For instance, the 1 https://kanji.jitenon.jp/cat/joyo.htmlkanji character 山 'mountain' can be read as either the native Japanese word yama or the Chinesederived term san.The native Japanese pronunciations of the kanji character 文 'letter, sentence, writings' are humi, aya, and kaza, while Chinese borrowed pronunciations are bun and mon.Because a kanji character has multiple pronunciations, to predict the appropriate pronunciation for each kanji character, a text-tospeech synthesis engine must select the appropriate reading.This is a form of homograph disambiguation.This research is a computational study of Japanese kanji homograph disambiguation.Recent research in homograph disambiguation in Japanese is limited because of the lack of extensive data sets that include comprehensive pronunciations for the most commonly used kanji characters.The goal of this research is to fill this void, make new data sets to conduct the analysis of kanji characters with multiple pronunciations, and use the computational methodology to test the data set to lay a foundation for computational research on Japanese kanji homographs in the future. Japanese writing scriptsThe Japanese writing system uses three different scripts, Chinese characters (kanji), and two kana systems: hiragana and katakana, which are derivatives of Chinese characters.Hiragana resulted from the cursive style of writing Chinese characters, while katakana developed from the abbreviation of Chinese characters.Roughly speaking, kanji are used for content words such as nouns, stems of adjectives,and verbs, whereas hiragana is used for writing grammatical words

Read the paper · More papers on PaperTik