Assessing creation methods of word embedding models for analyzing and repairing classical Japanese literature
Shota Kusabiraki, Y. Shimmura, K. Yamamoto, K. Iwata · 2024
This study not only proposes a valuable model for predicting missing words in classical Japanese literature but also suggests the potential of this model to be instrumental in repairing the literature. It could significantly advance the field of Natural Language Processing research in the context of historical literature. In recent years, Natural Language Processing has been applied to artificial intelligence programs such as ChatGPT and literary works. However, Natural Language Processing research in Japan has mainly focused on modern Japanese, and research in Japanese classical literature has yet to progress enough. Our research takes a novel approach by attempting to forecast missing words in classical Japanese literature. It creates several language models based on three pieces of classical literature using the Skip-gram of fastText. We employ LOOCV (leave-one-out cross-validation) to validate each model&s;s accuracy. The results highlight significant differences between the modern language model and our proposed models, which we attribute to the historical context. Next, the experiment demonstrates the efficiency of our model creation method in predicting a missing word. The results of the experiments show that our proposed method can predict words similar to a missing word.