Advanced Subword Segmentation and Interdependent Regularization Mechanisms for Korean Language Understanding
Mansu Kim, Yunkon Kim, Yeonsoo Lim, Eui‐Nam Huh · 2019
State-of-the-art neural language models highly rely on fixed-size subword vocabulary based pre-training process to improve the performance of language understanding. Existing subword segmentation algorithms have been researched to generate fixed-size vocabulary by frequency in large text pool without considering the linguistic structures. Also, most of research is focused on widely-used languages such as English and Chinese. However, for Korean, a segmentation algorithm considering the linguistic structure is required, and we thus propose the algorithm considering the characteristics of Korean. For example, Korean words include the special type of suffix called “Josa” that adds grammatical meaning. We also propose subword regularization algorithm based on mutual information, which is interdependence between words. The regularization algorithm customizes the size of vocabulary on demand. In addition, we present an experiment analysis of subword segmentation and interdependent regularization by testing neural language model. It can achieve better performance by small changes of vocabulary.