A comparative research of different granularities in Korean text classification
Chun Liu, Yahui Zhao, Xu Cui, Yitong Zhao · 2022 IEEE International Conference on Advances in Electrical Engineering and Computer Applications (AEECA) · 2022
Text classification is a process, which can make the specified documents group into several categories, predefined at the beginning through learning a series of rules or under the guidance of the goal function. This paper compared the subword-level, spacing-level of Korean and the word-level, then analyzed the influence of the preprocessing of different granularities on the text classification task of Korean. After that, analyzed the results of classification linguistically. Thus we can choose the proper granularity as the input to improve the classification effect. Firstly, cut the corpus according to different granularities; then, used Glove word embedding. Finally, used self-attention classification mechanism to verify the effect of the corpus which are preprocessed through different granularities. And using the accuracy and loss as the index. After experimental and comparison, the best result is selected when the accuracy is 83.46% through the spacing-level. Experiments show that the research in this paper can improve the effect of text classification.