Performance Comparison of Passage Retrieval Models according to Korean Language Tokenization Methods

WonJune Seo, V. Sakthivel, Dugki Min, Jae Wook Lee, Enumi Choi · 2023

Recently, NLP (Natural Language Process) is required in question-answering technology in many fields such as chatbots and AI speakers. Accordingly, companies and communities are conducting research on Open Domain Question Answering (ODQA), which is required to identify which question and answer questions should be used in a vast amount of data sets. In fact, in the case of question-and-answer learning, there are cases in which multilingual methods are included in most cases, but performance can be further improved by using a Korean-only model. Therefore, this paper focuses on information retrieval, which is a method that includes answers to questions in the ODQA Framework and tries to find a method unique to Korean. In this paper, we propose to apply the token for each query to the model using the Korean method using the BM25 (Best Matching) and DPR (Dense Passage Retrieval) models. To verify the above evaluations, KorQuAD v1 dataset is used. Accordingly, Korean BM25 can obtain the highest score for morpheme-based tokenization, and additionally, adding and maintaining other tokens according to the model performed well, and as a result of learning with the same model, DPR has more sub-tokens than the conventional method.

Read the paper · More papers on PaperTik