Chinese Short Text Matching Model Based on WoBERT Word Embedding Representation and Priori Knowledge

Jiaming Dong · 2023

Chinese short text matching is an important task in natural language processing. It is widely used in information retrieval, question and answer systems and dialogue systems. A model (WHCR) based on WoBERT word embedding representation and priori knowledge is proposed to address the problems that traditional Chinese short text matching models cannot solve multiple meanings of words, poor extraction of similar text features and insufficient extraction of granularity information. The pre-training model WoBERT is introduced instead of traditional Word2Vec, which combines contextual information in word vector representation. It introduces OpenHowNet as priori knowledge to guide word-sense interactions, and one-dimensional convolution to extract multi-granularity information, so that the model can obtain multi-level semantic information. Finally, for the problem of inconsistent training and inference due to random deactivation of neurons by dropout, R-drop is introduced to improve the robustness of the model. The model achieves good results on the Chinese datasets BQ and BZQMC, with Acc, AUC and F1-score reaching 85.57%, 92.68%, 85.44% and 84.04%, 93.19%, 85.71% respectively.

Read the paper · More papers on PaperTik