V-SBERT: A Mixture Model for Closed-Domain Question-Answering Systems Based on Natural Language Processing and Deep Learning
Yuan Zhuang, Boyan Liu, Xiaotao Lin, Canhao Xu · 2023
Smart question answering (QA) systems have been the focus of research for many years. This is known to be a challenging task due to the randomness of user input. As deep learning technology is evolving fast, closed-domain QA systems based on frequently asked questions (FAQs) have already become a mature field. In the project, the team designed a mixture model based on deep learning and natural language processing (NLP) for closed-domain QA tasks. We construct a novel dataset containing admission FAQs for a Chinese university1. Then we propose an ensemble system using the average voting method based on TF-IDF and fine-tuned versions of SBERT variation models (V-SBERT). Because both the query and questions in the knowledge base are formulated in natural language, it is important to use NLP approaches to represent the messages. Then, the matches between the user query and standard questions are scored. The accuracy and performance of several different encoding models are tested on the collected dataset, including Term Frequency - Inverse Document Frequency (TF-IDF), Bidirectional Encoder Representations from Transformers (BERT), Sentence-BERT (SBERT), Robustly Optimized BERT pre-training Approach (RoBERTa), Decoding-enhanced BERT with disentangled attention (DeBERTa), and OpenAI’s text-embedding-ada-002. The evaluation shows our V-SBERT system based on TF-IDF and SBERT variations outperforms other algorithms for this task.1https://github.com/PurpleFlow22/UIC-FAQ-chatbot