A Study of DistilBERT-Based Answer Extraction Machine Reading Comprehension Algorithm

Bo Li · 2024

The task of answer extraction for machine reading comprehension has been one of the most popular and emphasized directions of research in recent years, which requires a machine to be able to accurately predict the answer to a given question across the span of answer intervals. Generally speaking, increasing the model size usually improves the performance of the downstream task when pre-training the language model, however, in some cases, due to the limitations of computational resources and training time, further increasing the model's size will be not suitable for specific operations. In order to effectively solve the above problems, based on the full consideration of the target task, this paper proposes a model named Parallel-DistilBERT-QA for the answer extraction machine reading comprehension task. The model combines the methods of knowledge distillation and parameter sharing, and uses DistilBERT as well as the dual-stream network structure to construct a feature extraction module, which uses three parallel DistilBERTs to process question features, question and document joint features, and document features. Between the first one and the second one as well as the second one and the third one of DistilBERTs, the dual-stream network structure is used, which realizes the sharing of weight parameters; at the same time, in order to improve the computational accuracy, an information interaction module is used after the feature extraction module, which calculates the correlation of the local features in order to capture the blind spots that are omitted by the global features. The experiments show that the computational performance of the model in this paper outperforms BERT on SQuAD1.1 and SQuAD2.0 in all aspects, and the amount of parameters is reduced by about 31% compared with that of BERT-large, and the EM and F1 scores on the test set of SQuAD2.0 are ahead of BERT-large, with the leading scores of 1.91 points and 1.73 points.

Read the paper · More papers on PaperTik