Knowledge Distillation Scheme for Named Entity Recognition Model Based on BERT
Chengqiong Ye, Alexander Arcenio Hernandez · 2023
In order to improve the applicability of named entity recognition model in industrial application scenarios, this paper studies the BERT named entity recognition task. From the perspective of lightweight model, an adaptive weight distillation method is designed to compress the BERT named entity recognition model. By adjusting the weight hyperparameter of distillation loss to an adaptive function, this method enables the model to automatically adjust the weight value according to the training process, avoiding the influence of manual parameter adjustment. Furthermore, by decreasing the middle hidden layer’s dimension, this method compresses the teacher model from the width and incorporates the distillation of the middle Transformer layer of the model into the distillation process, allowing it to fully utilize the rich semantic information in the deep structure of the model. Ultimately, a light-weight, well-functioning model with a straightforward structure and few parameters is trained. By employing the suggested technique, the model’s inference speed can be increased by seven times while reducing the number of model parameters to 15% of the original model with little loss of accuracy.