Compacting Language Model for Natural Language Understanding on English Datasets
Nicholaus Hendrik Jeremy, Derwin Suhartono · 2024
Researchers strive to create a language model that is capable of doing various natural language understanding tasks. To do so language models have been improved significantly in terms of method and architecture. However, trends on improving also comes with the drawback of significantly increasing parameter amount. Knowledge distillation is one of many parameters reduction method that is popular to use by utilizing teacher-student relationship. In this paper we attempt to implement knowledge distillation from BERT using technique used in DistilBERT and investigate how far the distillation can go. The research is evaluated on GLUE, comparing the result to BERT and DistilBERT when trained on the same dataset prior to finetuning to target task. Our research discovers that the least amount of remaining layers possible for reduction is 4, pushing from 6. Our paper is also the first that investigates the effect of sequential distillation on language model.