Joint Goal for Word Embedding Compression Based on Word Frequency
Yunyi Kang, Defu Lian · 2022 5th International Conference on Advanced Electronic Materials, Computers and Software Engineering (AEMCSE) · 2022
With the rapid progress in natural language processing many models grow larger. However deploying these models in small devices such as mobile devices requires smaller models. Thus techniques of model compression should be used. We consider the problem of compressing word embedding of these models. Existing methods usually treat all the words as the same. These words often have the same dimension and each word cost the same memory. During our work,first we show the feasibility of further compressing word embedding by considering different impact of each word. The idea comes from the famous Zipf’s law which says few words cover most of the document. We divide the words into groups based on word frequency and encode each group differently. Second we propose a joint goal for automatically choosing the split points between different word groups and determining the size of each group. The joint goal between task performance and model compression ratio solves the problem of how to balance between both goals. By combining both goals into to an explicit formula, it’s easy to find a good solution using an end to end training strategy.It turns to an architecture search problem and we solve it using AutoML methods. The experiment show that the word embedding layer could be further compressed to 22.6% of the baseline. We also get hints on how to choose number of word groups and to balance between the model performance and compression ratio on the experimental result.