Enhancing Scalability of Pre-trained Language Models via Efficient Parameter Sharing
Peiyu Liu, Zefeng Gao, Yushuo Chen, Xin Zhao, Ji-Rong Wen · 2023
In this paper, we propose a highly parameterefficient approach to scaling pre-trained language models (PLMs) to a deeper model depth.Unlike prior work that shares all parameters or uses extra blocks, we design a more capable parameter-sharing architecture based on matrix product operator (MPO), an efficient tensor decomposition method to factorize the parameter matrix into a set of local tensors.Based on such a decomposition, we share the important local tensor across all layers for reducing the model size and meanwhile keep layerspecific tensors (also using Adapters) for enhancing the adaptation flexibility.To improve the model training, we further propose a stable initialization algorithm tailored for the MPObased architecture.Extensive experiments have demonstrated the effectiveness of our proposed model in enhancing scalability and achieving higher performance (i.e., with fewer parameters than BERT BASE , we successfully scale the model depth by a factor of 4× and even achieve 0.1 points higher than BERT LARGE for GLUE score).The code to reproduce the results of this paper can be found at https: //github.com/RUCAIBox/MPOBERT-code.