Aligning Word Embeddings from BERT to Vocabulary-Free Representations
Alejandro Rodríguez Pérez, Korn Sooksatra, Pablo Rivas, Ernesto Quevedo, Javier S. Turek, Gisela Bichler, Tomas Cerny, Laurie Giddens, Stacie Petter · 2023
This paper investigates the limitations of transformer-based models in handling a fixed vocabulary, which can lead to poor generalization of out-of-vocabulary words and domains. To address this, we explore the use of transfer learning from a vocabulary-rigid transformer to a vocabulary-free one by aligning the word-embedding layer. Our approach trains a CNN to mimic the word embeddings layer of a BERT model, using a sequence of byte tokens as input. By replacing the word embeddings layer of the baseline BERT model with the aligned CNN network, we evaluate the model's generalization performance and ability to handle a broader range of linguistic inputs. Our results demonstrate the advantages of using cosine-based loss functions in the alignment process. Our approach makes important contributions toward developing more flexible and robust NLP models.