Multilingual BERT Cross-Lingual Transferability with Pre-trained Representations on Tangut: A Survey
Xiao-Ming Lu, Wenjian Liu, Shengyi Jiang, Changqing Liu · 2023
Natural Language Processing (NLP) systems have three main components including tokenization, embedding, and model architectures (top deep learning models such as BERT, GPT-2, or GPT-3). In this paper, the authors attempt to explore and sum up possible ways of fine-tuning the Multilingual BERT (mBERT) model and feeding it with effective encodings of Tangut characters. Tangut is an extinct low-resource language. We expect to introduce a tailored embedding layer into Tangut as part of the fine-tuning procedure without altering mBERT internal structure. The initial work is listed on. By reviewing existing State of the Art (SOTA) approaches, we hope to further analyze the performance boost of mBERT when applied to low-resource languages.