On the Utility of Pretraining Language Models on Synthetic Data

Alcides Alcoba Inciarte, Sang Yun Kwon, El Moatez Billah Nagoudi, Muhammad Abdul-Mageed · 2024

Development of pre-trained language models has predominantly relied on large amounts of datasets.However, this dependence on abundant data has limited the applicability of these models in low-resource settings.In this work, we investigate the utility of exploiting synthetic datasets acquired from different sources to pretrain language models for Arabic.Namely, we leverage data derived based on four different methods: optical character recognition (OCR), automatic speech recognition (ASR), machine translation (MT), and generative language models.We use these datasets to pre-train models in three different architectures: encoderonly (BERT Base ), encoder-decoder (T5), and decoder-only (GPT-2).We test the capabilities of resulting models on Arabic natural language understanding (NLU) tasks using the ORCA benchmark.Our results show that utilizing synthetic data can achieve performance comparable to, or even surpassing, those trained on gold data.For example, our model based on a GPT-2 architecture trained on a combined synthetic dataset surpasses the baseline model ARBERT v2 .Overall, our models pre-trained on synthetic data demonstrate robust performance across various tasks.This highlights the potential of synthetic datasets in augmenting language model training in low-resource settings.

Read the paper · More papers on PaperTik