Trained on 100 million words and still in shape: BERT meets British National Corpus

David Adebisi Samuel, Andrey Kutuzov, Lilja Øvrelid, Erik Velldal · 2023

While modern masked language models (LMs) are trained on ever larger corpora, we here explore the effects of down-scaling training to a modestly-sized but representative, wellbalanced, and publicly available English text source -the British National Corpus.We show that pre-training on this carefully curated corpus can reach better performance than the original BERT model.We argue that this type of corpora has great potential as a language modeling benchmark.To showcase this potential, we present fair, reproducible and data-efficient comparative studies of LMs, in which we evaluate several training objectives and model architectures and replicate previous empirical results in a systematic way.We propose an optimized LM architecture called LTG-BERT.

Read the paper · More papers on PaperTik