AliBERT: A Pre-trained Language Model for French Biomedical Text

Aman Berhe, Guillaume Draznieks, Vincent Martenot, Valentin Masdeu, Lucas Davy, Jean‐Daniel Zucker · 2023

Over the past few years, domain-specific pretrained language models have been investigated and have shown remarkable achievements in different downstream tasks, especially in biomedical domain.These achievements stem on the well-known BERT architecture which uses an attention-based selfsupervision for context learning of textual documents.However, these domain-specific biomedical pre-trained language models mainly use English corpora.Therefore, non-English, domain-specific pre-trained models remain quite rare, both of these requirements being hard to achieve.In this work, we proposed AliBERT, a biomedical pre-trained language model for French and investigated different learning strategies.AliBERT is trained using regularized Unigram based tokenizer trained for this purpose.AliBERT has achieved stateof-the-art F1 and accuracy scores in different down-stream biomedical tasks.Our pre-trained model manages to outperform some French non domain-specific models such as Camem-BERT and FlauBERT on diverse down-stream tasks, with less pre-training and training time and with much smaller corpora.

Read the paper · More papers on PaperTik