French regional languages identification using custom dataset and Naive Bayes model
Stephane Maillard, Mohammad Safayet Khan · 2024
Identifying regional languages or dialects can be challenging due to the scarcity of data. There can be a huge contrast between smaller languages and more globally prevalent languages like English, where data is abundant. A lack of available resource in a specific language can contribute to smaller performance of a model in term of accuracy or other important metrics. This paper focuses on the identification of five written French regional languages, specifically Alsatian, Basque, Corsican, Breton, and Occitan, spoken by more than one million people combined. The main contribution of this study resides in the building of a custom dataset (http://dx.doi.org/10.13140/RG.2.2.31088.06406) for these regional languages using new data sources. Another contribution is the identification of these specific regional languages, that was relatively unexplored in previous studies, and the usage of a custom Naive Bayes model to achieve this goal. The motivation lies in achieving comparable performance in language identification for regional languages or dialects but also contribute significantly to their preservation and study. The custom Naive Bayes model in combination with the built dataset demonstrates improved performance in accuracy, precision, recall and F1 score compared to other models. It achieved $\mathbf{0. 9 9 2}$ on four different metrics (accuracy, precision, recall and F1 score) and 0.042 for $\log$ loss. Despite minor disparities, especially for French and Breton, the study contributes valuable insights to regional language and dialect detection, emphasizing the importance of quality data and architectural choices.