Multilingual Text Classification Based On Deep Learning Models
Lanxin Lin · 2023
In an era characterized by global interconnectedness, the imperative for seamless communication across diverse languages is more pronounced than ever. The present research introduces a sophisticated pre-trained model, denoted as Multilingual BERT (mBERT), meticulously engineered to execute text classification with efficacy across 21 European languages. This model is pivotal in augmenting performance uniformity amongst a myriad of European language families. Significantly, mBERT showcases enhanced parallelization and a notable reduction in processing time for extensive documents, standing out in comparison to its contemporaries. In addition, a distilled model is cultivated under the tutelage of mBERT, aiming to delve into the intricacies of knowledge transfer between the two entities. The empirical results corroborate the superior proficiency of mBERT in the realm of language type classification, thereby furnishing a robust methodology for language categorization. Concurrently, the distilled model adeptly assimilates knowledge from its precursor, mBERT, thereby validating the efficacy of the proposed methodology.