Language Identification Models for Short Medical Texts
Daniela Gîfu, Radu Adrian Ciora · 2022 E-Health and Bioengineering Conference (EHB) · 2022
In the context of SARS-CoV-2 transmission prevention, the short texts on the social networks are full of abbreviations, technical medical terms, slang words that are outside the native languages or pejorative jargon. In other words, the language identification issues are not trivial. In fact, always this task remained a challenge for short texts originating from social media, which abounds in hard-to-understand sequences of letters. Moreover, when the whole world is facing a “medical crisis”, the situation must somehow be controlled in order for the panic to not degenerate even more. As a result, we need enhanced tools that can help us speed up the process of detection of linguistic aspects of the online content. This paper presents a new method intended to identify the language of a Twitter collection related to COVID-19 subject, based on AI algorithms. The aim of this work is to optimally determine the main language in which a text is written, with the constrains given by the diversity of the short text language found in tweets. Therefore, to find an optimal solution for language detection, we evaluated the impact of several language detection algorithms for short texts. Moreover, the results of this analysis were used for the implementation of a new language detection model, focused primarily on the detection of the Romanian language of the gathered tweets. The results suggest that a hybrid approach which combines classic techniques looks like a realistic direction of research.