Predicting Language Endangerment: A Machine Learning Approach
Pankaj Dwivedi, C Shraddha, Shreyas Mathews, Sudipta Majumder, R Madhumathi, M R Vasundhara · 2020
Some languages flourish, while others may decline or eventually turn extinct in due course of time. When there are fewer and fewer people claim a language as their own, the language falls to endangerment; they neither use it nor pass it on to next generation [1]. Identifying actual vitality of a language is subject to many factors, which interact in dynamic ways and; therefore, they are not entirely predictable. However, recognizable trends may be found on careful observation. This study attempts to create a predictive model based on regression using machine learning approach to forecast the timeline as to when a language may become extinct following its ongoing vitality trend. The main datasets have been used from Census of India surveys of 1991, 2001, and 2011. Presently, this model is tested on the languages of Sikkim. However, the model may forecast the vitality prediction of any language.