Issues of Homonyms/Homoforms Automatic Recognition in Applied Linguistics
Akerke Meirbekova, Anar Fazylzhanova, Aiman Zhanabekova, Aigul Baydebekkyzy Amirbekova, Гүлнар Талғатқызы · Cuadernos de Investigación Filológica · 2026
The phenomenon of homonymy creates additional difficulties for automatic text recognition processes, requiring the use of more complex algorithms and processing methods. The research aims to address the issues of automatic homonym recognition in Kazakh, Russian, English, Turkish and Tatar. The study identified the main problems that arise in the automatic detection and identification of homonyms/homoforms, in particular, the lack of a clear sentence structure (Russian, Turkish), significant morphological diversity of the language (the presence of a large number of grammatical categories, high level of affixation), contextual ambiguity, insufficient development of theoretical issues of homonymy study. The most effective methods of distinguishing homonyms in national corpora of Russian (Markov model with maximum entropy), English (Embeddings from Language Model), Turkish (hybrid method) and Tatar (method of removing homonyms from a corpus marked with homonyms) were considered. The present study is more aimed at solving the problem from the point of view of analysing morphology and syntax and requires a more detailed study in semantic and contextual aspects.