A NEURAL NETWORK APPROACH FOR DETERMINING THE SEMANTICS OF HOMONYMS IN RUSSIAN-LANGUAGE TEXTS
Svetlana Anatol'evna Nikitina, A. A. Korzunina · Vestnik komp iuternykh i informatsionnykh tekhnologii · 2025
The phenomenon of homonymy can be found in many natural languages. Its essence consists in the sound coincidence of various linguistic units that have unrelated concepts. The presence of homonyms in the text can become an obstacle to its correct computer processing. Therefore, currently, the removal of homonymy is often considered as a separate stage of text analysis in machine translation tasks, extracting the main semantic content from information, improving the accuracy of query processing, and others. The process of automatic detection of homonyms, as well as determining their semantic meaning, is an important task in the field of artificial intelligence. A person is able to define homonymy based on context. The rules for computer processing of texts containing homonyms are based on a similar approach, that is, first a search for a homonymous word takes place, and then its semantics is predicted according to a given context. The most urgent task of removing homonymy is for languages with complex word formation and inflection, including the Russian language. The article discusses the use of neural networks for processing Russian-language text in order to determine the semantic meaning of homonyms. To solve this problem, a special neural network architecture was created, a dataset from texts from the National Corpus of the Russian Language website was collected and marked up, and special data preprocessing (tokenization and lemmatization) was carried out. The resolution of homonymy in this case boils down to the task of classification. The trained network showed acceptable recognition accuracy on test examples.