Programming analysis of the lexical unit context

Alexey Ivanovich Gorozhanov · Current Issues in Philology and Pedagogical Linguistics · 2024

Modern natural language processing tools make it possible to operate with big textual data, but the linguists have not yet fully used these tools to solve their (often highly specialized) issues. This applied research is devoted to the development of the original method of (automatic) analysis of the contextual environment of a given lexical unit in a coherent text based on the corpus approach. The relevance of the study is due to the need to create software for linguistic purposes in Russian Federation, as well as the significant interest shown in the topic of natural language processing and corpus research in our country and abroad. The research is based on the original text of the novel “The Castle” by F. Kafka as well as texts from the online version of the “Spiegel” magazine, collected in July, 2024. Methods of modeling, experiment, corpus analysis, as well as the method of professionally oriented programming are used. As a result, two functions were written for the software complex “Balanced Linguistic Corpus Generator and Corpus Manager”, created at the Laboratory for Fundamental and Applied Issues of Virtual Education at Moscow State Linguistic University, providing the operator with data for evaluating given lexical units, which build a frequency list of the connected tokens. The nature of the connection in the first case is determined by being within the same sentence, and in the second case – by taking into account the multi-level hierarchy “main word – subordinate word” in the sentence. The author comes to the conclusion that the proposed method provides more informative data when analyzing media texts. The algorithms did not reveal errors in the programming code, but in the future they can be improved in terms of the efficiency of selecting the tokens that are significant for the analysis.

Read the paper · More papers on PaperTik