Terminological subsystems of modern Russian school textbooks: A study based on Word2Vec and neural networks
Sergei Monakhov, Vladimir Vladimirovich Turchanenko, Ekaterina A. Fedyukova, Dmitry N. Cherdakov · Journal of Applied Linguistics and Lexicography · 2021
The article reports the results of the study that explored the inventory and functioning of scientific terms and special lexemes in textbooks for Russia’s secondary schools. The toolset included modern methods of natural language processing and deep learning. The number of terms from different fields of knowledge that a secondary school student should learn has never been evaluated. According to the preliminary evaluations based on the Model Basic Curriculum for General and Secondary Education 2015, a secondary school leaver is supposed to be able to understand, recognise and use about 1,000 terms and terminological combinations in the subject Russian Language alone. Thus, taking into account the number of school subjects, the total number of special vocabulary studied in general education schools is measured in thousands. At the same time, the comparative characteristics of the inventory and functioning of terms in textbooks for different school subjects are under-scrutinized and remain unknown. Besides, it is unclear how the terminological density of school textbooks for different subjects correlates with the place occupied by these subjects in the curriculum. The traditional way of compiling lists of special terms is simply to glean them from special texts and write down manually. This method is reliable to gain insights into the best selection practices, however, it cannot be applied to large data sets and does not reflect the term frequency, the specificity of their syntagmatic connections, or the systemic relationship between them. Our project is aimed at filling this gap through: 1) creating a full-text corpus of school textbooks approved by the Ministry of Education for grades 5–11, 2) automatic extraction, stratification, and mapping of terms with the help of distribution semantics algorithms, 3) creation and training of a deep neural network capable of predicting the subject, level of education and educational topic given a group of vector theoretical development of terminology science. They may also find practical application, e. g., in the development of different types of educational literature.