DISTRIBUTIONAL MODELS AND AUXILIARY METHODS FOR DETERMINING THE HYPERNYMS OF WORDS IN RUSSIAN

Vasiliy Yadrintsev, Anastasiia Ryzhova, Илья Владимирович Соченков · Computational Linguistics and Intellectual Technologies · 2020

This paper describes our participation in the first shared task on Automatic Taxonomy Construction for the Russian language RUSSE’2020. The goal of this task is the following: input words (neologisms that are not yet included in the taxonomy) need to be associated with the appropriate hypernyms from an existing taxonomy. For example, for the input word “duck”, it is expected that participants will provide a list of its ten hypernyms-synsets to which the word can most likely be attributed, such as “animal,” “bird” and so on. An input word can refer to one, two, or more “parents” at the same time. In this article we are trying to answer the following question: what results can be achieved using only “raw” vectors from distributional models without additional training? The article presents the results for several pre-trained models that are based on fastText, Elmo, and BERT algorithms. Also, an outof-vocabulary analysis was performed for the models under consideration. Taking into account all public scores from the leaderboards, we showed the results corresponding to the following places in the ranking: the 3rd place on public nouns, the 2nd on private nouns, the 4th on public verbs, and the 4th on private verbs.

Read the paper · More papers on PaperTik