Uuden suomenkielisen Wikipedia-pohjaisen yksikäsitteistämistehtäviin soveltuvan tietokannan kehittäminen
Einari Tuukkanen · Aaltodoc (Aalto University) · 2022
This paper provides a first sense-annotated encyclopedia-based Finnish glossary for evaluating disambiguation tasks and a simple yet transparent application-ready NED solution for Finnish NLP pipelines. We aim to develop a Finnish WSD solution to be used on top of an existing NLP pipeline to improve text indexing results by disambiguating similarly named entities. To do this, we gather a data dump of Finnish Wikipedia and process it with TurkuNLP's public models to achieve lemmatization, segmentation, and POS and NER taggings. We utilize Wikipedia's disambiguation pages to link ambiguous and disambiguous entities and create a glossary with articles as the sense definitions. We use the glossary to generate disambiguation tasks and perform a user test with humans solving them. The test gathers 1523 answers. We measure a human disambiguation accuracy of 85.6 %, which we define as the golden standard. We then define criteria for an application-ready WSD algorithm based on our original goal. We then examine different algorithms using an extended Lesk algorithm with POS-based word filtering, glossary document length normalization, and TF-IDF score weighting. This algorithm reaches 69.3 % accuracy on the same subset of tasks the users solved and 59.7 % accuracy on the entire dataset. Finally, we suggest improving the dataset by using an external source for generating context for the tasks and using word embeddings to store the data. In the future, database vectorization would enable the use of more precise algorithms, such as cosine similarity and state-of-the-art neural models.