Vietnamese Antonyms Detection Based on Specialized Word Embeddings using Semantic Knowledge and Distributional Information
Van-Tan Bui, Khac-Quy Dinht, Phuong-Thai Nguyen · 2020
Antonymy is one of the fundamental relations shaping the organization of the semantic lexicon. Therefore, automatic detection of antonymy can be leveraged to make contributions to different NLP tasks, such as Machine Translation, Sentiment Analysis, and Information Retrieval. Currently, most prior studies just focus on discriminating between antonyms and synonyms. However, not only synonymy but other semantic relations, such as hypernymy, co-hyponyms, which also get high similarities thereby making it hard to discriminate. Therefore, it is necessary to make a thorough research on identifying antonyms from a wide variety of other semantic relations. In this paper, we aim to identify Vietnamese antonyms pairs according to the vector semantics approach. Specifically, we build up specialized word embedding models by incorporating lexical-semantic resource and distributional information. In addition, we propose specialized Vietnamese features and utilize mutual information between words in order to integrate with word embedding vectors. This aims to generate more meaningful feature vectors for supervised classifiers solving antonym detection problems. Furthermore, we construct three reliable Vietnamese testing datasets consisting of AntSynlOOO, AntHyplOOO, and AntMixlOOO, for this task. Experimental results conducted on the datasets demonstrated that our model performs effectively.