A Synonym-based Approach for the Semantic Indexing of Texts
Georgios Grigoriadis Kotsalis, Stavros D. Nikolopoulos, Iosif Polenakis · 2022
In this work, we present an algorithmic technique for text indexing based on the utilization of classes of synonyms. The method proposed in this study utilizes a set of synonym classes in order to develop a more abstract representation of a given text focusing on the indexing of texts that express semantic similarity, according to the terms utilized. The content of the texts under consideration is represented by a set of terms that correspond to the class of synonyms substituting each term of the sentences of the text. In the proposed approach the terms are stored into vectors where the uniqueness and the multiplicity of their appearance inside the text are considered to deploy a corresponding similarity metric. Through the development of our model, we omit words that consist of monograms, di-grams and tri-grams, where a novel approach is deployed considering the optimally discriminating words over each class of synonyms that characterize each thematic area on which a text is indexed according to its relevance with semantically similar texts. We describe thoroughly the proposed approach and perform a series of evaluation experiments utilizing an adequate number of text samples from specific thematic areas, such as business, politics, sports, entertainment and technology, intending to attest the potentials of our proposed model to index texts from specific areas.