Dynamic Text Modeling and Categorization Framework based on Semantics Extraction and Similarity Checking

Gaukhar Dauzhan, Alisher Iskakov, Miras Kenzhegaliyev, Dinh-Mao Bui · 2022

Semantics extraction is necessary to obtain relevant knowledge from general data. Recently, many researchers have developed numerous techniques to retrieve the semantics and use this information for categorization purposes. Primarily, the methodologies can be grouped into two main approaches: knowledge-based and corpus-based approaches. In this work, we studied how to achieve the same purpose using semantics extraction and matching methods. In order to prove the effectiveness of our approach, we utilized the problem of news article categorization as a use case. In fact, we developed a dynamic tool that adopts the proposed technique to crawl and classify articles from many well-known newspapers. Furthermore, the topic of articles can be grouped by their meta-information or by user-defined categories flawlessly. For generalization, the Universal Sentence Encoder model has been used for distributed semantic text representation and trained on a large processed Wikipedia corpus to achieve better accuracy. Besides, there is a possibility to enhance the results further using multi-label categorization or optimizing the meta-categories from the training dataset shortly.

Read the paper · More papers on PaperTik