Sensitive Keyword Detection on Textual Product Data: An Approximate Dictionary Matching and Context-score Approach
Duc-Hong Pham · Indian Journal of Computer Science and Engineering · 2021
In this paper, we define and study a new problem in the field of natural language processing and data science called Sensitive Keyword Detection, which aims at detecting variants of keyword in each textual product.We propose a method using approximate dictionary matching algorithm and contextscore to solve this new text mining problem in a general way.Given a product data set, for each document will be dectected keyword-variant pairs, then to determine if they are similar in semantic, we compare the context score between them.We conduct experiments on 1.189.690textual sentences extracted from descriptions and titles of 300.000 products, each textual sentence contains an original keyword or a keyword variant.Experimental results show that important role of the context-score approach and the proposed method is effective when using context-score based on character and word embedding.