A Novel Text Features Jointing Model for Review Spam Filtering of Chinese
Faxin Zhang, Lichen Qiu, Peng Qi, Hong Luo · 2020
Review spam filtering of Chinese is a hot research topic in the field of natural language processing. In recent years, there have aroused a lot of neural network models for review spam filtering of Chinese, but these models mainly focus on utilizing text semantics or part-of-speech of a review without considering the variants of a sensitive word. In this paper, based on TextCNN, we propose a novel review spam filtering model of Chinese that embeds various features of the text. Considering the common types of variant words, we first extract three sets of vectors for each Chinese character of the review, including the single character vector, the pinyin vector, and the pinyin vector of each single character. Then, we use TextCNN to learn the features of these vectors, respectively. After that, the learned features are added as the comprehensive features of review. Finally, we input it to the softmax layer to get the final classification results. The experimental results show that the joint model performs better than classic classification methods in review spam filtering of Chinese. The recall and F1-score on review spam reach 92.4% and 93.6%.