Multilingual Inappropriate Text Content Detection System Based on Doc2vec
Kazuki Aikawa, Shin Kawai, Hajime Nobuhara · 2019
In this paper, an inappropriate multilingual text content detection method is proposed based on Neural Machine Translation - Doc2Vec (NMT-D2V) is proposed.. NMT-D2V has three features, i.e., it extends an existing detection system (e.g. based on English) to another language using a multilingual scheme, it improves detection system transparency based on similar text presentation, and it does not required text translation into the original language for each input. An experimental comparison using a Japanese dataset demonstrated that the evaluation index area under the curve of the Receiver Operating Characteristic of NMT-D2V (0.717) is better than that of an existing translation method(0.703).