Sentiment Analysis for Chinese Dataset with Tsetlin Machine
Xuanyu Zhang, Zhou Hao, Ke Yu, Xuan Zhang, Xiaofei Wu, Anis Yazidi · 2022
With the increasing popularity of deep learning, researchers and practitioners seem to prefer deep neural networks (DNN) in Natural Language Processing (NLP) owing to the superior performance compared with classical machine learning techniques. However, the black box nature of deep learning poses transparency and explainability barriers and reduces their trustworthiness. Recently, a newly mentioned model, Tsetlin Machine (TM), offered reliable performance and human-level interpretability in many natural language processing (NLP) tasks. However, the related work is concentrated on English language, while the research on Chinese datasets is still open. In this paper, we employ the TM model for sentiment analysis for Chinese datasets, where the learning process is transparent and easily-understandable. More specifically, the clauses of TM make it capable of learning semantic information of Chinese vocabulary. Experiment results have shown that TM can provide similar or even higher accuracy and F1 score than more complex but non-transparent deep learning models.