Cross-Language Aspect Extraction for Opinion Mining

Nguyễn Thị Thanh Thủy, Ngo Xuan Bach, Tu Minh Phuong · 2018

Aspect extraction is a subtask of aspect-based opinion mining, which aims to identify and categorize opinion targets such as product features in opinionated text. This is a challenging problem since users often use different words to express the same aspect or even mention about aspects implicitly. Supervised methods, while accurate, are not widely used for this problem mainly because they require annotated data which is costly to obtain, especially for new product and aspect categories. In this paper, we present a supervised aspect extraction method that utilizes annotated data from another language (English), which is translated to the target language (Vietnamese) by a machine translation tool (Google Translate). To alleviate the negative effect of the difference in expressions and words in different languages, we propose using word embedding features to capture the similarity and the relationships between words in order to increase the effectiveness of the extraction process. We also introduce an annotated corpus of aspect categories extracted from restaurant reviews in Vietnamese, in which we conducted experiments to evaluate our method. Experimental results demonstrate the effectiveness of the proposed approach.

Read the paper · More papers on PaperTik