Subword-based Sentence Representation Model for Sentiment Classification
Danbi Cho, Hyunyoung Lee, Seung-Shik Kang · 2020
While most embedding methods in the Korean language focus on morpheme unit to alleviate the out of vocabulary problem, recent researches in the English use the subword unit for embedding. Considering that a word is composed of subwords, which have a partial role in a word, we hypothesize that a sequence of subwords enriches the meaning of a sentence than a sequence of words or morphemes. We propose a sentence embedding method based on a sequence of subwords in the Korean language. We evaluate the effectiveness of our sentence embedding method on binary sentiment classification using Naver Sentiment Movie Corpus. By comparing the performance of sentence embedding based on a sequence of words, morphemes, and subwords, we verify that sentence embedding based on a sequence of subwords is more robust to the out of vocabulary problem than the others.