Multi-prototype Morpheme Embedding for Text Classification

Hye‐Jin Won, Hyunyoung Lee, Seung-Shik Kang · 2020

Representing a word into a continuous space, also known as a word vector, has been successful in various NLP tasks. The word-based embedding has two problems; one is the out-of-vocabulary problem and the other is does not take into account the context of a word, which is homonymy, and polysemy since word vector is represented as a single vector. To against out-of-vocabulary problem, the previous researches handle it as splitting smaller unit than word unit, in particular, mainly morphologically rich language is decomposed into morpheme unit as Korean. However, morpheme embedding has also a problem that doesn't take into account multiple senses of a morpheme although the morpheme mitigates the out-of-vocabulary problem. Therefore, we propose the Korean multi-prototype morpheme embedding method representing multiple senses of a morpheme. Connecting morphemes and POS vectors to handle multi-prototype morpheme. In the experiment, we found that our multi-prototype morpheme embedding makes morpheme in a similar context closer in the vector space than the previous morpheme embedding. Our method outperforms the previous morpheme embedding as well as a baseline.

Read the paper · More papers on PaperTik