Classification of Code-Mixed Text Using Capsule Networks

Shanaka Chaturanga, Surangika Ranathunga · 2021

A major challenge in analysing social media data belonging to languages that use non-English script is its code-mixed nature.Recent research has presented state-of-the-art contextual embedding models (both monolingual s.a.BERT and multilingual s.a.XLM-R) as a promising approach.In this paper, we show that the performance of such embedding models depends on multiple factors, such as the level of code-mixing in the dataset, and the size of the training dataset.We empirically show that a newly introduced Capsule+biGRU classifier could outperform a classifier built on the English-BERT as well as XLM-R just with a training dataset of about 6500 samples for the Sinhala-English code-mixed data.

Read the paper · More papers on PaperTik