Analysis of Subword based Word Representations Case Study: Fasttext Malayalam

Mannava Vivek, Priya Chandran · 2022 IEEE 19th India Council International Conference (INDICON) · 2022

Representation learning has played a crucial role in various natural language processing tasks since the advent of deep neural networks. Recent research shows that word representation learnings perform more accurately by incorporating subword information. With increased parameter sharing, subword-based representations learn more semantic and syntactic knowledge, thus gaining the ability to generate reliable representations for rare and out-of-vocabulary words. This paper reviews some popular techniques involved in subword-based representations and the publicly available pre-trained subword embedding models for Indian languages. The greater morphological richness and higher inflections make Indian languages differ from other languages commonly studied in NLP. Selected fastText models in Malayalam are analyzed in-depth in this paper. Our experiments show that intrinsic and extrinsic evaluations of some of the currently available models produce low accuracy for Malayalam compared to high-resource languages like English. It is hypothesized that the data quality, morphological richness, and higher inflections play a significant role in the low accuracy.

Read the paper · More papers on PaperTik