Pre-trained Word Embeddings for Malayalam Language: A Review

K Reji Rahmath, P. C. Reghu Raj, P C Rafeeque · 2021

Word embeddings are used to convert human language into a numerical form by encoding the semantic properties of words. Using it each word can be transformed to a set of N-dimensional vectors. It plays a vital role in processing of linguistic applications like natural language inference, information retrieval, sentiment analysis, etc. The goal of word embedding is to capture the meaning of words in their context. And it also find the semantic relationships and similarities between words. The aim of this work is to summarize the existing embedding techniques for words and available corpus for Malayalam language. Since Malayalam is a resource-constrained Indian language, this paper is expected to help NLP researchers in Malayalam to identify the existing resources and to improve the current research trend.

Read the paper · More papers on PaperTik