Automatic Text Summarization for Hindi Language Using Word Embeddings: A Critical Review

Showket Ahmad Khan, Mohd Mudasir, Hilal Ahmad Khanday · 2025

For languages with an abundance of resources like English, Chinese, Arabic, French, etc., a number of Automatic Text Summarization (ATS) techniques have been proposed. However, researchers relatively devoted minimal focus on lowresource languages such as Hindi and other regional languages of India. Taking into account that the Hindi language has limited resources available, Hindi ATS is a challenging task. The two key qualities of a useful summary are the ability to capture semantics and hidden relationships between the text units. In recent years, word embeddings have been used to capture text semantics. This review provides an in-depth examination of word embeddings by investigating conventional, contextual, cross-lingual, multilingual, and Indian language-specific models. In contrast to resourcerich languages, there are not many comprehensive datasets for the Hindi language. Most Hindi datasets are either very small for practical purposes or unavailable to the masses. This review also explores the effect of word embeddings on extractive and abstractive summarization approaches, emphasizing their performance on the Hindi corpus. Lastly, this comprehensive review highlights the open challenges, recent developments, and future research directions in the constantly evolving Hindi ATS field.

Read the paper · More papers on PaperTik