BERT-based ensemble model for Hindi summarization

Dharam Buddhi, Digvijay Singh · 2022

A growing percentage of internet users report Hindi as their primary language of communication. Hindi, one of the world's most widely spoken languages, is India's national language. Because of its proliferation, Hindi data poses a considerable problem in organising, analysing, and summarising its vast stores of information for a variety of purposes. However, the scope of language modelling and NLP efforts with this audience in mind is extremely limited. Even the most sophisticated multilingual models struggle with the intricacies of the language. To bridge this gap, the MuRIL [37] language model was developed and trained using sizable Indian text corpora. Document summarization in Hindi is the topic of this research. Leverage the power of the MuRIL model by embedding the language model for a novel extractive summarization-based solution. The model developed in this study outperforms previous baselines on the accuracy metric by using a wide range of newspaper articles from different categories as training data.

Read the paper · More papers on PaperTik