Sentence Continuation Inference of Urdu Text by BERT Technique

Saad Bin Ahmed · Research Square · 2022

Abstract The transformer-based BERT pre-trained language model has helped to improve state-of-the-art performance on a number of natural language processing (NLP) tasks.The BERT for Arabic has been well researched and has shown to be effective.However, the BERT for Urdu has yet to be tested.In this study, we show how pre-trained BERT models may be used to recognise cursive text such as Urdu.We employed pre-trained models because a dataset for synthetic Urdu was not accessible for research, thus we provided a synthetic Urdu dataset of 53k characters and got good results. If we could provide the BERT with a large amount of data, it would deliver impressive results.Initial results for Urdu text show that if we provide more data to the BERT pre-trained model, accuracy will improve even further.

Read the paper · More papers on PaperTik