Impact of Ligature Coverage on Training Practical Urdu OCR Systems
Muhammad Ferjad Naeem, Noor ul Sehr Zia, Aqsa Ahmed Awan, Faisal Shafait, Adnan Ul Hasan · 2017
A major hurdle in the development of practical Urdu Nastaleeq script OCR is the lack of transcribed data, which is a pre-requisite for training machine learning algorithms. Most of the previous research has focused on UPTI, a publicly available data set with no particular focus on performance on real world images. UPTI contains only 6000 of the most probable 26,000 ligatures of Urdu. We build upon UPTI with a new data set, UPTI 2.0 that covers over 18,000 ligatures of Urdu Nastaleeq, hence covering over 70% of the ligatures that can practically occur. We further train a system on UPTI 2.0 and compare its performance against the only commercial Urdu Nastaleeq OCR system to date. Bidirectional Long Short-Term Memory (BDLSTM) network is employed with Connectionist Temporal Classification (CTC) layer as the recognizer. We show that systems trained on UPTI 2.0 outperform the commercial system.