Semi-Automated OCR Database Generation for Nabataean Scripts
Adnan Ul-Hasan, Syed Saqib Bukhari, Sheikh Faisal Rashid, Faisal Shafait, Thomas Michael Breuel · UWA Profiles and Research Repository (UWA) · 2012
A large amount of real-world data is required to train and benchmark any character recognition algo-rithm. Developing a page-level ground-truth database for this purpose is overwhelmingly laborious, as it in-volves a lot of manual efforts to produce a reason-able database that covers all possible words of a lan-guage. Moreover, generating such a database for his-torical (degraded) documents or for a cursive script like Urdu1 is even more complex and grueling. The pre-sented work attempts to solve this problem by propos-ing a semi-automated technique for generating ground-truth database. It is believed that the proposed automa-tion will greatly reduce the manual efforts for devel-oping any OCR database. The basic idea is to apply ligature-clustering prior to manual labeling. Two pro-totype datasets for Urdu script have been developed us-ing the proposed technique and the results are also pre-sented. 1