Dissected Urdu Dots Recognition Using Image Compression and KNN Classifier
Shivani Wadhwa, Deepak Kumar, Shivani Gupta, Vinay Kukreja · 2022 International Conference on Data Analytics for Business and Industry (ICDABI) · 2022
Character recognition for the Urdu language is always a tedious task because of its complexity as Urdu words are composed of ligatures that further consist of primary and secondary components. The Secondary components further consist of Urdu diacritical marks which play an important role in recognition of the Urdu language. This paper presents the technique for recognizing these secondary components to improve the recognition accuracy of the Urdu language. For this, three types of features of secondary components are extracted i.e. invariant moments, DCT, and DFT. The trained data set consists of a total of 228 features out of which 28 are invariant moment features, 100 are DCT features and 100 are DFT features. These extracted features are classified using the KNN classifier for recognition. After extraction of variant and invariant features of characters, the KNN is evaluated with different threshold values. When the threshold value of KNN is considered as 1, then the KNN achieves 89.6% recognition rate for variants and invariants features. But, if the value of k is set to be 2 the recognition rate of KNN (97.57%) increases. Thus, the value of k improves the character recognition rate.