Automatic Thai and English fonts identification without character recognition

Boontee Kruatrachue, P. Piyatrakul · 2002

This paper describes a simple and fast algorithm to detect Thai and English characters in a document without doing actual characters recognition. The document is segmented into strings of letters separated by a blank, then each string is identified using characters features and their writing positions. This method achieves 100% accuracy if the characters have clear head feature. But if this feature is not used 90% of the strings still can be identified. This identification provides more information about the character set so that OCR can recognize faster with better accuracy.

Read the paper · More papers on PaperTik