Extracting generic text information from images
Chen Zeng · UTS ePRESS (University of Technology Sydney) · 2013
Another algorithm is designed for detecting texts from natural scene images.Maximally Stable Extremal Regions (MSERs) as character candidates are classified into character MSERs and non-character MSERs based on geometry-based, stroke-based, HOG-based and colour-based features.Two types of misclassified character MSERs are retrieved by two different schemes respectively.A false alarm elimination step is performed for increasing the text detection precision and the bootstrap strategy is used to enhance the power of suppressing false positives.Both promising recall rate and precision rate are achieved.In the aspect of text binarisation research, the combination of the selected colour channel image and graph-based technique are explored firstly.The colour channel image with the histogram having the biggest distance, estimated by mean-shift procedure, between the two main peaks is selected before the graph model is constructed.Then, Normalised cut is employed on the graph to get the binarisation result.For circumventing the drawbacks of the grayscale-based method, a colour-based text binarisation method is proposed.A modified Connected Component (CC)-based validation measurement and a new objective segmentation evaluation criterion are applied as sequential processing.The experimental results show the effectiveness of our text binarisation algorithms.