Text Area Extraction Method for Color Images Based on Labeling and Gradient Difference Method

Jongkil Won, Hye-Young Kim, Jinsoo Cho · The Journal of the Korea Contents Association · 2011

영상 입출력 장치 사용이 증가함에 따라 컬러영상 내 문자영역 추출의 중요성 또한 높아지고 있다. 본 논문은 이러한 영상 내 문자영역을 효과적으로 추출하기 위해 레이블링 기법과 화소 단위의 밝기값 변화에 기반한 문자영역 추출 방법을 제안한다. 제안하는 방법은 레이블링 및 필터링 과정을 통해 비문자 영역을 미리 제거하고, 밝기값의 변화가 큰 문자영역의 특성을 이용하여 문자영역 후보군을 추출한 후 노이즈 제거 및 문자영역 병합의 후처리 과정을 통해 문자영역을 추출한다. 제안한 방법의 강점은 기존 방법보다 단순하면서도 높은 정확성에 있다. 실험 결과 제안한 방법의 정확도와 재현율, 비문자 추출의 역 비율(IRNTE)은 각각 99.59%, 98.65%, 82.30%로 측정되었다. As the use of image input and output devices increases, the importance of extracting text area in color images is also increasing. In this paper, in order to extract text area of the images efficiently, we present a text area extraction method for color images based on labeling and gradient difference method. The proposed method first eliminates non-text area using the processes of labeling and filtering. After generating the candidates of text area by using the property that is high gradient difference in text area, text area is extracted using the post-processing of noise removal and text area merging. The benefits of the proposed method are its simplicity and high accuracy that is better than the conventional methods. Experimental results show that precision, recall and inverse ratio of non-text extraction (IRNTE) of the proposed method are 99.59%, 98.65% and 82.30%, respectively.

Read the paper · More papers on PaperTik