An Unsupervised Learning Approach to Text Line Detection in Complex Illuminated Medieval Manuscripts

Lizeth Gonzalez-Carabarin, Lisandra S. Costiner · Digital Humanities in the Nordic and Baltic Countries Publications · 2021

This paper outlines a simple and effective clustering and filtering approach for text line detection in challenging illuminated medieval manuscripts from Western Europe. As illuminated medieval manuscripts were copied and illustrated by hand and do not have regular formats, they pose particular difficulties for traditional methods of text line extraction, which have been designed for printed books. This paper introduces an unsupervised learning approach to text line detection in challenging manuscripts based on clustering, using a k-means algorithm with a combination of three salient features that are well-suited for image processing. These are: the gradient in the y-direction, mean values of rows, and grayscale values. The strength of the method lies in its reduced number of features, its computational lightness, low-memory use, transparency at every step of the process, and versatility. It can also be used on a single image, being perpage based, unlike supervised learning approaches which require large training datasets. This stands as an alternative to computationally-heavy algorithms such as neural networks, which have been increasingly used in recent years to solve such task.

Read the paper · More papers on PaperTik