From Extraction to Spotting for Cuneiform Script Analysis

Bartosz Bogacz, Hubert Mara · 2018

Cuneiform tablets appertain to the oldest textual artifacts used for more than three millennia and are comparable in amount and relevance to texts written in Latin or ancient Greek. We present a complete digital analysis workflow enabling modern text processing on the complex and non-linear script. Our tools encompass the whole pipeline, from digitization and wedge extraction to word-spotting and frequent pattern mining facilities. Tablets are being acquired from different sources requiring different methods for digitalization. Each representation is typically processed with its own tool-set. To homogenize these data sources, we introduce an unifying minimal wedge constellation description. For this representation, we develop similarity metrics based on the optimal assignment of wedge configurations. We combine our wedge features with work on segmentation-free word spotting using part-structured models. The presented search and similarity facilities enable the development of advanced linguistic tools for cuneiform sign indexing and spatial n-gram mining of signs.

Read the paper · More papers on PaperTik