A System for Identifying and Exploring Text Repetition in Large Historical Document Corpora

Aleksi Vesanto, Filip Ginter, Hannu Salmi, Asko Nivala, Tapio Salakoski · DSpace repository (University of Tartu) · 2017

We present a software for retrieving and exploring duplicated text passages in low quality OCR historical text corpora.The system combines NCBI BLAST, a software created for comparing and aligning biological sequences, with the Solr search and indexing engine, providing a web interface to easily query and browse the clusters of duplicated texts.We demonstrate the system on a corpus of scanned and OCR-recognized Finnish newspapers and journals from years 1771 to 1910.

Read the paper · More papers on PaperTik