SPELLING BASED RANKED CLUSTERING ALGORITHM TO CLEAN AND NORMALIZE EARLY MODERN EUROPEAN BOOK TITLES

Evan Bryer, Theppatorn Rhujittawiwat, John R. Rose and Colin F. Wilder · 2021

The goal of this paper is to modify an existing clustering algorithm with the use of the Hunspell spell checker tospecialize it for the use of cleaning early modern European book title data. Duplicate and corrupted data is a constantconcern for

Read the paper · More papers on PaperTik