SPELLING BASED RANKED CLUSTERING ALGORITHM TO CLEAN AND NORMALIZE EARLY MODERN EUROPEAN BOOK TITLES
Evan Bryer, Theppatorn Rhujittawiwat, John R. Rose and Colin F. Wilder · 2021
The goal of this paper is to modify an existing clustering algorithm with the use of the Hunspell spell checker tospecialize it for the use of cleaning early modern European book title data. Duplicate and corrupted data is a constantconcern for