Corpus-based Name Standardization

Gerrit Bloothooft · History and Computing · 1994

A method is described to standardize nominal data on the basis of a combination of rules and a probabilistic similarity measure. Onomastic corpora are used to estimate the probability of spelling variations automatically. These corpora are also the basis for finding the most likely standard for a name not encountered before.

Read the paper · More papers on PaperTik