Probabilistic Name and Address Cleaning and Standardisation
Peter Christen, Tim Churches, Justin Xi Zhu · 2002
In the absence of a shared unique key, an ensemble of nonunique personal attributes such as names and addresses is often used to link data from disparate sources. Such data matching is widely used when assembling data warehouses and business mailing lists, and is a foundation of many longitudinal epidemiological and other health related studies. Unfortunately,