Probabilistic linkage

William E. Winkler · Wiley series in probability and statistics · 2015

Probabilistic linkage (or record linkage) consists of methods for matching duplicate records within or across files using non-unique identifiers such as first name, last name, date of birth and address, or a national health identification code with typographical error. This chapter provides an introductory overview of some well-established methods of record linkage that have been used for different types of lists for more than 40 years. It first gives a background on the record linkage model of Fellegi and Sunter, methods of parameter estimation without training data, some brief comments on training data, and string comparators for dealing with typographical error and blocking criteria. Next, the chapter provides details of the difficulties with the preparation of data for linkage. Then, it describes methods for error rate estimation without training data and methods for adjusting statistical analyses of merged files for linkage error. The chapter ends with some concluding remarks.

Read the paper · More papers on PaperTik