Rigorous, systematic approach to automatic data editing and its statistical basis

Gunar E. Liepins · 1980

Automatic data editing is the computerized identification and correction (optional) of data errors. These techniques can provide error statistics that indicate the frequency of various types of data errors, diagnostic information that aids in identifying inadequacies in the data collection system, and a clean data base appropriate for use in further decision making, in modeling, and for inferential purposes. However, before these numerous benefits can be fully realized, certain research problems need to be resolved, and the linkage between statistical error analysis and extreme-value programing needs to be carefully determined. The linkage is provide here for the special case that certain independence and symmetry conditions obtain; also provided are rigorous proofs of results central to the functioning of the Boolean approach to automatic data editing of coded (categorical) data. In particular, sufficient collections of edits are defined, and it is shown that for a fixed objective function the solution to the fields to impute problem is obtainable simply from knowing which edits of the sufficient collection are failed, and this solution is invariant of the particular sufficient collection of edits identified. Similarly, disjoint-sufficient collections of edits are defined, and it is shown that, if the objective function of the fields to impute problem is determined by what Freund and Hartley call the number of involvements in unsatisfied consistency checks, then the objective function will be independent of the disjoint-sufficient collection of edits used.

Read the paper · More papers on PaperTik