Notice the noise: detecting misclassifications in register data

V. Oosterveen · Utrecht University Repository (Utrecht University) · 2020

Label noise is present in many registers and databases. In this paper, we propose a method to identify misclassified instances in large, noisy registers. The proposed method utilizes the observed class as found in the register, a class predicted by a machine learning algorithm based on additional data sources, the agreement between those, and background characteristics of a unit that affect the misclassification probability. An expectation-maximization (EM) algorithm is then used to estimate a unit's classification error probability. We illustrate the method with the NACE classification of enterprises in the general business register (GBR). We conducted experiments to demonstrate under which circumstances the proposed method is (not) able to find the misclassified units. In the experiments, we took non-random label noise and the possible interchangeability between classes into account. Experimental results are promising: the proposed method identifies the misclassified units accurately.

Read the paper · More papers on PaperTik