Improving Medical Predictions with Label Noise Tolerant Classification

Shivam Vernekar, Arun B. Ayyar, Arun Rajagopalan, Bijendra Kumar, Vivek Kumar Mishra · 2022

Supervised machine learning algorithms depend crucially on availability of a large pool of correctly labelled instances. However many real world datasets may contain noisy features and label values. This is even more problematic for medical datasets since noise can be asymmetric (class dependent) due to imprecise class definitions or subjective human biases during data collection. There is easy availability of a large pool of unlabelled instances which can’t be directly used for training, for example raw data from medical devices like fitness trackers, data from health surveys etc. Labelling them (with respect to certain task) with the help of human experts might be a time consuming and a costly process. Besides, learning algorithms trained on noisy datasets can reduce the quality of future inferences making them unsuitable for business use-cases. In this paper we propose a semi-supervised learning based approach where we combine information embedded in unlabelled and incorrectly labelled instances with a very small set of correctly labelled instances. To the best of our knowledge this is the first work based on semi-supervised learning for reducing the noise in generated predictions using imperfect datasets. It is demonstrated through extensive experimentation using a medical diagnostic diabetes dataset, that by using only a fraction of correctly labelled instances, we are able to generate high quality predictions on unseen data.

Read the paper · More papers on PaperTik