Mitigating Bias in Image Classification - Pretraining on Diverse Data for Improved Generalization

Rigved Shirvalkar, M. Kalaiselvi Geetha · 2023

Deep learning models have demonstrated remarkable abilities to learn complex patterns and concepts from the data they are trained on. However, recent studies have also revealed that these models can inherit and amplify the biases that exist in the data, such as those related to gender or race. These biases can pose serious ethical issues and need to be addressed. One possible solution to reduce these biases is to pretrain the models on large and diverse datasets that cover a broad spectrum of scenarios and features. In the current study,, we investigate the effect of pretraining and transfer learning on a biased dataset that we created using a subset of the celebA dataset. The dataset is biased towards people with black hair smiling and people with blonde hair not smiling. We use the new YOLO V8 model to classify the images based on whether the person in the image is smiling or not. The paper conducts a comparative analysis between a model that underwent no prior pretraining and a model pretrained on the ImageNet dataset. We find that the pretrained model has a lower bias and also a higher accuracy than the non-pretrained model. Our findings support that pretraining has the potential to reduce the bias in computer vision tasks while enhancing the model’s generalization capabilities.

Read the paper · More papers on PaperTik