Data Augmentation with Diffusion Model for Hand Detection
Genta Matsukawa, Atsuo Yoshitaka · 2024
In this paper, we propose a method of data augmentation for object detection using diffusion model. Specifically, we focus on human hands as our target objects, and our model generates images of human hands with gloves from bare hands for data augmentation. The reason for focusing on the human hands is based on the observation in the study of human behavior understanding. Additionally, the reason for focusing on augmentation of human hands with gloves is that there are many situations where gloves are worn in industrial fields, such as food or manufacturing factory work. However, it was difficult to perform gloves data augmentation by means of typical data augmentation methods. Therefore, in the experiments of this paper, we first prepare a fine-tuned image generation model combining Stable Diffusion and Low-Rank Adaptation of large language models (LoRA). For natural image generation, we combine ControlNet for contour extraction as preprocessing and feed the preprocessed images into the image generation model to create natural images for training data. Finally, we evaluated the recognition performance of object detection after training with generated human hands with gloves images. Our results indicated that our method of data augmentation improves recognition performance of human hands with gloves.