Synthetic Data Generation for Biomedical Deep Learning: Methods, Challenges, and Opportunities
Zlatan Car, Sandi Baressi Šegota, Branka Dobraš · 2023
With the proliferation of deep learning (DL) techniques in biomedical applications, the need for large-scale and diverse datasets has become more and more apparent. However, obtaining labeled biomedical data is often challenging. Concerns such as patient privacy, data sharing issues, ethical questions, and the lack of data – either due to the bias towards patients affected by an illness in medical examinations or due to the rarity of the investigated disease, can cause significant issues in the data collection process. In addition, manual annotation of collected data is time-intensive and requires trained personnel. One of the potential solutions discussed in the area of data science is the application of synthetically generated data, with the goal of creating artificial data points, based on previously collected data, which can aid in model training. A look into the existing synthetic data applications and generation methods, for both numeric and image data is provided by the authors. This paper explores the potential of synthetic data generation as a solution to this data scarcity, with the focus given on current state-of-the-art methods, standard approaches and challenges introduced by the application of the synthetic data in DL methodologies, and future opportunities in the field.