Addressing Small Sample Size and Improving Interpretability in Biomedical Applications with Adaptive Complexity Deep Neural Networks
William I. Baskett, Chi‐Ren Shyu · 2025
Biomedical research often faces challenges with limited dataset sizes, primarily due to privacy concerns limiting aggregation and the high costs associated with data collection. These limitations make it difficult to train effective machine learning models, especially those designed to work with complex, unstructured data such as time-series sequences from medical records, or images and 3D volumes from CT and MRI scans. To address this issue, we propose Adaptive Complexity Deep Neural Networks (ACDNNs), a novel approach designed for small dataset scenarios. ACDNNs tackle the critical challenge of selecting the right model complexity to avoid overfitting in limited data settings. They simultaneously learn and solve problems across a hierarchy of complexity levels, enabling the model to adjust its complexity during evaluation, rather than relying on pre-set assumptions before training begins. We demonstrate that ACDNNs can be trained on small datasets without significant performance loss due to overfitting. In a time-series AKI prediction dataset derived from MIMIC-IV EHR data, ACDNNs show lower training error compared to conventional CNNs on datasets with fewer than 10,000 samples. Additionally, on the BreastMNIST and AdrenalMNIST medical image and 3D volume benchmarks, ACDNNs achieve accuracies of 90.1% and 82.4%, respectively, outperforming modern methods tested by the dataset authors which achieve maximum scores of 86.3% and 80.3% respectively. In exploratory evaluations on structured synthetic and genetic data we observe that ACDNNs excel in problems with large numbers of confounders. We demonstrate that this property can be exploited to aid identification of key predictive objects in medical images.