Descriptor: Not-A-DAtabase of Synthetic Shapes Benchmarking Dataset (NADA-SynShapes)
Giulio Del Corso, Federico Volpini, Claudia Caudai, Davide Moroni, Sara Colantonio · IEEE data descriptions. · 2025
NADA is a synthetic dataset of 1,500,000 (128x128 pixels) images of geometric shapes with non-elementary multivariate parameter distributions designed to benchmark and test novel probabilistic deep learning models. Benchmarking uncertainty-aware techniques is critical in real-world scenarios, especially in high-stakes domains such as automated driving or health data analysis. This is because the most robust and reliable AI methods are rooted in Bayesian reasoning and uncertainty analysis, yet there is no consensus on how to test or compare them. In particular, the public, synthetic, and real datasets currently available in the literature are inadequate for benchmarking most probabilistic methodologies: first, due to a lack of control over population variability; and second, due to an oversimplification of the distributional properties of latent variables. From this perspective, NADA is organized into three main From this perspective, NADA is organized into three main repositories, specifically developed to challenge uncertainty-aware methods and provide a unified benchmark reference dataset across three key areas: (1) characterization of complex latent space and evaluation of disentangling ability (NADA_Dis: 300,000 images), (2) identification of different types of aleatoric and epistemic uncertainties (NADA_AlEp: 500,000 images), and (3) detection of out-of-distribution elements for reliable AI (NADA_OOD: 700,000 images). Each repository includes the dataframe describing the image parameters (e.g., rotation, position, shape, color, deformation, noise), the dataset-generating hyperparameters (e.g., marginal distributions and correlation matrix), and evaluation plots to assess data quality. In addition, the dataset is coupled with an open-source Python synthetic generator, allowing easy modification and adaptation to specific research questions.