Explaining and Visualizing Synthetic Data Quality Using Statistical Distances
Juko Yamamoto, Takayuki Miura, Rina Okada, Masanobu Kii, Atsunori Ichikawa · 2025
The quality of synthetic data is crucial for ensuring its usability and reliability in various applications, yet evaluating its utility remains a challenge. Statistical distances quantify deviations between synthetic and real data, but their effectiveness varies. We analyze eight statistical distance measures for categorical attributes through empirical evaluation. To better interpret these differences, we introduce the concept of a quality explanation distribution, which provides a structured probabilistic view of how synthetic data deviates under a given statistical distance constraint. A gradientbased optimization approach is implemented to explore these distributions, revealing statistical distance-specific trends. Our experiments, conducted on the Adult Dataset, show that Kullback-Leibler divergence (KLD) emphasizes lowfrequency attributes more than Total Variation Distance (TVD) and$L_{2}$distance, making it preferable for analyzing rare categories, whereas TVD and$L_{2}$better capture overall distributional trends. Additionally, we observe that in highdimensional distributions,$L_{\infty}$distance is less effective in capturing distributional characteristics. These findings emphasize the importance of selecting appropriate statistical distances and show that quality explanation distributions improve interpretability.