Full-Element Analysis of Side-Channel Leakage Dataset on Symmetric Cryptographic Advanced Encryption Standard
Weifeng Liu, Wenchang Li, Xiaodong Cao, Yihao Fu, J. F. Wu, Jian Liu, Aidong Chen, Yanlong Zhang, Shuo Wang, Jing Zhou · Symmetry · 2025
The application of deep learning in side-channel analysis faces critical challenges arising from dispersed public datasets—i.e., datasets collected from heterogeneous sources and platforms with varying formats, labeling schemes, and sampling settings—and insufficient sample distribution uniformity, characterized by imbalanced class distributions and long-tailed label samples. This paper presents a systematic analysis of symmetric cryptographic AES side-channel leakage datasets, examining how these issues impact the performance of deep learning-based side-channel analysis (DL-SCA) models. We analyze over 10 widely used datasets, including DPA Contest and ASCAD, and highlight key inconsistencies via visualization, statistical metrics, and model performance evaluations. For instance, the DPA_v4 dataset exhibits extreme label imbalance with a long-tailed distribution, while the ASCAD datasets demonstrate missing leakage features. Experiments conducted using CNN and Transformer models show that such imbalances lead to high accuracy for a few labels (e.g., label 14 in DPA_v4) but also extremely poor accuracy (<0.5%) for others, severely degrading generalization. We propose targeted improvements through enhanced data collection protocols, training strategies, and feature alignment techniques. Our findings emphasize that constructing balanced datasets covering the full key space is vital to achieving robust and generalizable DL-SCA performance. This work contributes both empirical insights and methodological guidance for standardizing the design of side-channel datasets.