Machine Learning VLSI CAD Experiments Should Consider Atomic Data Groups
Andrew David Gunter, Steven J. E. Wilton · 2024
Machine learning (ML) has proved useful across a wide range of applications in the very-large-scale integration computer-aided design (VLSI CAD) domain. To avoid overestimating ML models' generalization capabilities for real-world deployments, best practices utilize realistic data and avoid test set information leakage during ML model preparation. In this paper we identify a further consideration, atomic data groups, which are sets of very highly correlated data that may also lead to such overestimation if not accounted for in train-test splits during model evaluation. We investigate the potential impact of atomic data groups in experimental design through a case study of field-programmable gate array (FPGA) routing. Our investigations show that model performance in deployment is overestimated by 38% in this case study when atomic data groups are ignored. We hope that these results motivate other ML CAD practitioners to be critical of their train-test splits and identify when atomic data groups are relevant to their model evaluations.