Data-Centric Machine Learning Pipeline for Hardware Verification
Hongsup Shin · 2022
In hardware verification, constrained random testing is used to cover the vast verification space efficiently. Verification engineers adjust testbench settings to guide test behavior. Since this method produces many tests, we can train a machine learning (ML) model with test outcome to identify tests with high probability of having bugs, leading to increase in testing efficiency. However, maintaining good model performance after deployment is challenging. The data types are heterogeneous, change over time, and have multiple interpretations. I developed a data-centric ML pipeline that preprocesses custom verification-specific data types automatically and enables tuning data preprocessing methods ("data tuning"). This ML pipeline is adaptive to changes in future data and thus increases deployment scalability and robustness. Real-world benchmark testing shows that data tuning with this pipeline improves model performance.