Exploring model generalizability from a novel perspective

Shuheng Wang, baodi yu, Fanyong Meng · Machine Learning Engineering · 2025

The black-box nature is one of the bottlenecks preventing machine learning (ML) models, especially neural networks, from playing a more important role in the field of engineering. Thus, explaining the generalizability of ML models is a crucial topic in the field of artificial intelligence. Since ML model and its training dataset constitute a complex system, the training process cannot be precisely described by mathematical or physical formulas. Therefore, a unified understanding of this issue has not yet been established. This study introduces the concept of compromise in competition (CIC) from mesoscience, which originated in the field of chemical engineering, to elucidate ML model generalizability. In this work, a scale decomposition method is proposed from the perspective of training samples, and the CIC between memorizing and forgetting, refined as dominant mechanisms, is studied. Empirical studies on computer vision and natural language processing datasets demonstrate that the CIC affects model generalizability significantly. Furthermore, techniques such as dropout and L2 regularization, which aim to mitigate overfitting, can be reinterpreted in terms of the CIC between memorizing and forgetting. This study offers a novel perspective on the interpretability of the generalizability of ML models.

Read the paper · More papers on PaperTik