Capacity-aware learning by rejecting complex samples
Ali Khudiyev, Anne Jeannin‐Girardon · Procedia Computer Science · 2025
Deep learning (DL) has achieved significant success in tackling complex tasks, largely due to insights from the universal approximation theorem, which emphasizes the importance of appropriate model scaling. Scaling laws further show a direct relationship between the size of the models and their capacity to learn difficult tasks. Although DL models work well with substantial computational resources, this incurs significant costs. Firstly, many models do not deliver satisfactory results until extensive experimentation is conducted to find out an optimal model architecture and size. Secondly, models that work well for a certain task fail when faced with a new or more complex task, however, this does not imply they should be discarded as merely “experimental studies” never to be used again. In either case, a significant amount of computational power is wasted before making it to the final inference stage. To address this issue and enable incremental learning in DL models, we propose the Capacity-Aware Learning (CAL) framework. In the CAL framework, models are trained sequentially and separately to capture the decision-making aspects overlooked by previous models. The final decision-making process aggregates the outcomes from all models. Our experimental results in image reconstruction tasks demonstrate that models trained in the capacity-aware learning setting are incremental learners that are also prone to catastrophic forgetting. In the CAL framework, no model is left behind.