Direct vs. Cross-Validated Stacking in Ensemble Learning: Evaluating the Trade-Off between Inference Time and Generalizability on Fashion-MNIST
Raad Bin Tareaf, Alex Maximilian Korga, Sebastian Wefers, Keno Hanken · 2024
This study undertakes a detailed evaluation of two prominent ensemble learning techniques – direct stacking and cross-validated stacking – using the Fashion-MNIST dataset. The focus is on exploring the trade-offs between generalizability, accuracy, and inference time inherent in these methodologies. Through comprehensive experimentation, we compare the performance of various machine learning models, including LightGBM, XGBoost, Logistic Regression, and different neural network architectures, within the stacking frameworks. Our results reveal that cross-validated stacking typically offers superior accuracy and generalizability at the expense of increased computational complexity and inference time, whereas direct stacking, though computationally more efficient, is less robust in performance. The study also examines the importance of the architecture in stacking models, emphasizing that increased complexity does not necessarily lead to enhanced performance. Our findings provide essential insights for machine learning practitioners in selecting suitable ensemble techniques based on their specific needs. This research also opens avenues for future investigations into the application of these ensemble techniques across various domains and the integration of advanced model architectures. Our results show that on average the stacking models took four times longer than the respective best base models, with an average increase in accuracy of 0.875%.