Cognitive Horses of Deep Learning: Regularization and Optimization Practices
Asst. Prof., CSE, KMIT, Hyderabad, India, Ajeet K. Jain, PVRD Prasad Rao, Professor, CSE KLEF, Vaddeswaram, AP India, K. Venkatesh Sharma, Karan Jain, Scholar, VIT, Vellore India · 2023
Deep learning networks incorporate various regularization and optiization techniques motivated by effective parameters tuning.The arduous task of training a deep network invariably requires lot of parameter setting along with different architectural design in order to deal with problems of overfitting and convergence speed.These two cognitive working horses play the critical role in convergence for better performance!In recent past, new modifications in these two avatar horses being put into practice and their suitability and applicability has been reported on various domains.Though technique like batchnormalization, data augmentation, dropout and skip connections have promising results in abating overfitting, each method reported have peculiarities and sometime not well addressed and have dithering understanding towards their use.Similarly optimizing algorithms use bias correction second order moment estimates for faster convergence (ADAM, NADAM, YOGI and others).Though this fine, but findings have reported that a simple SGD based optimizers sometime outperform than advanced version of the same implantation.This contradictory and agnostic behavior suggests that, one intuitive way to get into their appropriateness is empirically apply them on different data set and measure their relative performances.There is no single trending horse practices which can either perform better or outperform over another variant for a given application domain.Eventually, it is a difficult practitioner's choice to foresee and get a feeling of which regularizer | optimizer is a best choice!This is indeed a challenging task and is an issue of cognitive deep learning research-to choose an intelligent best-fitting regularizer | optimizer performing optimally for a given model.Additionally, increasing the number of deep layers with hyperparameter tuning introduces higher complexity and a tradeoff is one way to find a niche fit.The paper delves deeper into the best practices offered by various parameters tuning and reports the agnostic behavior on various datasets.