Dynamic Binary Cross Entropy: An effective and quick method for model convergence

Chinmay Kulkarni, Mohith Rajesh, S. S. Shylaja · 2022

The current deep learning era has huge models of over a million parameters. Training such huge models can take a significant amount of time. The recent advancements in hardware like CUDA-enabled GPUs to train the machine and deep learning models have become extremely popular. Claims suggest that the learning times can often are reduced from days to hours. However, the training time can further be reduced by optimizing the software component of the machine and deep learning models. This is the opportunity we take to introduce our loss function, the Dynamic Binary Cross Entropy (DBCE). Recent studies have shown that dynamizing the parameters of the loss functions facilitates the training of deep learning models. We aim to perform minimal calculations to update the parameters of the loss function with no extra hyper-parameters, yet not compromising on the performance. We claim that by the end of an epoch, the DBCE updates its weights such that the model’s performance is better than the model trained by its vanilla counterpart. This can help in the early stopping of the model’s training, as with DBCE, the model can achieve higher metric values earlier than the model trained with standard Binary Cross Entropy.

Read the paper · More papers on PaperTik