Training Large Kernel Convolutions with Resized Filters and Smaller Images
Shota Fukuzaki, Masaaki Ikehara · 2023
Convolution is an essential component in neural networks for vision tasks. Convolutions on large receptive fields are more suitable than small kernel convolutions when aggregating global features in convolutional neural networks. However, large kernel convolutions require much computation and memory usage, slowing training neural networks. Then, we propose to train convolution weights with small images, resizing the convolution filters. While this idea shortens the time for training filters, simply applying this causes profound degradation. In this paper, we introduce four techniques that suppress degradation; weight scaling, removing Batch Normalization, defining a minimum resolution, and training with various-size images. In our experiment, we apply our proposals to train an image classification model based on RepLKNet-B on the image classification task of the CIFAR-100 dataset. Training with our proposals is approximately eight times faster than conventional training on the target spatial scale, keeping its accuracy.