Stochasticity and Skip Connection Improve Knowledge Transfer
Luong Trung Nguyen, Kwang-Jin Lee, Byonghyo Shim · 2020
Deep neural networks have achieved state-of-the-art performance in various fields. However, DNNs need to be scaled down to fit real-word applications where memory and computation resources are limited. As a means to compress the network yet still maintain the performance of the network, knowledge distillation has brought a lot of attention. This technique is based on the idea to train a student network using the provided output of a teacher network. Deploying multiple teacher networks facilitates learning of the student network, however, it causes to some extent waste of resources. In the proposed approach, we generate multiple teacher networks from a teacher network by exploiting stochastic block and skip connection. Thus, they can play the role of multiple teacher networks and provide sufficient knowledge to the student network without additional resources. We observe the improved performance of student network with the proposed approach using several dataset.