Model Distillation on Multi-domained Datasets
Alexey Orlov, Kamil Bayazitov, Andrey Grabovoy · 2023
Traditionally, machine learning models are developed for a specific task. Collecting and processing datasets for each new task are extremely expensive and time-consuming processes, and sufficient data for training may not always be available. In such cases, knowledge transfer from a teacher model to a student model can be used. The paper considers the problem of model complexity reduction when transferred to new data of lower cardinality. We use a teacher model pre-trained on a large dataset and a student model trained on a small dataset. Methods based on the distillation of machine learning models are considered. We consider the assumption that solving the optimization problem from the parameters of both models and domains improves the quality of the student's model. We present a method to train the student model using the teacher model and domain connections, which enhances the approximating ability of the student model. A computational experiment is being carried out on real datasets for computer vision and text processing tasks.