BestOf: an online implementation selector for the training and inference of deep neural networks

Sergio Barrachina, Adrián Castelló, Manuel F. Dolz, Andrés E. Tomás · The Journal of Supercomputing · 2022

Abstract Tuning and optimising the operations executed in deep learning frameworks is a fundamental task in accelerating the processing of deep neural networks (DNNs). However, this optimisation usually requires extensive manual efforts in order to obtain the best performance for each combination of tensor input size, layer type, and hardware platform. In this work, we present , a novel online auto-tuner that optimises the training and inference phases of DNNs. automatically selects at run time, and among the provided alternatives, the best performing implementation in each layer according to gathered profiling data. The evaluation of is performed on multi-core architectures for different DNNs using , a lightweight library for distributed training and inference. The experimental results reveal that the auto-tuner delivers the same or higher performance than that achieved using a static selection approach.

Read the paper · More papers on PaperTik