AI layer reference generator engine for early enablement of workloads over multiple frameworks
Mansi Agarwal, Chandan Kumar‐Sinha, Pramod Kumar, Swadesh Verma · 2022
Compute operators/kernels (like Convolution, Relu, Gemm etc.) form the backbone of Deep Neural Network (DNN) models. With the advent of AI accelerators, there is a need to have specialized and optimized kernels for these operators to take full advantage of the features provided by the hardware. Kernels are mostly developed at High level languages like C, C++. They should work with almost the same accuracy and functionality as they would work on CPU/GPU machine. There are various ways listed below to check kernel accuracy and functionality. One of the common ways is to use Python interface provided by the different AI/ML frameworks. However, there are certain disadvantages like script level exposure which requires full stack integration and lack of multithreading because of python global interpreter lock. Another way is to develop a reference generator engine based on understanding of operators from frameworks. This strategy requires validating the correctness of the reference generator engine itself. Also, the generated reference could be at times in incompatible format and gets difficult to consume in the kernel validation code. In this paper we propose PyTeNet++ which stands for AI framework Pytorch[1], TensorFlow[2], Mxnet[3] using C++ APIs. The proposed solution is one of its kind and no such reference generator framework exists at the high-level language utilizing listed frameworks as backends. PyTeNet++ sits on top of C++ APIs exposed by these frameworks and helps to generate reference which could be seamlessly integrated with the kernel development/validation stack of the operators. It also helps to achieve maturity while kernel development. It exploits various benefits C++ provides like multithreading, low latency, high performance and OpenMP[4]. All these features are useful for faster kernel maturity. The smarter way of doing validation close to development environment helps in the early enablement (without full stack integration) of the DNN models, leading to quick maturity of the product and improved time-to-market.