VTensor: Using Virtual Tensors to Build a Layout-Oblivious AI Programming Framework
Feng Yu, Huimin Cui, Xiaobing Feng · 2019 IEEE International Conference on Signal, Information and Data Processing (ICSIDP) · 2019
With neural networks have achieved great successes in a number of applications such as image classification, visual recognition. Researchers have proposed serval AI programming frameworks such as TensorFlow, Caffe that allows users to train neural network models or deploy them for inference using different accelerate libraries. However, when adding new libraries or new operators to existing frameworks, developers need to spends a lot of effort on layout-dependent code because tensor layout and framework are tightly coupled. Furthermore, inserting a new tuning algorithm into the framework is also difficult. In this paper, we propose VTensor, a novel programming model for developing neural network operators. The key insight is that, layout-dependent can divide into three categories and tuning algorithms have fixed pattern. Specifically, layout-dependent code can split into: accessing tensor information, memory allocation, data preparing and tuning. Based on this observation, we provide layout-oblivious APIs for first and second-class code. For the three-class code, we let developers use description language to describe layout-related information of library, and use code generators to generate dispatch code and tuning code automatically. We evaluate VTensor with a variety of neural network operators in the TensorFlow. Results show that VTensor can effectively decouple the tensor layout from the framework and reduce code size by 50.85% on average. The performance overhead is smaller than 2% with ten typical Convolution Neural Networks (CNNs) execute on Nvidia Titan V.