Runtime Prediction of AI Model Operations Using a GRU-Based Neural Network
Phothilimthana, Phitchaya Mangpo, Sami Abu-El-Haija, Kaidi Cao, Bahare Fatemi, Burrows, Mike, Charith Mendis, Bryan Perozzi · arXiv (Cornell University) · 2023
Modern AI models can be represented as computational graphs, where eachnode corresponds to a tensor operation (e.g., matrix multiplication, convolution),and edges represent tensor data flows. Optimizing the executionof these graphs on hardware accelerators such as Tensor Processing Units(TPUs) requires careful selection of compiler configurations that control layoutand tiling strategies.The compilation configuration involves two key types of optimizations:• Layout Configuration: Controls how tensors are arranged in physicalmemory by specifying the dimension order for inputs and outputs ofeach operation node.• Tile Configuration: Controls the tile size of each fused subgraph,impacting data locality and parallelism. Accurately predicting the runtime of AI model graphs under various configurationscan automate and improve the selection of optimal compiler settings,reducing execution time and resource consumption. The Kaggle competition dataset “Google - Fast or Slow? Predict AIModel Runtime” provides runtime data for XLA High Level Optimizer (HLO)graphs running on TPU v3 hardware. This dataset, called TPUGraphs,comprises multiple collections with diverse layouts and tiling configurations,posing a challenging performance prediction task. This work proposes a GRU-based runtime prediction pipeline leveragingopcode runtime features, graph structure dependencies, and configurablenode embeddings. The method consolidates node features, integrates configurationconvolutions, and trains a neural network to predict runtime withmean squared error loss. The approach captures both static graph propertiesand dynamic configuration effects, enabling enhanced runtime estimation toguide compiler heuristics.