A similarity-aware MOE-based method for optimizing tensor programs across diverse GPUs

Haowen Hou, Zihan Wang, Yining Song, Lei Gong, Chao Wang, Xi Li, Xuehai Zhou · 2025

There is an increasing demand to bring the neural network to a wide diversity of hardware devices. AI compiler can automatically optimize tensor programs, saving significant manual effort in optimizing tensor programs. However, existing methods lack hardware modeling, and knowledge-sharing mechanisms, resulting in significant re-search cost or re-training costs. We propose a hybrid analytical-empirical similarity-based hardware representation and an MoE-like model architecture that fully leverages hardware similarity. Its core idea involves representation and utilization of similarity between different hardware. By evaluation, our model compiles 61x faster with very little performance decrease compared to MetaSchedule. The compilation time aligns with TLM, and its performance is 1.25x better on average.

Read the paper · More papers on PaperTik