GPU Instruction Hotspots Detection Based on Binary Instrumentation Approach
Anton V. Gorshkov, Michael Berezalsky, Julia Fedorova, Konstantin Levit-Gurevich, Noam Itzhaki · IEEE Transactions on Computers · 2019
The problem of profiling a compute kernel running on the CPU is mostly solved with the help of technologies that explore a code behavior in detail. But with a last decade trend when computation spreads to other devices, more power and performance efficient, we face a high need for fine-grain code profiling on such devices. Traditional methods are not always sufficient: program counter sampling requires special hardware support, and performance simulation works slowly. In this paper, we introduce a novel method for instruction hotspots detection based on the binary instrumentation approach and demonstrate it on Intel® Graphics. This method relies on three key principles: measurement with instruction block granularity, conscious placement of probes, and combination of static and runtime information. We demonstrate its ability to highlight the hottest instructions and lines of code, and its relatively low runtime overhead comparable to a native run. The method is applicable to GPUs and accelerators with in-order architecture, and could be used in a rapidly growing segment of accelerator solutions for computer vision and artificial intelligence.