A Standalone GCC-based Machine-Dependent Speed Optimizer for Embedded Applications
S. Aguirre, Vikas A. Aggarwal, D. Mlynek · 2004
Abstract. Architecture-dependent optimizations in GNU C Compiler (GCC) have been studied in order to increase superscalar processor per-formance. However, for cost and power consumption reasons, embedded systems are essentially based on scalar processors. We present a software solution at the compilation level that improves the performance of single instruction issue, in-order processors at no cost and also helps save energy. More precisely, this work highlights the GCC compiler toolchain's weaknesses in terms of eciency for scalar proces-sors and proposes a standalone software tool to reduce program execution time by avoiding useless stall cycles thanks to improved scheduling. A custom proler based on the VMIPS simulator has shown that many clock cycles associated with the delay slots of branch and load instruc-tions are lost. This is due to the GCC algorithm, and more precisely the one implemented in the assembler (GAS). To avoid some of these wasted clock cycles, an independent software pro-gram, easily integrated into the compilation ow, attempts to ll in as many load delay slots as it can, in addition to those lled by the GAS algorithm. After brie y describing the gas algorithm, we detail the one implemented in our tool. Experimental results on the MIPS architecture have been performed with gains of up to 9 % at no cost. 1