Retargetable instruction scheduling for pipelined processors
David G. Bradlee · 1991
Retargetable code generators for complex instruction set computers (CISCs) have focused on sophisticated pattern matching code selection, because CISCs provide many machine instruction sequence choices. Recent pipelined processors, known as reduced instruction set computers (RISCs), provide fewer instruction sequence choices, but expose pipeline and functional unit costs to the compiler. For RISCs the compiler's emphasis must be shifted from code selection to instruction scheduling, resulting in code generation issues that are different than those for CISCs. In particular, the machine description language for a retargetable RISC compiler must contain scheduling requirements. Also, the interaction between register allocation and instruction scheduling is significant. This dissertation comprises three components. The first component discusses the Marion retargetable code generator system, which includes a machine description language that contains instruction scheduling requirements, along with other code generation information. Using Marion, code generators have been constructed for the MIPS R2000, Motorola 88000, Intel i860. The second component compares three code generation strategies for handling the interaction between instruction scheduling and register allocation, including one that I developed, RASE, that integrates the two phases. On a computation-intensive workload for three RISCs, RASE produces significantly better code than the Postpass strategy, which does not integrate the two phases, and slightly better code than IPS, which integrates the two phases to a lesser degree than RASE. The third component investigates the interaction between code generation strategies and architectural features, including register set size and structure, and operation and load latencies. On a computation-intensive workload, 64 registers yields a significant improvement over 32 registers for Postpass, but little improvement for IPS or RASE. This study also shows that, on architectures with long latencies or small register sets, or on programs with large basic blocks, RASE produces significantly better code than IPS.