Scheduling for Clustered VLIW Architectures

Jesus SBnchez, Antonio GonzBlez · 2000

nowadays a common trend in the design of embedde&DSP processors. In this work we propose a novel niodulo scheduling approach for such architectures. The proposed technique performs the cluster assignment and the instruction scheduling in a single pass, which is more effective than doingflrst the assignment and latter the scheduling. We also show that loop unrolling signijicantly enhances the performance of the proposed schedule< especially when the communication chunriel among clusters is the main perjiormance bottleneck. By selectively unrolling some loops, we can obtain the best performance with the minimum increase in code size. Performance evaluation for the SPECfp95 shows that the clustered architecture achieves about the same IPC (Instructions Per Cycle) as a unified architecture with the same resources. MoreoveK when the cycle time is taken into account, a 4-cluster conjguration is 3.6 times faster than the uniped architecture.

Read the paper · More papers on PaperTik