Scheduling coarse-grain operations for VLIW processors
N.G. Busa, A. van der Werf, M. Bekooij · 2002
In order to speed up current DSP applications, complex hardware accelerators may be added in DSP architectures. This means that "coarse-grain" operations, characterized by a long latency and by a complex input-output timeshape, may be available to implement the given application. In a traditional scheduling approach, coarse-grain operations are treated as bulky atomic multi-cycle operations, under the worst case assumption that inputs and output are confined at the beginning and at the end of the operation itself. We propose a novel scheduling method for VLIW processors, where coarse-grain operations are decomposed into a number of fine input and output operations. Therefore, each I/O operation is scheduled separately in order to synchronize data communication among operations in a "just in time" fashion. This leads to a higher instruction level parallelism (ILP) in the processor, and decreases the number of registers needed in the architecture. The experiments show that embedding custom hardware accelerators in a VLIW datapath, as proposed in this paper, enhances performance keeping the VLIW controller's microcode width small.