Architecture and compiler design issues in programmable media processors
Wayne H. Wolf, Jason E. Fritts · 2000
The processing demands for multimedia applications are rapidly escalating. Many current applications are pushing the limits of existing microprocessors, and the next generation of multimedia promises considerably greater demands. Adequate support for future multimedia requires the flexibility and computing power of high-level language (HLL) programmable media processors. This thesis examines the architecture and compiler design issues for programmable media processors. Design of the architecture requires an accurate understanding of multimedia characteristics. Using the MediaBench benchmark suite and the Impact compiler, workload and architecture evaluations were performed to define the essential architecture for programmable media processors. The workload evaluation examines various processing aspects, including functional necessities, data types and sizes, branch performance, loop characteristics, memory statistics, and instruction level parallelism. The architecture evaluation examines the performance of dynamic versus static architecture features. Most existing media processors use static architectures, but as processors progress to higher frequencies, the dynamic aspects become more prominent and dynamic hardware may be necessary to minimize stall penalties. The architecture evaluation examines static versus dynamic scheduling, dynamic aspects of instruction fetch, and performance effects in higher frequency processors. Finally, an investigation of the memory hierarchy identifies the most significant bottlenecks in memory performance. The high degree of parallelism available in multimedia applications is well researched, but less well understood is how a compiler extracts and schedules that parallelism to highly parallel architectures. Evaluation of the compiler issues begins with an investigation of the available parallelism in multimedia. While instruction level parallelism unfortunately provides only modest performance, data parallelism offers a promising avenue for increased parallelism. However, data parallelism is of a coarser level of granularity than instruction level parallelism, so conventional compiler methods do not prove very effective. Parallel compiler methods are necessary to realize the benefits of data parallelism. Unfortunately, parallel compilation requires complex dependence analysis that is often unable to identify all available parallelism. Consequently, we propose a speculative run-time technique for data parallelism that executes loop iterations in parallel across a multi-clustered architecture. This method speculatively executes several loop iterations in parallel and provides architecture support for identifying and recovering from misspeculations.