Architecture and arithmetic for multimedia-enhanced processors

Daniel F. Zucker · 1998

In the past, displaying video on desktop systems has required high cost special purpose hardware to handle the computationally intensive task of video compression. Recently, special purpose multimedia instruction sets have allowed software-only real time video decompression without extra hardware. This means video functionality can now be handled by a general purpose CPU. With this low cost video capability, the video data type is becoming truly ubiquitous. Most major CPU vendors have adopted this strategy of enhancing a general purpose processor for multimedia applications. Examples include Hewlett Packard's MAX instruction set, Sun's VIS (Visual Instruction Set), Digital's MVI instructions, MIPS MDMX instructions, and Intel's MMX instructions. This work investigates similar techniques for applying cost-effective enhancements to a general purpose processor. Using public domain MPEG implementations as benchmarks and trace based simulation, we investigate performance for typical MPEG video decompression applications. Beginning with a system level breakdown of execution time, we propose techniques to improve execution time in three separate architectural components: I/O, arithmetic, and cache memory. For I/O, we show how applying traditional techniques of I/O cache prefetching can reduce the time for reading compressed video data. For arithmetic, we propose software-only techniques to pack multiple data words into a single floating point operand to achieve SIMD type parallelism. A theoretical framework is presented to define the technique's capabilities and limitations. We also propose hardware extensions to increase the robustness of this technique. For cache memory accesses, we compare several hardware based prefetching techniques v and then define a stream cache that e...

Read the paper · More papers on PaperTik