Design and implementation of a Multimedia Extension for a RISC Processor

Martínez Montes, Eduardo Jonathan · 2015

Nowadays, multimedia content is everywhere. Many applications use audio, images, and video. Over the last 20 years significant advances have been made in compilation and microarchitecture technologies to improve the instruction-level parallelism (ILP). As more demanding applications have appeared, a number of different microarchitectures have emerged to deal with them. A number of microarchitectural techniques have emerged to deal with them: pipelining, superscalar, out-of-order, multithreading, etc. All these techniques increase the ILP at core level. An alternative to increase the performance is to exploit data level parallelism. This technique improves performance in applications where highly repetitive operations need to be performed. It performs the same operation on multiple pieces of data. The first approach was vector processors and most modern implementations use short fixed size vectors in what is known as Single Instruction Multiple Data (SIMD) extensions. Present and next generation of mobile and embedded devices require having multimedia support to be competitive. The most popular technique at the microarchitecture level to deal with these requirements is to use SIMD extensions to the Instruction Set Architecture. It improves performance by processing vector operations in parallel. Their high compute power and hardware simplicity improve overall performance in an energy efficient manner. Furthermore, their replicated functional units and simple control mechanisms make them manageable to scaling to higher vector lengths. MIPS born as an academic research and also it has been the favorite architecture used to teach Computer Architecture Design. Other architectures as x86 or ARM are quite complex to understand for university students. Most popular microarchitectures today have at least a SIMD unit implementation, x86-64 (MMX, SSE, AVX), PowerPC (AltiVec), ARM (Neon), MIPS (MDMX, MSA), and so on. This work focuses on the MIPS microarchitecture because it follows the RISC (Reduced Instruction Set Computer) philosophy quite closely, which is desirable since it simplifies implementation and it is easy to understand by students. First SIMD implementation for MIPS was MIPS Digital Media Extension (MDMX) that supported video, audio and graphics pixel processing by introducing two vectors formats of small integers. These vectors have a width of 64-bits, in the form of signed 16-bit or 8 unsigned 8-bit integers. MDMX is quite old, and it never reached production. Latest SIMD implementation on MIPS appeared in 2014 as an add-on of MIPS32/64 Release 5. It is called MIPS SIMD Architecture (MSA). It is designed to support 128-bit vectors of 8, 16, 32 and 64bit integer vectors; 16 and 32-bit fixed-point, or 32 and 64-bit floating-point elements. We have decided to implement the MSA ISA on an FPGA together with a MIPS-like implementation (soft-core). Since it is not intended for general purpose computing we take a MIPS32 core from opencores.org, upgraded it with the needed MIPS Release 6 instructions, microarchitecture and control to use MSA coprocessor. We have created a MSA coprocessor, a testing bench and adapted some embedded micro benchmarks to test the MSA coprocessor.

Read the paper · More papers on PaperTik