Scalable dual-instruction multiple-data processing on an efficient systolic-array architecture

Yuxi Tan, Riadh Ben Abdelhamid, Bingjie Guo, Qixiang Gao, Masaru Nishimura, Yoshiki Yamaguchi · 2024

For half a century, advancements in semiconductor manufacturing and processor design have propelled the evolution of computer systems, yet power consumption remains a critical challenge for future innovation. Essential for sustainable development and performance gains in AI, big data, and autonomous driving, energy-efficient computing now demands a shift from boosting processing element (PE) performance through frequency increases to improving computational efficiency. This shift has prompted a move from Single- Instruction Multiple-Data (SIMD) to Single-Instruction Multiple- Thread (SIMT) architectures, like GPU s, offering computational speed and energy savings [1]–[3]. Despite SIMD's effectiveness, its limitations necessitate further architectural innovations, such as predicate bits and advanced thread management in GPU s. Our research introduces a dual-instruction multiple-data (DIMD) architecture to enhance PE parallelism efficiently for FPGA overlay architecture. U sing the Fast Fourier Transform (FFT) [4], [5] and AI convolution [6] as benchmarks, we demonstrate the proposed architecture's potential for effectively handling complex parallel computations.

Read the paper · More papers on PaperTik