Programmable Architecture for Thread Level Parallel Computing
Zhenfu Feng, Huan Zhang, Lidong Xing, Hao Chang, Ang Li · 2024
A power-efficient Single Instruction Multiple Threads (SIMT) processor is proposed to address the increasing demands of high-performance computing in graphics rendering media processing, artificial intelligence etc... The processor supports thread-level parallelism by managing multiple threads into WARPS. The specific work includes the following aspects: 1) merge a common five-stage pipeline structure with an emission stage to form a low latency memory access structure; 2) an independent RISC instruction set; 3) two-level WARP scheduling algorithm for highly efficient thread management; 4) functional validation and performance evaluation were conducted on the overall implementation. The results indicate that the processor achieves optimal performance when the number of PEs equals to the matrix dimension on GEMM task. Furthermore, the architecture shows super linear speedup capability when the number of PE cores is less than 32. The performance grows continuously as the number of PE cores, but the decreased growth rate is observed.