Pipelined Algorithm and Modular Architecture for Matrix Transposition
Yunxiang Wang, Zhenguo Ma, Feng Nan Yu · IEEE Transactions on Circuits & Systems II Express Briefs · 2018
This brief presents a novel pipelined algorithm for transposing an N × N matrix, as well as a modular architecture for this algorithm. The architecture is optimal in both memory, using the minimum number of registers necessary to transpose an N × N matrix, and latency, achieving the theoretical minimum. The architecture is composed of a series of identical cascaded basic circuits and has a simple control strategy. Furthermore, the algorithm and architecture can be easily extended to p-parallel where p is any factor of N.