Balanced loop retiming to effectively architect STT-RAM-based hybrid cache for VLIW processors
Keni Qiu, Weigong Zhang, Xiaoqiang Wu, Xiaoyan Zhu, Jing Wang, Yuanchao Xu, Chun Jason Xue · 2016
Loop retiming has been extensively studied to maximize instruction-level parallelism (ILP) of multiple function units by rearranging the dependence delays in a uniform loop. Recently loop retiming technique has been proposed to mitigate the migration overhead of STT-RAM-based hybrid cache by changing the interleaved read and write memory access pattern. However, the previous ILP-aware loop retiming is unaware of its impact on the hybrid cache's migration while the migration-aware loop retiming has not fully considered the parallelism of arithmetic and logical units (ALUs) in VLIW processors.