Freeflow Core: Enhancing Performance of In-Order Cores with Energy Efficiency

Raj Kumar Choudhary, Newton Singh, Harideep Nair, Rishabh Rawat, Virendra Pal Singh · 2019

The Out-of-Order (OoO) superscalar core design has been widely adopted for high performance computing. It exploits both instruction level parallelism (ILP) and memory level parallelism (MLP) to speedup the program's execution. However, due to unresolved data dependencies among instructions, exploiting ILP becomes at times difficult and make the OoO core idle, thus reducing its energy-efficiency. Our study focuses on selective exploitation of inherent ILP present in the program. Our proposed architecture, Freeflow Core (FC), focuses on discovering the selective opportunities for conversion of instruction guided execution to data guided execution. Such selective mechanism improves the performance without incurring substantial energy budget overheads. Memory-related instructions are handled in FC by memory-access pipeline and compute instructions are handled by compute pipeline. Giving priority to load/store instructions has been one of the known techniques to improve performance. However, less importance has been given to non-ready instructions which may block the head of the compute pipeline. We observe that such instructions are mainly those whose producers are unresolved in the memory-access pipeline. Hence, FC detects instructions that are dependent on unresolved memory instructions and guides them to a dedicated independent execution path. This segregation enables the younger ready instructions to free flow through the compute pipeline. Our evaluations show that FC outperforms InO and state-of-the-art Load Slice Core (LSC) both in performance and energy efficiency metrics.

Read the paper · More papers on PaperTik