An Effective Architecture for Trace-Driven Emulation of Networks-on-Chip on FPGAs
Thiem Van Chu, Kenji Kise · 2018
Modern many-core systems use Networks-on-Chip (NoCs) to move data around their cores. As the number of cores increases, the overall performance becomes highly sensitive to the NoC performance. Research and development of NoCs thus play a key role in designing future systems with hundreds to thousands of cores. However, current methodologies for evaluating NoCs are not scalable with respect to the system complexity. Conventional software simulators are too slow for evaluating middle-and large-scale NoCs. Recent FPGA-based emulators provide promising emulation speedups over software simulators. However, emulating large-scale NoCs with hundreds to thousands of nodes on FPGAs is a challenging problem because of the FPGA capacity constraints. Moreover, supporting trace-driven emulation is not trivial because trace data must be stored outside of the FPGA (usually in off-chip DRAM). Most of the existing FPGA-based NoC emulators rely on soft processors like Microblaze or hard processors on SoC FPGAs for loading trace data from the off-chip memory, generating messages, injecting the messages to the target NoC, manipulating the emulation, and making sure that there is no timing error. This approach makes the implementation easy but drastically degrades the emulation speed. This paper proposes an effective architecture for trace-driven emulation of NoCs on FPGAs. We present methods to scale to large NoCs and effectively hide the off-chip memory access latency. Our evaluation results show that (1) the proposal achieves a speedup of 260x compared to BookSim, one of the most widely used NoC simulators, when emulating an 8x8 NoC with trace data collected from full-system simulation of the PARSEC benchmark suite; and (2) the speedup is increased to three orders of magnitude when emulating a 64x64 NoC.