The design of a latency constrained, power optimized NoC for a 4G SoC

Rudy Beraha, Isask’har Walter, Israel Cidon, Avinoam Kolodny · 2009

Network on-Chip (NoC) is being adopted by chip architects as a means to improve design productivity. As the number of modules connected to a bus increase, its physical implementation becomes very complex, and achieving the desired throughput and latency requires time consuming custom modifications. Conversely, NoCs are designed separately from the functional units of the system to handle all foreseen inter-module communication needs. Their inherent scalable architecture facilitates the integration of the system and shortens the time-to-market of complex products. In this work, we discuss and evaluate the design process of a NoC for a state-of-the-art system on-chip (SoC). More specifically, we describe our experience in designing a cost optimized NoC interconnect for a high-performance, power constrained 4G wireless modem. We focus on the power and performance aspects of various module mapping schemes, looking for a tradeoff that is characterized by a minimal power consumption that still meets the timing requirements of all targeted applications. Using a simulated annealing based mapping process, we place the system's modules on a grid, minimizing the dynamic energy consumed by the transmission of packets over the NoC. In many of the studies where network latency was used as a performance goal (either as the optimized cost function or as a constraint), the average delay of all packets over all communicating pairs was considered. However, in a practical SoC, different streams of communication may require different delays and therefore the overall average latency is an inappropriate measure. Consequently, the individual per-flow, point-to-point (source-destination) latencies should be accounted for to get better results. In this paper, we go further to suggest a third, improved approach: knowing the application that is to be used in the SoC, we utilize its functional timing requirements, which are defined by the application end-to-end latency constraints. Each of those end-to-end traversal delay requirements is composed of the cumulative requirement of a sequence (or a "chain") of point-to- point flows. For example, the application may require that a block of data which is generated by module A is sent to module B, and then to module C. By observing that the performance of the application is subject to the total time it would take the data to get from module A to module C, we can use this delay as the targeted performance measure, rather than specifying the two separate latency constraints (for the flow from module A to module B and from module B to module C). Since pair-wise delays may be traded, the timing constraints are relaxed and the optimization program can use more freedom in its operation. To the best of our knowledge, this paper is the first to discuss and quantify the benefits of specifying the end-to-end traversal requirements during the design process.In order to quantify and evaluate different alternatives, we report the actual throughput and timing requirements of the commercial SoC as well as the synthesis results. We evaluate three mapping schemes: a power optimized mapping; a power optimized mapping with point-to-point timing requirements; and a power optimized mapping with end-to-end timing requirements. For each of these mappings, we use simulations to find a uniform assignment of link capacities so that the run-time latency requirements of all flows are met. We then further optimize the network by tuning the capacity of links, reducing the bandwidth of links that operate faster than necessary. Synthesis results reveal that considering end-to-end requirements during the mapping phase of the design results in an improved implementation, even if only a limited number of discrete port configurations and link bandwidths are supported. According to our findings, the proposed mapping and link tuning techniques offer up to 40% savings in the total router area and a reduction of up to 49% in the inter-router wiring area. As part of this work, we present the bandwidth and timing requirement of the high-performance, state-of- the-art 4G application we examine. This information can be used by the NoC community as a benchmark for future research.

Read the paper · More papers on PaperTik