for High-Radix Crossbar Schedulers
Manolis G. H. Katevenis, Dionisios N. Pnevmatikatos · 2011
We study the scaling of parallel-matchin g crossbar sched ulers to radices above 100. First, we examine a traditional microarchitectur e that implements the matching decision of each input and each output of the crossbar in a sepa rate arbiter block and communicates the matching decisions between the input and the output arbiters through global point-to-point links. Using simple models and experimenta tion with 90nm CMOS layouts, we show that this architec ture is expensive because the global point-to-point links take up O(N4) area, where N the radix of the crossbar. Next, by observing that the wiring of an arbiter fits in a mini mal O(NlogN) area, we propose a novel microarchitectur e that inverts the locality of wires by orthogonally interleav ing the input with the output arbiters, thus lowering the wiring area of the scheduler down to O(N2log 2 N). Using this architecture, the scheduler for a radix-128 FIFO, VOQ, or 2-VC crossbar becomes gate limited, fitting in 3.6, 7.2, and 7.2mm 2 respectively, which is a 40, 50, and 70% im provement compared to the traditional. Moreover, the pro posed schedulers find a new match in less than IOns, thus allowing a minimum packet below 30Bytes at 24Gb/s line rate. Based on these findings, we conclude that crossbar schedulers are feasible even for radices above 100.