A 300MB SRAM, 20Tb/s Bandwidth Scalable Heterogenous 2.5D System Inferencing Simultaneous Streams Across 20 Chiplets with Workload-Dependent Configurations

Srivatsa Srinivasa, Dileep John Kurian, Paolo A. Aseron, Prerna Budhkar, Arun Radhakrishnan, Alejandro Cardenas Lopez, Jainaveen Sundaram, Vinayak Honkote, Leonid Azarenkov, Dan Lake, Jaykant B. Timbadiya, Mikhail Moiseev, Brando Perez Esparza, Ronald Kalim, Erika Ramirez Lozano, Mukesh Bhartiya, Sriram K. Muthukumar, Satish Yada, Suresh Kadavakollu, Saransh Chhabra · 2025

Disaggregating large systems has shown multifold advantages especially with current application trends prompting a shift towards chiplet-based architectures [1]–[3]. To meet increasing computing demands, 2.5D systems should allow greater interoperability across advanced technology nodes from multiple foundries, higher system memory capacity, higher 10 counts and scalable interconnect pitches. To further address the escalating memory capacity demands and mitigate the bandwidth bottleneck prevalent in AI applications, chiplet systems should be capable of workload-tailored configurations at assembly time. This adaptability enables optimal resource allocation and facilitates processing of voluminous datasets and complex AI computations.

Read the paper · More papers on PaperTik