The Virtuous Cycles of Determinism: Programming Groq's Tensor Streaming Processor
Satnam Singh · 2022
FPGAs and other 2D and 3D spatial computing fabrics share several common characteristics e.g. a deterministic computing model with distributed memories, but also differ along important dimensions e.g. granularity and communication infrastructure. This talk will position Groq's Tensor Streaming Processing (TSP) chip relative to other spatial computing fabrics and highlight common aspects of programming models for such spatial computing engines as well as highlighting some of the unique characteristics of the TSP architecture and its programming model One common characteristic this talk will focus on is \em determinism and specifically the use of static scheduling to know at compile time exactly how many cycles a program will take to execute regardless of the specific values of the input data to the TSP chip. Groq's TSP architecture is based on a parallel collection of coarse grain processing units including matrix multiplication or convolution (MXM), vector operations (VXM), arbitrary vector transformations (SXM) and various numerical operations. Data originates from a memory reads (MEM) and is then chained through a tensor pipeline through one or more computing blocks and terminated by a memory write (MEM). The chip can operate on in-flight tensors as they are produced and consumed, aggressively exploiting data-flow locality. A variety of deterministic computing models will be surveyed as we seek abstractions that support rapid multiplexing between multiple workloads which can be co-resident on TSP chips, with scheduling and coordination facilitated by determinism.