Advances in Adaptable Computing

Amit Kumar Gupta · 2019

Recent technical challenges have forced the industry to explore options beyond the conventional "one size fits all" CPU scalar processing solution. Very large vector processing (DSP, GPU) solves some problems, but it runs into traditional scaling challenges due to inflexible, inefficient memory bandwidth usage. Traditional FPGA solutions provide programmable memory hierarchy, but the traditional hardware development flow has been a barrier to broad, high-volume adoption in application spaces like the Data Center market. Recent technical challenges in the semiconductor process prevent scaling of the traditional "one size fits all" CPU scalar compute engine. Changes in semiconductor process frequency scaling with the end of Dennard scaling have forced the standard computing elements to become increasingly multicore [Ref 1]. As a result, the semiconductor industry is exploring alternate domain-specific architectures, including ones previously relegated to specific extreme performance segments such as vector-based processing (DSPs, GPUs) and fully parallel programmable hardware (FPGAs). The question becomes: Which architecture is best for which task? Scalar processing elements (e.g., CPUs) are very efficient at complex algorithms with diverse decision trees and a broad set of libraries but are limited in performance scaling. Vector processing elements (e.g., DSPs, GPUs) are more efficient at a narrower set of parallelizable compute functions, but they experience latency and efficiency penalties because of inflexible memory hierarchy. Programmable logic (e.g., FPGAs) can be precisely customized to a particular compute function, which makes them best at latency-critical real-time applications (e.g., automotive driver-assist, ADAS) and irregular data structures (e.g., genomic sequencing), but algorithmic changes have traditionally taken hours to compile versus minutes.

Read the paper · More papers on PaperTik