Keynote2: Global-Scale FPGA-Accelerated Deep Learning Inference with Microsoft's Project Brainwave

2019

The computational challenges involved in accelerating deep learning have attracted the attention of computer architects across academia and industry. While many deep learning accelerators are currently theoretical or exist only in prototype form, Microsoft's Project Brainwave is in massive-scale production in data centers across the world. Project Brainwave runs on our Catapult networked FPGAs, which provide latency that is low enough to enable "Real-time Al" - deep learning inference that is fast enough for interactive services and achieves peak throughput at a batch size of 1. Project Brainwave powers the latest version of Microsoft's Bing search engine, which uses cutting-edge neural network models that are much larger than the neural networks used in typical benchmarks. In this talk, I'll discuss how Project Brainwave's FPGA-based and software components work together to accelerate both first-party workloads - like Bing search - and third-party applications using neural network models, like high energy physics and manufacturing quality control. I'll also talk about how FPGAs are the perfect platform for the fast-changing world of deep neural networks, since their reconfigurability allows us to update our accelerator in place to keep up with the state of the art.

Read the paper · More papers on PaperTik