COnfigurable Network Protocol Accelerator (COPA): An Integrated Networking/Accelerator Hardware/Software Framework

Venkata Krishnan, Olivier Serres, Michael A Blocksome · 2020

Modern FPGAs are more than just “gate arrays” - they offer hardware acceleration capabilities and advanced features that include High Bandwidth Memory (HBM), Cache/Memory coherent/IO interconnect (CXL/PCIe), high speed transceivers, and embedded processors. Unfortunately, lack of a standardized hardware/software infrastructure has relegated FPGAs to being second-class citizens under the control of a host platform rather than being compute/network accelerator nodes with SmartNIC capabilities. This stands in the way of broader deployment of FPGAs in a wide variety of distributed/heterogenous platforms.Intel's COnfigurable Network Protocol Accelerator (COPA) addresses this challenge and provides a customizable framework that integrates communication with computation on an FPGA platform. The FPGA incorporates SmartNIC capabilities and functions as an “autonomous” node that attaches to the network. In this paper, we provide an overview of COPA architecture and the acceleration modes that it supports. The hardware component provides the necessary networking/accelerator infrastructure while the software component of COPA abstracts the underlying FPGA from the application or middleware. The API is based on an open standard network API (OFI) that has been extended for COPA to expose the various acceleration modes to software. Given the lack of a standard API for SmartNICs, the extended OFI can also potentially serve that purpose and efforts are underway to upstream the changes.COPA has been implemented on different variants of Stratix10 FPGAs. Multiple COPA FPGAs can autonomously connect to a standard 100GigE switching network with the ability to mix-and-match FPGA variants to this network. Both inline and lookaside accelerator functions are also supported. The framework has been validated by microbenchmarks as well as proxy benchmarks that mimic the behavior of client/server flows - and achieves bandwidths close to 100Gbps when acceleration is enabled.

Read the paper · More papers on PaperTik