A hybrid CPU/GPU Scheme for Optimizing ChaCha20 Stream Cipher
Ziheng Wang, Heng Chen, Weiling Cai · 2021
The secure transmission of large-scale data has attracted more and more attention. In the widely recognized security protocol TLSv1.3, the only algorithms that support large-scale data en-/decryption are ChaCha20 and the Advanced Encryption Standard (AES). Although AES has a higher usage rate, ChaCha20 still has the advantage of speed and security on many platforms, and has a better performance against post-quantum attacks. However, for a CPU/GPU platform, compared to the AES algorithm, no work has fully described the application scheme of ChaCha. This paper proposes an optimization scheme to optimize the performance of the ChaCha20 algorithm on a CPU/GPU platform. On a CPU platform, we provide a parallelization implementation that is better than that of OpenSSL. On a single GPU, our implementation of ChaCha20 achieves peak throughput of 211.41GB/s, which is better than any previous implementation of ChaCha20 and AES algorithms on GPU. More importantly, we are the first to detail the optimization of ChaCha on GPU. When considering the interconnection between CPU and GPU, we use the 87.76% peak bidirectional bandwidth of a PCIe channel. Finally, we also provide a scheme for the application of ChaCha20 on a CPU/GPU platform.