Optimizing Fast Trigonometric Functions on Modern CPUs

Jie Shen, Biao Long, Chun Huang · 2022

Traditional math libraries in high performance computing (HPC) are designed with accuracy as the first priority. With the development of modern hardware processors and the expansion of HPC application domains, it is highly desirable to develop fast, approximate math function implementations for performance-hungry and error-tolerable applications. In this paper, we propose an acceleration method for trigonometric functions (sine and cosine) based on specialized hardware instructions. We implement vector versions of the math functions which utilize single instruction multiple data (SIMD) vectorization to enable data-level parallelism. We apply Estrin's scheme to compute the polynomial approximation and customize the hardware instruction with subtle changes to the original one to maximize instruction-level parallelism. We also identify key performance impact factors for vector trigonometric functions through a thorough empirical evaluation. Experimental results show that our proposed implementation outperforms the state-of-the-art fast implementation by up to$\mathbf{1}.\mathbf{35}\times$without visible loss of accuracy.

Read the paper · More papers on PaperTik