Implementation and Optimization of Double-Precision Floating-Point Exponential Functions on ARMv8 NEON Architecture
Rong Luo, Cheng Xu, Dan Hu, Xiaozhe Zhang, Тао Чен, Chunye Gong · 2023
The natural exponential function is an important basic arithmetic function with a wide application in many fields such as artificial intelligence. In order to satisfy the demands of high-performance computing on Armv8 architecture, we use Neon assembly instructions to implement and optimize the double-precision floating-point exponential functions. Firstly, we implemented sine, cosine and exponential function for real numbers to computing two arguments simultaneously with Neon assembly instructions based on efficient algorithms. Then the computations of sine and cosine are combined to improve the performance of exponential function for complex numbers. Finally, assembly optimization methods of instruction selection, loop unrolling and instructions re-ordering are implemented to efficiently compute for more arguments at once. The experimental results show that, on the premise of ensuring the accuracy, the exponential functions we implemented have excellent performance, with the fastest achieving a speed up of 2.29 to 2.90 and 1.35 to 1.67 over the corresponding functions in the GCC standard math library and the Arm Optimized-routines respectively.