Ph.D. Project AIM: Accelerating Arbitrary-Precision Integer Multiplication on Heterogeneous Reconfigurable Computing Platform Versal ACAP
Zhuoping Yang, Peipei Zhou · 2025
Arbitrary-precision integer multiplication serves as the core kernel in many applications such as cryptographic algorithms, scientific computing, and etc. To compute arbitrary-precision integer multiplication using low-bit function units (32/64-bit) on existing hardware, decomposition methods like Karatsuba and Schoolbook are usually adopted. In general, the decomposition methods use two steps to finish the calculation. First, it decomposes the two large integers into many smaller integers and generates a group of low-bit multiplications that can be calculated in a spatial or sequential manner. Second, the results of low-bit multiplications are shifted and added together to get the final result. The first step involves massive parallel byte-level processing, while the second step requires a long propagation chain, which involves bit-level processing. Prior works have leveraged vector instructions on CPUs, CUDA cores on GPUs, and DSPs on FPGAs to accelerate arbitrary-precision multiplication. We use the state-of-the-art FPGA accelerator and libraries on GPUs and CPUs, and find that the FPGA has the lowest energy efficiency. We identify that the dedicated vector units on CPUs and GPUs bring the biggest energy efficiency in the first computation step. Although DSPs and LUTs on FPGAs introduce extra energy overhead in the first step compared to dedicated vector units but are more suitable for the second computation step. To benefit both two steps, we propose the AIM framework to generate efficient arbitrary-precision integer multiplication accelerator on AMD Versal adaptive compute acceleration platform (ACAP) VCK190, which comprises 400 AI Engine (AIE) ASIC processors, an FPGA, and a ARM CPU. AIM uses 400 AIEs to compute the first step and the FPGA to process the second step. Our experimental results show that AIM achieves up to 12.6x and 2.1x energy efficiency gain over the Intel Xeon Ice Lake 6346 CPU, and Nvidia A5000 GPU respectively with the respect to the multiplication kernel. We open-sourced AIM on GitHub: https://github.com/arc-research-lab/AIM. Also, we use three different applications including large integer multiplication, RSA, Mandelbrot to demonstrate the usability of AIM.