Modified VLSI Architecture for Montgomery Modular Multiplication
S. Bhagyamma, A L Choodarathnakara, N D Vindya, Y T Swamy, H Meghana · Zenodo (CERN European Organization for Nuclear Research) · 2017
This paper proposes a simple and efficient Montgomery multiplication algorithm such that the low-cost and high-performance Montgomery modular multiplier can be implemented accordingly. The proposed multiplier receives and outputs the data with binary representation and uses only one-level carry-save adder (CSA) to avoid the carry propagation at each addition operation. This CSA is also used to perform operand pre computation and format conversion from the carry save format to the binary representation, leading to a short critical path delay at the expense of extra clock cycles for completing one modular multiplication. To overcome the weakness of extra clock cycle, a configurable CSA (CCSA), which could be one full-adder or two serial half-adders, is proposed to reduce the extra clock cycles for operand pre computation and format conversion by half. In addition, a mechanism that can detect and skip the unnecessary carry-save addition operations in the one-level CCSA architecture while maintaining the short critical path delay is developed. As a result, the extra clock cycles for operand pre computation and format conversion can be hidden and high throughput can be obtained. The experimental results show that the proposed Semi Carry Save Montgomery Modular Multiplier New (SCS-MM New) Architecture utilizes Delay of 2.5 ns, Clock cycle of 10, Area of 245 slices and achieves moderate Throughput of 387.78 Mbps.