An Improved Unified Scalable Radix-2 Montgomery Multiplier
David Money Harris, Ram Kumar Krishnamurthy, Mark A Anders, Sanu K. Mathew, Steven Hsu · 2005
This paper describes an improved version of the Tenca-Koc unified scalable radix-2 Montgomery multiplier with half the latency for small and moderate precision operands and half the queue memory requirement. Like the Tenca-Koc multiplier, this design is reconfigurable to accept any input precision in either GF(p) or GF(2/sup n/) up to the size of the on-chip memory. An FPGA implementation can perform 1024-bit modular exponentiation in 16 ms using 5598 4-input lookup tables, making it the fastest unified scalable design yet reported.