OISMA: On-the-Fly In-Memory Stochastic Multiplication Architecture for Approximate Matrix Multiplication
Shady O. Agwa, Yihan Pan, Georgios Papandroulidakis, Themis Prodromakis · IEEE Journal on Exploratory Solid-State Computational Devices and Circuits · 2026
Artificial Intelligence models are currently driven by a significant up-scaling of their complexity, with massive matrix-multiplication workloads representing the major computational bottleneck. In-memory computing architectures are proposed to avoid the Von Neumann bottleneck. However, both digital/binary-based and analogue in-memory computing architectures suffer from various limitations, which significantly degrade the performance and energy efficiency gains. This work proposes OISMA, an energy-efficient in-memory computing architecture that utilizes the computational simplicity of a quasi-stochastic computing domain (Bent-Pyramid system), while keeping the same efficiency, scalability, and productivity of digital memories. OISMA converts normal memory read operations into in-situ stochastic multiplication operations with a negligible cost. An accumulation periphery then accumulates the output multiplication bitstreams, achieving the matrix multiplication functionality. A 4KB 1T1R OISMA array was implemented using a commercial 180nm technology node and in-house RRAM technology. At 50 MHz, it achieves 0.789 TOPS/W and 3.98 GOPS/mm2for energy and area efficiency, respectively, occupying an effective computing area of 0.804241 mm2. Scaling OISMA to 22nm technology shows a significant improvement of two orders of magnitude in energy efficiency and one order of magnitude in area efficiency, compared to dense matrix multiplication in-memory computing architectures.