Design of Concurrent Error-Detecting Systolic Amy s
James L. Olivier, F. Özgüner · 1989
This paper addresses the problem of detection and iden- tification of a faulty processing element in a systolic array. A method for designing processing elements with concurrent error detection is presented. The I RAN(,,, code ( 11 is shown to be an effective code for encoding the operands in a systolic array. It is shown that the 1 g3N I,+, code is equivalent to a residue code with the check and information bits interchanged, for odd number of information bits. This allows arithmetic to be performed separately on the information and check bits while the output can be checked by an ANchecker. An architecture and rules for designing a self-checking processing element (PE) for sys- tolic arrays are presented. Both redundancy and extra delay of the self- checking PE are shown to be low. This paper first investigates the effectiveness of differ- ent classes of arithmetic codes for concurrent error detec- tion in systolic arrays by analyzing the implications of these codes on the size of the multiplier circuits. The 1 @NIM code ( 13 is shown to be an effective code for en- coding the operands in a systolic array. Here, (gANIM denotes gAN modulo M. In Section 111, multiplication is defined for the I gAN code. It is shown that the 1 g3N IM code is equivalent to a residue code with the check and information bits interchanged, for an odd number of in- formation bits. When the number of information bits is even, the check bits can easily be obtained from the res- idue of the information bits. Using this interesting prop- erty, an architecture and rules for designing a self-check- ing processing element (PE) for systolic arrays are presented in Section V. A parallel multiplier array design that performs modulo 2k - 1 multiplication is also devel- oped. Pepnnance of the self-checking PE is evaluated based on a gate level design. It is shown that the redun- dancy is very low, especially for larger wordlengths. The delay is due to the checkers only since the circuits per- forming arithmetic on the check bits operate in parallel with the circuits performing arithmetic on the information bits. The single stuck-at fault model is used in the dis- cussion of hardware faults. The self-checking design pre- sented detects both permanent and transient faults. The class of systolic arrays where each PE performs the function ci + = ci + ab is addressed. An example of such a systolic array for matrix multiplication is shown in Fig. 1.