Implementation of a parallel extended Kalman filter using a bit-serial silicon compiler

Phillip L. Shaffer · Fall joint computer conference · 1987

Two different architectures to compute the extended Kalman filter have been designed using a bit-serial silicon compiler. For both systems, the Kalman filter is to be used for target tracking, with 9 states, 3 observations, and 3 process noise sources. The first system is a systolic-type architecture. Data matrices are presented by columns and data vectors presented by element, and recursive computations are used to triangularize matrices. The feedback necessary to use recursion limits the throughput of this architecture. For this application, with a 20 MHz clock rate, the throughput is 1736 samples/sec, and the latency (delay from data availability to production of results) is 0.859 msec. This circuit requires about 51 chips of 13 different types.The second system is a pipelined synchronous dataflow architecture. All data values are presented simultaneously, and recursion is not used. Because most of the matrices involved are triangular, block triangular, band, or sparse, the number of elements dealt with is manageable. The throughput of this system is simply the clock rate divided by the word length, or 625000 samples/sec. The circuit requires about 112 chips of 10 different types, and has a latency of 0.842 msec. This system could thus track 526 different targets at 1188 samples/sec each. Thus, by using about twice as many chips and avoiding recursion, the throughput was increased by 360 times.

Read the paper · More papers on PaperTik