Processor element architecture for nonshared memory parallel computers
Thomas J. Holman, Lawrence Snyder · 1988
The architecture of the processor elements in a parallel computer can have a significant impact on overall performance. However, because the hardware used for architectural features can alternatively be used to add more processor elements, and hence increase parallelism, adding architectural features is not necessarily the best choice. In this dissertation, a combination of analytical techniques, program analysis, and simulations is used to explore this trade-off between processor element architecture and added parallelism in a non-shared memory parallel computer. The impact of an architectural feature cannot, in general, be assessed by comparing program execution times. This is because execution time can be changed independently of architecture by changing either the number of processor elements or the total data memory. We argue that a fair comparison based on execution time can only be performed when total hardware and total memory remain constant. These constraints are used to explore the relationship between processor element architecture and overall performance. The key insight gained from this analysis is that features with the greatest potential for improving performance are those that affect all or most of the operations. Consequently, a substantial hardware investment for a specific class of operations, such as floating point, does not give the best performance. The above observations are substantiated by the analysis of several specific architectural features. This study utilizes dynamic frequency and cost distributions of seven parallel programs representative of numerical computations and physical simulations. Performance estimates for these programs are combined with area estimates for the features to obtain the following results: (1) With a given amount of hardware, a 32-bit data path outperforms a serial path by a factor of five. (2) The MIMD architecture studied requires 3.4 times more hardware than an equivalent SIMD architecture. (3) Though code size increases, a reduced instruction set gives better performance. (4) Small instruction caches show promise; a data cache does not. (5) Partial support for floating point operations is more beneficial than a full floating point unit. (6) DMA for communication channels does not improve performance. (7) A few additional registers can significantly reduce the overhead of providing virtual machine facilities.