Design and simulation of an MIMD shared memory multiprocessor with interleaved instruction streams
Thomas R. Stiemerling · ERA · 1991
The design of the Eppi MIMD shared memory multiprocessor is described, and its performance evaluated by simulation.The Eppi has a dancehall architecture with p instruction interleaved RISC processors connected to p shared memories by a packet switched, combining, indirect binary n-cube multistage network composed of P 2 10 92 p 2 x 2 crossbar switches.There is no processor cache or local memory, and no paged virtual memory.Memory addresses are low order interleaved across the memories.The fetch-and-add instruction is used for inter-process synchronisation, and the switches support the combining of load and fetch-and-add memory requests.Simulation results of a single Eppi processor with varying interleaving level and instruction mix are presented, and of an isolated network with varying queue size and network load.A distributed time-driven, instruction level simulator of the Eppi design has been implemented in Occam, and runs on a transputer based, distributed memory multiprocessor.Three parallel benchmark programs: matrix multiply, bitonic merge sort and Moore shortest path, have been written in the processor assembly language, and are used as workloads in the simulations.The programs use the fetch-and-add instruction to implement process control primitives.A number of simulation experiments have been carried out using the Eppi simulator which investigate the effect on performance of increasing the system size (speed-up), varying the switch queue and wait-buffer size, increasing the combining level, increasing the interleaving level, and varying the network and memory speed relative to the processor.These experiments are repeated for each benchmark program, and detailed execution statistics are presented for each simulation.A dynamic execution profile for each benchmark program is also presented.