Portable high-performance supercomputing: high-level platform-dependent optimization
Eric Brewer, William E. Weihl · 1994
Although there is some amount of portability across today’s supercomputers, current systems cannot adapt to the wide variance in basic costs, such as communication overhead, bandwidth and synchronization. Such costs vary by orders of magnitude from platforms like the Alewife multiprocessor to networks of workstations connected via an ATM network. The huge range of costs implies that for many applications, no single algorithm or data layout is optimal across all platforms. The goal of this work is to provide high-level scientific libraries that provide portability with near optimal performance across the full range of scalable parallel computers. Towards this end, we have built a prototype high-level library “compiler” that automatically selects and optimizes the best implementation for a library among a predefined set of parameterized implementations. The selection and optimization are based on simple models that are statistically fitted to profiling data for the target platform. These models encapsulate platform performance in a compact form, and are thus useful by themselves. The library designer provides the implementations and the structure of the models, but not the coefficients. Model calibration is automated and occurs only when the environment changes (such as for a new platform). We look at applications on four platforms with varying costs: the CM-5 and three simulated platforms. We use PROTEUS to simulate Alewife, Paragon, and a network of workstations connected via an ATM network. For a PDE application with more than 40,000 runs, the model-based selection correctly picks the best data layout more than 99% of the time on each platform. For a parallel sorting library, it achieves similar results, correctly selecting among sample sort and several versions of radix sort more than 99% of the time on all platforms. When it picks a suboptimal choice, the average penalty for the error is only about 2%. The benefit of the correct choice is often a factor of two, and in general can be an order of magnitude. The system can also determine the optimal value for implementation parameters, even though these values depend on the platform. In particular, for the stencil library, we show that the models can predict the optimal number of extra gridpoints to allocate to boundary processors, which eliminates the load imbalance due to their reduced communication. For radix sort, the optimizer reliably determines the best radix for the target platform and workload. Because the instantiation of the libraries is completely automatic, end users get portability with near optimal performance on each platform: that is, they get the best implementation with the best parameter settings for their target platform and workload. By automatically capturing the performance of the underlying system, we ensure that the selection and optimization decisions are robust across platforms and over time. Finally, we also present a high-performance communication layer, called Strata, that forms the implementation base for the libraries. Strata exploits several novel techniques that increase the performance and predictability of communication: these techniques achieve the full bandwidth of the CM-5 and can improve application performance by up to a factor of four. Strata also supports split-phase synchronization operations and provides substantial support for debugging.