D7.2.1 A Report on the Survey of HPC Tools and Techniques

Michael Lysaght, Bjørn Lindi, Vít Vondrák, John Donners, Marc Tajchman · Zenodo (CERN European Organization for Nuclear Research) · 2013

he objective of PRACE-3IP Work Package 7 (WP7) ‘Application Enabling and Support’ is to provide applications enabling support for HPC applications codes which are important for European researchers to ensure that these applications can effectively exploit multi-petaflop systems. This applications enabling activity will use the most promising tools, algorithms and standards for optimisation and parallel scaling that have recently been developed through research and experience in PRACE and other projects. This deliverable contains a comprehensive survey of the research activity undertaken within PRACE to date so as to better understand what HPC tools and techniques have been developed that could be successfully applied to help other applications within WP7 effectively exploit multi-petaflop systems and so see how the various applications communities have progressed over the last few years within PRACE. As well as surveying the tools and techniques that have been employed in PRACE, this deliverable also reports on how tools and techniques are being used in various exascale projects and initiatives outside of PRACE. This perspective is presented so as to inspire new “forward looking” approaches to enable European applications on the road to exascale computing. The survey covers four separate topics that we consider relevant to enable applications on current multi-petascale systems. We summarize our findings separately by topic: Programming Interfaces and Standards, Debuggers and Profilers, Scalable Libraries and Algorithms and I/O Management Techniques. Programming Interfaces and Standards As part of this report we have surveyed thirteen individual programming languages and standards and report on how they have been used in PRACE to date. While we have found that, unsurprisingly, the MPI model still dominates within PRACE, evidence suggests that the most recent version of the standard provides features that are starting to confront the challenges of exascale computing and which have not yet been exploited within PRACE to any considerable extent. As a result, we recommend that the latest features of the standard be exploited during the enablement of applications on multi-petascale systems in WP7. We have also assessed the wide range of programming models for exploiting heterogeneous architectures and conclude that the entry of new competitors to the many-core space has increased the relevance of open standards on the road to exascale (where many-core typically implies > 50 cores). Indeed, even for GPUs we have found considerable evidence that an open standards approach, of which OpenACC represents the strongest offering to date, is becoming more popular both within and outside PRACE and should be considered for enabling applications within WP7. In terms of more novel approaches to exploiting multi-petascale systems, we have drawn rich information from the European exascale projects as well as work being carried out in the US, which includes interesting findings on novel extensions to OpenMP and the use of Partitioned Global Address Space (PGAS) languages in real applications, which should inspire WP7. Debuggers and Profilers As part of this report, we have surveyed fourteen debugging and profiling tools. We have found that all of the European exascale projects are concentrating effort into tools for debugging and performance analyses. This is deemed a necessity for efficient use of multi-petascale and future exascale systems: If we are to enable applications on such systems, then we need to have as clear a view as possible of the barriers to achieving performance. In some respect, we feel that the European exascale project, DEEP, provides a model for how the profiling tools should enable applications in WP7. It is worth noting that this type of rigorous assessment of debugging and profiling tools has rarely been seen in PRACE reports or whitepapers to date. A substantial effort of training on tools for debugging and performance analysis has been carried out within PRACE. However, very little is documented on how successfully these tools have been employed within enabling projects. One of our missions within WP7 is to fix this discrepancy and to work more closely with tool developers to understand the full benefits and limits of such tools in extreme cases. Scalable Libraries and Algorithms As part of this report we have surveyed a representative collection of libraries and techniques that currently garner much interest both within and outside PRACE. As a consequence of the move towards large multi-petascale heterogeneous systems, there is an increasing demand for new and improved scalable, efficient, and reliable numerical algorithms and libraries that confront existing and upcoming complexities associated with such systems, including complex memory hierarchies, the overhead of data movement and fault tolerance. In particular, we have surveyed very interesting exploratory work that has recently been carried out in WP12 PRACE-2IP on libraries and algorithms, which we feel should be exploited further on real applications within WP7. As well as surveying research within PRACE, we have also looked to the work being carried out in European exascale projects and further afield to find out more about how such projects are tackling the challenges confronting libraries and algorithms at extreme scales. I/O Management Techniques As part of this report we have surveyed five I/O management techniques. The increasing data needs of scientific and engineering applications mean that the problems associated with reading, writing, analysing, storing and sharing large amounts of data are becoming more relevant to a wider user community within PRACE. While the performance gap between file systems and compute systems is well known, during our surveying we have found that users within PRACE have in general not been able to squeeze as much performance from existing parallel file systems as they have from computational hardware, particularly for the case of high-level I/O libraries. Deeper investigations into extracting performance from parallel file systems (with WP7 applications) will be the main focus of our enablement work within WP7.

Read the paper · More papers on PaperTik