D12.5: Summary of Novel Programming Techniques Results

Carlos Sanchez, Jose, Christian Pérez, Cevdet Aykanat, R. Oguz Selvitopi, Alberto Miranda, Ramón Nou, Toni Cortés, Eric Boyer · Zenodo (CERN European Organization for Nuclear Research) · 2014

This Work Package performed research and development on the programmability of future multi-petascale and exascale systems. In particular, it was focused on main four areas, auto-tuned runtime environments, scalable numerical algorithms, fault tolerant tools, and file system optimization for exascale systems. A total of 16 research projects are reporting in this document covering multiple different techniques. On auto-tuned runtimes environments a total of six projects were exploring different auto-tuning techniques at different granularities. For example, it goes from a coarse technique that auto-tunes the job scheduler down to the finest granularity when auto-tuning Intel vector instructions. Specifically, on SLURM job scheduler we explored the case capability to do topologically aware mappings of jobs on hierarchically interconnected systems. On the other hand, three projects were focused mostly on the programming language exploring one these projects the case of controlling dynamically the number of threads allocated to OpenMP parallel regions to decrease the overall wall-times; and the rest were focusing on improving of workflow executions through meta-model methods and the other using a component based approach to improve the 3D FFT. Additionally, on project was focused on compilation utilities to improve the vectorization of their codes. And finally, the last project was focused on improving the collective communications operations, specifically the MPI_All_to_All. On scalable numerical algorithms, it was explored new algorithms or techniques that improve the scalability of existing algorithms. Eight projects were also focused on this research line. with several diverse parallel techniques such as utilization of GPU and MIC accelerators, message-passing paradigm and shared memory constructs. There are algorithmic approaches to increase the efficiency of widely used numerical algorithms (such as Conjugate Gradient (CG)) as well as adaptive parameter determination and utilization. The target applications and libraries include HYDRO, PETSc, CP2K and FETI. On the other hand, on fault tolerant tools, it was shown the performance of the fault tolerant tool called FTI based on application-based checkpoint/restart. The overhead of using this tool could be very low at large scale. It was reporting less than 6% on HYDRO at 9,600 cores. And finally, on file system optimization it was showed the performance of user hint guided I/O prefetching which is seen a scalable technique for accessing I/O at exascale systems. In fact, it was shown that anticipating user reads, with small information from the user, can produce great benefits in terms of performance. In addition, it was shown than the memory usage could be reduced with the use of an advanced filtering technique.

Read the paper · More papers on PaperTik