Special Issue of the Cray User Group (CUG 2018)

Abhinav Thota, Yun He · Concurrency and Computation Practice and Experience · 2019

The Cray User Group (CUG) is an independent, international corporation of member organizations that own Cray Inc computer systems.1 The CUG holds an annual conference every year, and this year's conference, CUG 2018, the 61st meeting of the group, was held in Stockholm, Sweden in May 2018. There were 215 attendees from 54 participating sites. The conference received a total of 76 paper submissions, out of which 58 were accepted for inclusion in the conference proceedings. Out of the 58, 13 of the top-rated papers were selected for inclusion in the Concurrency and Computation: Practice and Experience Special Issue of the CUG proceedings. The papers in this Special Issue cover a diverse set of topics, including deep learning, data analytics, containers, and Advanced RISC Machine (ARM) processor performance characteristics. As always, there are papers touching on the latest developments in Cray architectures, operations, and application scaling and performance studies. The papers about containers and ARM performance characteristics both were highly attended at the conference, which speaks to the continuing interest and progress the supercomputing community is making in the area of containers, whereas ARMs are an emerging trend in the community that people are closely watching. The papers accepted at CUG 2018 covered a wide ranging collection of topics. A few of the themes that highlight the emerging trends in supercomputing are discussed below. Traditional topics at the CUG conferences include information exchange among vendors and Cray site staff and users on system management, software, and applications in the format of technical talks, Birds of a Feather sessions, tutorials, and panels. It is not an exception this year. A subset of papers in this Special Issue talks about the advances made in the Cray architecture, operations, and management,2-4 a very relevant topic to a large majority of sites that operate Cray supercomputers. Another subset of the papers touched on the latest developments and techniques in performance analysis and benchmarking in supercomputing.5-10 Deep learning and data analytics continue to be areas of strong interest in the supercomputing domain. There have been continuing efforts to mold the supercomputing environment to make it suitable for deep learning and data analytics frameworks. Multiple papers investigate new approaches to scaling TensorFlow and other frameworks on Cray supercomputers.11, 12 The “Alchemist” interface described in the work of Rothauge et al13 bridges the gap between Apache Spark and Message Passing Interface and tries to improve the performance and scalability of data analytics applications on supercomputers. Containers are used very differently in the supercomputing field compared to how they are used in enterprise computing. In many instances, containers are used to extend the capabilities and expand the number of applications that are supported on supercomputers, in addition to giving users the ability to bring their own containers from Docker Hub and other public repositories. There is only one paper in this Special Issue that deals with containers on supercomputers, but this was a topic of great interest at the CUG conference this year and the last couple of years. Sparks14 does a review of the state of the art of container tools and technologies as relevant for the supercomputing domain. ARM processors have traditionally, in recent years, been used in mobile devices. This processor architecture is known for its energy efficiency and memory bandwidth capabilities. The ARM processors introduce a new architecture to the supercomputing landscape, which has traditionally been dominated by x86 processors with optional accelerators such as graphics processing units. However, given the cost per FLOP ratio that ARM processors offer, they are very compelling. McIntosh-Smith et al7 presented the performance results of what is likely the first Cray supercomputer based on ARM central processing units. This machine is located at the GW4 Alliance in the United Kingdom and is named “Isambard.” The Cray XC50 “Scout” form factor–based machine features Cavium ThunderX2 ARM 64-bit central processing units and the Aries interconnect. The initial performance benchmarking results and ease of use with respect to software and library support seem promising. We saw that some of the big trends such as deep learning efforts in the supercomputing domain and containers continued this year at the CUG, whereas some of them have emphatically ended, such as Intel's Xeon Phi coprocessors. An ARM-based supercomputer was presented to the attendees this year, and we will find out if they become a trend next year. In addition to articles shedding light on new trends and developments, a robust set of articles in the areas of system management and application performance rounds up this year's Concurrency and Computation: Practice and Experience journal's Special Issue of selected proceedings from the CUG 2018 conference. Abhinav Thota is the manager of the Scientific Applications team in the Research Technologies center at the Pervasive Technology Institute, Indiana University. The Scientific Applications team helps users efficiently use the supercomputing resources at Indiana University. Abhinav has a master's degree in Systems Science from Louisiana State University. Dr Yun (Helen) He is a High Performance Computing Consultant at the National Energy Research Scientific Computing Center, Lawrence Berkeley National Laboratory. She specializes in the software programming environment, parallel programming models, scientific application porting and benchmarking, and climate models. Helen is the Program Chair for CUG 2018. We would like to thank all of the authors and all of the CUG participants who provided valuable contributions to this Special Issue. We would also like to thank the members of the CUG Review Committee for the feedback provided to the editors and authors, which was essential to selecting papers for publication and further improved many of the submissions. Finally, we would like to thank Professors Geoffrey Fox and David Walker, the editors, for providing us with this opportunity to present our works in the international journal of Concurrency and Computation: Practice and Experience.

Read the paper · More papers on PaperTik