Foreword to the Special Issue on Processors, Interconnects, Storage, and Caches for Exascale Systems
Manuel E. Acacio, Julio Sahuquillo · Concurrency and Computation Practice and Experience · 2019
Exascale computing constitutes nowadays a significant challenge both for the academia and the industry. Although traditional computer systems continue to make important advances, achieving exascale computing requires mass customization. With this aim, several ongoing research projects are focusing on different architectural (computing boards or nodes, interconnects, storage, etc) issues of future exascale systems. Most of them devise heterogeneous computing boards consisting of CPUs (high performance and/or low power), FPGAs, GPUs, etc, sharing a common memory hierarchy. In this context, efficient intra- and inter-board within the same rack and inter-rack interconnect with the memory hierarchies are required. Also, performance and reliability design constraints for exascale storage systems rise significant challenges for HPC system designers. High performance I/O must be also faced because storing and retrieving such large amounts of data can greatly affect the overall performance of applications. Finally, it is important to characterize the demands that exascale applications exert on the different components of an exascale system. The goal of this special issue is to promote research on all of these aspects related to exascale computing. Six papers that address several components of exascale systems were carefully selected from open submissions. The six manuscripts included in this special issue cover different aspects of an exascale system. Lant et al1 present a network interface architecture and networking infrastructure, designed to sit inside the FPGA fabric of a cutting-edge heterogeneous MPSoC (Multi-Processor System-on-Chip), enabling networks of these devices to communicate within both a distributed and shared memory context, with reduced need for costly software networking system calls. This work presents in detail the factors that influenced the implementation and system prototype-based upon the use of Xilinx Zynq Ultrascale+ and discusses the main design decisions and implementation challenges. Crespo et al2 emphasize the need of interconnect technologies alternative to the classical electrical one. This work focuses on silicon photonics and highlights practical challenges that must be met to enable the adoption of this technology in building efficient, extreme-scale interconnection networks. In particular, they show that signal loss sources, suffered mainly due to waveguide crossings and propagation, play a critical role in photonic exascale network designs as they constrain the ability to perform data transmission in an effective as network size increases. Also focused on the interconnection network of an exascale system, Duro et al3 conduct an extensive simulation study using realistic photonic network configurations with synthetic and realistic traffic and show that, compared to electrical networks, optical networks can reduce the execution time of the studied real workloads in almost one order of magnitude. The study is performed from an architectural perspective, and the authors state that the photonic configuration highly impacts on the network performance, being the bandwidth per channel and the message length the most important parameters. Piernas and González-Férez4 address the scalability of file systems aimed to exascale systems. In particular, they describe how they have implemented the support for data objects in their previously proposed Fusion Parallel File System (FPFS). They show that the utilization of a unified data and metadata server (an enhanced object-based storage device or OSD+) provides FPFS with a competitive advantage over other file systems like Lustre or OrangeF, which brings higher performance in some file operations. Metadata-intensive workloads are used to stress the network traffic and analyze the scalability of the file systems. Pascual et al5 investigate alternatives for the storage subsystem of a novel exascale-capable system with special emphasis on how allocation strategies would affect the overall performance. They consider several aspects of data-aware allocation (such as the effect of spatial and temporal locality, the affinity of data to storage sources, and the network-level traffic prioritization for different types of flows) and show that scheduling policies exposing data-locality information can be essential for the appropriate utilization of future large-scale systems. They also found that the distributed storage system they implement can outperform traditional SAN architectures, even with a much smaller (in terms of I/O servers) back-end. Finally, Castro et al6 present an energy study on critical parameters for the deployment of CNNs on flagship image and video applications, ie, object recognition and people identification by gait, respectively. Their experimental results on a multi-GPU server endowed with twin Maxwell and twin Pascal Titan X GPUs demonstrate that energy correlates with performance and that Pascal may have up to 40% gains versus Maxwell. Larger batch sizes extend performance gains and energy savings but accuracy must be watched, which sometimes shows a preference for small batches. The manuscripts presented in this special issue provide insights into several cutting-edge aspects of exascale computing. We believe that the main contributions presented in these manuscripts are timely and important. We hope that readers can benefit from these research manuscripts and contribute to these rapidly growing areas. Manuel E. Acacio is a Full Professor of computer architecture and technology at the University of Murcia, Spain. He obtained his PhD degree in Computer Science in March 2003. Before, in the summer of 2002, he worked as a summer intern at IBM TJ Watson, Yorktown Heights (NY). Currently, Prof. Acacio leads the Computer Architecture & Parallel Systems (CAPS) research group at the University of Murcia. He is author of more than 100 papers in refereed international conferences and journals. As well, he has served as a committee member of important conferences, ICPP and IPDPS among others. His research interests are focused on the architecture of multiprocessor systems. From April 2011 to April 2015, Prof. Acacio served as an associate editor of IEEE Transactions on Parallel and Distributed Systems International Journal; since August 2016, he is a member of the editorial board of MPDI Computers Int'l Journal; and more recently, since September 2018, he serves as an academic editor in the editorial board of Hindawi Scientific Programming journal. He is also a member of the board of distinguished reviewers of ACM Transactions on Architecture and Code Optimization Int'l Journal since May 2014. Julio Sahuquillo is a Full Professor with the Department of Computer Engineering at the Universitat Politècnica de València. He has enjoyed a postdoctoral research stay with Prof Antonio Gónzalez, former director of Intel Barcelona. He has taught several courses on computer organization and architecture. He has co-authored more than 150 refereed conference and journal papers. His current research interests include multi- and manycore processors, memory hierarchy design, cache coherence, GPU architecture, resource management in the cloud, and architecture-aware scheduling. In these topics, he has advised more than 10 PhD Theses and has been the principal investigator of competitive Spanish domestic projects, international European projects, and projects with international companies. He has participated in the organization of about 30 conferences (HPCC, Euro-Par, HPCS, etc.) in different positions: Publicity Chair, Local Organizing Committee, Workshop Co-Chair, and Special Session Co-Chair. He participates assiduously in the PC of major Computer Architecture conferences. He is a member of the IEEE and the IEEE Computer Society. The guess editors would like to thank all the authors who made valuable contributions to this special issue. We also thank the reviewers for their detailed review reports that have helped to further enhance the manuscripts originally submitted. Finally, we would like to express our sincere gratitude to Prof Geoffrey Fox, the editor-in-chief, for having provided us with the opportunity to edit this special issue in the international journal of Concurrency and Computation: Practice and Experience, as well as for his assistance throughout all the review process.