Guest Editor's Introduction: Special Section on Challenges and Solutions in Multicore and Many‐Core Computing

Shujia Zhou, Judy Qiu, Ken A. Hawick · Concurrency and Computation Practice and Experience · 2011

It is our honor to serve as guest editors of this special section of the journal of Concurrency and Computation: Practice and Experience on Frontiers of GPU, Multi- and Many-Core Systems (FGMMS). We are pleased to present nine high-quality contributions in this special issue, where they were first presented at the Frontiers of GPU, Multi- and Many-Core Systems Workshop in conjunction with the 10 th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGrid 2010), held from May 17 to 20, 2010, in Melbourne, Victoria, Australia. The invited papers in this special issue represent augmented works drafted at the beginning of 2010 that addressed the issues below. Multicore and many-core microprocessors are being deployed in a broad spectrum of applications including clusters, clouds, and grids. Both conventional multicore and many-core processors, such as Intel Nehalem and IBM Power7 processors, and unconventional many-core processors, such as NVIDIA Tesla and AMD FireStream graphics processing units (GPUs), hold the promise of increasing performance through parallelism. However, GPU approaches in parallelism are distinctly different from those of conventional multicore and many-core processors, which raises new challenges. For example, how do we optimize applications for conventional multicore and many-core processors? How do we re-engineer applications to take advantage of GPUs’ tremendous computing power in a reasonable cost–benefit ratio? What are effective ways of using GPUs as accelerators? Enormous and rapid progress has been made in accelerator computing over the last two years, but nevertheless we believe the themes developed for the FGMMS workshop and that are represented here in this special issue are still very important ones. In the last two years we have seen the continued rise of GPU computing and indeed its uptake as multi-GPU systems now dominates the top 10 entries within the Top 500 supercomputing systems list 1. Although there has been something of a shakeout of the accelerator technologies that were prevalent three years ago, we have also seen the continued and steady rise of uptake of multicore conventional CPU devices, and an exciting future with combined multicore CPU and GPU devices seems likely. As the articles in this special issue suggest, there are still challenges ahead for application developers to make best use of these future and hybrid highly concurrent systems. There are however still many really important applications that will continue to drive interest, investment, and development of these technologies. The goals of this special issue are to discuss these and other issues and bring together developers of application algorithms and experts in utilizing multicore and many-core processors. We briefly introduce the articles as follows. El Zein and Rendell 2 explore the effect of different GPU programming options (e.g., memory type, memory access methods, and data types) on the performance of routine evaluating sparse matrix vector products and discuss the method for optimal performance. Playne and Hawick 3 report on their approach in accelerating finite-differencing applications using multiple GPU devices with a single CPU host and the asynchronous CPU/GPU communication. Kato and Hosino 4 present their algorithms for speeding up a k-nearest neighbor problem in the recommendation system through multiple GPUs. Ino et al. 5 discuss a cooperative multitasking method for concurrent execution of scientific and graphics applications on GPU and acceleration of compute unified device architecture-based applications using idle GPU cycles in the office. Gillan et al. 6 present a case study on how the instruction-level parallelism offered by three accelerator technologies — field-programmable gate array, GPU, and ClearSpeed — can be exploited in atomic physics with considerable differences in the implementation strategies. Barhen et al. 7 present an unconventional fast Fourier transform implementation scheme for the IBM Cell B.E. processors, named transverse vectorization, and provide the first results for multifast Fourier transform implementation and application on the novel, ultralow power Coherent LogixHyperX processor. Zhou et al. 8 investigate the software and system issues in accelerating climate and weather models in a prototype hybrid computing system, which comprises Intel blades and IBM Cell B.E. blades, connected with both InfiniBand and 1-Gigabit Ethernet and communicate with IBM's Dynamic Application Virtualization software. Qiu and Bae 9 present performance results of two significant bioinformatics applications, gene clustering and dimension reduction, on a Microsoft Windows cluster with up to 768 cores using Message Passing Interface and two variants of threading — Concurrency and Coordination Runtime and Task Parallel Library. Hackenberg et al. 10 present a tool for the graphical program flow analysis of hardware accelerated parallel programs. The tool monitors the hybrid program execution to record and visualize many performance relevant events along the way, and is exemplified through representative real-world applications written for both IBM Cell B.E. processor and NVIDIA Compute Unified Device Architecture (CUDA) API. This work was supported in part by a Microsoft CRMC grant. The guest editors of this special issue would like to express their deep gratitude to all authors, external reviewers, and Geoffrey Fox for their efforts in making this issue possible.

Read the paper · More papers on PaperTik