Programming models and applications for multicores and manycores

Pavan Balaji, Zhiyi Huang · Concurrency and Computation Practice and Experience · 2015

Rapid advancements in multicore and manycore chips have been a revolution within chip manufacturing, almost eradicating single-core processors. From high-end servers to mobile phones, multicore and manycore chips are steadily entering every single aspect of information technology. However, programming multicore and manycore architectures remains challenging today. To fully utilize these chips, parallel programming models that allow sequential programs and programs utilizing limited parallelism to transition to architectures with massive parallelism, while maintaining good performance and productive development, are urgently needed. This special issue contains six articles selected from the 2014 International Workshop on Programming Models and Applications for Multicores and Manycores (PMAM 2014). These articles cover issues in parallel programming languages and models, work-stealing in parallel runtime environments, heterogeneous computing, and their implications on computational patterns. In ‘Adaptive Demand-aware Work-Stealing in Multi-programmed Multi-core Architectures’ 1, Chen et al. discuss how work-stealing models can be extended to support environments where multiple parallel programs might be concurrently executing. They propose a new work-stealing algorithm called demand-aware work-stealing, which alleviates this issue by allowing programs to donate and steal idle cores as needed. In ‘Palirria: Accurate On-line Parallelism Estimation for Adaptive Work-Stealing’ 2, Varisteas and Brorsson present a self-adapting work-stealing scheduling method for nested fork/join parallelism called ‘Palirria’. The proposed approach can be used to estimate the number of utilizable workers and self-adapt accordingly. In ‘Efficient CPU-GPU cooperative computing for solving the subset-sum problem’ 3, Wan et al. study the impact of heterogeneous computing with CPUs and Graphics Processing Unit (GPUs) for solving the subset-sum problem. The authors observe that while heterogeneous computing is prominent, not much study has been performed in the simultaneous usage of both CPU and GPU resources during computation. This paper proposes an efficient CPU–GPU cooperative computing scheme for solving the subset-sum problem, which enables the full utilization of all the computing power of both CPUs and GPUs. In ‘Dynamic Partitioning-based JPEG Decompression on Heterogeneous Multicore Architectures’ 4, Sodsong et al. introduce a novel JPEG decoding scheme for heterogeneous architectures consisting of a CPU and a general-purpose GPU. They employ an offline profiling step to determine the performance of a system's CPU and GPU with respect to JPEG decoding. Then the runtime partitioning and scheduling scheme exploits task, data, and pipeline parallelism by scheduling the non-parallelizable entropy decoding task on the CPU, whereas inverse discrete cosine transformations, color conversions, and upsampling are conducted on both the CPU and the GPU. In ‘Compiler Transformation of Nested Loops for GPGPUs’ 5, Tian et al. present their experiences in creating an open-source OpenACC compiler in an industrial framework (OpenUH as a branch of Open64). They also discuss in detail the techniques that they developed for loop-scheduling reduction operations on General Purpose Graphics Processing Unit (GPGPUs). In ‘Vectorizing Unstructured Mesh Computations for Many-core Architectures’ 6, Reguly et al. present results on achieving high performance through vectorization on CPUs and the Xeon-Phi on a key class of irregular applications: unstructured mesh computations. Using Single Instruction Multiple Threads (SIMT) and Single Instruction Multiple Data (SIMD) programming models, they show how unstructured mesh computations map to OpenCL or vector intrinsics through the use of code generation techniques in the OP2 Domain Specific Library and explore how irregular memory accesses and race conditions can be organized on different hardware. We hope that the articles in this special issue will provide readers with relevant insights into the emerging parallel programming models for multicore and manycore systems. Pavan Balaji holds appointments as a Computer Scientist and Group Lead at the Argonne National Laboratory, as an Institute Fellow of the Northwestern-Argonne Institute of Science and Engineering at Northwestern University, and as a Research Fellow of the Computation Institute at the University of Chicago. He leads the Programming Models and Runtime Systems group at Argonne. His research interests include parallel programming models and runtime systems for communication and I/O on extreme-scale supercomputing systems, modern system architecture, cloud computing systems, data-intensive computing, and big-data sciences. He has nearly 150 publications in these areas and has delivered nearly 150 talks and tutorials at various conferences and research institutes. Dr Balaji is a recipient of several awards including the U.S. Department of Energy Early Career award in 2012, TEDxMidwest Emerging Leader award in 2013, Crain's Chicago 40 under 40 award in 2012, Los Alamos National Laboratory Director's Technical Achievement award in 2005, Ohio State University Outstanding Researcher award in 2005, six best paper awards, one best paper finalist, and one best poster finalist. He has served as a chair or editor for nearly 50 journals, conferences, and workshops and as a technical program committee member in numerous conferences and workshops. He is a senior member of the IEEE and a professional member of the ACM. More details about Dr Balaji are available at http://www.mcs.anl.gov/balaji. Contact him at [email protected] Zhiyi Huang is an Associate Professor at the University of Otago, New Zealand. His research interests include parallel and distributed computing, multicore architectures, parallel programming models and environments, task scheduling, operating systems, green computing, and computer networks. More details about Prof. Zhiyi Huang are available at http://www.cs.otago.ac.nz/staffpriv/hzy. Contact him at [email protected]

Read the paper · More papers on PaperTik