Foreword to the special issue of the workshop on data‐intensive computing in the clouds
Tonglin Li, Bing Xie, Boyu Zhang · Concurrency and Computation Practice and Experience · 2020
The purpose of this special issue is to collect a selection of representative research articles that were primarily presented at the Eighth Workshop on Data-Intensive Computing in the Clouds, held in conjunction with SC'17. In particular, this annual workshop brings together domain scientists, researchers, scholars, vendors, and practitioners from the complementary fields of data science, cloud computing, and high-performance computing, in order to promote an exchange of ideas, discuss future collaborations, and develop new research directions. Data scientists increasingly rely on high-performance computers and cloud infrastructures to analyze high volumes of scientific data, automatically process data, and manage data privacy and performance. As scientific data continues to grow in volume and complexity, computational capabilities also increase at both supercomputing facilities and industry data center. Processing scientific data on emerging hardware with high performance-efficiency and privacy requires a knowledge combination from specific scientific domains and computer system techniques. This special issue presents examples of the successful collaboration from domain scientists, researchers, scholars, vendors and practitioners to address the research challenges on processing scientific data on the infrastructures of high-performance computers and cloud. The scope of this special issue is representative of the multidisciplinary nature of scientific computing in high-performance computing and cloud computing. This special issue addresses the challenges on practical experiences on processing scientific data in different domains and different platforms in high-performance computers and cloud. In particular, Zamani et al.1 show how to automatically integrate large-scale facilities with cyberinfrastructure services for automated data processing. Peng and Plale2 identify the requirements for managing computational analysis among candidate storage solutions. Koulouzis et al.3 describe their experiences that identify the time-critical requirements of environmental scientists making use of computational research support environments and provide a case study whereby their software suite is used to optimize runtime service quality for a data subscription service. Dayarathna and Suzumura4 demonstrate their approach on producing optimized stream query performance, and further compare the solution to naive deployments using two real-world stream processing applications in the domains of health care and search advertising. We encourage the readers to review the aforementioned articles to gain insight into the breadth and depth of problems and innovative solutions in the multidisciplinary field of scientific data in high-performance computing and cloud computing.