Algorithms and applications towards the convergence of high‐end data‐intensive and computing systems
Jesús Carretero, Javier Garcia‐Blas, Koji Nakano, Peter R. Müeller · Concurrency and Computation Practice and Experience · 2017
With the increasing availability of data generated by scientific instruments and simulations, today, solving many of our most important scientific and engineering problems requires high-end computing systems (HECS)1 that may be able to process and storage a huge amount of data.2 With this landscape, many synergies between extreme-scale computing, simulations, and data intensive applications might arise.(3, 4) However, the high-performance computing and data analysis platforms, paradigms, and tools have evolved in many cases in different fields, having their own specific methodologies, tools, and techniques. We need to evolve systems and paradigms to create High-End Data-Intensive Computing Systems (HEDICS) to create high-end resources that must be powerful enough in a broad sense (computation, storage, I/O capacity, communications, etc), but at the same time have to provide utilities from the Big Data computing (BDC) space to satisfy the data management and analytics needs of near future applications. Future HECS platforms will be likely characterized by a three to four orders of magnitude, increasing in concurrency, a substantially larger storage capacity, and a deepening of the storage hierarchy. Moreover, the advent of the Big Data challenges5 has generated new initiatives closely related to ultrascale computing systems in large scale distributed systems. The current uncoordinated development model of independently applying optimizations at each layer of the system software I/O software stack will not scale to the required levels of distribution, concurrency, storage hierarchy, and capacity.6 Thus, we need reusable, modular, and scalable frameworks for designing high-end reconfigurable computers, including novel data processing building block and innovative programming models. In those aspects, many new topics are open to research: parallel and distributed algorithms for HEDICS; algorithms for aggressive management of information and knowledge from massive data sources; resource management and scheduling in high-end data and computing systems; tools and environments for parallel/distributed high-end software development; new programming models, as well as machine and application abstractions; resilience issues in HEDICS; adaptive software; architectures, networks, and systems suited for extreme-scale and Big Data; massive distributed and parallel data analytics and feature extraction; new I/O and storage systems valid for HEDICS; and novel and redesigned high-end scientific and engineering computing. This special issue is intended to provide an overview of some key topics and state-of-the-art of recent advances in subjects relevant to High-End Data-Intensive Computing Systems. The general objectives are to address, explore, and exchange information on the challenges and current state-of-the-art in HEDICS, new programming models, run-times, and data facilities design and performance, and their application in various science and engineering domains. This special issue includes research papers addressing the state-of-the-art in high-end data-intensive computing systems. A set of carefully selected works was invited based on the original presentations at the 16th International Conference on Algorithms and Architectures for Parallel Processing (ICA3PP 2016),7 which was held in Granada, Spain, December 2016 and the Third International Workshop of Sustainable Ultrascale Network (NESUS 2016),8 held in Sofia, Bulgaria, October 2016. The extended works have been thoroughly reviewed by an international technical reviewing committee, and only nine papers covering a wide range of relevant challenges in HEDICS were selected for this special issue. The manuscripts present research works showing the convergence of High-End Data and Computing Systems, including new frameworks and platforms, system software enhancements, algorithm design and optimization, programming paradigms and techniques, data processing support in high-end computing systems, and run-time support for HEDICS and performance simulations, measurement, and evaluations. The set of accepted papers can be organized under the following key subjects and subsections and are briefly described in the remaining parts of this section. Current parallel and distributed programming frameworks aid developers to a great extent in implementing applications that exploit homogeneous resources. Nevertheless, it is generally accepted that the ability to develop large-scale distributed applications has lagged seriously behind other developments in cyber-infrastructure.9 Thus, developers strongly require additional expertise to properly port and tune their applications to operate efficiently on specific parallel and distributed platforms, which is not straightforward and demands considerable efforts and specific knowledge. One important cause is the lack of high-level parallel pattern abstractions in the existing frameworks. Dolz et al,10 in their paper A Generic Parallel Pattern Interface for Stream and Data Processing, propose GRPPI, a generic and reusable parallel pattern interface for both stream processing and data-intensive C++ applications available for high-end nodes. GRPPI accommodates a layer between developers and existing parallel programming back-ends targeting multi-core processors, such as C++ threads, OpenMP and Intel TBB, and accelerators back-end like CUDA Thrust. Furthermore, thanks to its high-level C++ API and pattern composability features, GRPPI enables users to easily expose parallelism via stand-alone patterns or pattern compositions matching in sequential applications. The authors evaluate this interface using an image processing use case and demonstrate its benefits from the usability, flexibility, and performance points of views. Furthermore, they analyse the impact of using stream and data pattern compositions on CPUs, GPUs, and heterogeneous configurations. To scale to the next level, as high-end data intensive computing systems become more widespread for scientific applications, there is a necessity of simplifying the development, deployment, and execution of complex data analysis applications for scientific discovery. The scientific workflow model is the leading approach for designing and executing data-intensive applications in high-performance computing infrastructures. Commonly, scientific workflows are built by a set of connected tasks generally arranged in a directed acyclic graph style, which communicate through storage abstractions. Regarding the paper A Data-aware Scheduling Strategy for Workflow Execution in Clouds, Marozzo et al11 present the integration between DMCF and Hercules solutions by using a data-aware scheduling strategy for exploiting data locality in data-intensive workflows. The Data Mining Cloud Framework (DMCF) is a system allowing users to design and execute data analysis workflows on cloud platforms, relying on cloud storage services for every I/O operation, while Hercules is an in-memory I/O solution that can be used in DMCF as an alternative to cloud storage services, providing additional performance and flexibility features. The experimental results demonstrate the performance improvements achieved using the proposed data-aware scheduling strategy in the Microsoft Azure cloud platform. In particular, with the new proposed scheduling strategy, the I/O overhead has been reduced by 55% with respect to the Azure storage, leading to a 20% reduction of the total execution time. In spite of former solutions, network traffic is always a major problem in HEDICS due to data movements. In their paper A scalable synthetic traffic model of Graph500 for computer networks analysis, Fuentes et al12 provide a simulation tool for network architects that need to evaluate the suitability of their interconnect for Big Data applications. Their development is a low computation- and memory-demanding synthetic traffic model that emulates the behaviour of the Graph500 communications and is publicly available in an open-source network simulator. The characterization of network traffic is inferred from a profile of several executions of the benchmark with different input parameters, and the equations in the model have been validated against an execution of benchmarks with a different set of parameters to measure also the impact of the node computation capabilities and network characteristics in the execution time of the model. To cope with huge jobs, some organizations use volunteer computing to get computing resources to scientific projects, so that organizations can be able to attain large computing power from volunteer clients instead of making a high investment in infrastructure. However, there are projects, like the ATLAS@Home project,13 in which the number of running jobs has reached a plateau, due to a high load on data servers and networks caused by file transfers. Alonso et al,14 in the paper A New Volunteer Computing Model for Data-Intensive Applications, provide an alternative, named ComBoS, to improve the performance of volunteer computing projects that have reached their limit due to the I/O bottleneck in data servers by having a percentage of the volunteer clients running as data servers, called data volunteers, to reduce the load on data servers. This solution also improves data locality, leveraging the network latencies of closer machines, as shown by the performance increase provides by their solution, applied to three different BOINC projects. Two current trends in Big Data processing have made the usage of GPGPUs very popular in HEDICS: information discovery and deep-learning techniques15 and collective video games.16 In both cases, there is an increasing trend to discharge client nodes by sending bulk computing to heterogeneous high-end computing nodes for data processing. Data compression is an important area in many data management applications, like training of deep learning, where data must be decompressed many times. Nakano et al,17 in their paper Adaptive Loss-Less Data Compression Method Optimized for GPU Decompression, present a novel lossless data compression method, called Adaptive LossLess (ALL) data compression, designed with the objective of performing decompression very efficiently on the GPU. Evaluations of the ALL data compression method against published lossless data compression methods implemented in GPU show improvements between 1.22 and 23.5 times running on the same GPU. Due to the massive extension of many mobile applications, such as sensors and smart phones, it is crucial for HEDICS to offload applications to high-end nodes so that low-power devices can be used as clients. One possible approach to deal with this problem is the solution proposed in the paper Accelerating Linux and Android applications on low-power devices through remote GPGPU offloading by Montella et al.18 They describe the architecture and integration of RAPID, a complete framework suite for computation offloading to help low-powered devices overcome these limitations. RAPID supports CPU and GPGPU computation offloading on Linux and Android devices, providing lightweight secure data transmission of the offloading operations. The proposed framework is highly modular and exposes a rich Application Programming Interface (API) to developers, making it highly versatile while hiding the complexity of the underlying networking layer. The evaluation results show that Java/Android GPGPU code offloading is possible, through a BioSurveillance application, a commercial real-time face recognition application. High-end networked scientific and engineering applications requires usually HPC for numerical computing and large storage capabilities at end nodes. As the problem grows, the scientific community, in its never-ending road of larger and more efficient computational resources, is in need of more efficient implementations that can adapt better to the current parallel platforms and in need of more new solutions for memory problems that are now memory bound. The memory problem is addressed by Valero19 in the paper Reducing Memory Requirements for Large Size LBM Simulations on GPUs, where he proposes some initiatives to minimize the memory requirements of the Lattice- Boltzmann Method for its usage on GPGPUs to run large scale simulations. The proposed approach allows the author to execute bigger simulations on the same platform without additional memory transfers, those achieving a high performance. In particular, the paper presents two new implementations, LBM-Ghost and LBM-Swap, which are deeply analysed, presenting the pros and cons of each of them. The need of parallelization at high-end nodes is addressed in the paper Parallel solvers for fractional power diffusion problems by Starikoviius et al. 20 The authors construct and investigate parallel solvers for problems described by fractional powers of elliptic operators, like fractional diffusion. Three state-of-the-art approaches are used to transform the non-local fractional-order differential problem into local partial differential equation problems formulated in a space of higher dimension. Scalability of the developed parallel algorithms is investigated, and their parallel performance is compared in the paper. Finally, the problem of accuracy and efficiency for statistic distributions is addressed by Monni et al21 in the paper Fitting Long-Tailed Distribution to Empirical Data. The authors discuss about the limits of the analysis of empirical fat-tailed distributions, which can describe a variety of evolving systems, both natural and man-made. An algorithm to fit fat-tailed distributions is presented and tested against samplings of the power law, the Yule, the log-normal, and Weibull distributions. The algorithm is general and can be applied to any numerical dataset. Thus, the authors compute the parameters defining the shape of each distribution and test the results against simulations. Their method with another state-of- the-art technique to estimate the parameters of empirical distributions. The accuracy of the estimations is discussed, and they conclude that their method based on a weighted iterated χ2 test performs better than the other. Power laws can fit a variety of distributions coming from real data, so a systematic approach to the measurement of the accuracy of fitting algorithms is essential. Articles presented in this special issue provide recent advances in some fields related to high-end data-intensive computing systems. They were selected by invitation of best ranked from two conferences and peer reviewed by journal selection. Acceptance rate for the special issue was below 50% of the invited papers. We hope that the ideas presented in this special issue can contribute to this strategically important, exciting, and fast growing research area and will be of interest for readers of the journal. As guest editors of this special issue, we would like to express our gratitude to all of the authors who submitted their papers to this special issue, and to the Reviewers that helped us with their hard work and the feedback provided to the authors. We also wish to express our gratitude to the Editor-in-Chief Geoffrey C. Fox for the opportunity to edit this special issue and his assistance during the special issue preparation. We acknowledge the following Reviewing Committee members: Pawe Czarnul (Poland), Guilherme Dinis (Sweden), Ece Guran Schmidt (Turkey), Massimiliano Ferrara (Italy), Shih-Hao Hung (Taiwan), Florin Isaila (Spain), Dingde Jiang (China), Amin Khan (Portugal), Marcin Kostur (Poland), Kenli Li (China), Francesco Longo (italy), Francesc Lordan (Spain), Najme Mansouri (Iran), Panagiotis Michailidis (Greece), Eike Mueller (UK), Tomas Potuzak (Cezch Republic), Philipp Reinecke (Germany), Francisco Rodrigo (Spain), Gopal Shyam (India), Shengen Yan (China), Wenwu Tang (USA), and Peng Zhang (USA).