Latest advances in distributed, parallel, and graphic processing unit accelerated approaches to computational biology
Ivan Merelli, Horacio Pérez‐Sánchez, Sandra Gesing, Daniele D’Agostino · Concurrency and Computation Practice and Experience · 2013
Bioinformatics is a discipline that performs analyses, modeling, and simulations of complex biological systems by using a computer science approach, which typically means dealing with huge amounts of data. Although the most powerful supercomputers in the world are heavily involved in computational biology research, scalability, portability, integration, and usability of bioinformatics software still represent open issues. Currently, the possibility of parallelizing algorithms and analysis techniques exploiting various high-performance computing (HPC) techniques and platforms is receiving an even increasing interest. Examples include the porting of legacy applications to clusters, such as those for genome analysis, and the use of distributed technologies like grid and cloud computing for large, embarrassingly parallel computations. Also, performance acceleration using on-chip supercomputing, such as graphic processing units (GPUs) and massively parallel architectures for the processing of large data sets, is becoming largely exploited. This trend is motivated by the lightening improvement of novel molecular biology high-throughput technologies, such as next generation sequencing, which allow the analysis of inter personal variations in genomics and transcriptomics, but also to the development of mass spectrometry techniques for proteomics and metabolomics profiles. This clearly calls for novel solutions in the field of HPC. Also in the field of structural biology, in silico molecular dynamic simulations, ligand screening projects for neglected and complex diseases, and drug discovery campaigns are examples of the great advantages that HPC can provide to medicine and healthcare. Many are the projects aiming at developing technologies in the field of high-performance computational biology worldwide. Examples are the EU-funded FP7 projects Venus-C 1 and scientific gateway based user interface (SCI-BUS) 2, and the project Extreme Science and Engineering Discovery Environment (XSEDE) 3 funded by the US National Science Foundation. These projects aim to improve the quality of HPC services for diverse research areas and support a large number of applications and projects in computational biology, for example, Biodrugscore 4 in XSEDE and the Swiss Proteomics Gateway 5 in SCI-BUS. Furthermore, the FP7 projects BioHPC 6 and Elixir 7 are specifically dedicated to the life sciences. Whereas BioHPC offers a suite of applications in a science gateway, Elixir functions as data hub for life sciences organizations. As regards for the most important national initiatives, the Flagship project Interomics 8, Laboratory for Interdisciplinary Technologies in Bioinformatics (LITBIO) 9, and Italian Bioinformatics Network (Italbionet)10 in Italy and IdeeB 11, Grid Support for Bioinformatics 12 and Renabi 13 in France are concerned with the development of bioinformatics HPC infrastructures. Also, in the UK, there are very important initiatives in this sense, such as MyGrid 14, which is very active in the development of tools for e-Science. In Germany, the D-Grid project MediGRID 15 was focused especially on grid services for biomedical research. Its work is continued via Technologie und Methodenplattform für die vernetzte medizinische Forschung e.V. 16, which is the umbrella organization for networked medical research in Germany. The D-Grid project MoSGrid 17 offers a complete solution for the molecular simulation community supporting HPC infrastructures via a web-based science gateway. It is being further developed via SCI-BUS. WeNMR 18 follows a similar approach like MoSGrid and it supplies a worldwide e-Infrastructure for NMR and structural biology. All these projects have prompted the diffusion of parallel and distributed solutions for bioinformatics, which also promoted the discussion of these topics in a number of conferences. In particular, this special issue of Concurrency and Computation: Practice and Experience is partially a follow-up to some special sessions/workshops held in the context of international conferences in the field of HPC, such as the special session ‘Grid, Parallel and Distributed Bioinformatics Applications’ of the ‘Euromicro International Conference on Parallel, Distributed and Network-Based Processing (PDP)’ 19. European Grid Infrastructure 20 has organized each year two forums as successors of the Enabling Grids for e-Science in Europe conferences 21 for the HPC community with dedicated sessions to the life sciences since 2009. Also, smaller events are launched very successfully in this field like the workshop series International Workshop on Science Gateways for Life Sciences (IWSG) 22 extended to IWSG for a wider community or the Black Forest Grid Workshop23. Accelerating applications using specific paradigms of computations, in particular for scientific computations, is a long-standing task and usually involves the use of custom-designed solutions. Traditionally, parallel computing has been employed for addressing bioinformatics problems that would otherwise be impossible to solve. A first memorable example has been the computational assembly of the human genome as post analysis of the whole genome shotgun sequencing proposed by Craig Venter, which challenged the long-standing Human Genome Project, arriving very close to an unpredictable victory 24. The point is that using parallel computing implies rethinking the whole application to exploit multiple processors, shared or distributed memory resources, and the network itself. The role of software architect is therefore of primary importance to maximize performance beside, and of course, the possibility of accessing a large computational facility 25. In this context, the cost of buying and maintaining an in-house cluster is very important, and this explains why the grid computing paradigm gained a great success in the mid-1990s 26. The term grid was used in analogy to the electric power grid to indicate the main goal to make the access to computing power as easy as the access to electricity. Ian Foster extended his definition of grid computing by a three-point checklist, which emphasizes that a grid manages distributed resources, uses standard, open, and generic-purpose protocols and interfaces, and carries out nontrivial quality of services 27. Grid middlewares are responsible for all the aspects related to the efficient management of the available computing power: the authentication of users, the submission and monitoring of jobs, and the data movement. Accordingly, the underlying hardware, the involved operating systems, batch systems, and file systems are hidden from the users. Grid was a very innovative paradigm of distributed and collaborative computing, but it often requires users to adapt their code for being able to run in this environment. Ten years later, cloud computing was presented as a more flexible solution 28. Cloud computing overcomes the idea of volunteer computing for resource sharing by proposing an ‘on-demand’ paradigm in which users pay for what they use. Cloud computing providers offer their services according to several fundamental models: infrastructure as a service (IaaS), platform as a service, and software as a service, where IaaS is the most basic model whereas the other provide higher level of abstraction 29. Furthermore, in the last decade, new paradigms have emerged in parallel computing, in particular concerning the exploitation of massively parallel architectures (GPUs) 30 such as CUDA 31 and OpenCL 32. Users can presently exploit up to 10 TFlops on a single workstation equipped with multiple CPUs and accelerators, with a cost of less than $5000 33. This allows also low-budget research groups to achieve very high computational performance for many compute-intensive applications in the field of computational biology 34. It is widely recognized that the biology research field is data driven because of the presence of high-throughput acquisition techniques and the increased level of details of simulated biological complex systems. The consequence is that there is the need to apply the most advanced HPC techniques and platforms to process the data and to turn them into real knowledge. This explains why all the available HPC solutions have been exploited to deal with the different kinds of analyses. Several kinds of analysis algorithms have been developed with very different behaviors: some of them are mainly CPU-intensive, whereas others require to access a large amount of data; some require a high number of communications between the parallel processes for each step of the computation, and some analysis require a stochastic approach where the accuracy of results is proportional to the number of runs. This is the reason why different parallelization paradigms, parallel, and distributed architectures have been considered in order to be able to achieve the highest possible performance figures. Salah and Kenli 35 proposed PAR-3D-BLAST, a parallel tool for protein structure comparison, where a two-level parallelism is exploited to improve the performance. Meier-Kolthoff et al. 36 describe a parallelized method to infer phylogenetic relationship using Genome Blast Distance Phylogeny on a reference set of microbial genome from the GEBA project. In particular, they developed an infrastructure in order to save computational time and data storage able to scale up the generation of phylogenetic tree. Even if the low-level details of the grid infrastructures are hidden via middlewares, often the application of bioinformatics methods on HPC facilities require specialized knowledge. Thus, researchers in this field currently need to deal especially with usability issues. On the one hand, these issues derive from the applied methods: the usability of many tools is limited and the implemented methods are very complex, reflecting the underlying complex theory. They thus require a lot of experience, also because the lack of user interfaces and of pre-configured settings deter novice users from them. It is clear that users have to become acquainted with the domain-related features, but the presence of a mere command line interface represents a big issue for a wide audience. On the other hand, the heterogeneity of the underlying distributed computing infrastructure augments the complexity, especially for users who do not have an information and communication technology background. A suitable solution to offer easy-to-use and intuitive access to applications are science gateways, which offer specific services tailored to the users' needs. Additionally, the users mostly do not only analyze and process data via single jobs but via workflows in computational biology. Thus, science gateways capable of managing workflows are essential for processing all the necessary steps. Kertesz et al. 37 introduced a case study on generating conformers via unconstrained molecular dynamics parallelizing single steps and creating a reusable workflow in the grid portal WS-PGRADE. The workflow can be performed significantly faster on European grid infrastructures than on single CPU machines. Grunzke et al. 38 addressed a further aspect of usability in workflow-enabled science gateways: metadata management via standards. They present Molecular Simulation Markup Language for the area of molecular simulations of small and large molecules. Molecular Simulation Markup Language allows for workflow-interoperability on data level and is supported in the grid workflows of the MoSGrid portal. Cloud computing is a model in which users access computational resources and storage facilities from a vendor over Internet, that is, the commercial Amazon Elastic Compute Cloud and Simple Storage Server 39. The user can exploit computers and storage for any task, such as serving websites or running computationally intensive parallel bioinformatics pipelines, being an administrator of its services and paying just for the time of effective usage. By instantiating many virtual resources, a parallel cluster can be deployed on demand, where common libraries such as the Message Passing Interface can be exploited. Also, batch-processing systems can be used to manage the different computations in a queue. Moreover, frameworks for distributed access to files such as Hadoop can be adapted to distributed programming paradigms such as MapReduce 40. The flexibility and the cost-effectiveness provided by cloud computing is extremely appealing for computational biology, in particular for small-medium biotechnology laboratories which need to perform bioinformatics analysis without coping with all the issues of having an in-house information and communication technology infrastructure 41. An intermediate solution is represented by Hybrid Clouds that couple the scalability offered by general-purpose public clouds with the greater control and ad hoc customizations supplied by the private ones 42. Kiss et al. 43 investigated and analyzed how Windows Azure cloud can be applied for a virtual screening experiment relying on docking simulations, building a framework that exploits the generic worker concept. Kavasidis et al. 44 presented a bioinformatics knowledge discovery tool, BioWizard-C, for extracting and validating implicit associations between biological entities. The aim of the work is to demonstrate how porting a data-intensive application to the Cloud, affects positively its efficiency. Guerrero et al. 45 stated the importance of GPUs in Cloud Computing environments but also highlighted that the efficiency of such infrastructure, with respect to the use of local resources, should be evaluated on the size of the studied problem. In 2007, NVIDIA released the first version of the CUDA 31 programming tools, which enormously facilitated the exploitation of GPU hardware. Afterwards, other possibilities were presented (i.e. OpenCL 32 and OpenACC 46) aiming to provide users a simplified interface to the great computing capabilities offered by present graphics cards. In the field of bioinformatics, a huge number of applications has been ported to CUDA, and depending on the algorithm, the performance can be very competitive when compared with other platforms 47. In general terms, compute-intensive algorithms with simple mathematical operations are the ones that benefit the most from these computational architectures. NVIDIA cards can be installed in workstations quite easily, but they can be also exploited as part of distributed infrastructures such as grid and cloud platforms. Exploiting simple procedures, such as the Peripheral Component Interconnect Passthrough approach, NVIDIA cards can be mounted on virtual machines to provide the user more complete and customizable solutions, whose costs can be very competitive with respect to in-house implementations. Chessa and Pasquale 48 show in another problem of great relevance for biomedicine, how visual coding in the primary visual cortex can be modeled thanks to a well-designed parallel implementation, conveniently tuned to the GPU architecture. D'Agostino et al. 49 report dramatic speedup improvements obtained in local GPU machines for calculation of the molecular surface, problem of great relevance, which appears in many contexts of molecular sciences. Guerrero et al. 50 demonstrate how computational kernels for virtual screening applications can be ported on GPU architectures achieving both interesting energy efficiency and performance figures. We would like to thank the authors for contributing papers on their research on latest advances in Distributed, Parallel, and GPU-accelerated Approaches to Computational Biology for this special issue and all the reviewers for providing constructive reviews and in helping to shape this special issue. Finally, we would like to thank Prof. Geoffrey Fox for providing us an opportunity to bring this special issue to the research community.