Survey: Translational Bioinformatics embraces Big Data

Nigam Haresh Shah · Yearbook of Medical Informatics · 2012

We review the latest trends and major developments in translational bioinformatics in the year 2011–2012. Our emphasis is on highlighting the key events in the field and pointing at promising research areas for the future. The key take-home points are: Translational informatics is ready to revolutionize human health and healthcare using large-scale measurements on individuals. Data–centric approaches that compute on massive amounts of data (often called “Big Data”) to discover patterns and to make clinically relevant predictions will gain adoption. Research that bridges the latest multimodal measurement technologies with large amounts of electronic healthcare data is increasing; and is where new breakthroughs will occur. Introduction Summarizing an entire research field is an intrinsically hard problem and for the purpose of this survey, I rely on discussions among the Scientific Program Committee of the 2012 AMIA Summit on Translational Bioinformatics (TBI), the focus areas of the excellent submissions received at the 2012 Summit [1] and the year-in-review presentations of the past two years at the TBI Summit [2]. The key areas of activity at the 2012 Summit were focused on research that take us from base pairs to the bedside [3], with a particular emphasis on clinical implications of mining massive data-sets, and bridging the latest multimodal measurement technologies with large amounts of electronic healthcare data that are increasingly available. Among the submissions to TBI, those that stood out for their innovation were invited into a special issue of the Journal of the American Medical Informatics Association. These capture some the trends underway in translational bioinformatics. For example, Liu et al [4] demonstrated how the ability to predict Adverse Drug Reactions (ADRs) can be increased by integrating chemical, biological, and phenotypic properties of drugs. They demonstrated that data fusion approaches are promising for large-scale ADR predictions in both preclinical and post-marketing phases. Similarly, for advancing the state of the art on interpreting GWAS data, Russu et al. [5] introduced a novel Bayesian model search algorithm, Binary Outcome Stochastic Search (BOSS), for model selection when the number of predictors (e.g. SNPs) far exceeds the number of observations. Finally, advancing the science on using the genome-in-the-clinic, Morgan et al [6] constructed genomic disease risk summaries for 55 common diseases using reported gene-disease associations in the research literature. They constructed risk profiles based on the SNPs as well as based on 187 whole genome sequences and show that risk predictions derived from sequencing differ substantially from those obtained from the SNPs for several non-monogenic diseases—by as much as a factor of 20 times in some instances. Beyond this year's conference papers, in the larger informatics community, the following significant themes emerge over the past two years: 1) the genome has arrived at the door of the clinic [7, 8]. 2) “Big Data” approaches that compute on massive amounts of data to make clinically relevant predictions are poised for breakthroughs [1, 9–11]. 3) Efforts to bridge the latest multimodal measurement technologies with large amounts of electronic healthcare data are increasing. We refer to this emerging focus area as research on mass phenotyping. Genome in the clinic Researchers from the eMERGE project recently demonstrated that GWAS can now be performed by leveraging large amounts of EMR data [12]. For example, Kho et al showed that by using commonly available data from five different EMRs it is possible to accurately identify T2D cases and controls for genetic study across multiple institutions [13]; although in some instances the algorithms need some local tweaking [13, 14]. In parallel, genomic sequencing has moved out of the research realm and established itself in the clinic. For example, at the Medical College of Wisconsin, Dr. Howard Jacob's team used exome sequencing to identify a novel casual mutation that led to successful treatment of a 6-year-old boy with an extreme form of inflammatory bowel disease [7, 8]. In this landmark study, the authors used the patient's medical history, genetic and functional data, to diagnose an X-linked inhibitor of apoptosis deficiency. Going a step ahead, they performed an allogeneic hematopoietic progenitor cell transplant based on this finding to prevent the development of life-threatening hemophagocytic lymphohistiocytosis. Since treatment, there has been no recurrence of gastrointestinal disease, suggesting this mutation drove the gastrointestinal disease. This report demonstrates the power of exome sequencing to arrive at a molecular diagnosis in an individual patient in the setting of a novel disease, and illustrates clinical use of genomic sequencing. In recognition of the importance of such systematic clinical use of genomic information, the team's activities were recently the focus of a PBS NOVA episode titled “Cracking your genetic code”. With the increasing use of genomic information in the clinic, we are bound to be faced with having to interpret sequence variations that have not been observed and cataloged before. The problem is particularly acute in the case of multigenic diseases where known variants only contribute a small amount of risk. In research that attempts to find disease-causing variants in whole genome sequences, Yandell et al, develop a Bayesian method for prioritization of coding and non-coding variants combining several sequence features. They demonstrate the ability to detect rare variants in key genes in small cohorts, and common multigenic diseases. Given the complexity of interpreting genomic information and the lack of comprehensiveness of current genomic variation databases—which are based on a small number of individuals, cover mostly Caucasian population, and where the variant-to-disease correlations don't generalize well—patients look to the scientific community to reliably detect disease and to predict their likelihood of responding to specific drugs. As a testimonial to the importance of this task, the Institute of Medicine recently published a consensus report on the Evolution of Translational Omics: Lessons Learned and the Path Forward. Inspite of the challenges ahead, it is now clear that the era of using the Genome in the clinic has arrived. It remains to be seen if the actual benefit obtained from using genomic information for patient care lives up to the promise genomic data has been hyped up to.

Read the paper · More papers on PaperTik