“Big Data” in Laboratory Medicine

Nicole V. Tolan, M. Laura Parnas, Linnea M. Baudhuin, Mark A Cervinski, Albert Solomon Chan, Daniel Thomas Holmes, Gary L. Horowitz, Eric W. Klee, Rajiv B Kumar, Stephen R. Master · Clinical Chemistry · 2015

Depending upon who you ask, you may get very different answers to the question of “what does ‘big data’ mean to you?” Most obviously, the term “big data” applies to the high-resolution omics data for which we rely on various bioinformatics tools to make conclusions on how to improve patient care. However, “big data” also readily refers to the data reported every day as a part of the clinical laboratory testing environment, and more broadly to the information generated in electronic health records (EHRs).11 There are several practical IT solutions for handling day-to-day “big data” that enable millions of test results to be reported per year. Informatics is changing the processes behind laboratory medicine. With ever-growing demands on laboratory medicine professionals not only to collect and interpret omics data in the era of the Precision Medicine Initiative, but also to ensure high-quality, low-cost patient management in the structure of accountable care organizations, we have invited several experts to discuss their take on “big data.” Our experts highlight how to ensure that the data analyzed are high quality, so that the conclusions we make will translate to effective clinical management and optimal patient care. They review a number of IT solutions they rely on to gain efficiency in the clinical laboratory to benefit clinical practice. Also, our experts discuss the ability to query the clinical laboratory database in an effort to improve test utilization and how “big data” analytics allows for a more effective means of quality management. These 8 experts, with diverse backgrounds and interests, highlight various IT solutions to tackle our “big data.” How do you define “big data” and what does it mean to you in your clinical practice? Eric Klee: In my opinion, the term “big data” has different meanings depending on the context being considered. From an IT perspective, “big data” is anything that challenges an institution's computational infrastructure and requires application-specific modifications. A clear example is providing sufficient compute nodes, memory, and storage to account for the demands of whole genome sequencing. The size of the data sets generated often requires high-performance computer clusters and specialized storage infrastructure. I think “big data” has a different meaning from the context of a cyto- or molecular geneticist. From that perspective, I would assert that “big data” refers to any data set that challenges or exceeds an individual's ability to manually evaluate all data points for clinical relevance. This does not necessarily require the data to be whole genome sequencing, but any targeted next-generation sequencing (NGS) panel of sufficient size to require informatics solutions to enable data reduction before clinical interpretation. For example, a molecular geneticist might be capable of reviewing all variants called on a 10-gene panel (approximately 30–40 variants per case) without additional informatics support; however, they would be overwhelmed in trying to do this for a 60–100-gene panel. Linnea Baudhuin: “Big data” is a broad term that relates to data sets that are so large, diverse, and/or complex that traditional data processing applications are inadequate to analyze, capture, curate, share, visualize, and store. Thus, “big data” requires innovative bioinformatics solutions for processing, to make the data meaningful, and to derive usable information. In the arena of clinical molecular genetic testing, NGS has required us to develop complex bioinformatics solutions that require raw data mapping, alignment, filtering, and variant calling. Stephen Master: There are several different definitions of “big data” that get thrown around. One school of thought says that we should only use the term “big data” when we have too much information to store or process on a single computer. Another way to define “big data” is by the characteristics of volume (amount of data), velocity (how quickly we acquire data), and variety (different kinds of data). We also talk about “big data” in the context of analyzing large, complex data sets such as those derived from omics experiments. The important thing to recognize is that many of the same analytical approaches can be applied to laboratory medicine data regardless of the precise definition that we choose. As a result, I think that it's appropriate to use “big data” to describe the information that we get from our large numbers of patients, samples, and analytes in the clinical laboratory. Dan Holmes: The term “big data” is frequently used by my laboratory medicine colleagues, clinicians, and health administrators in various settings. From the context of these conversations, I (somewhat jokingly) would be prepared to say that most of us use the term to mean “I can't do the analysis in Excel.” Restricting my thinking to healthcare environments, I would define the term as follows: Big data (in clinical medicine): (n) Extremely large data sets obtained from demographic, clinical (medical, nursing, paramedical, pharmacy), diagnostic, and public health records used to direct decisions about diagnostics, patient care, resource allocation, and epidemiological trends. The connotation of the term suggests that the data itself might have to be pooled from disparate sources, usually databases of diverse structure, and that the data might require substantial “wrangling” to prepare it for analysis using preexisting or custom computational tools. Mark Cervinski: All of the high-volume test result data produced by automated instruments in the clinical laboratory could qualify as “big data.” A medium-sized laboratory such as ours can generate 3 to 4 million patient test results a year and in addition, each one of those results has associated data that never make it to the chart. For every test performed, we track the time when, and where, a sample was drawn, transit time, processing time, analyzer time, resulting time, and specimen integrity (hemolysis, icterus, and lipemia indices). However, the only data to accompany the test result are typically the time the sample was drawn and when it was resulted. All of these nonresult data are valuable and mineable. With the proper tools and questions in hand, a laboratorian can dig into these data to query whether their reference ranges are appropriate, to look for test utilization patterns, to curb test overutilization, or to monitor preanalytic and analytic quality. Gary Horowitz: Although my laboratory generates over 5 million patient test results each year, I don't usually think of the work I do as involving “big data,” but maybe I should. To me, “big data” relates to the kind of analytics that Google does—e.g., using the frequency of search terms to track influenza epidemics almost in real time. The work I do with large amounts of patient data reflects practices at my institution, which is only a small piece of the “big data” our laboratory produces each day. My goal in analyzing these clinical data is to see whether I can uncover ways to improve not only laboratory practice, but overall clinical practice. That is, in addition to ensuring the accuracy, precision, and turnaround times (TATs) of laboratory results, I try to see whether the results themselves can be used to monitor and improve clinical care. For example, if many of the vancomycin levels we report are outside the therapeutic range, it's not enough that our assays are accurate and TATs are good. I'd like to know what we can do, as an institution, to help ensure that patients' vancomycin levels are therapeutic. Rajiv Kumar: “Big data” is the assessment of massive amounts of information from multiple electronic sources in unison, by sophisticated analytic tools to reveal otherwise unrecognized patterns. As a pediatric endocrinologist, to me “big data” means the method to enhance the care of human disease, such as insulin-dependent diabetes mellitus. Multiple and fluctuating factors affect blood sugar control, and patients/parents are asked to make real-time insulin multidaily dosing decisions without truly knowing if a given dose will have the intended effect. They have the benefit of their own experience and healthcare provider guidance based on intermittent retrospective review of available data. However, each decision point represents a new combination of variables with subtleties in patterns that may not be readily identified. As the diabetes research community moves closer to realization of an automated closed-loop artificial pancreas in clinical use, “big data” will be the backbone that facilitates optimal glycemic control. Albert Chan: Our world today is geared to improve the lives of consumers. From Amazon to Google to our local grocery store, we have the same expectation: a product that meets our consumer expectation of utmost quality and convenience. This is made possible by “big data.” On the continuum from creepiness to utility, our expectations have shifted. We now readily provide the personalized data that can simplify our transactions or personalize our experiences for the better. How do you tackle your “big data” and what do you make out of it in your practice? Linnea Baudhuin: Most laboratories utilize multiple different software programs and home-brewed IT solutions to analyze extremely large NGS output files. There are 3 major types of NGS tests in current clinical use: (1) cancer genetic variant testing for diagnosis, prognosis, or therapeutic response, (2) gene panel testing for diagnosis of inherited disorders, and (3) whole exome (or genome) sequencing for rare inherited disorders. In the future, we can NGS tests for testing as as and For all of these we the ability to out large numbers of or we frequently variants that are of with the more that we the more we testing of and can help to are not available for For these and we require a to help us tackle NGS data. Most clinical laboratories now at one bioinformatics with the laboratories a of such Eric Klee: The “big data” data sets that I work with are all NGS and most of what we to make out of these data are variant The variant types from variants and small or variants to number variants to and The to make these filtering, to the appropriate reference and a of that are These are into a with the variant to provide the appropriate context to the data at the time of interpretation. Stephen Master: From the of software the of our analytic work is being in the There are a that also have for analytics in a data but we now have a of clinical who are approaches for analyzing laboratory data in Holmes: I use the for processing, and the large data sets I have to with in laboratory medicine. I use it has a very large of from many disparate This means that tools for almost any analysis are available in the to clinical and this tools for database data and and sophisticated and data real-time is available and and has a broad community that and Our with have quality management of quality in new and in For example, for all tests on all for the are in the of These are and to by a single with and to all for review when they for As for in we an analyzer panel results on on a very rare With we to a to these 5 million over the 3 to the patient records and make and Mark Cervinski: For the using our “big data” to a to monitor the mean patient for a number of analytes in real time. is not a new quality but it has to to the in and analyzing the data. We only to develop we to million test results in a this “big data” data we to the process in a software This of “big data” us to develop to analytical the we would only be to if our would be enough to a only to this data in our laboratory and only on the results and when and the sample is to the care that outside of the Gary Horowitz: We make use of and We are in that we have to a database of laboratory data to with many clinical of and We use to to data of and we do the of our using To the number of variables in the future, more sophisticated analytical tools will be such as data from this database using Rajiv Kumar: when reviewing data for a patient with we to recognize patterns that may benefit from of insulin dosing and blood sugar characteristics of insulin and in and These data and are to and are not readily in a for In response, my has an infrastructure goal of more patient more without effort or provider resource of with our patient we are now to to blood per day for using of this data in the of laboratory test results, and provider work data the and data of this patient health information. With using this will be important in clinical decision more precise diabetes care. Albert Chan: have made decisions about patient care based on data. For example, may a based on a blood obtained in the of a outside of the This is the of and providing our care with a more of the For example, a blood to our electronic health our care with a more of blood to make clinical our with this data can new “big data” analytics you use to gain for quality or to provide clinical Eric Klee: We use all the same of NGS analytics that the community on at and variant number of variants per or variant variant frequency In addition, we use a specialized database to enable variant storage with us to quickly generate and Linnea Baudhuin: We utilize bioinformatics tools to analyze NGS data quality and out data that are of quality or require that are number of per of of of variant frequency for and and for analytical to be test quality are for each test and bioinformatics will with quality before To gain for variant we utilize multiple sources to if a variant has and reported in the in and/or in or with These sources our own variant the and and the We utilize information we from these sources, with in and to help us make decisions on variant We also variants test and store this information in a to help of variants when the test is This information is for or Stephen Master: now we with our “big data” it has from the laboratory information in a This is for that are not time like of testing or but it provide a way for us to our analysis into real-time Our at an is that we have much more to the raw results data. In terms of analytic my has the use of analyzer data to However, we to the to take of these analytics in our clinical practice. Holmes: is a for the of quality in of traditional of and the of of in of testing of clinical all of this was manually in which is for a number of it is not is of the in the it is not automated the same each to generate the tools in programs are and are automated report or real-time data tools. For these my and I are tools to the traditional of quality We to and utilization using the of query and and to We may to a using the for a the will be that the what has and we can use the same to laboratory quality the large in our Mark Cervinski: In addition to the we use tools available in our software to analyze sample work and tests per to our and test We also monitor and specimen quality in real time, as from the could and On a the data we collect could be “big data” but as we monitor these on a and we to to this as data.” These data” are to that laboratory and to on those that A set of and can the and the per that more important as for laboratory testing to Gary Horowitz: We generate and test with a goal of analyzing and clinical practice. As an example, we our so many they for or they of should be very in the all and have Our analysis that the test was being in numbers and almost in the of Our data that it an clinical 3 of out of over an We we generate results in a we would not to do the test at In a we can look at how often a single in for how often therapeutic levels are their and how often to In all of these our goal is to try to where, with our clinical colleagues, we can improve patient care and accurate test results in a Rajiv Kumar: In of an in data from our patients, we an analytic report and data in the to retrospective data review without available The automated report is generated at to by glycemic control. This allows a diabetes provider to time and they are For a patient data the provider the and the to review and quickly data trends. questions or are to the patient and/or using the patient in the We are now this provider work to analysis of health data for additional are from your experience with real-time data Rajiv Kumar: A major of data is expectations about data. Our current of using is not to take over real-time but to of In review of to data in the we that may think their provider is their data and get when they are not for an we use and to appropriate expectations only intermittent provider To we have only as we are the expectations we a the diabetes provider with questions or provider to the data with additional effort is on There is a for patients/parents to their and to data this is not a major for most we use of a with to be for How do you see these healthcare in the and way in the should be in to have a at the from Rajiv Kumar: In the of health data to the to this information in the context of laboratory and variables in the chart. This of in the of provider work may improve care for a given In the this medicine will to and clinical decision tools for and and with to of data sets in a given will health to provide on and approaches for human In of the and of health and should be in data infrastructure an patient We to that the data we generate and have are to and share, and can be to questions that we have not thought to we also to for healthcare that use of data in care without provider Albert Chan: In our experience with a personalized healthcare for almost of with who their blood data with such as are at control. such as a test will be with that our of From and tests that provide of disease, our will have to with an ability to take complex data and the and for I my and the of computer This is not a that they will to be it is based on a that all of those of us to in clinical care, will to as for our Our healthcare will on and that we have these to with to make the decisions that their be of to ensure we make accurate conclusions from the analysis of our “big Eric Klee: is important that all of the that have into any “big data” analysis these of per informatics solutions that will not necessarily the A example is the or frequency of a variant that would be called and These are often set with the that the is analyzing data in a test and will when thinking about or complex are of the made complex variant or in to is important that the laboratorian is with the of variant quality that is being one is with extremely large data automated and data reduction is required for interpretation. A laboratorian be to review all possible variants for each but take the time and a of the data reduction and in a “big data” test to ensure the proper are being Linnea Baudhuin: NGS has us to from targeted or single gene analysis to whole and whole genome with our analysis of the data has from software solutions to the to a set of bioinformatics software that are a combination of and A bioinformatics us to testing with high and In the world of this means that we can as many variants and types of variants as are ensuring that the data being reported are quality we to this with being about a test that is in that more is not necessarily better. In the more we the more variants we and the more variant that to be in to by the laboratory time by the trying to and the results, a for of the report by and testing on Thus, the have a to provide NGS tests that are and as as We also to what the of testing are what is not with the and we variants in a and Stephen Master: I think that the for the laboratory is not enough time thinking about data management. There have several the complex data and to possible to or The is that about “big data” it can be very to we have ways of and processing data. Another important is the of in our with to use “big data” approaches in clinical and laboratory we to be to and review each laboratories the This has important not only for the way that we our use of data for computational but also for the way that we the of clinical Mark Cervinski: it is or data can be to a The of questions and of all analysis what or data or from the is all the results of our analysis of “big data” be to be by our our data sets may not be possible of health information or of the size of the I would the of the tools used so that they can be and upon by Holmes: The data out of our are not as as we For example, in analytical tools for we have that data with results, and review of the quality of the data and a for data are to ensure that the results are and We usually with small on to the of we the into and it In medicine we often a test if it does not clinical With “big data” we are tests on our data to see if we can help direct patient care, and However, we are in of our in custom analytics only to with that are or to which the appropriate is the analysis does not or clinical or laboratory practice, we that have is that the analysis have a of clinical and/or laboratory medicine they are to that may to an but don't from a practical For this of quality management and from the clinical or clinical laboratory does not a to improve patient care. A professionals from and from Gary Horowitz: the most important for us to relates to the of local on the the of the As an example, we an analysis of extremely high from our of the traditional disease, our we and most this our patient our and/or we our by that a is not in those A example relates to to derive reference from laboratory which are they massive amounts of information. as as sophisticated have used in these the results depending on whether one all patients, or the to or the to with the of the use of these to the data can be We analyzed the of with diabetes at our institution, on in the database to the The of was very so high in that we to dig into the at which point we that many of the being for diabetes an diagnosis of In an may not and one be in how much data one Albert Chan: “Big data” does not with of data it more to the from the To truly the we to develop and computational approaches to “big data” into to for clinical With these new we will be to our to be in ways not possible without data. electronic health records next-generation sequencing laboratory information turnaround time variant of variant or variant number variant quality of and laboratory information health information for This for a number of IT in this The of the for the was a one and a result of the by the and and of and of and of of of of and of Mark of of and of and of and of of and

Read the paper · More papers on PaperTik