Dealing With Binary Repeated Measures Data
Christophe Toukam Tchakoute, Kristin Lynn Sainani · PM&R · 2018
Many longitudinal studies and randomized trials have binary outcomes that can repeat. For example, in a study evaluating the safety of a new drug, researchers might record whether a subject developed an adverse reaction at multiple time points during follow-up. These observations are correlated and thus require statistical tests that handle correlated data. Several different statistical models can be applied; this article reviews the advantages and disadvantages of each approach. In a study of 749 infants from Cape Town, South Africa, and Jos, Nigeria, researchers assessed the effects of in utero human immunodeficiency virus (HIV) exposure and breastfeeding practices on the occurrence of infections during the first year of life 1. Infants were HIV-negative but were considered HIV-exposed if born to an HIV-positive mother. Researchers recorded whether or not infants had a lower respiratory tract or other infection, such as diarrhea or sepsis, at 8 time points from birth till 1 year of age. It was hypothesized that in settings with good antiretrovirals coverage and high exclusive breastfeeding rates, HIV-exposed infants will have similar infectious morbidity as HIV-unexposed infants. Before running any statistical models, it is important to graph the data to identify patterns or anomalies with the data, thus informing which statistical test(s) to apply. Plots for repeated binary events can be divided into 2 main types: mean plots, which treat time as a discrete variable, and cumulative hazard plots, which treat time as continuous. Outcomes are typically assessed and collected at discrete time points. In our case study, infants were monitored for lower respiratory tract infections or gastroenteritis at 1, 4, 7, 11, 15, 24, 36, and 52 weeks of age. Mean plots show the percentage of individuals with the event at each time point. For example, Figure 1 plots the percentage of infants with an infection in the HIV-exposed and HIV-unexposed groups at each time point. Mean plot of infections among HIV-exposed and HIV-unexposed infants during the first year of life. (HIV, human immunodeficiency virus.) Mean plots are easy to create and understand. However, they are affected by losses to follow-up. For example, the percentage of infected infants may seem to increase or decrease with time as a result of significant missing data. These plots also assume that outcomes were measured at the exact same time point for all study participants. One can also treat time as a continuous variable using a cumulative hazard plot. The plot ticks up every time events occur. In this case study, events were only recorded at 8 discrete time periods, so the graph has just 7 steps (Figure 2). But the graph could theoretically incorporate events that were measured at any time. Cumulative infectious morbidity incidence according to HIV-exposure status during the first year of life. (HIV, human immunodeficiency virus.) Reprinted with permission from Tchakoute CT, Sainani KL, Osawe S, Datong P, Kiravu A, Rosenthal KL, et al. Breastfeeding mitigates the effects of maternal HIV on infant infectious morbidity in the Option B+ era: A multicenter prospective cohort study. AIDS 2018;32:2383-2391. Each time an event occurs, the cumulative hazard is estimated as the cumulative number of events divided by the cumulative risk set. For example, the HIV-exposed group had 9 infections at week 1, at which time 489 infants were still being followed; the cumulative hazard was thus 9/489=0.018 (1.8%). At week 4, there were 20 new infections among 468 infants; the cumulative hazard was thus (9+20)/(489+468) = 0.03 (3%). Note that the graph accounts for losses to follow-up because infants stop contributing to the denominator at the point at which they leave the study. Examining both graphs (1, 2) reveals that infection rates are similar in the HIV-exposed and HIV-unexposed infants and that infection rates tend to be higher in older infants than in younger infants. It is important to understand these patterns before doing any statistical modeling so that one can recognize if the statistical model yields results that do not make sense. Three main statistical methods are used to analyze repeated binary outcomes: logistic regression for correlated data, Poisson regression for correlated data, and Cox regression for correlated data. We will review the properties of each approach. Table 1 summarizes the results we obtain when applying each method to our case study data set to assess the effects of HIV exposure and breastfeeding status in infectious disease morbidity during the first year of life. For this data set, all 3 models yield similar results. Logistic regression is the most common statistical model used for binary outcomes 2. Logistic regression models the natural log of the odds, or logit, of an outcome as a function of a linear combination of predictors, and is part of a larger class of models known as generalized linear models (see In-Depth Box for more details). Logistic regression can only be used for data in which outcomes were measured at discrete time intervals. Logistic regression also yields odds ratio, which can be difficult to interpret 3. When the outcome of interest is common—typically defined as occurring in more than 10% of the reference group—the odds ratio should not be interpreted as a risk ratio, as doing so may exaggerate the size of the effect. The odds ratio for formula feeding versus exclusive breastfeeding was 1.70, indicating a 70% increase in the odds of infection in formula-fed babies. Because infections were relatively rare in our case study, the increase in odds is only slightly higher than the increase in risk (Table 1). This model does not account for the correlated nature of the outcomes. But an extension of the generalized linear models, called a generalized estimating equation (GEE), does. GEEs extend the generalized linear model to account for the correlated nature of the data by correcting the error, or residual, term. A GEE model requires the user to specify the type of regression (eg, logistic) as well as a correlation structure. However, the model is robust to the choice of correlation structure so long as one uses robust standard errors, also known as sandwich estimators of the variance. To run the model, the data must be formatted such that there is 1 row of data per discrete time period per participant. Predictors may change across each row of data for the same person. In this way, GEE models can incorporate predictors that change over time, such as breastfeeding status. Poisson regression is typically used to analyze count outcomes, such as the number of falls that an individual has during a given time period. Poisson regression models the natural log of the count as a function of a linear combination of predictors. When applied to count data, Poisson regression yields incidence-rate ratios. Beyond count data, Poisson regression can also be used to analyze binary outcomes, because a binary outcome can be viewed as a limited count variable—with only counts of 0 and 1 possible 4. Using a Poisson model on binary data causes the standard errors to be overestimated, but this can be corrected by the use of a robust standard error. Robust standard errors are estimated empirically and they correct misspecifications of the error term (in this case, arising from assuming that the data are Poisson distributed when they are actually binomially distributed). Poisson regression for binary outcomes gives risk ratios rather than odds ratios, which is its main advantage over logistic regression. For repeated binary events, one can use a GEE version of Poisson regression 4. Because infections were relatively rare, logistic regression and Poisson regression give similar results for these data (Table 1). The risk ratio from Poisson regression for formula feeding versus exclusive breastfeeding was 1.68, indicating that the risk of infections was 68% higher among formula-fed infants when compared with exclusively breastfed infants. As for logistic regression, the data must be formatted such that there is 1 row of data per discrete time period per participant. Alternatively, one could collapse the data to 1 observation per person by summing up the total number of events per individual. The outcome would then be a count variable rather than a binary variable, and one could apply standard Poisson regression. However, this would preclude the inclusion of time-changing predictors such as breastfeeding status. One drawback of both logistic and Poisson regression is that they can only be applied to data with discrete time points. Outcomes and dropouts can only be recorded at certain times, and time-changing predictors can only be updated at certain times. Thus, these models sacrifice some precision around the exact time at which events occur and individuals drop out of the study. Cox regression models (also referred to as proportional hazard models) are part of survival analysis methods and are used to model time-to-event outcomes 5. Cox regression models the natural log of the hazard rate of the outcome as a linear combination of predictors. The main advantage of Cox regression over logistic and Poisson is that it can be applied to data with a continuous time variable. In studies with a continuous time variable, the study's time points are not restricted to specific values. Events and drop outs are recorded whenever they occur rather than at limited measurement times. In our case study, measurements were limited to 8 discrete time periods, so Cox regression is likely to give similar results to Poisson. Note also that when dealing with data with very small discrete time intervals, Poisson regression is almost equivalent to Cox regression. Cox regression is typically used to analyze only the time to the first event, such as the first infection. But the model can also be extended to recurrent events by allowing individuals to return to the risk set after having an event. There are several different approaches for modeling recurrent time-to-event outcomes 6. We will detail the counting process model, which is the simplest choice and also the model that was applied in our case study. In the counting process model, individuals rejoin the risk set after each event, and their time-to-event clock restarts at 0. For example, if an infant had infections at 4 weeks, 24 weeks, and 36 weeks, and was followed until the end of the study at 52 weeks, then this infant would have 4 observations with 4 time-to-event values in the data set: 4 weeks (ending in infection), 20 weeks (ending in infection), 12 weeks (ending in infection), and 16 weeks (ending in censoring). Note that the counting process model makes no distinction between an infant's first, second, and third (or more) infections. The implicit assumption is that the relationship between the predictor and the outcome is independent of event number. This may be a reasonable assumption for acute infections, but may not be for other binary outcomes. The multiple observations per individual will be correlated, so a robust standard error is used. Note that in Cox regression the intercept term, which represents the baseline rate of events, is not estimated. Cox regression also has a "proportional hazards assumption," which says that the relative difference in the hazard rates between the comparator groups is the same at all time points. For example, the effect of breastfeeding on infection rates must be the same at 4 weeks as at 52 weeks. Cox regression models are more precise than other choices, but can involve tricky data reformatting. In this study, the data were organized as 8 observations per infant. These 8 observations had to be changed to 1 observation per event (or final censoring), with an added variable capturing the time between events. In addition, breastfeeding data had to be reformatted into an 8-variable array that indicated breastfeeding status at each of the 8 time points when breastfeeding was recorded. Incorporating time-changing predictors in Cox regression is also trickier because it involves comparing event times to the dates at which time-changing predictors were measured. Cox regression models yield hazard ratios. In our study, the hazard ratio for formula feeding versus exclusive breastfeeding was 1.64, which indicates that formula-fed babies had a 64% increase in the rate of infections. Repeated binary outcomes are common in longitudinal studies. Multiple analytic approaches can be used to analyze these data. Logistic regression for correlated data is the simplest approach, but yields odds ratios, which can be misleading if the outcome is common. Alternatively, Poisson regression for correlated data can be applied to binary outcomes to give risk ratios. Both logistic regression Poisson regression assume discrete time intervals. Cox regression, in contrast, can be applied to data in which events were measured continuously. A drawback of Cox regression is that it often requires tedious data reformatting, especially when dealing with time-changing variables. In many cases, the 3 modeling techniques will give similar results. In our case study, our initial hypothesis was confirmed. Formula-feeding and not HIV-exposure was the major predictor for morbidity during the first year of life in this cohort.