A Random-effects Regression Specification Using a Local Intercept Term and a Global Mean for Forecasting Malarial Prevalance

Benjamin George Jacob, Ranjit de Alwiss, Semiha Caliskan, Daniel A. Griffith, Dissanayake Gunawardena, Robert John Novak · American Journal of Computational and Applied Mathematics · 2013

Historically, malaria disease mapping has involved the analysis of disease incidence using a prevalence responsible variable often available as aggregate counts over a geographical region subdivided by admin istrative boundaries (e.g., districts). Thereafter, co mmonly, univariate statistics and regression models have been generated fro m the data to determine covariates (e.g., rainfall) related to monthly prevalence rates. Specific district-level p revalence measures however, can be forecasted using autoregressive specifications and spatiotemporal data collections for targeting districts that have higher prevalence rates. In this research, initially, case, as counts, were used as a response variable in a Poisson probability model framewo rk for quantifying datasets of district-level covariates (i.e., meteorological data, densities and distribution of health centers, etc.) sampled fro m 2006 to 2010 in Uganda. Results from both a Poisson and a negative binomial (i.e., a Poisson random variable with a gamma d istrusted mean) revealed that the covariates rendered fro m the model were significant, but furnished virtually no predict ive power. Inclusion of indicator variables denoting the time sequence and the district location spatial structure was then articulated with Thiessen polygons which also failed to reveal mean ingful covariates. Thereafter, an Autoregressive Integrated Moving Average (ARIMA) model was constructed which revealed a conspicuous but not very prominent first-order temporal autoregressive structure in the individual d istrict-level time-series dependent data. A random effects term was then specified using monthly time-series dependent data. This specification included a district-specific intercept term that was a rando m deviation fro m the overall intercept term which was based on a draw fro m a normal frequency distribution. The random effects specification revealed a non-constant mean across the districts. This random intercept represented the combined effect of all o mitted covariates that caused districts to be more prone to the malaria prevalence than other districts. Additionally, inclusion of a random intercept assumed random heterogeneity in the districts' propensity or, underlying risk of malaria prevalence which persisted throughout the entire duration of the time sequence under study. This random effects term displayed no spatial autocorrelation, and failed to closely conform to a bell-shaped curve. The model's variance, however, implied a substantial variab ility in the prevalence of malaria across districts. The estimated model contained considerable overdispersion (i.e., excess Poisson variability): quasi-likelihood scale = 76.565. The following equation was then employed to forecast the expected value of the prevalence of malaria at the district-level: prevalence = e xp (-3.1876 + (random effect)i) . Co mpilation of additional and accurate data can allow continual updating of the random effects term estimates allowing research intervention teams to bolster the quality of the forecasts for future district-level malarial risk modelling efforts.

Read the paper · More papers on PaperTik