A COMPARISON OF CLUSTER DETECTION METHODS APPLIED TO MASSACHUSETTS CANCER DATA
Al Ozonoff, Thomas F. Webster, Verónica M. Vieira, David M. Ozonoff, Janice Weinberg, Ann Aschengrau · Epidemiology · 2004
ISEE-515 Introduction: Various statistical methods have been developed to assess the degree and/or the location of spatial clustering of disease cases. Determining whether or not there is an unusual geographical pattern of disease is a matter of public concern and can also direct and focus epidemiological studies into potential environmental and other factors associated with the disease in question. However, there is relatively little in the literature devoted to comparison and critique of the various cluster detection methods, and most of the existing studies rely on simulated data rather than real data sets. Using three methods with very different approaches to the problem of cluster detection, we analyzed breast, lung, and colorectal cancer data from a large case-control study in Southeastern Massachusetts. In this paper we discuss the methodology and compare the results of the analyses, with the goal of improving our understanding of this important problem as it applies to real data. Methods: In this paper we use three methods -- Kulldorff’s scan statistic; Bonetti and Pagano’s M-statistic; and a statistic based on a generalized additive model (GAM) smoothing approach by Webster et al. -- were applied to three different cancers, each with two different latency criteria, for a total of six data sets. For each of the six data sets, smoothed rate maps were produced based on the smoothed additive model, to give a visual rendering of the data. Results: There was general concordance between the three statistics, but instances where the results did not agree. Discussion: There are relatively few comparative studies of cluster detection methodology that perform analyses on real data sets, hence these results offer an opportunity to evaluate the performance and differences of various test statistics in a more natural context than that of synthetic data. Areas in the real data sets where different cluster methods give different results illustrate their potential strengths and weaknesses.