Clustering in the Service of the Public's Health
Leslie Andrew Lenert, Alfred Lin, Richard A. Olshen, Catherine A. Sugar · 2013
Our purpose here is to provide an overview of some recent developments in cluster analysis and their applications to health services research. To illustrate, we show how k-means clustering can be used to partition a population of patients with depression into clinically relevant health states and to analyze how they change over time. Among the issues discussed are preprocessing of data, identifying the "right " number of clusters, implementation via the Lloyd algorithm, reverse engineering of health state descriptions from clusters, and longitudinal modeling of patients' transitions. 2. Introduction and the Depression Panel of the Medical Outcomes Study Because resources for medical care are increasingly expensive and in some cases scarce, governments and societies must confront the problem of allocating them equitably. As an aid to setting policy it can be worthwhile to measure the quality of life for patients in different health states. By the latter we mean a partition of the space of "attributes, " or "dimensions of health, " that