Differentiating Effective Data Mining From Fishing, Trapping, and Cruelty to Numbers
Kathleen A. Fairman · Journal of Managed Care Pharmacy · 2007
nn Differentiating Effective Data Mining From Fishing, Trapping, and Cruelty to Numbers Just Right or TooMuch of aGood Thing?It is said that politicians use statistics the way that an inebriated person uses al amppost-"for support, not illumination."1 In managed carep harmacy today,s ome would argue that the same has become true of analyses of medical or pharmacy administrative claims data.The reasoning goes that, given a claims dataset and enough time to massage the data, one can set out to prove nearly anything and produce the desired answer.Is the accusation justified?Compared with other types of research such as randomized controlled trials or patient surveys, retrospective analyses of administrative claims present greater potential for violations of ethical research standards.With atypical database and minimal effort, it is possible (not appropriate, but possible) to recalculate study results post hoc using seemingly endless combinations of methodological decisions.Some of the opportunities to revise study results, either for manipulation or legitimate scientific inquiry, include decisions about these questions: •H ow many claims during what period of time constitute a drug user?•H ow long is the washout period to define a"new start" with the medication?•W hich diagnosis codes in which positions (primary, secondary, tertiary) on how many medical claims constitute the appropriate inclusion criteria?•F or how long should patients be followed and continuous eligibility be required for inclusion in the sample?•H ow should the researcher translate ab road concept, such as noncompliance or treatment success, into measurable decision rules?So many study design changes arepossible, all at the push of acomputer key.This ability to create multiple scenarios so easily has precipitated the lureo ft he "fishing expedition" in which repeated attempts arem ade to produce ap articular desired finding.Unfortunately,this approach poses asubstantial risk of generating incorrect information; while the resulting finding might be appealing, it might also represent nothing moret han sampling error.As tatistical significance standardo f P <0.05 refers to a1in20probability of "Type 1" error,falsely detecting as tatistically significant result when outcomes area ctually due to chance.After just 10 attempts using as tatistical significance standardof P <0.05, the probability of obtaining at least 1false positive result is 40%.After 20 attempts, that probability increases to 65%. 2 However,t he veryf eatureo fc laims database research that is am ajor source of ethical and statistical shortcomings-the ready ability to perform post hoc analysis-is also ak ey tool in avoiding or mitigating those shortcomings.Used properly, for legitimate scientific inquirya nd not to support ap redeter-