Choosing the Resampling Scheme when Bootstrapping: A Case Study in Reliability
Hani Doss, Yuang-Chin Chiang · Journal of the American Statistical Association · 1994
Often when dealing with complex data structures there is no unique way to bootstrap. If the data can be viewed as U 1, …, U n iid from some distribution P, then one can bootstrap by resampling the U's. Alternately, one can resample in a more model-based way; that is, by making use of the structure of the model P. A typical example of this is linear regression, in which the data is (Y i , X i ), i = 1, …, n. One can resample the pairs (Y i , X i ), or one can resample the residuals from a fitted model. This phenomenon arises over quite a wide spectrum of problems, and in many cases the different methods of bootstrapping can give substantially different results. It seems hopeless to come up with a general theory that compares the different ways of bootstrapping. In this article we study in some detail a certain model that arises in reliability theory in which there are two natural ways to bootstrap. This model is described as follows. Available for testing is a sample of n iid systems each having the same structure of m independent components. Each system is continuously observed until it fails. For every component in each system, either a failure time or a censoring time is recorded. A failure time is recorded if the component fails before or at the time of system failure; otherwise a censoring time is recorded. Thus the system failure acts as a censoring mechanism on the component lifelengths. In this model bootstrapping can be carried out in two ways: resample n systems at random from the original n systems or formally compute the Kaplan-Meier estimates [Fcirc] 1, …, [Fcirc] m of the component life distributions F 1, …, F m . One then generates artificial lifelengths from these Kaplan-Meier estimates and from those form artificial data. We show that asymptotically, bootstrapping by either method yields correct answers. Intuitively, one expects the model-based method to outperform the other method. The results of an extensive Monte Carlo study show that this is usually true, with substantial gains possible. But there are also some cases when the model-based method does worse than the “naive” model-free method.