Estimation and Prediction for a Simple Software Reliability Model
David E. Wright, Claire E. Hazelhurst · Journal of the Royal Statistical Society Series D (The Statistician) · 1988
The Jelinski-Moranda model is a simple model for software failure time data. The likelihood function for this model sometimes exhibits pathological behaviour and fails to provide sensible estimates of the number of bugs in the software. Despite this the likelihood is informative about the time to the next failure. This is illustrated using real data and looking at the predictive distribution of the time to the next failure. Jelinski & Moranda (1972) use a simple model for software with an initial N bugs. The times to failure resulting from these bugs are assumed to be independent exponential random variables with common failure rate A. The model assumes that on failure the bug responsible is removed and that no other bugs are affected by this corrective action. As the program is run successive failure times thus constitute the order statistic of N random variables from an exponential distribution with failure rate A. This so called Jelinski-Moranda model has received considerable attention in the literature. Blumenthal & Marcus (1975), considering the model in a different context, present tables facilitating the calculation of maximum likelihood estimates of N. They provide a condition for the existence of a finite maximum likelihood estimator. Littlewood & Verrall (1981) present an equivalent condition and describe how, if the data exhibit reliability decay, infinite estimates occur. Foreman & Singpurwalla (1977) apply the model and use asymptotic results to develop a rule for debugging software. Meinhold & Singpurwalla (1983) apply a Bayesian approach assuming prior indepen- dence with a gamma prior for A and Poisson prior for N. Bendell & Mellor (1986) give a comprehensive review and classification of various models including the Jelinski- Moranda model. The model has been criticised on two levels. With its assumptions of a fixed number of bugs, independence and the same constant failure rate for each bug it can be criticised for being over simplistic. On a more fundamental level the notion of failure as an observable event may be considered a naive representation of software behav- iour (see Bendell & Mellor, 1986, Chapter 3). Despite these criticisms we believe that the inferential problems associated with the model are of interest in their own right. Moreover they provide insight into the behaviour which might be expected if some of the assumptions are removed. We focus attention on difficulties associated with estimation of N which occur when the failure history can be explained by a large number of bugs, each with a small failure rate, or a small number of bugs, each with a large failure rate. We argue that in these situations, unless firm prior information is available, there is little information to be gained about the number of bugs N. Despite this the likelihood function is informative about the time to the next failure of the software. This is demonstrated by looking at the predictive reliability function of the time to the next failure which decays to a positive limit, representing the probability that the software is bug free. In practical applications this reliability function may be more useful than an estimate of