A Simple Method for Monitoring Routine Statistics

M. J. R. Healy · Journal of the Royal Statistical Society Series D (The Statistician) · 1983

It is proposed to assess the surprisingness of the latest observation in a time series, when only a very few earlier observations are available, by comparing it with a prediction based on weighted linear regression with weights decreasing exponentially into the past. An enormous quantity of statistical material is collected on a routine basis and published at more or less regular intervals. The problem of presenting such figures to best advantage is not a simple one; almost by definition, the figures are not collected in order to answer a few well-defined questions but rather so as to be available when new questions arise or, better, in order to stimulate questions otherwise would not be asked. This last role requires a new reading in an on-going series should be scrutinized so it can be labelled as surprising or non-surprising according as it invites the comment that seems peculiar or much as would have been expected. This is the monitoring problem, and the second comment indicates the solution lies in comparing the new observation with some kind of expected result based on the previous values in the series. There is a very large literature on the subject and many fairly elaborate methods (notably those based on the work of Box and Jenkins (1970)) have been developed. It is however quite common for the series concerned to be very short, no more perhaps than five to ten readings in all, and then there is a requirement for very simple methods whose assumptions are both few and appar- ent. This note describes one such method which is mainly designed for annual figures; modifications to cope with seasonal variations will be briefly mentioned below. The method supposes, then, we have a small number of equally spaced observations and we wish to assess the surprisingness of the next observation in the series. It will usually be too restrictive to assume the readings are varying about a constant mean, but with a short series it may be adequate to suppose any trend is roughly linear. The proposal is then merely to fit a straight line to the data points, to derive from the line an expected or predicted value for the new point, and to judge the new value by the size oI its departure from expectation. The line can be fitted by simple regression, but it is possible to recognize the time-series aspect of the data by doing a weighted regression with the recent past weighted more heavily than more remote points. One possibility suggested by exponentially weighted moving averages (Chatfield, 1975, section 5.2.2) is to use weights which decrease by a constant factor with each successive step backwards in time. A difficulty occurs with the assessment of the statistical significance (the surprisingness) of the calculated deviation. It is of course technically possible to estimate the scatter about the fitted line, but with the short series envisaged here such an estimate will itself be exceed- ingly imprecise. The proposed method will really be satisfactory only when the precision

Read the paper · More papers on PaperTik