A Note on the Reliability of Simple and Residualized Differences
Richard Hays Williams, Donald W. Zimmerman · The Journal of Experimental Education · 1992
We want to express our gratitude to Bruno Zumbo for pointing out the error in our article Comparative Reliability of Simple and Residualized Difference Scores, published in the Journal of Experimental Education in 1983. This error has escaped our notice during the decade since the article originally appeared. Fortunately, the main conclusions are unaffected by this oversight. Zumbo 's systematic approach in posing a series of questions in the form of decision rules goes a long way toward clarifying the concepts developed in the article, and we will quibble with only one feature of his presentation. He sug gested that these decision rules can help a researcher decide whether to use the simple or residualized difference score. We want to emphasize that the equations provided a comparison of the reliability of these two types of scores, and that reliability is only one property, psychometric or otherwise, that determines the usefulness of scores in research. The validity of scores is also a major consideration. See, for example, Wil liams and Zimmerman (1982), which contains a comparison of the validity of simple and residualized differences. There are still other considerations. The sta tistical power of a measure is important in research, and readers of the statistical papers in the Psychological Bulletin are undoubtedly aware that power does not always increase monotonically with reliability. The sort of paradox originally noted by Overall and Woodward (1975) could occur in the present context. That is, it is possible for simple differences to be less reliable than residualized differ ences and yet have more power to detect a false null hypothesis when used in significance tests (see also Zimmerman & Williams, 1986; Williams & Zimmer man, 1989). The equations in our 1983 article reveal which measure has the greater relia bility under specified conditions, everything else being equal. This comparison has theoretical interest for psychometricians, and sometimes, perhaps, practical implications as well. But it does not dictate the choice of the measure to be employed in a research setting. However, we probably worry needlessly that any