Estimating interrater agreement
W. Holmes Finch, Brian F. French, Jason C. Immekus, Shenghai Dai · Educational and Psychological Measurement · 2025
Chapters 4 and 5 were devoted to the topic of reliability estimation and the closely related issue of generalizability theory (GT). We saw that the goal of these methods was to estimate the level of precision or consistency in a set of items. We described in some detail how the concepts underlying classical test theory (CTT), as described in Chapter 3 , could be brought to bear in understanding the degree of consistency in scale scores. In Chapter 6 , we consider the issue of estimating agreement among raters who provide scores for some set of behaviors. For example, a group of five raters may score teachers on various aspects of their classroom practice or students on the quality of their science projects. In other contexts, a team of three psychologists may rate children on their level of behavioral aggression while on the playground, or on the severity of depression symptoms for patients taking part in a medical clinical trial. In each of these situations, we want to quantify the extent to which the raters agree (or do not agree) with one another on their scores.