A Comparison of Two Methods for Determining Difficulty in a Multiple Choice Test
Margery V. Hawes, Gerald C. Helmstadter · The Journal of Educational Research · 1963
TRADITIONALLY, item difficulty has been measured by determining the percentage of exam inees in a group who correctly answered the ques tion. Although most persons have used this clas sical method without question, some dissatisfac tion with it has been expressed. See, for example, Davis (1951). The major limitation of using the proportion who get the item correct is that is does not reveal differences among individuals with re spect to the degree of difficulty they have on each single item. Thus, for one person, an item may be troublesome because he has difficulty in making a choice between only two of the alternatives pre sented; for another, all of the alternatives of the same item may appear as reasonable answers. The rationale for ignoring such differences in the usual measure of difficulty assumes that, when the size of the group is large enough, the laws of chance will operate in such a way that the expected value for the number of persons who will get the item correct will occur. Even so, it is possible to have items with identical proportions of persons who get the correct response which, intuitively, are not of the same difficulty. Consider, for ex ample, the hypothetical situation for the three 5 choice items represented in Table I. While the ex pected proportion of persons who mark the item correctly is identical for all three, the breakdown, showing the different ways in which that proportion was achieved, suggests that the items are definite ly not of equal difficulty. Item B, where 15% of the group could eliminate none of the alternatives and an additional 16% only one alternative, would ap pear to be the most difficult. Item C, where every person was able to eliminate at least two alterna tives and where over half of the group eliminated three of the five choices, would seem to be the eas iest. In spite of this limitation of the classical meas ure, no one seems to have suggested any other way of determining item difficulty. The problem of this study, therefore, was that of devising and ?valua -