Relative effectiveness of generalized Mantel-Haenszel, simultaneous item bias test and logistic discriminant function analysis for detecting differential item functioning in ordinal test items

Abdul-Wahab Ibrahim · IFE PsychologIA · 2017

IntroductionAn important procedure in the development of psychological instruments is ensuring that no individual or group responding to the instrument is disadvantaged in any way. For instance, Differential Item Functioning (DIF) has an important impact on the fairness of psychological and educational testing. This is because one of the important factors which should be taken into account in ensuring the validity of any test is the issue of fairness. A test that shows valid differences is fair; a test that shows invalid differences is not fair. Hence, DIF forms a major threat to the fairness and validity of psychometric measures (Ibrahim, 2013).Test fairness is a crucial issue in testing, which according to Afolabi (2012), a test that is not fair is a malfunctioning test. The process for developing instruments that are fair for all test takers requires the removal or revision of potentially malfunctioned items. In practice, this implies that before any instrument is ready for use, all malfunctioned items are first detected, and either eliminated or revised. Questions of test malfunctioning are closely related to questions of test validity. A test possesses validity if it measures what it purports to and invalidity if it does not (Afolabi, 2012; Pedrajita, 2015). DIF is a kind of invalidity that arises relative to groups. Validity is an essential requirement of all tests. A valid test produces outcomes that are based only on the trait being measured rather than irrelevant characteristics. When test scores depend on irrelevant characteristics such as group membership (i.e., gender, age social; status) then the test is considered as potentially functioning differently.One way to investigate potentially malfunctioned items especially at the item level is through Differential Item Functioning (DIF) analysis. DIF analysis is a means of statistically identifying unexpected differences in performance across matched groups of examinees. It compares the performance of matched majority (or reference) and minority (or focal) group examinees. Differential item functioning (DIF) is said to be present in a test item when, despite controls for overall test performance, examinees from different groups have a different probability of answering an item correctly or when examinees from two subpopulations with the same trait level have different expected scores on the same item (Osterlind and Everson, 2009).As a psychometric technique, DIF was developed to counteract test item bias, and DIF tells whether a particular test item functions differently to different groups. Based on Item Response Theory (IRT), DIF equips the instrument developer to become aware of situations where examinees of the same ability but from different groups have different probabilities of success on an item. It is expected under equivalent testing conditions, that individuals from different groups (but with comparable ability levels) exhibit similar probability of responding correctly to a given item. Therefore, DIF represents a modern psychometric approach to the investigation of between-group score discrepancies. An advantage of DIF over Classical Test Theory (CTT) is that reliability is not constrained to a single co-efficient, but instead can be measured continuously over the entire ability spectrumcontinuum of variation. Therefore, total score is used as a reference to classify testees into high and low ability groups (Ibrahim, 2013).One of the pioneering methods used to detect DIF is known as the Generalized Mantel-Haenszel procedure (GMH) (Mantel & Haenszel, 1959). This method is based on contingency table analysis and was first used to detect DIF by Holland and Thayer (1988). The GMH procedure compares the item performance of the reference and focal groups, which were previously matched on the trait measured by the test; the observed total test score is normally used as the matching criterion. In the standard GMH procedure, an item shows DIF if the odd of correctly answering the item is different for the two groups at a given level of the matching variable (Stephen-Bounty, 2005). …

Read the paper · More papers on PaperTik