Investigating the validity of rating scales using many-faceted rasch model
So Young Jang · Studies in English Education · 2019
The characteristics of a test instrument, particularly the frame of a scoring system, would be a strong candidate to increase or decrease rater reliability due to individual raters’ different perceptions on each rating scale. This indicates that scoring reliability may be affected by two parts, both the rater and the scoring system. The purpose of this study is to quantitatively investigate the optimized scoring system in order to acquire the validity of a rating scale by reducing systematic errors from the test instrument. The validation process of an analytical rating scale was conducted and the utility of the newly proposed rating scales were evaluated using the Many-Faceted Rasch Measurement (MFRM). The rating data were derived from multiple ratings of ninety essays. Four scoring models were proposed and evaluated. The results showed that the rating scale seemed to show some problems at the initial analysis. As a remedy to consider this matter, the number of the rating categories was reduced and re-evaluated in terms of reliability. It was found that a scoring system would function more reliably and consistently when evaluating rater reliability. The findings are discussed in reference to their implications for the development of statistically solid rating scales. The validation process of the rating scale is essential in order to identify the source of errors, for instance, whether it comes from the test instrument or randomness in rater performance.