Cross-Loadings in Scale Development: Monte Carlo Study of Structural Item-Total Correlation Analysis with Small Samples

Gordon P. Brooks, James Pokoo, Nina Adjanin, George A. Johanson · General Linear Model Journal · 2023

This study investigated the possibility of using item correlations with subscales as a tool to diagnose crossloadings in the scale development process using smaller sample sizes than required for factor analysis.A Monte Carlo simulation using R examined sample sizes from 30-120 under several conditions of correlations, numbers of items , and numbers of factors (2-3).Within each condition, most items were generated to load on their expected dimensions, but some items were generated that (a) cross-loaded on multiple dimensions or (b) loaded on no dimensions.Based on its consistently larger average loading differences between loading power (loading correctly on the right dimension) and loading error (wrongly loading on an incorrect dimension), combined especially with its consistently lower loading errors, Structural Item-Total Correlation Analysis (SITCA) diagnosed cross-loading and non-loading items most effectively across most of the conditions when sample sizes were approximately 50-60.he importance of quality items in scale development cannot be overestimated.However, in the current world, samples are increasingly difficult to obtain, making the process of scale development more difficult.Potentially, a single new scale could conceivably require multiple samples for pilot studies and validity studies to improve items and provide evidence of validity and reliability, respectively, depending on how many changes are made to the scale during the process (ideally, item analyses and validity studies continue after any items are changed before the scale is used for applied research purposes).Having the ability to weed out and improve items that need to be repaired or replaced (or removed) as early as possible in the process, with typically smaller pilot study samples, allows for fewer more extensive and more expensive validity study samples later.Applied researchers perform a number of statistical analyses as they develop tests and scales, in order to provide evidence of both reliability and validity.For example, they perform item analyses (e.g., alpha, item-total correlations, alpha-if-item-deleted) for unidimensional scales and subscales to verify that all items are contributing positively to reliability.Previous authors have suggested that preliminary or pilot studies for these issues in scale development can be performed successfully with small sample sizes of 24-36 cases based on the precision of correlations (standard errors, confidence intervals) used in item analyses (Johanson & Brooks, 2010).Evidence (often from experts) for content validity arguments is also needed.Scale developers also typically perform factor analyses for structural validity to provide evidence that the underlying dimensional structure of an instrument supports construct validity.These factor analytic methods-for example, exploratory factor analysis (EFA), confirmatory factor analysis (CFA), or principal components analysis (PCA)-are also the best methods to verify (a) that items are contributing to the measurement of their designated subscale and (b) that items are not contributing to multiple subscales (item unidimensionality).Items that contribute significantly to more than one subscale (or factor or component) are commonly called cross-loaded items.Therefore, in the scale development process, while still creating and revising items, researchers often desire to use factor analysis to help examine the quality of items: whether items load statistically on any of the dimensions of the scale, whether items load on the correct dimensions, and whether items load on just one dimension.This evidence is useful as the scale and items are still being developed so the items can be improved as much as possible before the more extensive construct validity evidence is collected.Unfortunately, these factor analytic methods generally require hundreds of cases for most scales of decent size and with multiple subscales.For example, some say minimum sample sizes of at least 150 should be used even in the very best circumstances (Bandalos & Finney,

Read the paper · More papers on PaperTik