Regression-based normative data in neuropsychology: Using raw scores as observed response variable outperforms transforming for normality.

Javier Oltra‐Cucarella, Rubén Pérez-Elvira, Beatriz Bonete-López, Clara Iñesta, Esther Sitgés Maciá, Rafael de Andrade Moral · Psychological Assessment · 2025

= 100, 200, 5,000, 1,000, 10,000) for seven different scenarios, analyzed the percentage of individuals scoring in the lowest 5%, and analyzed the agreement between models with the Cohen's κ statistic. Linear models for raw scores and for scaled scores were similar when the model included all the covariates, but barely identified low scores when scaled scores were corrected with covariates taken from different regressions (κ = 0.58). Models with raw scores showed that the expected number of individuals scoring low was close to the expected 5%, whereas models with scaled scores with covariates taken from different regressions were close to 0%. The two models agreed only when the response variable was random symmetrical and uncorrelated with the covariates. When calculating normative data using linear regressions, raw scores should be the preferred choice. If residuals analysis shows that the model does not fit the data well, researchers should consider using nonlinear models. Transforming data for normality of the observed response is discouraged. (PsycInfo Database Record (c) 2025 APA, all rights reserved).

Read the paper · More papers on PaperTik