ERRORS IN THE USE OF CORRELATION AND DETERMINATION COEFFICIENTS

Alexander Ivanovich Orlov · Industrial laboratory Diagnostics of materials · 2018

Coefficients of correlation and determination are widely used in statistical analysis of data. Some of the errors attributed to their use are considered in this article. We confine ourselves to the case of two variables. The linear Pearson correlation coefficient and nonparametric rank coefficients of Spearman and Kendall are used most commonly. According to the theory of measurements, the Pearson correlation coefficient can be applied to variables measured in the interval scale (and in scales with a narrower group of permissible transformations, for example, in the ratio scale) but it cannot be used in analysis of ordinal data. Spearman and Kendall’s nonparametric rank coefficients are designed to evaluate the relationship of ordinal variables. They can also be used in scales with a narrower group of permissible transformations, for example, in the scales of intervals or ratios. The critical value in testing the significance of the difference in the correlation coefficient from zero depends on the sample size and approaches zero as the sample size grows. Therefore, the use of the «Cheddock scale» is incorrect. When using a passive experiment, the correlation coefficients can be reasonably used only for forecasting, but not for control. To obtain the statistical models valid for control, an active experiment is required. S. N. Bernshtein has shown that the effect of outliers on the Pearson correlation coefficient is very large. The effect of «inflation» of the correlation coefficient is that with increasing number of analyzed sets of predictors, the maximum of the corresponding correlation coefficients, the quality of approximation, increases noticeably. A common mistake is to use the determination coefficient to estimate the quality of the least-squares recovery.

Read the paper · More papers on PaperTik