Computer-Feasible Method for Handling Incomplete Data in Regression Analysis

John Evangelist Walsh · Journal of the ACM · 1961

The data ordinarily available for determining the regression function of a variable to be estimated on the variables used for the estimation consist of a number of multivariate observations, where each observation contains values for the estimation variables and the corresponding value for the estimated variable.However, in many situations involving biological, medical, and other types of data, some of the values for the variables are missing.This can happen among the observations used in determining the regression function and also in the application of the regression function for estimation purposes.This paper presents a method for handling these two problems.The underlying procedure involves the estimation of the missing value for an estimation variable from its regression function on the estimation variables with known values.A scheme which is feasible for application on a high speed computer is developed for determining the huge number of regression relations that can arise among subsets of the estimation variables.This scheme consists in first establishing a basic set of regression relations among sufficiently small subsets of the estimation variables and then determining the remaining regression relations in terms of appropriately weighted sums of functions in the basic relations.Also, for cases where the forms specified for the regression functions are linear in unknown constants, a special type of least-squares curve fitting technique is developed for determining the constants on the basis of incomplete data The method is aimed at cases where there exist reasonably high correlations among some of the estimation variables and is based on common sense rather than a theoretical analysis.In some cases, the method presented may be inefficient or not meaningful.A discussion of such cases is given at the end of the paper.

Read the paper · More papers on PaperTik