Appendix A: Data Set Descriptions
Carl J. Huberty, Stephen F. Olejnik · Wiley series in probability and statistics · 2005
The following five data sets are available at the Wiley website. Data Set A1 (5GED)This data set is based on a sample of community college students.Over 700 students responded to 150 (or fewer) items on the Community College Student Experience Questionnaire (CCSEQ) (Ethington et al. 2001).Items of interest to us were those that were scored with numerical values for two to four categories.Nine response variables were defined for our use; see Table A.1.CCSEQ validity and reliability information for the last six effort scales is reported by Ethington and Polizzi (1996); validity and reliability index values are judged to be respectable.An inspection of the 700 × 9 data matrix was made.We ended up with a 545 × 9 complete data matrix.The 545 (squared) Euclidean distances for each student (represented by a vector of 9 variable scores) to the "typical student" (represented by variable means) were calculated (via SAS).The 545 distances ranged from 5.60 to 37.92 with no appreciable gaps; therefore, it was judged that no outliers were present.Two grouping variables are considered.These are Race with three levels (Black, Hispanic, White) and Grade with five levels (A, A-or B+, B, B-or C+, and C or C-).To obtain the Grade variable, students were asked to report the grade they typically earned in their classes.The number of students in each Race-Grade combination are given in Table A.2.In the covariance analysis context, a hypothetical covariate termed Time is included. Data Set A2 (3GED)This data set is a subset of the 5GED data set.Here, we are using three Grade levels A, B, and C or C-, and does not include Race. Applied MANOVA and Discriminant Analysis, Second Edition, by