Basic Features of Statistical Packages and

Data Documentation · 2012

On its own, this data file is simply a series of numbers. To interpret these numbers requires information about what each value represents. The column labels in Table 3.1 gave us some information, but the information needed to interpret the numbers fully would more generally be found in the data’s documentation. To prepare the data for analysis by a statistical package, we would need to instruct the software on what each value means. For example, we would tell the software which variables are located in which position (e.g., the case identifier, then the mother’s education, then the mother’s age, then the distance away the mother lives, separated by spaces). And, we would choose variable names for each variable (such as caseid, momeduc, momage, and mommiles). We might want to label the variables to remind ourselves about important details found in the documentation (e.g., that momeduc and momage are recorded in years and mommiles is recorded in miles). Our instructions would convert these numbers into a data file format that the statistical package recognizes. Once converted, the data would be ready for analysis by that statistical package, and could be saved in that package’s such as these to “read” the data. This process was time-consuming and error-prone. As we will see in our NSFH example, these days, the raw data file can often be obtained directly from an archive in a format ready to be directly understood by the statistical package. But, it still can be viewed in a row and column format, similar to Table 3.1.

Read the paper · More papers on PaperTik