Basic Data Management in R
Nils B. Weidmann · Cambridge University Press eBooks · 2023
Basic Data Management in RAs we saw in the previous chapter, using spreadsheets to prepare data for analysis may be convenient at first, but entails a number of major drawbacks.In this chapter, we introduce the basics of data management using the R statistical toolkit.R is one of the most popular software tools for data analysis -it has numerous features and extension packages for statistics, visualization, machine learning, etc.Therefore, it is convenient to also use it to prepare your data before you actually analyze it.This way, you can stick to a single software package and one language to implement your entire research workflow from beginning to end.As you know, you interact with R not by pointing and clicking with your mouse, but by entering commands in the R programming language.This way, you can have R run statistical analyses for you, visualize your data, but also perform data management operations.While cumbersome at first, this mode of interaction is extremely powerful and has a number of advantages.Most importantly, the set of commands you send to R (which is typically called a "script") can be saved, such that you can later return to it, fix potential problems, or simply replicate the steps you carried out to arrive at a particular result.This resolves one of the main drawbacks in the spreadsheet-based data management approach we discussed in the previous chapter, where it is difficult -if not impossible -to keep track of the different modifications you made to your data.In this chapter, we focus on "base R," which is the set of commands and functions that are part of R's core functionality.We do this with a particular emphasis on R's features for data storage and processing, and how we can get data into R and back out.In the next chapter, I describe 74