Data Manipulation in R
Mario Cortina‐Borja · Journal of the Royal Statistical Society Series A (Statistics in Society) · 2009
The significance of the R system in the wider statistical community is manifested in the exponentially growing amounts of libraries, journal papers, international conferences and books covering many possibilities for this important resource. However, until a few years ago R’s ability to deal with large data sets, e.g. containing millions of records, was a drawback. The availability of considerable computing power and the steady flow of improvements in newer versions allow the analysis of very large databases in R with ease. Perhaps for historical reasons, as R was developed primarily as a computational tool, an aspect which has not received much attention is how to deal with all parts of data management in R. This is what this excellent text is about and there must be very few books whose title matches their content and style more precisely than this one’s. The book starts with a general chapter on data in R, introducing data modes, classes, structures and missing values. This is followed by two chapters on reading and writing data, and on R and databases. These two chapters are likely to be the most useful for experienced R users; they show how to develop the ‘open data base connectivity’ and ‘structural query language’ facilities within the context of R programming and how to exploit the flexibility that is afforded by R to solve complex data management problems. Chapters 4–7 discuss in a very clear way dates, factors, subscripting and character manipulation in R. Chapter 8 deals with data aggregation using well-known R functions such as lapply and tapply and also the use of loops in R. The final chapter is about reshaping data, in particular on how to modify data frames by recoding, reshaping and combining them. Though the ordering of the chapters might look arbitrary the book is well structured and is easy to read. Of the nine chapters three are very short (mean 11.5 pages) whereas Chapters 2, 3 and 8 are relatively long (mean 27 pages). This division shows the balance on which the book is built: it quickly covers the basics of data manipulation with R and it explains with more detail features that are specific to database functionality. Several packages that are available from the Comprehensive R Archive Network are used throughout the book.