Analyzing Generic Tabular Numeric Datasets in R
Edward Curry · 2020
This chapter discusses the the basic data-handling aspects of R, and how these can be used to work with tables of numerical data that arise very frequently in biological research. In the initial stage of data analysis, it can be very helpful to examine the data visually to check that there are no obvious anomalies. The cell viability measurements can simply be plotted as numeric values using the plot function. The chapter compares the z-scores following each transfection in MCF7 cells with the corresponding transfection in HeLa cells. Statistical correlation measures provide a means of assessing the similarity between trends in data. Clustering is generally the task of associating similar entities from a dataset: this could involve finding groups of genes with similar expression levels across a set of samples, or it could involve finding groups of samples with similar expression levels of certain sets of genes.