First Visualizations and Statistical Tests with Real Data

John M. Shea · 2024

This chapter shows how to take the techniques introduced in Chapter 2 and apply them to a real data set. It uses an example involving determining whether the early spread of the COVID-19 coronavirus might have been affected by difference in socioeconomic factors between states. To work with real data, the reader is introduced to the Python Pandas library. Pandas provides powerful tools for loading and manipulating tabular data, i.e., data that can be tabulated into rows and columns. The reader learns how to load data from comma-separate value (CSV) files, and how to view, access, and convert that data to other formats. More advanced techniques are introduced for making scatter plots and histograms involving multiple variables. The concept of a partition of a data set is introduced as a way to separate the data into two classes based on various socioeconomic factors. Summary statistics are introduced as a way to represent data in terms of a single numerical value. The concept of choosing a summary statistic to optimize some function of the error to the data is introduced, and it is shown that the average, median, and mode of a data set all optimize different functions. The concept of a Null Hypothesis Significance Test (NHST) is formalized, and a NHST is conducted using bootstrap resampling, a statistical test in which new samples are drawn from the existing data. A brief introduction to two-dimensional statistical techniques is provided.

Read the paper · More papers on PaperTik