Essential Statistics with Python and R
S. Rao Jammalamadaka · eScholarship (California Digital Library) · 2019
Statistical ideas have become indispensable not only for doing and understanding scientific research but just for being well-informed citizens.We deal with uncertainties around us in everyday life from weather forecasting to stock-market gyrations and are bombarded with polls and advertisements containing claims and counter-claims.So a certain level of "Statistical literacy" is desirable for all of us.The need for such statistical literacy in this modern age was foreseen by H.G. Wells, who said "Statistical thinking will one day be as necessary for efficient citizenship as the ability to read and write" and that day, we think, is already upon us!This book attempts to cover basic ideas of statistics in a direct and succinct way.The goal is to help develop familiarity with statistical concepts among students and others with varied backgrounds.Minimal mathematical background is assumed and the emphasis is on understanding concepts and how they apply to data.Statistical ideas are explained and then illustrated with one or two simple examples.We use practical and realistic examples in many places, and believe a statistical idea is more easily explained through a simple illustrative example.Besides teaching basic statistics, this book can also serve students and novices as an introduction to modern computational packages that are being commonly used nowadays for statistical analysis, namely Python and R. Two Appendices at the end of the book provide basic introduction to these two packages, and demonstrate their use by working out sample exercises taken from the book, in a follow-up Appendix.Further introduction to these packages comes in the form of worked out Examples in various chapters throughout the book.Clearly many more complex data analyses can be done using Python and R, and we restrict ourselves to their use in connection with the basic topics covered in this book on "Essential Statistics".R code: 1 p i e ( d a t a s e t , r a d i u s =1.0 , l a b e l s=c ( " Heart " , " Cancer " , " S t r o k e " , " Pulmonary " , " A c c i d e n t s " , " Others " ) , main=" P ie Chart o f c a u s e s o f death " ) Python code instruction: Color can be defined for each part using "colors= " inside of the plt.plot().The Syntax for plotting a Pie Chart is: 1 p l t .p i e ( ) Python code: 1 c o l o r s =( ' r ' , ' y ' , ' g ' , ' b ' , ' g r e y ' , ' p u r p l e ' ) 2 p l t .p i e ( d a t a s e t , l a b e l s=name , c o l o r s=c o l o r s , r a d i u s =1.0) 3 p l t .t i t l e ( ' P ie Chart o f c a u s e s o f death ' ) 4 p l t .show ( ) 3Quantitative variables on the other hand, can be represented in many ways.We will describe just two basic graphical methodshistograms and stem-and-leaf plots.A good way to summarize and make sense of a large set of numbers is to form a frequency distribution which tells us where the values are and how frequently they occur.To form a frequency table (or frequency distribution table), we proceed as follows:(i) Locate the minimum and maximum values among the data.(ii) Break this range of values into a small number of groups -called "class intervals" or "bins".(iii) Find the frequency in each bin i.e., how many data points fall into each of these class intervals.* Quantile function; Exact value of a certain percentile * p: specified percentile of interest * size: total number of trials * prob: probability of success in a single trial * Lower.tail: if TRUE then calculate the value x, area to the left of which is p, if FALSE then calculate x, the area to the right of which is p * e.g. the 25th percentile of bin(n, p): qbinom(0.25,size, prob, lower.tail= T RU E) -rbinom(n, size, prob) * Random Generation function * n: number of generated data * size: total number of trials * prob: probability of success in a single trial * e.g.randomly generate 1000 data following a bin(n, prob) rbinom(1000, size, prob) 5.2 Normal distribution dnorm(x, mean =, sd =) * Density function at a given point * x: critical value * mean: mean of normal distribution * sd: standard deviation of normal distribution * i.e. f X (x) = dnorm(x, mean =, sd =) -pnorm(q, mean =, sd =, lower.tail=) * Distribution function; probability after or before a certain value of r.v.X * q: critical value * mean: mean of normal distribution * sd: standard deviation of normal distribution * Lower.tail: if TRUE then calculate the area to the left of the critical value, and if FALSE then calculate the area on the right of critical value qnorm(p, mean =, sd =, lower.tail=) * Quantile function; Exact value of a certain percentile * p: critical value * mean: mean of normal distribution * sd: standard deviation of normal distribution * Lower.tail: if true then calculate the area on the left of critical value if false then calculate the area on the right of critical value rnorm(n, mean =, sd =) * Random variables generating function * n: number of samples generated * mean: mean of normal distribution * sd: standard deviation of normal distribution