Abstract
Charles Wheelan, Naked Statistics: Stripping the Dread from the Data, First Edition. New York: W.W. Norton and Company, 2013, xviii + 282 pp. US$ 26.95 (ISBN: 978-0-393-07195-5 [hardback])
The field of statistics is rapidly evolving into a discipline. From batting averages and political polls to game shows and medical research, the real-world applications of statistics continue to grow by leaps and bounds. The author strips away the superfluous outer garments and expresses the underlying beauty of the subject in a way that everyone can appreciate. The right data and a few well-chosen statistical tools can help us in answering many questions of practical life. The author focuses on the underlying intuition that drives statistical analysis. He explains the various statistical concepts such as inference, correlation and regression analysis in an easy and accessible way. He reveals that biased and careless parties can manipulate or misrepresent data and points out how brilliant and creative researchers can exploit the valuable data from natural experiments to tackle thorny problems.
The book is divided into thirteen chapters. The first chapter highlights the points of learning statistics. It helps in summarizing huge quantities of data, making better decisions, answering important social questions, recognizing patterns that can refine how we do everything from selling diapers to catching criminals, catching cheaters and prosecuting criminals, evaluating the effectiveness of policies, programmes, drugs, medical procedures and other innovations and in spotting the scoundrels who use these very same powerful tools for nefarious ends.
The second chapter is about descriptive statistics that are often used to compare figures and quantities. They give us insight into phenomena that we care about through a manageable and meaningful summary of the underlying phenomenon. They help us in framing the issue(s).
The third chapter is on deceptive description. It discusses that even the most precise and accurate descriptive statistics can suffer from a fundamental problem: a lack of clarity over what exactly we are trying to define, describe or explain. It suggests that all calculations and measurements should be checked against common sense and further emphasizes the importance of judgement and integrity in the entire statistical analysis.
The fourth chapter is related to correlation. Correlation measures the degree to which two phenomena are related to one another. One of the crucial points discussed in the chapter is that correlation does not imply causation; a positive or negative association between two variables does not necessarily mean that a change in one of the variables is causing the change in the other.
The fifth chapter discusses basic probability. Probability is the study of events and outcomes involving an element of uncertainty. Probabilities do not tell us what will happen for sure; they tell us what is likely to happen and what is less likely to happen. Sensible people can make use of these kinds of numbers in business and life.
The sixth chapter enumerates some of the most common probability-related errors, misunderstandings, and ethical dilemmas. For example, one of the problems is assumption of events as independent when they are not. A different kind of mistake occurs when events that are independent are not treated as such.
The seventh chapter describes the importance of good data for any statistical analysis. First, good data are based on a sample that is representative of some larger group or population. Second, good data provide some source of comparison and are unbiased. Behind every important study there are good data that make the analysis possible; but getting good data is harder than it seems.
The eighth chapter is about how we are able to draw sweeping and powerful conclusions from relatively little data. It is through central limit theorem. To apply this theorem, the sample sizes need to be relatively large.
In the ninth chapter, the author highlights the importance of inference. The power of statistical inference is derived from observing some pattern or outcome and then using probability to determine the most likely explanation for that outcome. We can gain great insight into many life phenomena just by determining the most likely explanation. Most of us do this all the time; statistical influence merely formalizes the process.
The author in the tenth chapter talks about polling. As compared to other forms of sampling, a poll is a percentage or proportion. The real challenge of polling is twofold: finding and reaching that proper sample, and eliciting information from that representative group in a way that accurately reflects what its members believe.
The eleventh chapter outlines regression analysis. Specifically, regression analysis allows us to quantify the relationship between a particular variable and an outcome that we care about while controlling for other factors. In other words, we can isolate the effect of one variable, while holding the effects of other variables constant.
The twelfth chapter enumerates ‘top seven’ common regression mistakes: using regression to analyse a nonlinear relationship, demonstrating causation between two variables, establishing reverse causality, omitted variable bias, multicollinearity, extrapolating beyond the data and data mining.
The thirteenth chapter is on programme evaluation, which is the process by which we seek to measure the causal effect of some intervention, typically called the ‘treatment’ in statistical context. In other words, the purpose of any programme evaluation is to provide some kind of counterfactual against which a treatment or intervention can be measured. In the case of a randomized, controlled experiment, the control group is the counterfactual. In cases where a controlled experiment is impractical or immoral, we need to find some other way of approximating the counterfactual. Our understanding of the world depends on finding clever ways to do that.
To conclude, the field of statistics like fire, knives, automobiles etc. serves an important purpose. It makes our lives better and can cause serious problems when abused. Hence, we should use data wisely and well.
On the whole, the author has presented the apparently ‘dry’ discipline of statistics in a conversational and easy to understand style. He has minimized the technical jargon and illustrated everything with day-to-day problems. It is really a fun while going through the book. It is going to be useful to all those who are interested in academics, sports, politics, business or any other areas in which statistics rule the roost. In my opinion, this book gives value to the buyer for the money he or she pays for buying this book.
