Abstract

Aside from its technical architectural aspect, the massive framing of our early houses is a thing to delight by anyone possessed of the smallest amount of architectural sense. A feeling of boundless strength, of security and steadfastness, as well as a notable kind of dignity, is inseparable from the ponderous timbers which go to make up these mighty frames
The scientific method requires crafting hypotheses, testing them empirically, analyzing the results, drawing conclusions, and, finally, communicating all the above as clearly, thoroughly, and accurately as possible so that replications and extensions are feasible. Logic and reasoning enable manipulating theories and concepts in order to craft sensible hypotheses; coherent methodological steps enable those hypotheses to be tested rigorously and without confounds; appropriate statistical methods provide a researcher with mathematical tools to analyze the data at hand and draw firm conclusions.
Testing a scientific hypothesis is assessing the appropriateness of a theory with regard to an observed phenomenon that has been measured qualitatively or quantitatively. Testing a statistical hypothesis is assessing, via statistical methods, the relationship, the difference, or the lack of fit of a prediction with regard to an observed phenomenon affected by random sources of error. The boundary between these two inferential methods is blurry in the scientific method as both aim at carrying out inferences, making arguments, and reaching conclusions. But how strong is an argument made on the basis of a statistical method? Are all statistical methods equally apt to draw conclusions? This special issue, “Perspectives on the Use of Null Hypothesis Statistical Testing (NHST),” focuses on the use of statistical methods to test scientific hypotheses and uncover central properties of data.
This special issue is divided into three subsections. The first, titled, “The Scientific Basis of Null Hypothesis Statistical Testing,” is presented in this issue of Educational and Psychological Measurement. The second, titled, “The Proper Meaning and Correct Use of Null Hypothesis Statistical Testing,” will follow in the next issue of Educational and Psychological Measurement. The last subsection, “On Some Alternatives to Null Hypothesis Statistical Testing,” will be published in a subsequent issue. In total, 14 articles comprise this special issue.
In the first part, the scientific foundation of NHST is discussed. The link between the scientific method and hypothesis testing is discussed by Chang, whereas the role of hypothesis testing in science is the topic of Haig’s article. The p value is ubiquitous in scientific research but its interpretation and usage has been questioned and challenged. For instance, while some researchers focus on its correct interpretation (e.g., Held, 2010), others question its reliability (e.g., Nuzzo, 2014) and rather stress the importance of research steps prior to the estimation of such a value (e.g., Leek & Peng, 2015). As is well known, Fisher’s p value is meant to measure how consistent the sample data are given the null hypothesis (Reid, 2015). Others extended the meaning of this value such that “the smaller the p-value, the stronger the sample evidence that the alternative hypothesis is true” (Berger, 2015, p. 493). Patriota questions the logic of p values obtained from NHST theory and Bayesian theory and proposes an alternative measure of evidence. Marsman and Wagenmakers describe the reasoning behind p values from the point of view of Bayesian theories.
How hypotheses should be tested (e.g., via frequentist, likelihood, or Bayesian approaches) is a matter of much debate. In particular, the NHST procedure in which no association or no difference between two measured phenomena is assumed at the very onset of the process has been highly attacked. Yet, in the second part of this special issue, some articles defend the usefulness of this position. Häggström presents arguments to embrace such an approach and indicates ways for its appropriate use. García-Pérez argues that current alternatives to the null hypothesis testing do not solve issues this type of testing is said to have. Miller reinforces the idea that null hypothesis testing is an appropriate method to do science.
Alternatives to NHST have also been proposed, the Bayesian approach receiving much visibility. In the final part of the special issue, a few alternative approaches are examined. Jamil, Marsman, Ly, Morey, and Wagenmakers feature the Bayesian approach in the comparison of proportions. Trafimow, on the other hand, proposes the coefficient of confidence as an a priori inferential statistic to be used as an alternative to both Bayesian and NHST procedures. Grice, Yepes, Wilson, and Shoda propose a method called observation-oriented modelling, a method that relies less on estimators of location and scale and instead focuses on the visual examination of data to detect and explain dominant patterns within a set of observations. Another alternative to the NHST that has received much attention is that of inferences based on confidence intervals (see Cumming, 2012). Work in this area is highly active, as evidenced among others by the proposal of a method to assess statistical significance via the (non-) overlap of two confidence intervals (Noguchi & Marmolejo-Ramos, 2016). Wiens and Nilsson illustrated this approach with contrast analyses in factorial designs where confidence intervals are rarely given.
Breiman (2001) argues that most researchers and statisticians perform data modelling (as is the case, e.g., in linear regression where the population distribution of the residuals is given). Model-based inference is but one approach to data analyses (Smith, 1994), one that has been criticized for its lack of information about the experimental design (Little, 2004). Breiman (2001) additionally contends that more should be done in the area of algorithmic modelling (e.g., decision trees). He states that “the goal of statistics is to extract information from the data about the underlying mechanisms producing the data” (p. 203), but data modelling is too often restricted to dichotomic yes–no answers. Statistical inference can instead be seen as a form of algorithmic modelling. Bzdok, Varoquaux, and Thirion advocate the use of such an approach in the analysis of neuroimaging data. Campitelli, Macbeth, Ospina, and Marmolejo-Ramos propose the combination of statistical graphics with modelling techniques to understand the phenomena of interest and exploit all the parameters of the data’s distribution. Another growing field of research in statistics is how to reconcile parameter estimation and inference with the presence of outliers and non-Gaussian distributions (see Portnoy & He, 2000). Wilcox and Serang demonstrate how robust statistics can address problems that null hypothesis, confidence intervals, and Bayesian approaches have.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
