Abstract
This article presents a simple approach to making quick sample size estimates for basic hypothesis tests. Although there are many sources available for estimating sample sizes, methods are not often integrated across statistical tests, levels of measurement of variables, or effect sizes. A few parameters are required to estimate sample sizes and by holding the error probabilities constant (α = .05 and β = .20), an investigator can focus on effect size. The effect size can be thought of as a measure of association, such as the correlation coefficient. Here, effect size is linked across three of the most commonly used bivariate analyses (simple linear regression, the two-group analysis of variance [ANOVA] or t-test, and the comparison of proportions or χ2 test) with a correlation coefficient or equivalent measure of association. Tabled values and examples are provided.
This article describes how to estimate sample size. When designing a study, it is important to plan for an adequate sample size so that effort is not wasted. Project proposals typically include a section on analysis, detailing how the research questions will be tested, and a rationale for the proposed sample size. This article can assist with the formulation of a sample size rationale for those unfamiliar with the process but who have some background in statistics. The focus is on the case where there is one independent and one dependent variable. Three of the most common situations are considered: (1) an interval- or ratio-scaled (continuous) independent variable and a continuous dependent variable (simple linear regression); (2) a categorical independent variable (with two categories) and a continuous dependent variable (analysis of variance [ANOVA] or t-test); and (3) two dichotomous categorical variables (contingency tables or χ2 analysis). Although the focus is on these basic comparisons, suggestions are included on how to estimate increased sample size needs with additional independent variables.
When testing hypotheses, three pieces of information are needed to estimate sample size: (1) the tolerable probability for incorrectly concluding that an effect is present when there is really no effect (a Type I error, α); (2) the probability of concluding there is no effect when there really is an effect (a Type II error, β); and (3) the size of the effect you would like to detect. In this article, the two error probability values are set to customary values (α = .05 and β = .20, respectively) to simplify estimation procedures and so that the discussion can focus on effect size.
Although many sources are available on sample size estimation, measures of association to assess effect size are rarely linked across different analyses, so changes in how to operationalize and measure a particular construct or variable (as continuous or categorical) can inadvertently lead to very different sample size estimations. Here, sample size estimates are obtained within the context of the size of the association the researcher considers meaningful and then translated across measurement and analysis options. Effect size is linked across analyses using the Pearson correlation coefficient or an equivalent measure of association: Pearson r in linear regression, eta (η) for ANOVA, and phi (φ) for 2 × 2 contingency tables. A single metric for effect size may make it easier to move between analyses and a single sample size table may be a useful fieldwork tool.
The Process of Statistical Inference
Sample information can be used to test specific hypotheses about the population. Probabilistic statements can be made about population values or the likelihood that a hypothesis may or may not be true. For hypothesis testing, an interval is constructed around a hypothesized value using a theoretical probability distribution. This procedure begins by stating an alternative or null hypothesis, a statement that there is no difference between two values or that there is no association between two variables. Assuming no association or no difference, the test estimates the likelihood of observing the sample value. If the sample value lies outside the test interval, then the null hypothesis is rejected and we say that the finding is statistically significant. It means that the occurrence is rare enough that most likely the null hypothesis is not true. If the sample value lies inside the test interval, then the null hypothesis is not rejected. Hypothesis tests work from the logic of trying to disprove the null hypothesis.
Hypothesis testing, then, is a decision guide. The “truth” is that there is a difference (or an association) or there is not. The hypothesis test allows for a conclusion that a difference exists or it does not. When a hypothesis test concludes that a difference exists and it really does, this is a correct conclusion. Similarly, when a hypothesis test fails to reject the null hypothesis and concludes that there is no difference, and there really isn’t a difference, this is also correct. However, a null hypothesis can be rejected when it is in fact true (Type I error), and you can fail to reject it when it is false (Type II error).
The probability of a Type I error (α) is the probability that the investigator is comfortable with being wrong in concluding that a difference exists, when there really is no difference (α). This probability is often set low, as it is considered a fairly serious error. The probability level may be anything, although it is often .05 (α = .05) meaning that 5/100 you may be wrong, simply by chance. This value can also be .10 or .01; it is not a fixed value. In fact, for a pilot study, it may be defensible to use α = .10. However, to simplify this discussion, the probability for a Type I error will be assumed to be .05 (α = .05, for a two-tail or nondirectional test), a commonly used value for hypothesis testing.
The probability of a Type II error (β) is the probability that the investigator is comfortable with being wrong in concluding that no difference exists, when there really is a difference. This error is not as serious as the Type 1 error and is often set higher. Again, this value may be anything, although it is often .20 (β = .20), it is not a fixed value. A Type II error is of particular concern when interpreting the results of a negative study—without statistical significance (Young et al. 1983). A study without a statistically significant finding means that (1) there is no difference or (2) there is a difference, but the test did not detect it. The latter error can occur when a small sample size results in low statistical power. If a study has a statistically significant finding, then the study clearly has adequate power to detect it. When planning a study, the study should have a sufficient sample size to be able to detect a meaningful effect if one is present. For this discussion, the probability of a Type II error is assumed to be .20, which also sets the statistical power (1 – β) at .80.
By holding α and β probabilities constant, we can focus on the third issue, namely effect size. The question of what is an important effect size is subjective, needs a rationale, and may reasonably be set at various levels (Young et al. 1983). Here, effect size is estimated with an equivalent metric across analyses, the Pearson correlation coefficient (r) for two interval/continuous variables, eta (η) for a categorical variable and an interval/continuous variable, and phi (φ) for two dichotomous variables.
Estimating Sample Size for Two Interval-scale or Continuous Variables: Simple Linear Regression and the Pearson Correlation Coefficient r
Variables may be observed and recorded as categorical, ordinal, interval, or ratio scaled. In statistics, interval- and ratio-scaled variables are considered together and sometimes called “continuous” variables. With categorical variables, for example, gender, numbers represent the names of categories and do not have any true numerical value. With ordinal or ranked variables, numeric values only indicate the relative order of size and do not have information about the differences between ranks. With interval- and ratio-scaled level variables, both the values and the intervals between values have numeric meaning; these are true numeric variables.
Regression or simple linear regression is used to examine the relationship between two interval- or ratio-scaled (continuous) variables. For example, we can test whether maternal smoking (the number of cigarettes smoked per day by a pregnant woman) is associated with infant birth weight (in grams). Or we can test whether blood glucose is associated with the number of depressive symptoms. Regression models the relationship between two interval-scaled variables (an independent variable, X, and a dependent variable, Y). A visual display of the relationship can be seen with a scatterplot of the independent variable as the horizontal axis and the dependent variable as the vertical axis. The measure of association between the two variables is the Pearson correlation coefficient (r). It indicates the degree that the relationship is linear and the degree to which the dependent variable can be predicted from the independent variable. It ranges from −1 to 1, and a larger absolute value indicates a stronger association. Inferential procedures (hypothesis tests or the construction of confidence intervals) involve the use of probability distributions that match the behavior of the sample statistic. For regression and the correlation coefficient, an F or t distribution is used.
The measure of association or effect size that we are trying to detect is the Pearson correlation coefficient, r. The question, then, is what is a meaningful size correlation? Height and weight are two naturally occurring and very strongly correlated variables (around .70). Father’s height and son’s height are strongly correlated (around .50) and probably reflects the 50% genetic sharing of fathers and sons. Smaller correlations may still be meaningful even though they do not predict as well. Another way to consider the effect size is that r 2 estimates the proportion of the explained variance, so a correlation of .50 indicates that the independent variable explains 25% of the total variance in the dependent variable. Cohen (1988) describes a .10 correlation as small, .30 as medium, and .50 as large.
When estimating sample size, it is important to estimate conservatively. So the choice of an effect size should be near the lower/smaller size that you consider meaningful. Although .30 is considered a moderate effect size, a study may be designed to detect a smaller correlation, around .20 or .25. Then, logistics come into play. What size sample can you afford to get? Here, a researcher might settle on an effect size of .25 because it would be too difficult to get the larger sample size needed to detect a correlation of .20. The necessary sample sizes to detect different levels of r appear in Table 1 (Gatsonis and Sampson 1989:520). The first column contains the minimum correlation that can be detected and the second column contains the minimum total sample size necessary to detect it.
Necessary Sample Size to Detect a Given Effect Size for Simple Linear Regression, ANOVA (t-test), and χ2 Analyses (α = .05 and β = .20).
Note: ANOVA = analysis of variance.
Returning to the two examples, predicting infant birth weight from maternal smoking and depression from blood glucose, the needed sample size can be estimated with the table regardless of the means and standard deviations of the variables. If I truly believe that a correlation of .20 is a minimum value for a meaningful correlation, each study needs at least 193 people. The final estimate, however, is determined by an acceptable minimum correlation and what I can afford to get (money and time). Say the second study is part of a student dissertation and a sample size of almost 200 is too large for one student. She thinks she can get 100 interviews, but probably not 200 interviews. Previous studies suggest a correlation between blood sugar and depression scale scores is around .30. A sample size of 84 would have adequate statistical power (1 − β = .80) to detect a correlation of .30 or greater between blood sugar and depressive symptoms (two-sided α = .05, β = .20) if one exists. With 84–100 interviews, the student would be able to detect a moderate correlation and possibly include an additional independent variable.
In another example, a researcher wants to test the idea that older men in a hunting-based society would have better hunting efficiency and better yields than younger men and that skill and strategy would win out over better physical conditioning of younger men. The main variables are a man’s age (10–50 years) and his hunting productivity. A man’s age can be measured in years, and measurement of hunting might be measured with the number of animals killed and/or the relative weight of kills over a hunting season (Koster 2010). Hunting yield can also be estimated by aggregating community member’s judgments about hunting ability. The null hypothesis is that a man’s age and his hunting yield are not correlated. What size of correlation should he accept as meaningful? He assumes that since age and hunting experience are almost synonymous, the correlation between age and hunting yield should be very strong. Thus, he considers a sample size to detect a correlation of .70 or greater. At α = .05 and β = .20 (two-sided test), only 13 men are required to detect r ≥.70. Since it is best to estimate on the conservative side, he notes that a sample size of 19 would be sufficient to detect a correlation of .60 or greater and that 29 would detect a correlation of .50 or greater (α = .05 and β = .20). He decides that if he can get a sample of 13–29 men, he will be able to test his hypothesis with adequate power and possibly include an additional independent variable in his analysis.
Estimating Sample Size with One Categorical Variable and One Interval-scale or Continuous Variable: The t-test and the Eta (η) Coefficient
An ANOVA compares a categorical independent variable with an interval-scaled dependent variable. This is also called a one-way ANOVA, indicating only one independent variable. A special case of a one-way ANOVA occurs when the independent variable has only two categories. This comparison is often called a t-test, because the hypothesis test for difference between the two means uses the t probability distribution. The general problem is to see if there is a statistically significant difference in means between two groups or, alternatively, whether the association is significantly greater than zero.
The effect size or measure of association can be estimated with the equivalent of a Pearson correlation coefficient for nonlinear associations, the eta (η) coefficient. To see a visual display of the association, a scatterplot is used with the independent variable as the horizontal axis and the dependent variable as the vertical axis. Just as the Pearson r 2 measures the explained variance in the dependent variable, eta squared (η2) estimates the explained variance between the independent variable (the groups) and the dependent variable. With one dichotomous variable and one interval-scaled variable, Pearson r and η are equivalent. The range of η is from 0 to 1, with a larger value indicating a stronger association. Either a Pearson r or an η can be calculated in this case and both are available in most standard software packages. Because ANOVA is a linear model (like regression), hypothesis tests concerning whether the association (the correlation coefficient η) is significantly different from 0 are tested with an F or t distribution.
Effect size is often expressed as the proportion of standard deviations that one would like to detect: Δ = (Ȳ 1 − Ȳ 2)/σ, where Ȳ 1 and Ȳ 2 indicate the two group means and σ is an estimate of the within-group standard deviation (Hays 1963:329–32). However, the within-group standard deviation (σ w ) or the standard deviation of the dependent variable (σ y ) can be used (Browner et al. 2007), as they are equal under the null hypothesis. 1 The effect size can also be expressed as η = Δ/2. 2
The choice of an effect size is guided by the same principles and magnitude as for the Pearson correlation coefficient. The table contains information on necessary sample sizes for a two-group comparison (t-test). In the table, the third numeric column contains η coefficients (equivalent to the Pearson r), the fourth column contains the effect size translated to Δ as 2η = Δ, and the fifth column contains the group sample sizes calculated from Δ (from Hays 1963:330). 3 Note that these estimates assume equal group sizes.
You do not need to know the standard deviation beforehand, but if information on means and standard deviations are available, the effect size can be expressed as the difference between means. With α = .05 and β = .20, a sample size of 16 in each group can detect a difference of 1.00 standard deviation in means (η = .50, a large difference). So, if a variable has a mean of 100 and a within-group standard deviation of 10, a difference of 10 units between the two group means can be detected with 16 people in each group (e.g., 90 vs. 100). If the mean of the variable is 1,000 and the standard deviation is 50, a sample size of 16 will still detect a difference of one standard deviation between group means (e.g., 950 vs. 1,000). The beauty of this approach is that it doesn’t matter what the variable is or what the mean and standard deviations are, you can estimate the necessary sample size based on a meaningful difference expressed in terms of a proportion of the standard deviation.
Reconsidering maternal smoking and infant birth weight, assume that women are classified as smokers or nonsmokers and that infant birth weight is measured in grams. How many mother–infant observations are needed? To detect a similar effect size as mentioned earlier (r ≥ .20), equivalent to η ≥ .20, a comparison of two means (α = .05 and β = .20) requires a minimum sample size of 99 smokers and 99 nonsmokers (total = 198). We can also go to the literature for guidance on effect size. A study by Wang et al. (2002) reports average birth weights of 3,110 gm to nonsmokers (within-group standard deviation = 775) and 2,830 gm to smokers (within-group standard deviation = 799), a difference of −262 gm. The two within-group standard deviations can be averaged with either a simple average (787) or an average weighted by the group sizes (781).
The Wang et al. (2002) study suggests that we might expect to find a .33 (262/787) standard deviation difference, equivalent to η = .17. To detect η ≥ .15 (α = .05 and β = .20), 176 smokers and 176 nonsmokers are necessary to detect a difference in means of 236 gm (e.g., 3,300 gm in nonsmokers and 3,064 gm in smokers). A meta-analysis by Kramer (1987) reports an average U.S. infant birth weight of 3,300 gm and about −150 gm difference in birth weight between smoking and nonsmoking mothers. The Kramer (1987) study suggests a smaller difference (150/787 standard deviations or η = .10) and would require 395 in each group to detect a 150-gm difference or larger between nonsmoking mothers (3,300 gm average birth weight) and smokers (3,150 gm; α = .05 and β = .20). So, although we began with assumptions suggesting 99 people per group would be adequate, previous research suggests the effect size is small and such a study would need at least 176–395 per group to detect a significant effect.
In a study testing for differences in plant knowledge between urban and rural residents, a researcher plans to sample people from both locations and ask them to name as many plants as they can. Previous studies used 18 or 24 observations per comparison group (Mathez-Stiefel et al. 2012; Reyes-Garcia et al. 2006). Time in the field is expensive, and the investigator realizes that a nonsignificant difference with such a small sample size will be difficult to interpret. Because such small sample sizes can only detect a fairly large effect (η ≥ .45 or r ≥ .45, and η ≥ .40, respectively), a nonsignificant difference might only imply that there are not large differences in plant knowledge, but there could still be moderate differences. So, she plans the study to detect a moderate effect size.
At α = .05 and β = .20, a sample of 44 people from each site will be able to detect a .60 standard deviation difference in average list length (η ≥ .30). This sample size will also allow her to test whether list length varies by age of the informant: A total sample size of 88 people at each site also will be able to detect r ≥ .30 between age and list length (α = .05 and β = .20). She also plans to report the observed effect size with 95% confidence intervals. If she cannot get more than 22 households per site, especially at the rural site, she will consider a slightly more complex design and interview two adults per household.
Estimating Sample Size for Two Categorical Variables: The χ2 Test and the Phi (φ) Coefficient
The analysis between two categorical variables is assessed with a contingency table or χ2 analysis. A contingency table shows the cross-classification of two variables, their joint frequency distribution, and the number of observations in each category. The association between two categorical variables can be measured in various ways. In the special case where both variables are dichotomous, the degree of association can be expressed as a correlation coefficient in the range from 0 to 1, with the phi (φ) coefficient. The φ coefficient is equivalent to the Pearson r and it is because of this equivalence that we focus on φ as a measure of effect size. With two dichotomous variables, the exact probability of an outcome is given with the hypergeometric probability distribution and the cumulative hypergeometric probability is used in the Fisher’s Exact test or can be estimated with a χ2 probability distribution. Thus, hypothesis tests concerning the association between two categorical variables are usually made with the χ2 distribution.
Effect size for sample size estimation often focuses on the difference between two proportions but also can be expressed as a φ coefficient. Hypothesis tests involve testing whether two proportions are equal or equivalently whether φ = 0. The sixth to eighth numeric columns in the table provide sample size information for detecting a difference in proportions for different levels of φ. The sixth column contains φ, the seventh column contains the sample size when the two proportions are centered on .50, and the eighth column contains the sample size when the smaller of the two proportions (P 2) is .10 (formula and tables in Fleiss 1981:38–44). The sample size requirement to detect a difference in sample proportions (P 1 – P 2) of .20 with α = .05 and β = .20 is 107 people in each group when the proportions are near .50 (e.g., .40 vs. .60); 71 people are necessary in each group when the smaller proportion is .10 (e.g., .10 vs. .30). When comparing a specific difference between two proportions and no information is available to guide the actual choice of proportions, estimates should center near or on .50, since the standard deviation and thus the sample size requirements are largest when the sample proportions are near .50.
For the maternal smoking and infant birth weight example, mothers can be categorized as smokers/nonsmokers and their infants as low birth weight (<2,500 gm) or normal birth weight (≥2,500 gm). The research question is whether maternal smoking is associated with infant birth weight and the null hypothesis is that there is no association (φ = 0). Without prior information on the rate of low birth weight babies among women who smoke, we first consider effect sizes that parallel the example mentioned previously, where φ ≥.15 and φ ≥ .10. To detect a φ ≥ .15 a sample size of 186–219 is needed in each group. A sample size of 186 would detect a difference of .15 or larger between the two sample proportions if they are near .50 (e.g., .40 vs. .55) and a sample size of 219 would detect a .10 difference in proportions (e.g., .10 vs. .20 or .80 vs. .90) if the smaller proportion is near zero or the larger proportion is near one. Say we knew that the rate of low birth weight births was about 7–10% and we expected the difference between groups to be small (as shown earlier). For φ = .15, the researcher would design the study to detect a .10 difference (.10 vs. .20) with 219 smokers and 219 nonsmokers (α = .05 and β = .20). But to detect a .05 difference (φ = .10), 725 women would be necessary in each group to detect a difference between 10% and 15% (1,450 total sample size).
In another example, an investigator is interested in testing whether the folk illness nervios occurs more in women or men. Nervios is a Latin American folk illness often translated as “nerves.” Symptoms include crying, difficulty sleeping, shaking/trembling, sadness (and depression), being easy to anger, and feeling hopeless; treatments include trying to relax, taking sedatives, praying, and sometimes seeing a psychologist or psychiatrist (Baer et al. 2003). In an urban Mexican setting, nervios was reported by 65% of the sample and having nervios was strongly associated with gender: 76% of women and 42% of men reported nervios (Weller et al. 2008). The new study in a rural Latin American setting will compare various aspects of gender roles and stress between men and women, and one outcome variable will be the experience of nervios. The urban Mexican study observed a large difference in proportions (.76 – .42 = .34) and their φ was .33. To detect a difference of at least .30 and φ ≥ .30 (α = .05 and β = .20), the study needs at least 49 men and 49 women. This sample size would detect a difference in nervios of 50% in men and 80% in women, or 20% in men and 50% in women (α = .05 and β = .20).
Discussion
An adequate sample size is necessary to detect a meaningful effect. In planning a new study, it is important to declare the size difference you want to detect and give the rationale for that choice. A meaningful effect size, however, is subjective and relies on judgment and previous research. The correlation coefficient offers a useful metric for considering effect sizes across different types of analyses. Although effect sizes are often expressed as the difference in means divided by the standard deviation for a t-test or the difference in proportions for a χ2 test, these can be translated into equivalent effect sizes using the correlation coefficient. Other measures have been proposed (Lachin 1981), but the correlation coefficient is familiar, interpretable, and translates readily across analyses. Also, using the correlation coefficient to link concepts across analyses shows that similar sample sizes detect similar effect sizes across analyses: For r ≥ .20, a sample size of 193 is needed (α = .05 and β = .20); for η ≥ .20, a total sample size of 198 is needed; and for a φ ≥ .20, a total sample size of 214 to 226 is needed.
Estimating sample size requirements is important for planning a study and determining feasibility. A general principle is that to detect smaller differences, a larger sample size is needed. Although this article only presents sample sizes for conducting basic tests between two variables, this information should be helpful for planning many types of research projects. Sample size can be estimated without any prior information for the mean and standard deviation of the variables involved. This is especially helpful when using derived variables like cultural competence or cultural consonance (Weller 2007) or other variables that are sample dependent and cannot be estimated prior to the study. Although sample size estimation does not need the mean and standard deviation of the variables, prior information can be useful in thinking about and expressing the difference one is trying to detect. This information can usually be found in prior publications or can be estimated from a pilot study or pretest.
In reality, projects often have more than one independent variable and additional variables increase sample size requirements. For example, a sample size of 84 would have adequate statistical power (1 – β = .80) to detect a correlation of .30 or greater between blood sugar and depressive symptoms (two-sided α = .05, β = .20). When a second independent variable is added (e.g., age), the minimum sample size increases to 101, and when a third independent variable is added (e.g., gender), the sample size increases to 115 (Elashoff 2007). Similarly, a study designed to detect a correlation of .20 or greater (n = 193) with one independent variable would need at least 292 people to accommodate four independent variables. For regression, the sample size needs to be increased for each additional independent variable by approximately five people when detecting a large effect (r = .50), 15 (10–20) people for a moderate effect (r = .30), and 30 (20–40) people for a smaller effect (r = .20).
Here, sample estimation only considered cases with equal group sizes for the t and χ2 tests. Total sample size is minimized and statistical power maximized with equal group sizes. When studying differences between men and women, a visible characteristic that is distributed 1:1 in the population, the sampling protocol can easily select equal numbers for interviewing. When the distribution of categories in the independent variable is not evident or one group is relatively rare, this may be more difficult. Since the smaller of the two groups limits the power of a statistical test, a quick estimate can be obtained from the size of the smaller group. If a two-group design with 176 per group (η = .15) were instead distributed at a 3:1 ratio (25% prevalence of smoking: 264 nonsmokers and 88 smokers), the detectable association would actually approach η ≥ .20 to η ≥ .25 (for ni = 88).
Sometimes sampling strategies will purposefully oversample one group by two to four times to gain statistical power. When testing for a difference between proportions (10% low birth weight for nonsmokers and 20% for smokers, α = .05, β = .20), the minimum sample size of 219 per group with equal group sizes changes to 316 nonsmokers and 158 smokers for a 2:1 group size, and 413 and 138 for a 3:1 group size (Elashoff 2007). Detailed information for estimating sample size with unequal group sizes is available in other sources (e.g., see Fleiss 1981:44–46 for the χ2 test).
This article offers an introduction to the process of sample size estimation and study planning with consideration for effect size. More advanced sources and an expert 4 should be consulted for more complicated study designs, including different values of α and β, more variables, unequal group sizes, and repeated measurements on the dependent variable (including a matched-pairs t-test). Sample size estimation is exactly that—an estimation. Although a few main factors must be considered (α, β, and effect size), the investigator must negotiate a final sample size from what is a meaningful effect size and what is the maximum sample size that can be obtained with available resources. Also, when reporting results, it is important to also report the effect size—the correlation or the observed difference between means or proportions.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The author acknowledges support from the University of Texas Medical Branch.
