Abstract
The purpose of the meta-analysis was to address the varied results within the area of singing ability research by statistically summarizing the data of related studies. Across 34 studies, analyses yielded an overall mean effect size for instruction of g = 0.43. Studies were limited to pretest-posttest or posttest-only quasi-experimental designs. The largest overall study effect size across categorical variables included the effects of same and different discrimination techniques on mean singing score gains. Overall mean effects by primary moderator variable ranged from trivial to moderate. Feedback yielded the largest effect regarding teaching condition, 8-year-old children yielded the largest effect regarding age, the Boardman assessment measure yielded the largest effect regarding measurement instrument, and song accuracy yielded the largest effect regarding measured task. Regarding gender, boys and girls improved similarly from singing interventions across studies. Implications for singing instruction pertain to the importance of intervention, especially between the ages of 5 and 8. Results from the meta-analysis have highlighted a tendency for singing interventions to improve singing ability more than traditional song-singing and more than no music instruction at all. Results from the meta-analysis have also highlighted the importance of self, teacher, and computer feedback in the development of singing.
Research regarding the effects of instruction on the development of singing abilities in childhood has yielded conflicting results. Several studies have indicated that instruction improved singing ability (Miyamoto, 2005; Paney & Kay, 2014; Roberts & Davies, 1975). Positive effects on singing ability have been found for feedback (Porter, 1977; Welch, Howard, & Rush, 1989), the presence and type of singing instruction (Apfelstadt, 1984; Miyamoto, 2005; Roberts & Davies, 1975; Rutkowski, 1996), and pattern instruction (Jarjisian, 1981; Porter, 1977). Conflicting effects, however, have been found for vocal model (Green, 1990; Small & McCachern, 1983; Yarbrough, Green, Benson, & Bowers, 1991; Yarbrough, Bowers, & Benson, 1992), use of singing voice (Persellin, 2006), use of accompaniment (Atterbury & Silcox, 1993; Guilbault, 2004), and performance context (Goetze & Horii, 1989; Moore, Chen, & Brotons, 2004; Yarbrough et al., 1991). Age and grade have yielded conflicting results regarding singing ability across grade levels (Moore, Fyk, Frega, & Brotons, 1996; Gault, 2002; Goetze & Horii, 1989; Hornbach & Taggart, 2005; Miyamoto, 2005; Welch, Sergeant, & White, 1997; Yarbrough et al., 1992; Yarbrough et al., 1991). Gender differences have also yielded conflicting results (Apfelstadt, 1984; Goetze & Horii, 1989; Hedden & Baker, 2010; Moore, 1994; Moore et al., 1996; Paney & Kay, 2014; Yarbrough et al., 1992).
Given the breadth of and disagreement within singing ability research, the meta-analytic process may provide an empirical synthesis that utilizes primary data to illustrate the extensive body of information. Statistically synthesizing past singing ability research might inform future singing ability research and encourage the dissemination of transparent results. Therefore, the purpose of the study was to address the varied results by statistically summarizing the data of related studies through meta-analysis. Toward this purpose, three research questions were developed: (1) what is the overall statistical effect of instructional characteristics on singing ability in children ages 5 through 11 years old?; (2) what are the statistical effects within and across the following primary moderator variables: instruction, age, gender, and measurement instrument?; and (3) what are the statistical effects of study characteristics within and across the following secondary moderator variables: publication source, publication year, population, research design, and treatment period?
Method
The population of the meta-analysis included 34 studies that measured the effects of instruction on singing ability. An initial search through literature yielded 168 possible studies. Application of the inclusion and exclusion criteria resulted in the exclusion of 134 studies. Refer to Tables 2, 3, 4, and 5 within the Supplementary Materials online for a selection of excluded article citations.
Three experts in the field were contacted to establish content validity of the inclusion and exclusion criteria list and initial list of studies. The experts provided suggestions regarding the overall content of the analysis and specific suggestions regarding inclusion and exclusion criteria. See Table 1 below for the final inclusion and exclusion criteria statements.
Inclusion/exclusion critera statements.
Inclusion and exclusion criteria
Inclusion and exclusion criteria were outlined to define and narrow the literature search. Inclusion criteria included nine statements, and exclusion criteria included seven statements. The justifications for inclusion and exclusion statements were based upon the purpose of the meta-analysis, existing singing ability literature, recommendations from published meta-analyses, and texts regarding meta-analytic procedures.
Three inclusion and exclusion statements were written to ensure generalizability to elementary music educators and classrooms. First, to be included, studies must have investigated singing ability of children within an age range typically found in elementary schools (5 through 11 or 12), excluding children younger than kindergarten and older than sixth grade. Second, because classrooms have a mixture of neurotypical and neuro-atypical children, the meta-analysis only included studies reflecting that population; it excluded special populations such as audition choirs and neuro-atypical children. Third, the singing analysis must have required participants to sing independently for at least part of the data collection because children may require opportunities to sing independently for singing ability development (Hedden, 2012); studies without independent singing as part of data collection were excluded.
Because the meta-analysis sought to address singing ability and instruction, included studies were limited to the effects of instruction on singing ability. Other constructs were permitted if instruction was also involved and if singing ability was measured through singing performance. Studies with children who spoke tonal languages were excluded due to variables that may attribute to singing development having more to do with language acquisition than with singing instruction. Studies regarding language acquisition were also excluded.
Studies qualified for inclusion if they reported necessary descriptive statistics. Means, standard deviations, and sample sizes are needed to calculate effect sizes (Card, 2012). F-test or t-test statistics are also acceptable when means, standard deviations, and sample sizes are not readily available. If the study satisfied all other criteria, the study author was contacted to request original data.
Meta-analyses, like most quantitative analyses, should be carried out (and reported) in a matter that would enable approximate replication. For replication purposes, and due to the researcher’s proximity to interpreters, only studies written in English (or translated to English) were included in the meta-analysis. For replication purposes, studies within the meta-analysis were limited to dissertations as well as articles published in peer-reviewed journals.
Limiting research designs may minimize variability of effects accounted for by research design. Therefore, the audience could focus on the effects of instruction, specifically. Three statements were written to indicate included or excluded research designs.
Literature search and coding
A literature search was performed to pool studies regarding instructional effects of singing ability. Key search phrases were derived from commonly observed terminology across published literature. Phrases included singing ability child, singing achievement child, vocal development child, teaching singing child, and singing development child. Initial searches utilized Summon. Additional articles were found through backward searches (gleaning articles from existing reference lists), forward searches (searching for articles that have cited seminal works or seminal researchers), and by contacting several seminal researchers in the field.
Articles were coded, and data within the categories were used for the moderator analyses and study quality. To check for intracoder reliability of coding procedures in the main study, the researcher recoded all of the studies two months after initial coding. The analysis yielded a tenable agreement rate of 91.32%. The agreement rate of 70% or higher has been commonly cited as acceptable across studies (Card, 2012).
In order to obtain an overall weighted mean effect size for the main study, the researcher first calculated effect sizes based on study subgroups (e.g., gender, grade level, and instruction). Hedges’ g effect sizes were calculated from the reported means and standard deviations of instruction effect for the following designs: (a) change from pretest to posttest, and (b) between-group effects. Cutoff numbers for small sample adjustments to effect sizes varied, therefore an equation recommended by Lipsey and Wilson (2001) was applied to all effect sizes calculated with groups smaller than 20 participants.
Alternate equations and Wilson’s Practical Meta-Analysis Effect Size Calculator (Lipsey & Wilson, 2001) were used to determine effect sizes in studies where means and/or standard deviations were not reported. In several studies, effect size was calculated using the F statistic, χ2 values, independent t-scores, or mean change between pretest and posttest scores. Attempts were made to obtain raw scores and standard deviations when sufficient data were not available.
Statistical methods
The current meta-analysis had three consecutive parts that encompassed multiple calculations. The first part yielded an overall weighted mean effect size across all studies within the current meta-analysis. The second part investigated heterogeneity among effect sizes. The third part investigated moderator variables.
SPSS 22 was used to compute group standard deviations when data were not provided. Excel was used to calculate equations in the current meta-analysis where warranted, specifically regarding calculating p-values and sums. All other calculations utilized a TI-83 graphing calculator.
Hedges’ g was used in the current meta-analysis due to inequivalent group sizes within studies. Cohen (1988) provided cutoffs for small, medium, and large d effect sizes of 0.2, 0.5, and 0.8, respectively. Effect sizes lower than 0.2 have been labeled as trivial (Ellis, 2010). The cutoffs have also been commonly generalized to Hedges’ g effect sizes (Card, 2012; Ellis, 2010).
The meta-analysis utilized a conditional inference model with fixed-effects procedures. Fixed-effects procedures, “assume that all studies in the meta-analysis share a common (true) effect size” (Borenstein, Hedges, Higgins, & Rothstein, 2009, p. 63), and result generalization should be limited to the studies included in the analysis (Card, 2012). The statistical model and moderator variables (instruction, age, gender, measurement instrument, publication source, publication year, population, research design, and treatment period) were selected a priori based on a conceptual assumption that the effects within the current meta-analysis would be heterogeneous. Therefore, results of the current study may be generalized to the population of the current studies and “to a population of studies with similar characteristics than those represented in the meta-analysis” (Huedo-Medina, Sanchez-Meca, Marin-Martinez, & Botella, 2006, p. 194). Primary moderator variables were selected a priori due to the variety of representation in literature. The secondary moderator variables were also selected a priori for the purpose of exploring validity and bias issues.
Results
The population of the current meta-analysis was comprised of 34 studies with a combined sample size of N = 5,497 participants who ranged in age from 5 to 11 years (M = 7.38, SD = 1.86). Studies represented all regions of the United States, Canada, and England. Refer to Table 6 in the online supplementary materials for information regarding each study grouped by teaching condition and instructional details listed for treatment and control groups. Refer to Tables 7 and 8 in the online supplementary materials for information regarding all other primary and secondary moderators by study.
Within the 34 studies, there were 433 unique effect sizes calculated between and within groups. In other words, there were 433 group comparisons made across the 34 studies. Unique group comparisons were also important when grouping effect sizes into moderator variables.
With regard to the first research question, the current meta-analysis yielded an overall small weighted mean effect size for instruction (ES—= 0.43, 34 studies, 433 unique effects, 95% CI [0.42, 0.44]). Mean effect sizes across the studies ranged from large (g = 0.71, p < .05, 95% CI [0.63, 0.79]) to trivial (g = 0.03, p > .05, 95% CI [−0.10, 0.17]). The largest overall study effect sizes across categorical variables included Jordan-DeCarbo’s (1982) investigation regarding the effects of same and different discrimination techniques on mean score gains, favoring the treatment group (gchange= 0.71, p < .05, 95% CI [0.63, 0.70]) and Porter’s (1977) investigation regarding the effects of multiple discrimination training on mean score gains, favoring the treatment group (gchange= 0.69, p < .05, 95% CI [0.46, 0.92]). The smallest overall study effect sizes across categorical variables included Rutkowski’s (1996) investigation regarding the effects of small group and individual participation in singing activities on mean scores, favoring posttest treatment scores (g = 0.03, p > .05, 95% CI [−0.10, 0.17]) and Klinger, Campbell, and Goolsby’s (1998) investigation regarding the effects of rote versus immersion song teaching on mean scores (g = 0.04, p < .05, 95% CI [−0.59, 0.67]). Refer to Table 1 in the online supplementary materials for a summary of overall mean effects, within-group effects (gch), and between-group effects (g).
A Q test for heterogeneity of effect sizes (Card, 2012) and an I2 index of the magnitude of heterogeneity indicated that the study effect sizes were significantly heterogeneous (Q = 1039.90, df = 33, I2 = 96.83%, p < .0001). Moderator variables selected a priori were evaluated to observe effect sizes by study characteristics.
With regard to the second research question, primary moderator variables included instruction, age, gender, and measurement instrument. The instruction moderator variable included seven subgroups in which the population of studies were categorized: (a) feedback, (b) vocal model, (c) presence and type of singing instruction, (d) use of accompaniment, (e) pattern instruction, (f) performance context, and (g) literature dissemination. The age moderator variable included seven subgroups in which the studies were categorized: (a) 5-year-olds, (b) 6-year-olds, (c) 7-year-olds, (d) 8-year-olds, (e) 9-year-olds, (f) 10-year-olds, and (g) 11-year-olds. The gender moderator variable included three subgroups in which the population of studies were categorized: (a) girls, (b) boys, and (c) between gender, favoring girls. The third category, between genders, favoring girls, referred to differences in group scores between girls and boys. Within the context of the current study, the girls’ socres were higher, therefore, the effect size favored the girls. The measurement instrument moderator variable included six subgroups in which the population of studies were categorized: (a) Boardman, (b) acoustic measures of pitch accuracy, (c) researcher as acoustic measurer of pitch accuracy, (d) researcher as acoustic measurer of song accuracy, (e) researcher as acoustic measurer of pattern and interval accuracy, and (f) Singing Voice Development Measure (SVDM). The measured task moderator variable included three subgroups in which the population of studies was categorized: (a) song accuracy, (b) pitch accuracy, and (c) pattern accuracy. See Table 2 for a detailed account of effect sizes by primary moderator variable subgroup.
Primary moderator variable within-group, between-group, and overall mean effect sizes.
Denotes statistical significance (p < .05).
With regard to the third research question, a secondary moderator variable analysis within each of the three categorical moderator variables yielded overall effect sizes within each categorical subgroup: (a) publication source, (b) population, and (c) research design. A secondary moderator variable analysis within each of the two continuous moderator variables yielded overall correlation coefficients and t-test statistics within each continuous subgroup: (a) publication year and (b) treatment period. Between-group effect sizes and within-group effect sizes were omitted from secondary moderator variable analyses as the secondary moderators described the study demographics as opposed to the instructional characteristics of the primary moderators. See Table 3 for a detailed account of effect sizes by secondary moderator variable subgroup. See the online supplementary materials for a more in depth discussion of secondary moderator variables.
Secondary moderator variable within-group, between-group, and overall mean effect sizes.
Denotes statistical significance (p < .05).
The two continuous moderator variables, publication year and treatment, period were analyzed to determine relationship strength across study weighted mean effect sizes. Publication years across the studies ranged from 1977 to 2014 (N = 34, SD = 11.14, Mdn = 1992, Mode = 1991), and treatment periods across the studies ranged from one week to 36 weeks (N = 34, M = 14.59, SD = 10.01, Mdn = 12, Mode = 10). A Pearson r correlation coefficient yielded a weak, negative relationship between publication years and corresponding effect sizes (r = −0.13). A Pearson r correlation coefficient also yielded a weak, negative relationship between treatment periods and corresponding effect sizes (r = −0.13). T-test results indicated that there were no significant relationships between effect sizes and variables (t = −0.74, p = .23).
Statistical significance
Statistical significance was calculated using the Wald test from standard errors, effect sizes, and z scores (Card, 2012). An effect size was significant if the product was greater than a z score of 1.96. Non-significant effects are still holistically meaningful. Therefore, interpretation based solely on statistical significance is strongly discouraged.
Publication bias
There may have been evidence from publication source subgroups indicating that the results were impacted by publication bias within the current meta-analysis. Effects across publication source yielded a larger effect for studies in the “other” category (g = 0.55, p < .05, 95% CI [0.53, 0.57]) than the subsequent three subgroups. Studies within Journal of Research in Music Education (g = 0.41, p < .05, 95% CI [0.37, 0.45]) and the Bulletin of the Council for Research in Music Education (g = 0.41, p < .05, 95% CI [0.32, 0.50]) yielded identical effect sizes and were both larger than the overall doctoral dissertation effects (g = 0.28, p < .05, 95% CI [0.27, 0.30]).
Discussion
Comparing the current study with other music-education-specific meta-analyses provided an external context that enabled the following conclusion: a “small” effect size can be regarded as effective. Therefore, instructional interventions may be effective to the development of singing ability with young children. Larger sample sizes and attention to a priori power analyses may benefit singing ability research and researchers, specifically regarding the detection of differences within and between groups.
Singing instruction may be seen as an overall effective means of improving singing ability scores in children between the ages of 5 and 11, specific to the studies and participants represented within the current meta-analysis. The difference between treatment group gain scores and control group gain scores may indicate that singing instruction is more effective than what has been commonly referred to as a traditional curriculum, or one which includes song-singing without vocal development guidance. Additionally, the difference may indicate that singing instruction is more effective than no music instruction at all.
Feedback was an overall effective means of improving singing ability scores across studies. Within the context of the current meta-analysis, feedback yielded the largest effect size of all teaching conditions. Several studies utilized feedback as a condition through computer programs. There was a greater effect for practicing singing accuracy with computer programs than receiving traditional instruction in third grade. Practicing fewer times (between one and five) with computer programs yielded a greater effect on mean scores than practicing more than five times with the computer programs. Additionally, teacher feedback combined with computer feedback had a greater effect than computer feedback alone. Specific feedback was also found to be effective without the use of a computer, especially when compared to generic, unspecific feedback.
Implications for teaching may be made with regard to feedback. Children are living in an age filled with technology and information. Children have access to information beyond that which was available when their teachers and parents were young. If children have access to information outside of the music classroom, perhaps there would be a benefit to having access to information within the music classroom. Information within the music classroom regarding the development of singing ability can be disseminated from teacher to student and from technology to student.
The type of vocal model influenced the improvement of singing ability scores across studies. Differences between vocal models yielded a medium overall and within-group effect size and a small effect size for between-group scores. Within the context of the current meta-analysis, the effect of a vocal model was nearly as large as the effect of feedback. Research regarding vocal models was represented within the current meta-analysis by comparisons of male models, female models, and whether the model sang for or with the students. Within the current meta-analysis, a female model yielded a greater effect on mean scores than a male model. The effect within the current meta-analysis aligns with previous results (Goetze, Cooper, & Brown, 1990; Hermanson, 1971; Simms, Moore, & Kuhn, 1982; Small & McCachern, 1983; Yarbrough et al., 1991). In male models, singing with either a comfortable range or falsetto voice both yielded similar effects on mean scores, although a comfortable range yielded higher mean scores than singing with a falsetto voice.
Whether or not a teacher sang for students, with students, or for and with students may make a difference regarding the development of singing ability. Although groups yielded similar effect sizes within the current meta-analysis, a combination of singing for and with students yielded a greater effect on mean scores than singing only with students. Additionally, singing only with students had a greater effect than singing only for students. These findings were based on one study; replications of the study are needed to strengthen the evidence of the finding given that they are contradictory to common pedagogical recommendations.
Overall, students may benefit more from a female vocal model than a male vocal model and more from a male model singing in his natural range than singing in falsetto. Singing abilities, however, should develop with proper instruction regardless of whether the vocal model is a male or female. Green (1990) suggested that neither adult model was as effective as a child vocal model. Regardless of model, consequences of quality instruction may surpass consequences of vocal modeling. Perhaps there is a need to encourage alternate views of instruction including gender-specific recommendations and gender-neutral recommendations.
The presence and type of singing instruction yielded a small effect size for overall, within-, and between-group scores. Pretest to posttest scores yielded the largest effects for instruction that included discrimination training. Discrimination training included an element of differentiating between tones that were the same or different. Vocal coordination instruction yielded a medium effect for vocal development in elementary music and a small effect regarding range and higher pitches.
The use of accompaniment yielded a small effect size for overall, within-, and between-group scores. Whether or not piano accompaniment influences singing ability was investigated across four studies in the current meta-analysis. Gain scores were similar for groups who received instruction with and without accompaniment, although a slightly greater effect was found for groups without accompaniment overall. Previous research aligns with this finding (Guilbault, 2004). Recommendations for teachers include limiting the use of accompaniment until vocal accuracy is consistently high for most students.
Pattern instruction yielded a small effect size overall and for between-group scores, but a trivial effect size for within-group scores. Pattern instruction has been used to provide children aural practice with the tonal vocabulary found in children’s literature while focusing on two to three pitches at a time. Although pattern instruction yielded a positive effect, pattern instruction did not yield greater effect sizes for treatment conditions with pattern instruction than a music curriculum without pattern instruction. There may be, however, differences regarding which patterns were most effective. A combination of pentatonic and diatonic patterns yielded a greater effect size than pattern instruction limited to diatonic patterns or limited to pentatonic patterns.
Pattern instruction can be a valuable tool for teachers who enjoy pattern instruction and who develop the skills to implement pattern instruction effectively. A combination of pentatonic and diatonic patterns may be preferable to pentatonic or diatonic alone. Therefore, it may be beneficial for the literature used to also include pentatonic and diatonic patterns rather than be limited to all diatonic songs or all pentatonic songs.
Performance context yielded a trivial effect size overall, and for between-group scores, as well as a small effect size for within-group scores. Within the current meta-analysis, Performance context included studies where how (or what) children sang was an instructional condition. Teaching conditions included individual versus group singing, response modes, movement while singing, and song range. There was a greater effect size for instruction that included individual and small-group singing than instruction limited to classroom singing. The use of songs with words in an elementary curriculum had the greatest effect for low-scoring kindergarten and first grade students when compared to singing most songs without words. The use of movement in a music curriculum had a greater effect on accuracy scores than without movement, and whether the movement was fluid or beat-centered yielded similar results. This finding aligns with previous research where 5- and 6-year-olds demonstrated improved vocal accuracy when associating movement with pitch levels (Liao, 2008). Regarding range, the song range in which literature was disseminated had a greater effect, favoring literature of a restricted singing range (from C3–B3) over literature of a higher singing range representative of music textbook literature.
Literature dissemination yielded a trivial effect size overall, and for between-group scores. Findings should be interpreted with caution regarding the literature dissemination subgroup. Two studies constituted the subgroup, and neither study included a pretest.
Regarding the current meta-analysis, singing instruction had the greatest overall effect on group scores for 8-year-old (third grade) children. After 8-year-old children, the second largest and subsequent effect sizes were with 9-year-olds, 10-year-olds, 7-year-olds, and 6-year-olds. Singing instruction had the lowest effect for 5-year-old (Kindergarten) children and 11-year-old children. This finding aligns with past research (Bentley, 1969; Gould, 1969). Children enter the public school system with varying background experiences (Welch, 2015) and generally spend much of Kindergarten exploring vocal timbres and foundational vocal development. The current meta-analysis also demonstrated a trend that aligns with previous research: accuracy tends to improve most across the ages of 5, 6, 7, and 8.
Children in Kindergarten through third grade may benefit from an elementary music program that prioritizes vocal development. If there is a social shift and decrease in accuracy scores after third grade, great care must be taken before third grade to build a respectful, safe classroom environment to negate or cushion the shift. Teachers may be doing so regardless, but doing so with the intention of vocal development in mind may offer a new perspective on building trust and skills in an elementary music program.
Within the current meta-analysis, singing instruction was approximately as effective for boys as it was for girls, and the difference between those effects was not statistically significant. It may be beneficial to encourage an alternate view regarding the existing trend of comparing male singers to female singers as an independent variable. Both groups may benefit from singing instruction. It seems that conflicting results may be attributed to other variables, such as student attitude toward music or student attitude toward the teacher. The interaction between a researcher’s interest in gender differences and a child’s vocal insecurities might be problematic for all involved. Rather than demonstrating concern regarding differences between genders, boys and girls alike may benefit from being compared to themselves. Having knowledge of one’s own singing development may encourage a sense of empowerment over one’s unique ability.
Implications for classroom teachers may be made with regard to research design. The pretest-posttest control group design (referred to by Campbell and Standley, 1963, as the nonequivalent control group design) controls for more threats to internal validity than when researching without a pretest or without a comparison group. When implementing new strategies for vocal development, teachers may see the most easily interpreted benefits by making a limited number of small changes in instruction with one or two intact classes across a predetermined time period. The other classes could serve as control groups and should receive similar instruction without the added strategies or small changes. A pretest and posttest should be chosen that are efficient, accessible, and will provide information specific to the informal research questions the teacher may be asking. The Singing Voice Development Measure has been most commonly used for investigating stages of vocal development, and computer programs are highly recommended if vocal accuracy is the primary goal.
An analysis of the moderator variable, measurement instruments, indicated a threat to construct validity by the frequency of tests and moderate heterogeneity of the subgroup. All of the effects, however, were generally similar with the exception of the Singing Voice Development Measure. The similar effect sizes may be indicative of differences between effects and can be attributed to variables other than what measurement instrument was used.
Across the 34 studies that measured vocal accuracy, nearly all were unique measurement instruments. Many more exist within research outside of the scope of the current meta-analysis. Measurement instruments were grouped by the designer of the instrument, who/what assessed the stimulus, and specific characteristics that were measured. Most were rating scales, although the acoustic measures evaluated pitches based on intonation. Of the 27 instruments, over a third included the use of singing voice as a variable, and about one-fifth included melodic contour. All measures assessed a specific type of accuracy including pitch, pattern, interval, or song.
The breadth of instruments created to measure singing ability has been addressed across literature reviews (Goetze et al., 1990; Salvador, 2010). The frequency of instruments may be indicative of the lack of consensus in defining singing ability or the lack of prior knowledge regarding established instruments at the onset of a study. The latter may be more relevant given the similarities across measures within the current meta-analysis.
Implications for future research
Original data are needed to continue meta-analyzing music education research accurately. Future meta-analyses may benefit from the publication of raw data sets within dissertations. Writing a dissertation provides young researchers with an opportunity to report specific information that might not otherwise be reported within journal articles; there are fewer publication limitations with dissertations.
Future research may benefit from consistently publishing descriptive statistics in music education research. Publishing overall descriptive statistics has been a common practice. When analyzing multiple independent variables, the inclusion of group sample sizes, group means, and group standard deviations with published results would benefit researchers involved in replication research and meta-analytic research.
Future research may also benefit from consistent reporting of effect sizes and confidence intervals. Out of the studies included in the current meta-analysis, one study reported d effect sizes, but only for groups with significant differences. To encourage growth of meta-analytic procedures within music education research, it would be beneficial for researchers to include effects for both significant and non-significant results. Without effect size and confidence intervals, information may be lost or go unreported within studies. Future research may also benefit from utilizing repeated-measures more often, since between-group comparisons yield higher threats to validity and smaller effect sizes.
Singing is important. Singing instruction is meaningful. The current meta-analysis has supported these beliefs and offered suggestions for future instruction. The current study has also suggested questions for future research and perhaps outlined problematic areas within singing ability research and research methodology within the field of music education.
Footnotes
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
