Abstract
Objective:
The objective was to conduct research synthesis for the U.S. Army on the effectiveness of two error prevention training strategies (training wheels and scaffolding) on the transfer of training.
Background:
Motivated as part of an ongoing program of research on training effectiveness, the current work presents some of the program’s research into the effects on transfer of error prevention strategies during training from a cognitive load perspective. Based on cognitive load theory, two training strategies were hypothesized to reduce intrinsic load by supporting learners early in acquisition during schema development.
Method:
A transfer ratio and Hedges’ g were used in the two meta-analyses conducted on transfer studies employing the two training strategies. Moderators relevant to cognitive load theory and specific to the implemented strategies were examined. The transfer ratio was the ratio of treatment transfer performance to control transfer. Hedges’ g was used in comparing treatment and control group standardized mean differences. Both effect sizes were analyzed with versions of sample weighted fixed effect models.
Results:
Analysis of the training wheels strategy suggests a transfer benefit. The observed benefit was strongest when the training wheels were a worked example coupled with a principle-based prompt. Analysis of the scaffolding data also suggests a transfer benefit for the strategy.
Conclusion:
Both training wheels and scaffolding demonstrated positive transfer as training strategies. As error prevention techniques, both support the intrinsic load–reducing implications of cognitive load theory.
Application:
The findings are applicable to the development of instructional design guidelines in professional skill-based organizations such as the military.
Keywords
Introduction
When one is learning difficult tasks and skills, it is often essential to be trained on easier versions of the task (Schneider, 1985). The reason for this is straightforward. When the learner is overwhelmed by the demands of trying to perform the complex task, he or she simply does not have the cognitive capacity or resources available to focus on learning: discerning the rules underlying the task, encoding and rehearsing the new information about the task for storage in long-term memory, and processing feedback about what worked and what did not. The notion that the demands for performance and for learning may be distinct (and in competition with each other) is the hallmark of cognitive load theory (CLT; Paas, Renkl, & Sweller, 2003; Paas & van Gog, 2009; Sweller, 1988).
In CLT, the resources associated with performance, demanded by task complexity, are called intrinsic load, which is closely related to the number of elements of the task to be learned and the interaction between those elements. Those cognitive resources associated with learning are called germane load (Kalyuga, 2011; Kirschner, Kester, & Corbalan, 2011; Paas, van Gog, & Sweller, 2010). Germane load may be closely related to the concept of deep processing in working memory (Craik & Lockhart, 1972). Within CLT, a third class of resource demand, extraneous load, defines those resources demanded by neither germane nor intrinsic load, for example, a poor interface or distracting characteristics of the learning environment (Sitzmann, Ely, Bell, & Bauer, 2010; Wickens, Hutchins, Carolan, & Cumming, 2012a).
An important prediction of CLT is that learning can be improved by removing sources of both intrinsic and extraneous load, to avail ample resources for germane load. For example, two related training strategies based on this logic are to train complex tasks in parts (i.e., part-task training) and to simplify the task early on and add complexity only as learning progresses (i.e., increasing difficulty) because the simplified components require fewer resources (Wickens, Hutchins, Carolan, & Cumming, 2012b; Wightman & Lintern, 1985).
CLT and Error Prevention During Training
In the current research, we examined two training strategies, training wheels (TW; e.g., J. Carroll, 1990; Carroll & Carrithers, 1984a, 1984b) and scaffolding (e.g., Pea, 2004), designed to reduce both intrinsic and extraneous load, in the hope of maximizing resources available for germane load. The common feature of both of these is the aim to prevent errors during skill acquisition. Such prevention is argued to have two benefits in the context of CLT: (a) It can reduce the complexity of learner choices of “what to do next” and the cognitive load of figuring out how to do it when the behavior of the learner is somewhat constrained and (b) it can prevent the commitment of catastrophic errors (such as deleting an entire file while learning a word-processing system) that can create a large amount of extraneous load, in which the learner is essentially thrashing or “floundering” (Kirschner, Ayres, & Chandler, 2011, p. 101) early in learning.
Although clearly two instances of the same functional goal (i.e., preventing errors), TW and scaffolding differed in the literature with the emphasis on the error-preventative mechanism, as we contrast next. In this sense, error prevention might be thought of as the umbrella strategy under which the methods of TW and scaffolding are two instances, each designed to reduce intrinsic and extraneous load, therefore maximizing resources available for germane load by preventing errors early in learning. Our approach for examining the transfer effects of these strategies during training is the meta-analysis to be described in more detail later.
Training Wheels
In the literature on error prevention strategies, TW has been exemplified as either lockouts or worked examples. As lockouts, or performance constraints, actions not relevant to the current phase of learning are “locked out,” or made unavailable. Consequently, learners do not invest resources in cognitive processes irrelevant to their current phase of learning. As phases are successfully completed by the learner, the relevant locked-out actions and processes are incrementally made available. Lockouts are often applied to computer tasks, in which interface access to advanced or irrelevant functions is restricted (J. Carroll, 1990; Carroll & Carrithers, 1984a, 1984b; van Merriënboer & Sweller, 2005). As worked examples, the basic idea is to provide learners with a worked-out solution state and/or partially completed steps toward the solution state. By providing the correct strategies early in learning, learners avoid using weak, inefficient, or otherwise inappropriate strategies to solve problems or perform tasks (van Gog & Rummel, 2010). Our review of the experimental literature on TW revealed an emphasis primarily focused on study manipulations of the degree of constraint in the error-preventative mechanism (e.g., full lockouts, worked-out examples, partial completion).
Scaffolding
Scaffolding is a term used for providing various types of instructional support to promote learning when concepts and skills are being first introduced. Scaffolding allows trainees to perform a task beyond the learner’s skill level by providing learning support that is removed as the learner becomes more proficient. Types of scaffolding include modeling advanced task solutions, simplifying or reducing the degrees of freedom of a task, and focusing the learner’s attention on relevant features of the task (Pea, 2004). In contrast to TW, the experimental literature on scaffolding revealed a common emphasis on manipulations involving the removal schedules of the error-preventative mechanism (i.e., fixed vs. adaptively driven schedules).
In spite of this slightly different emphasis, both TW and scaffolding strategies have much in common, and it is our intent not to contrast these two techniques but rather more to identify the convergence and commonalities between them in terms of their common implications for error prevention during learning. In particular, both prior research and consideration of the trade-off involving the benefits of errors suggest that the success of these techniques cannot be guaranteed (e.g., Keith & Frese, 2008). In particular, it has been argued that allowing the learner to make errors during training (rather than preventing them) can be an essential component of skill acquisition, and training can then focus on error management (Keith & Frese, 2008). If learners are forced to make the active choices within the task, they may commit occasional errors, but this engagement of active choice has been found to enhance learning in other environments (Roediger & Karpicke, 2006; Weinstein, McDermott, & Roediger, 2010). Generating such active choices forces the learner to engage in deeper thinking about the material in a way that is germane to knowledge acquisition. Slamecka and Graf (1978) refer to this benefit of active choices as the “generation effect.” This advantage in active learning (which allows errors to occur) could outweigh the cognitive load costs of errors, should they occur.
Moderators of Error Prevention Training
In addition to the aforementioned accounting of two forces in error prevention strategies, prior research can also point to moderator variables that may modulate the effects of each. First, given the possible benefits of allowing some, but not excessive, errors to occur, we examine the level of constraint on the transfer benefit of error prevention. We predict that preventing errors by guiding or encouraging the correct choice will be more effective than “locking out” the possibility of error by hard constraints that force the learner to perform only correct actions. The latter situation will be more likely to encourage the learner to simply passively follow the forced execution steps and therefore not need to make any active decisions about actions, a source of germane load (e.g., Roediger & Karpicke, 2006; Slamecka & Graf, 1978).
A second variable expected to influence the success of the two error prevention strategies is learner experience (van Merriënboer & Sweller, 2005). The inexperienced learner, encountering the task or equipment to be trained for the very first time, may find the task of overwhelming complexity (very high intrinsic load) and can benefit from any technique available to lower its intrinsic load. In contrast, a learner with some experience on the task will be able to perform components with reduced resource demands; can tolerate less aggressive implementation of load-reducing strategies, such as error prevention; and so allows the benefits of performing the task in its less protected version to emerge (as well as realize the potential benefits of errors during learning; Keith & Frese, 2008). This expertise effect, operationalized here by prior experience, is well documented in prior research (e.g., Kalyuga, 2011; van Merriënboer & Sweller, 2005; Rey & Buchwald, 2011) and indeed was observed in a meta-analysis of two other CLT-driven strategies: part-task training and increasing-difficulty training (Wickens et al., 2012b).
A third variable anticipated to modify the costs or benefits of error prevention strategies is the nature of the transfer test itself; the deeper the knowledge acquired, the more likely it will show up in a more distant-transfer test less similar in both time (e.g., delayed) and specifics than an immediate-transfer test on a task highly similar to that used in training. We hypothesize that the more freedom (e.g., to make errors) the learner is given during training, the more deeply he or she would understand the task and hence the more remotely this knowledge would transfer.
Method
The Meta-Analytic Approach
In the current research, we examined the relative training transfer effectiveness of two training strategies through meta-analysis of research. Narrative reviews of scaffolding (Pea, 2004) and TW (Carroll, 1990; van Gog, Paas, & Sweller, 2010; van Gog & Rummel, 2010; Shen & Tsai, 2009) were found, but no meta-analyses examining the techniques were located. The rationale of the current approach followed Borenstein’s recommendations arguing not for a single way to perform meta-analysis but rather for aligning the approach used with the synthesis purpose (Borenstein, Hedges, Higgins, & Rothstein, 2009; Lipsey & Wilson, 2001; Rosenthal, 1991). The approach developed for the meta-analyses emphasized discovery and inclusiveness in a manner consistent with ongoing training exploration (Carolan, McDermott, Hutchins, Wickens, & Belanich, 2011; Wickens et al., 2012a, 2012b; Wickens, Hutchins, Carolan, & Cumming, 2011). Our approach further views the meta-analysis as an analysis of analyses, whereby the analysis can be performed on any aggregate statistics (Rosenthal, 1991). In the current meta-analysis process, we examined transfer ratio (TR) data, representing percentage benefit versus cost of the treatment, as the primary effect measure and additionally examined Hedges’ g as a complementary measure.
The choice of TR as primary effect measure was guided by its perceived benefits (see also Lu, Wickens, Sarter, & Sebok, 2011; Wickens et al., 2012b; Wickens, Prinet, Hutchins, Sarter, & Sebok, 2011), that is, ease of interpretation (e.g., a ratio of 1.5 means 50% more effective) and available study data. Not all viable studies report the necessary statistical data required for computation of a standardized effect size measure, such as Hedges’ g, therefore excluding such studies from a meta-analysis. Use of TR allowed the meta-analysis to be more inclusive, increasing the number of studies and contrasts examined. An important issue of our approach was to analyze the degree of the concordance between the two meta-analytic techniques. On the basis of a review of the meta-analysis methodology literature, we considered the strengths and weaknesses of the fixed-effect, random-effects, and mixed-effects models, that is, assumptions of between-study variance, statistical power, independence, sensitivity, and so on (Berk, 2007a, 2007b; Briggs, 2005; Glass, 1999; Lipsey, 2007; Shadish, 2007). Ultimately, a fixed-effect model was chosen because of the discovery emphasis of the research and an analytical emphasis on avoiding Type II error (Wickens, 1998).
Inclusion and Exclusion Criteria
Demographics
The targeted participant population was normal-health adults representative of the army population. Consequently, searches focused on an age demographic that excluded any samples younger than high school juniors and any samples labeled as “older adults” or “elderly.” The primary demographic was vocational and undergraduate students. No date restrictions were imposed on the included studies.
Research design
For inclusion, all studies were required to have a transfer condition that was identical between the treatment (error prevention) and control group. Furthermore, the control group was required to involve either (a) no error prevention techniques at all or (b) substantially less error prevention than the experimental group (of the 39 studies included, 36 were in former category and 3 were in the latter).
Search Strategy
Databases
Keyword searches were applied to Defense Technical Information Center, the database services of EBSCO, the Web of Science, Wiley Interscience, and Science Direct, for a total of 76 unique databases.
Keywords
For the scaffolding searches, the keywords scaffolding and either training or learning in the title, abstract, subject terms, or author-provided keywords were used. For the TW searches, the keywords training wheels, error prevention, worked examples, or partial completion problems and either training or learning in the title, abstract, subject terms, or author-provided keywords were used.
Additional searches
As a manual check of comprehensiveness, manual reference searches and journal searches were performed. All were noted and set aside. The reference sections of high-value papers (i.e., literature reviews, meta-analyses, significant empirical studies, etc.) were reviewed for additional relevant papers. Candidate papers were retrieved and reviewed. Additionally, journals with training themes or volumes were reviewed to further ensure that no important studies had been unintentionally excluded, for example, Acta Psychologica, Volume 71 (1989).
Moderators
Detailed reviews of included articles were performed to identify the set of potential moderator variables. The moderator variables of interest were of essentially three classes:
Those driven by CLT theory and the theoretical foundations of error allowance described earlier. These variables included the expertise of the learner (defined as low vs. high prior experience on the basis of sampling description provided by authors), the degree of constraints imposed on error prevention, and the transfer “distance” in time and in elements, relative to the final training trials. Transfer was coded as near transfer (NT) identical (a test with a task or problem identical to that used in training), NT similar (a test with a task or problem different from but similar to that used in training), far transfer (FT) performance environment (a testing situation different or new from training that requires an approximate match in acquired skill or knowledge with the primary difference being performance environment), FT complexity (again, FT with the primary difference being greater complexity), or FT task (far transfer but to a task structurally different than that trained).
Variables that were not clearly identified in advance but that emerged as important both by the virtue that they were varied within the studies examined and by the fact that, in our parallel meta-analysis on increasing difficulty, they revealed a relatively robust effect. Such is the case of the “instructor present” effect (Wickens et al., 2012b), which revealed that having an instructor engaged in the training treatment reduced the benefit (or increased the cost) of the strategy in question. These moderators examined included the following:
Instructor role: Coded as instructor not present, individual instruction, or classroom instruction depending on whether a human instructor administered the training materials to participants. Delivery system for error prevention: Coded as manuals, lecture, or computer or web based depending on the actual system used to deliver the error preventative mechanism. Delivery categories were not mutually exclusive.
Variables specific to each of the two strategies, including the lockouts versus worked examples of TW and the adaptive versus fixed removal of scaffolding.
We note that several other moderator variables were examined that did not meet any of the aforementioned criteria. These effects are not reported here, but the reader is referred to Carolan et al. (2011) for more details.
Coding Procedures
Coders were team members who had been trained on operational definitions and had practiced parsing the full studies into moderator variables. Coding for moderators was performed independently by two of these team members. Final codes were checked for consensus. We resolved any discrepancies by collectively reviewing the original article and discussing the appropriate code.
Statistical Methods
Effect size metrics
TR was calculated with the study-reported measure of transfer for the treatment (error prevention) group as the numerator and that for the control group as the denominator. Some measures were converted so that the larger number always indicated better performance (e.g., error rate was converted to accuracy). Interpretation of the ratio centered on a deviation from 1, whereby values of 1.0 indicate no difference between treatment and control, those <1 indicate cost for training treatment, and those >1 indicate benefit.
For computation of standardized effect size, first Cohen’s d′ was calculated as follows:
where Mt = treatment mean, Mc = control mean, and SDpooled = pooled standard deviation, when mean, standard deviation, and sample size were provided for both control and treatment.
where t = t test statistic or t value, sqrt = square root, nt = treatment sample size, and nc = control sample size, when only t test results and sample size were provided.
where sqrt = square root, F = F test statistic or F value, nt = treatment sample size, and nc = control sample size, when only F test results and sample size were provided.
In the instances when only a p value was provided, the t value was approximated given sample size and p value. To correct for sample size bias in d′, all d′ statistics were transformed into Hedges’ g, a sample-adjusted measure of effect size. Hedges’ g was computed as
Hedges’ g values of 0 indicate no difference between treatment and control, negative values indicate a cost for training treatment, and positive values indicate a benefit.
Software packages and statistical models
We analyzed TR data in SPSS Version 19 using a fixed-effect model and weighted least squares regression. We analyzed Hedges’ g data in Comprehensive Meta Analysis Version 2 using a fixed-effect model.
Results
Reporting Convention
TR and Hedges’ g results are presented together in a manner consistent with Wickens et al. (2012b). Effects significant at the p ≤ .05 level are indicated with an asterisk. Note that a somewhat unconventional k† is used to indicate the number of data points used in analysis, and multiple effects per study were allowed. Our focus in reporting the results is on the differences between a given level of a moderator variable and the equivalence point (TR = 1.0 or g = 0) rather than on the differences between two or more levels of a moderator variable. This focus is because we believe the greatest value to the consumer of our results lies in determining the conditions in which specific variations of the training strategy in question are effective or not in contrast to the status quo rather than how their overall benefits or costs are moderated by variation in those conditions. Language referencing the magnitude of observed effects follows the small (0.2 < g ≤ 0.5), medium (0.5 < g ≤ 0.8), and large (0.8 < g) guidance of Cohen (1988). We do sometimes call attention to differences (between two levels of a moderator variable) in these categories of effect sizes while recognizing that doing so is not the same as a statistical test between them and hence suggest caution in the interpretation of such differences.
TW: Results
A total of 31 TW studies provided statistical information necessary for the analysis and yielded a total of 74 Hedges’ g estimates and 79 TR estimates, including multiple subgroups, treatment comparisons, and outcome measures. See Appendix A for the full table of study effects. Conditions were coded on the basis of the nature and presentation of the error prevention treatment. Overall, there was a benefit for TW relative to the unsupported (or less supported) control (TR = 1.3*, k† = 79; g = +0.21*, k† = 74). This benefit was moderated as follows (additional moderator variables were examined, and their analysis is found in Carolan et al., 2011).
Prior experience
Experience was coded for a subset of studies that clearly identified naive (no- or low-experience) learners and experienced learners; however, these studies yielded only data useful for the calculation of Hedges’ g. Naive learners benefited from TW, demonstrating a significant benefit compared with control (g = +0.44*, k† = 6). Experienced learners also, on average, benefited from TW training but not significantly differing from the control (g = +0.28, k† = 5).
Delivery environment
Collapsed across all instructor-present conditions, TW was beneficial when no instructor was present (TR = 1.51*, k† = 28; g = 0.37*, k† = 26) and still beneficial when there was an instructor present (TR = 1.18*, k† = 49; g = 0.14*, k† = 46); however, in the latter, g was not even a small effect by standard conventions (Cohen, 1988). With respect to the delivery system, only three combinations of systems were identified: computer or web based, lecture and manuals, or all combined (computer, lecture, and manuals). The delivery system did not differentially affect the overall TW benefit.
Type of transfer
NT findings demonstrated a TW benefit (NT identical, TR = 1.3*, k† = 54; g = 0.24*, k† = 51; NT similar, TR = 1.73*, k† = 9; g = 0.48*, k† = 9); however, FT findings revealed a neutral effect of TW (TR = 1.0, k† = 16; g = −0.02, k† = 14). Examination of the types of FT revealed that FT to different tasks requiring the same skill types benefited from TW (TR = 1.56*, k† = 5; g = 0.34*, k† = 3), but TW may not be effective for FT to a task of the same skill type but that is more complex (TR = .91, k† = 11; g = −0.08, k† = 11).
Constraints
The small significant benefit for the highly constraining lockouts was not consistent across effect measures (TR = 1.1, k† = 2; g = +0.3*, k† = 3), whereas a consistently significant benefit was observed for the softer constraints imposed by worked examples (TR = 1.29*, k† = 55; g = +0.31*, k† = 50). Any conclusions that the two types of TW differed or about lockout effectiveness (or lack thereof) must be drawn with extreme caution, both because a difference was manifest in only one of the transfer measures for lockouts and because of the extremely small sample of lockout studies.
Discussion: TW
The overall benefit of TW appeared to be mitigated or offset by four factors.
The CLT expectations (Kalyuga, 2011; van Gog & Rummel, 2010; van Merriënboer & Sweller, 2005) for learner experience were supported in the TW literature. Naive learners benefited from the TW support; learners with prior experience did not.
The “instructor effect” replicated findings in the part-task training and increasing-difficulty training meta-analyses (Wickens et al., 2012b), whereby the presence of an instructor mitigated the strategy’s transfer benefit.
The more remote the transfer task, the less effective TW became. Here NT and FT to a novel task requiring the same skills demonstrated a TW benefit, but FT to a more complex task did not. It is possible that the deeper processing of the material required to learn more abstract task rules for FT was better supported by the greater degree of error tolerance of the control groups, availed by learner choice.
Lockouts failed to demonstrate the clear benefit of worked examples, but lockouts were not a common form of error prevention in the literature examined. More research is needed to understand the effectiveness of lockouts whereby the wrong path is made inaccessible in contrast to worked examples in which the right path is made explicit.
Scaffolding: Results
Eight scaffolding studies provided statistical information necessary for the analysis and yielded a total of 21 Hedges’ g estimates and 23 TR estimates, including multiple subgroups, treatment comparisons, and outcome measures. Overall, there was a strong benefit for scaffolding relative to the unsupported (or less supported) control (TR = 1.58*, k† = 23; g = 0.46*, k† = 21). This benefit was moderated as follows.
Prior experience
The small sample notwithstanding, those with prior experience showed a large benefit of scaffolding (TR = 1.7, k† = 2; g = 1.34*, k† = 2). For naive participants, the magnitude of the beneficial effect of scaffolding was not as pronounced (TR = 1.3, k† = 15; g = 0.26*, k† = 15).
Delivery environment
The benefits of scaffolding were revealed by both effect measures whether the instructor was present or not. However, the effect size was quite large (Cohen, 1988) when the instructor was absent (TR = 1.81*, k† = 5; g = 1.09*, k† = 5) and small (Cohen, 1988) when the instructor was present (TR = 1.52*, k† = 18; g = 0.27*, k† = 16).
Type of transfer
Although both identical and similar types of transfer examined showed scaffolding benefit, the effect size was small for identical NT tasks (TR = 1.55*, k† = 20; g = 0.44*, k† = 18) yet medium (Cohen, 1988) for an unsupported transfer task similar but not identical to training (TR = 2.03, k† = 3; g = 0.66*, k† = 3). FT was not examined in the scaffolding studies.
Removal of scaffolding
Scaffolds were removed on either an adaptively driven or fixed schedule. The scaffolding effect size was medium (Cohen, 1988) with adaptive administration (TR = 1.63, k† = 7, p < .10; g = 0.72*, k† = 8) and small (Cohen, 1988), although still present, with fixed (TR = 1.56*, k† = 16; g = 0.29*, k† = 13).
Discussion: Scaffolding
Overall, the benefits of scaffolding as an error prevention strategy echoed those of TW, with a 58% benefit that was significant in both effect measures. Again, certain moderator variables modulated this benefit.
Prior experience revealed a pattern inconsistent with expectations of CLT (Kalyuga, 2011; van Gog & Rummel, 2010; van Merriënboer & Sweller, 2005). According to CLT, novice learners would benefit most from the additional support early in acquisition while schemas are forming; however, experienced learners showed greater benefit from scaffolding than did inexperienced learners. One possibility is that, as discussed with respect to worked-out examples, “for learners with more prior knowledge, imagining the solution steps can also impose a germane cognitive load, but not for novice learners” (Paas & van Gog, 2006, p. 88).
Again, as with TW, the scaffolding strategy appeared to be more effective when the instructor was absent than when the instructor was present.
Benefit or cost from the error prevention of scaffolding could not be assessed for FT.
The larger benefit of an adaptive removal schedule compared with a fixed schedule is consistent with the previous meta-analysis on adaptive versus fixed difficulty increases (Wickens et al., 2012b), which also revealed that cognitive load variables that are sensitive to individual learner differences are more effective.
Methodological Findings
Analysis of concordance revealed that TR and Hedges’ g provided convergent information. Interpretation of the analysis between the two effect size measures was guided by Table 1. Any cell in Blocks A or D indicates agreement in the direction of the effect (cost or benefit for the training strategy vs. control). Cells A.1, A.4, D.1, and D.4 indicate both agreement in direction and significance. Cells C.3 and D.2 represent counter indications in both direction and significance (no effects were found in these two cells). Cell values show the frequency of agreement and percentage of total between corresponding measures.
Concordance Table Cell Codes
For TW (Table 2), directional concordance was 98%, with a 74% agreement in direction and significance. There were no cases in which significant contraindications were found, and only 2% (one case) differed in direction, with neither measure indicating a significant effect. For scaffolding (Table 3), directional concordance was 97.5%, with a 62.5% agreement in direction and significance. Again, there were no cases in which significant contraindications were found and only one case (2.5%) in which the two effect size measures differed in direction, with one measure indicating a significant effect. Overall, the Hedges’ g measure was most powerful, yielding 35 and 39 significant effects for TW and scaffolding, respectively, whereas TR yielded 28 and 24 such effects.
Concordance Table for Training Wheels
Concordance Table for Scaffolding
General Discussion
Two error prevention strategies, TW and scaffolding, were examined in the context of CLT. Arguably, as is the case in the present work, both TW and scaffolding are examples of the overall error prevention strategy; however, the two were different as instances of the same higher-level strategy in the manner with which the error-preventative mechanism has been experimentally explored in the literature defined by the two methods. We hypothesized that reducing the intrinsic load of making error-prone choices in the skill, and avoiding the possible extraneous load resulting from catastrophic errors, would avail more resources for germane load, enhance learning, and hence, enhance transfer. However, we also considered the possible counterinfluence of a benefit for making errors during training, which might offset any benefit to reducing cognitive load.
Collectively, the two strategies reflected the common finding that the benefits of error prevention outweighed those of making errors whether via TW (30% transfer advantage) or scaffolding (60% transfer advantage). Also common to both literatures was the apparent advantage of the strategies when an instructor was not present. Having no instructor present for administration of the training itself (but not the training strategy itself, i.e., prompting) increased the transfer benefit of the error prevention strategies. Consistent with the hypothesized mechanisms in Wickens et al. (2012b), a possible explanation of this instructor effect for error prevention is increased extraneous load attributable to potential instructor–technology interactions or instruction inconsistency with the training environment. Further insights into the roles of human tutors and delivery systems can be found in the related body of literature on intelligent tutoring systems. In a review of tutoring systems, detailed comparisons revealed that human tutors may not be vastly superior to other, nonhuman guidance. Rather, the review supported a view that the interaction granularity of the tutoring mechanism itself, regardless of human or nonhuman delivery, can better explain observed benefits in tutoring systems (VanLehn, 2011).
Within the TW analysis, certain moderator variables supported the role of CLT in conferring an error prevention strategy advantage. First, the expertise effect (e.g., Rey & Buchwald, 2011) provided direct evidence: TW benefited inexperienced learners, for whom resources are scarcer (because intrinsic load is greater), but not experienced learners. However, at the same time, there was less clear evidence for the counterinfluence of a benefit to allowing errors (and hence a cost for error prevention). In the TW analyses, lockouts, the most extreme way of preventing errors, did not prove a less effective technique than did worked examples, a strategy within which errors, although discouraged, could still occur. However, because of low statistical power, any conclusion of no difference must be taken with extreme caution, awaiting further research. In addition, within the TW meta-analysis, the finding that more remote transfer benefited less from error prevention could be consistent with an errors-are-valuable effect in that remote transfer would seem to require a higher level of more abstract knowledge, the kind of knowledge that could be acquired by more exploratory learning.
Within the scaffolding analysis, a presumed key feature of error prevention strategies was the fading schedule. Adaptive and fixed fading schedules were compared, and both showed benefit for transfer; but adaptive had a larger benefit (medium effect size, g = +0.72) than fixed (small effect size, g = +0.29), consistent with the value of an adaptive schedule found with training difficulty increases (Wickens et al., 2012b). Overall, the scaffolding moderator variable analysis provided less clear-cut evidence for the balance between the two forces of CLT and error benefits. It did not contain an analogous variable, such as degree of constraint, and the learner expertise effect actually expressed itself in the opposite direction when we examined level of trainee experience, both from that found with TW and from prior research in this area (Rey & Buchwald, 2011; van Merriënboer & Sweller, 2005; Wickens et al., 2012b). The reasons for the contradiction remain unclear and a curious finding indeed but should be given more attention in future research, given the small sample and inconsistency with the majority of the CLT literature on expertise effects.
Concerning our methodological approach, analysis of concordance between the ratio and standardized effect size approaches not only provided evidence of the complementarity of the methods but also offers evidence for the validity of the novel ratio approach because of the high concordance with more widely accepted standardized measures of effect. In areas where quantitative literature synthesis is desired, but a lack of standardized reporting conventions negatively affect study inclusion, researchers should keep an open mind to the viability of lower-precision estimates, such as the TR, in the analysis of analyses.
Training Design Recommendations
The current research supports the concept that training designed to reduce intrinsic and extraneous loads with an instructional mechanism that prevents errors during learning improves transfer when compared with training designed with less or no error prevention. Further supported in the current research, as well as in prior research (Wickens et al., 2012b), designing training with adaptive removal of the error prevention mechanism should further enhance transfer. The curious influence of instructor presence on transfer effectiveness should highlight to training designers the potential for an increase in extraneous load during learning through unintended instructor interactions with design elements. Last, although inconclusive in the current research, training designers should not ignore individual differences, such as prior trainee experience, when making training design choices within the conceptual framework of CLT.
Limitations of Current Research
As with any meta-analysis, there is questionable internal validity, given the lack of true experimental manipulation; inference should be made accordingly. Another limitation is unintentional exclusion of studies (i.e., unpublished or missed). Even though in the current approach we attempted to maximize inclusiveness across the two methods, unintentional exclusion should be considered and, as noted for scaffolding, may have factored into sampling. Because the current research emphasized inclusiveness, there was less emphasis on estimate precision than would have been afforded with random- or mixed-effects models. Another source of limitation, inherent in most meta-analyses that are partially exploratory in nature, is that moderator variables are rarely “crossed” as they might be within traditional experimental design, so interactions between such variables cannot be easily discerned. Finally, as with all meta-analyses, our search for studies may have missed findings of negative results (0 transfer) because of the classic “file drawer problem” (Rosenthal, 1991), whereby researchers tend not to publish those findings with negative results.
Key Points
Error prevention strategies of training wheels (TW) and scaffolding are both successful, yielding a 30% to 60% transfer benefit.
The benefits of both strategies can be accounted for within the framework of cognitive load theory (CLT), although there are some offsetting benefits of error allowance strategies.
Consistent with CLT, in TW, novice learners benefit more from error prevention strategies than do experienced learners, and having instructors present may produce extraneous load to diminish the benefits of error prevention.
The two complementary methods of meta-analysis, transfer ratio and Hedges’ g, provided convergent results.
Footnotes
Appendix A
Table of Effect Sizes for Training Wheels Studies
| Study | Task | Contrast and Measure | TR | g (SE) |
|---|---|---|---|---|
| Bannert (2000) | Computer word-processing task | Unassisted task vs. lockouts: Accuracy | 0.99 | −0.06 (0.23) |
| Carroll (1984) | Computer word-processing task | Unassisted task vs. lockouts: Accuracy | 1.75 | 0.77 (0.55) |
| Carroll (1994) | Quantitative reasoning task | Unassisted task vs. worked examples with prompts (near transfer): Error | 2.5 | 0.75 (0.32) |
| Unassisted task vs. worked examples with prompts (delayed transfer): Error | 2.28 | 0.75 (0.35) | ||
| Unassisted task vs. worked examples with prompts (far transfer): Error | 1.61 | 0.5 (0.35) | ||
| Crippen & Earl (2007) | Chemistry content task | Unassisted task vs. worked examples with prompts: Accuracy | 1.05 | 0.53 (0.31) |
| Unassisted task vs. worked examples: Accuracy | 0.98 | −0.19 (0.31) | ||
| Halabi, Tuovinen, & Farley (2005) | Quantitative reasoning task | Unassisted task vs. worked examples (no prior experience): Accuracy | 1.1 | 0.1 (0.29) |
| Unassisted task vs. worked examples (no prior experience): Accuracy | 0.97 | −0.03 (0.28) | ||
| Hilbert & Renkl (2009), Experiment 1 | Computer concept mapping task | Unassisted task vs. worked examples: Accuracy | 0.83 | −0.29 (0.36) |
| Hilbert & Renkl (2009), Experiment 2 | Computer concept mapping task | Unassisted task vs. worked examples: Accuracy | 1.95 | 0.14 (0.28) |
| Unassisted task vs. worked examples with prompts: Accuracy | 2.4 | 0.16 (0.27) | ||
| Hilbert, Renkl, Kessler, & Reiss (2008) | Quantitative reasoning task | Unassisted task vs. worked examples (combined treatments, skill test): Accuracy | 1.32 | 0.65 (0.24) |
| Unassisted task vs. worked examples (combined treatments, knowledge test): Accuracy | 1.77 | 1.51 (0.26) | ||
| Unassisted task vs. worked examples (heuristic only): Accuracy | 1.46 | 0.14 (0.31) | ||
| Unassisted task vs. worked examples with prompts (self-explanation): Accuracy | 1.48 | 0.14 (0.3) | ||
| Hohn & Moraes (1997-1998) | Computer programming task | Unassisted task vs. worked examples (code grouping): Rating | 1.12 | 0.03 (0.24) |
| Unassisted task vs. worked examples with prompts (code grouping): Rating | 1.85 | 0.29 (0.24) | ||
| Unassisted task vs. worked examples (problem categorization): Rating | 1.33 | 0.05 (0.24) | ||
| Unassisted task vs. worked examples with prompts (problem categorization): Rating | 2.07 | 0.16 (0.24) | ||
| Lang (2007) | Computer video game problem-solving task | Unassisted task vs. worked examples (conventional worked examples): Accuracy | 1.02 | 0.15 (0.23) |
| Unassisted task vs. worked examples (just-in-time worked examples): Accuracy | 1.22 | 0.59 (0.23) | ||
| Leutner (2000) | Computer-aided design software task | Unassisted task vs. lockouts: Accuracy | 0.54 (0.22) | |
| Loring (2003) | Quantitative reasoning task | Unassisted task vs. worked examples (high ability): Accuracy | 1.13 | 0.02 (0.36) |
| Unassisted task vs. worked examples (moderate ability): Accuracy | 2.05 | 0.24 (0.41) | ||
| Unassisted task vs. worked examples (low ability): Accuracy | 1.25 | 0.18 (0.41) | ||
| Mahan (2007) | Quantitative reasoning task | Unassisted task vs. worked examples with prompt (near transfer, procedure prompt): Accuracy | 0.91 | −0.38 (0.12) |
| Unassisted task vs. worked examples with prompt (near transfer, principle prompt): Accuracy | 0.93 | −0.28 (0.12) | ||
| Unassisted task vs. worked examples with prompt (far transfer, procedure prompt): Accuracy | 0.9 | −0.15 (0.12) | ||
| Unassisted task vs. worked examples with prompt (far transfer, principle prompt): Accuracy | 0.94 | −0.06 (0.12) | ||
| McLaren, Lim, & Koedinger (2008), Experiment 1 | Quantitative reasoning task | Unassisted task vs. worked examples with prompt: Accuracy | 0.98 | 0.02 (0.25) |
| McLaren, Lim, & Koedinger (2008), Experiment 2 | Quantitative reasoning task | Unassisted task vs. worked examples with prompt: Accuracy | 1.05 | 0.04 (0.25) |
| McLaren, Lim, & Koedinger (2008), Experiment 3 | Quantitative reasoning task | Unassisted task vs. worked examples with prompt: Accuracy | 1.04 | 0.29 (0.22) |
| Moraes (1995) | Computer programming task | Unassisted task vs. worked examples (coding task): Accuracy | 1.12 | 0.31 (0.24) |
| Unassisted task vs. worked examples with prompt (coding task): Accuracy | 1.85 | 0.89 (0.25) | ||
| Unassisted task vs. worked examples (categorization task): Accuracy | 1.37 | 0.17 (0.24) | ||
| Unassisted task vs. worked examples with prompt (categorization task): Accuracy | 3.0 | 1.18 (0.26) | ||
| Ondrusek (1999) | Database search task | Unassisted task vs. worked examples (Search Problem 1): Accuracy | 1.26 | 0.58 (0.26) |
| Unassisted task vs. worked examples (Search Problem 2): Accuracy | 1.34 | 1.22 (0.32) | ||
| Paas & van Merriënboer (1994) | Computer numerically controlled spatial reasoning task | Unassisted task vs. worked examples (low difficulty): Time to completion | 1.08 | 0.31 (0.36) |
| Unassisted task vs. worked examples (low difficulty): Accuracy | 1.65 | 1.44 (0.4) | ||
| Unassisted task vs. worked examples (high difficulty): Time to completion | 0.94 | −0.29 (0.36) | ||
| Unassisted task vs. worked examples (high difficulty): Accuracy | 2.24 | 2.05 (0.44) | ||
| Rourke & Sweller (2009), Experiment 1 | Declarative design history recall task | Unassisted task vs. worked examples (near transfer): Accuracy | 1.32 | 0.43 (0.2) |
| Unassisted task vs. worked examples (far transfer): Accuracy | 0.15 (0.2) | |||
| Rourke & Sweller (2009), Experiment 2 | Declarative design history recall task | Unassisted task vs. worked examples (near transfer): Accuracy | 2.68 | 1.5 (0.51) |
| Unassisted task vs. worked examples (far transfer): Accuracy | 1.12 (0.48) | |||
| Shen & O’Neil (2006) | Computer video game-problem solving task | Unassisted task vs. worked examples (knowledge map): Accuracy | 3.56 | 0.63 (0.24) |
| Unassisted task vs. worked examples (Problem Solving 1): Accuracy | 1.18 | 4.22 (0.42) | ||
| Unassisted task vs. worked examples (Problem Solving 2): Accuracy | 1.2 | 2.83 (0.33) | ||
| Tookey (1994) | Quantitative reasoning task | Unassisted task vs. worked examples (near transfer): Accuracy | 0.99 | −0.11 (0.37) |
| Unassisted task vs. worked examples with prompt (near transfer): Accuracy | 1.05 | 0.41 (0.37) | ||
| Unassisted task vs. worked examples (far transfer): Accuracy | 0.98 | −0.04 (0.37) | ||
| Unassisted task vs. worked examples with prompt (far transfer): Accuracy | 0.84 | −0.34 (0.37) | ||
| Van Gerven, Paas, van Merriënboer, & Schmidt (2002) | Spatial reasoning and problem-solving task | Unassisted task vs. worked examples (near transfer): Accuracy | 0.99 | −0.39 (0.36) |
| Unassisted task vs. worked examples (far transfer): Accuracy | 0.97 | −0.23 (0.36) | ||
| van Gog, Jarodzka, Scheiter, Gerjets, & Paas (2009) | Problem-solving task | Unassisted task vs. worked examples: Accuracy | 2 | |
| Unassisted task vs. worked examples with prompt: Accuracy | 2 | |||
| van Gog, Paas, & van Merriënboer (2006) | Electric circuit problem-solving task | Unassisted task vs. worked examples (near transfer): Accuracy | 1.26 | 0.65 (0.26) |
| Unassisted task vs. worked examples (near transfer): Time to completion | 0.67 | −1.05 (0.27) | ||
| Unassisted task vs. worked examples (far transfer): Accuracy | 1.2 | 0.6 (0.26) | ||
| Unassisted task vs. worked examples (far transfer): Time to completion | 0.91 | −0.19 (0.25) | ||
| Ward & Sweller (1990), Experiment 1 | Physics problem-solving task | Unassisted task vs. worked examples: Accuracy | 1.33 | 1.05 (0.32) |
| Ward & Sweller (1990), Experiment 2 | Physics problem-solving task | Unassisted task vs. worked examples: Accuracy | 1.52 | 1.54 (0.39) |
| Ward & Sweller (1990), Experiment 3 | Physics problem-solving task | Unassisted task vs. worked examples (near transfer, horizontal linear motion): Accuracy | 1.04 | 0.12 (0.34) |
| Unassisted task vs. worked examples (far transfer, horizontal linear motion): Accuracy | 0.6 | −0.14 (0.34) | ||
| Unassisted task vs. worked examples (near transfer, vertical linear motion): Accuracy | 1.17 | 0.69 (0.34) | ||
| Unassisted task vs. worked examples (far transfer, vertical linear motion): Accuracy | 0.71 | −0.12 (0.34) | ||
| Unassisted task vs. worked examples (near transfer, projectile motion): Accuracy | 0.77 | −0.92 (0.35) | ||
| Unassisted task vs. worked examples (far transfer, projectile motion): Accuracy | 0.75 | −0.13 (0.34) | ||
| Unassisted task vs. worked examples (near transfer, collisions): Accuracy | 0.76 | −1.45 (0.38) | ||
| Unassisted task vs. worked examples (far transfer, collisions): Accuracy | 0.75 | −0.09 (0.34) | ||
| Ward & Sweller (1990), Experiment 4 | Physics problem-solving task | Unassisted task vs. worked examples (near transfer, one-move problem, conventional worked examples): Accuracy | 0.79 | −0.38 (0.36) |
| Unassisted task vs. worked examples (near transfer, two-move problem, conventional worked examples): Accuracy | 0.83 | |||
| Unassisted task vs. worked examples (far transfer, conventional worked examples): Accuracy | 1.4 | |||
| Unassisted task vs. worked examples (near transfer, one-move problem, integrated worked examples): Accuracy | 1.54 | 1.25 (0.39) | ||
| Unassisted task vs. worked examples (near transfer, two-move problem, integrated worked examples): Accuracy | 1.83 | |||
| Unassisted task vs. worked examples (far transfer, integrated worked examples): Accuracy | 2 | |||
| Ward & Sweller (1990), Experiment 5 | Physics problem-solving task | Unassisted task vs. worked examples (near transfer, conventional worked examples): Accuracy | 1.34 | 1.32 (0.39) |
| Unassisted task vs. worked examples (far transfer, conventional worked examples): Accuracy | 2 | |||
| Unassisted task vs. worked examples (near transfer, split-attention worked examples): Accuracy | 0.91 | −0.28 (0.36) | ||
| Unassisted task vs. worked examples (far transfer, split-attention worked examples): Accuracy | 0.78 |
Note. TR = transfer ratio.
Appendix B
Table of Effect Sizes for Scaffolding Studies
| Study | Task | Contrast and Measure | TR | g (SE) |
|---|---|---|---|---|
| Azevedo, Cromley, & Seibert (2004) | Declarative circulatory system task | Adaptive scaffolding vs. no scaffolding (mental model shift): Rating | 2 | 1.05 (0.36) |
| Fixed scaffolding vs. no scaffolding (mental model shift): Rating | 0.78 | -0.31 (0.34) | ||
| Adaptive scaffolding vs. no scaffolding (matching task): Accuracy | 2.27 | 0.48 (0.34) | ||
| Fixed scaffolding vs. no scaffolding (matching task): Accuracy | 3.43 | 1.09 (0.36) | ||
| Adaptive scaffolding vs. no scaffolding (labeling task): Accuracy | 1.66 | 1.62 (0.39) | ||
| Fixed scaffolding vs. no scaffolding (labeling task): Accuracy | 0.31 | −1.46 (0.38) | ||
| Azevedo, Cromley, Winters, Moos, & Greene (2005) | Declarative circulatory system task | Adaptive scaffolding vs. no scaffolding (mental model shift): Rating | 1.6 | |
| Fixed scaffolding vs. no scaffolding (mental model shift): Rating | 0.63 | |||
| Adaptive scaffolding vs. no scaffolding (matching task): Accuracy | 1.22 | 0.7 (0.24) | ||
| Fixed scaffolding vs. no scaffolding (matching task): Accuracy | 0.91 | -0.27 (0.23) | ||
| Adaptive scaffolding vs. no scaffolding (labeling task): Accuracy | 1.4 | 0.67 (0.23) | ||
| Fixed scaffolding vs. no scaffolding (labeling task): Accuracy | 0.47 | −0.89 (0.24) | ||
| Azevedo et al. (2002) | Declarative circulatory system task | Fixed scaffolding vs. no scaffolding (critical thinking prompt): Rating | 0.77 | −0.15 (0.43) |
| Fixed scaffolding vs. no scaffolding (self-regulation of learning prompt): Rating | 1.92 | 0.65 (0.44) | ||
| Adaptive scaffolding vs. no scaffolding (self-regulation of learning prompt): Rating | 2.35 | 0.93 (0.45) | ||
| Bixler (2008) | Web design problem-solving task | Fixed scaffolding vs. no scaffolding (analysis of problem): Rating | 1.77 | 1.5 (0.25) |
| Fixed scaffolding vs. no scaffolding (development of solution): Rating | 1.62 | 1.19 (0.24) | ||
| Chen & Bradshaw (2007) | Educational measurement problem-solving task | Fixed scaffolding vs. no scaffolding (critical thinking prompt): Accuracy | 2.7 | 1.07 (0.42) |
| Fixed scaffolding vs. no scaffolding (strategy-based prompt): Accuracy | 1.71 | 0.48 (0.4) | ||
| Fixed scaffolding vs. no scaffolding (critical thinking and strategy-based prompt): Accuracy | 1.69 | 0.46 (0.4) | ||
| Kao & Lehman (1997) | z test quantitative reasoning task | Adaptive scaffolding vs. no scaffolding (maximum support): Accuracy | 0.19 (0.28) | |
| Adaptive scaffolding vs. no scaffolding (incremental support): Accuracy | 0.78 (0.3) | |||
| Petsangsri (2002) | Declarative human–computer interaction task | Fixed scaffolding vs. no scaffolding: Accuracy | 1.08 | 0.35 (0.26) |
| Podolefsky (2008) | Physics-based quantitative reasoning task | Fixed scaffolding vs. no scaffolding (shift on correct answers): Accuracy | 3 | |
| Fixed scaffolding vs. no scaffolding (shift on partially correct answers): Accuracy | 1.33 |
Note. TR = transfer ratio.
Acknowledgements
This work is supported by the U.S. Army Research Institute under Contract No. W91WAW-09-C-0081 to Alion Science and Technology titled “Understanding the Impact of Training on Performance.” The view, opinions, and/or findings contained in this article are those of the authors and should not be construed as an official Department of the Army position, policy, or decision.
Shaun D. Hutchins is a senior human factors engineer at Alion Science and Technology in Boulder, Colorado. He is also a PhD student in the School of Education at Colorado State University. He received his MA in experimental psychology with a minor in experimental statistics at New Mexico State University and his BA in psychology from the University of Maine at Farmington.
Christopher D. Wickens is a senior scientist at Alion Science Corporation, Micro Analysis and Design Operation, in Boulder, Colorado, and professor emeritus at the University of Illinois at Urbana-Champaign. He received his PhD in psychology from the University of Michigan in 1974.
Thomas F. Carolan is a senior scientist and program manager at Alion Science and Technology in Boulder, Colorado. He received a PhD in psychology from the University of Connecticut.
John M. Cumming, PhD, is a recent graduate of the Research Methodology program within the School of Education at Colorado State University. He received his MA in educational psychology emphasizing in research and evaluation methodology from the University of Colorado at Denver and his BA in psychology from the Metropolitan State College of Denver.
