Abstract
Teacher professional development (PD) is seen as a promising intervention to improve teacher knowledge, instructional practice, and ultimately student learning. While research finds instances of significant program effects on teacher knowledge, little is known about how long these effects last. If teachers forget what is learned, the contribution of the intervention will be diminished. Using a large-scale data set, this study examines the sustainability of gains in teachers’ content knowledge for teaching mathematics (CKT-M). Results show that there is a negative rate of change in CKT after teachers complete the training, suggesting that the average score gain from the program is lost in just 37 days. There is, however, variation in how quickly knowledge is lost, with teachers participating in summer programs losing more rapidly than those who attend programs that occur during school years. The implications of these findings on designing and evaluating PD programs are discussed.
Introduction
Classroom teachers are one of the most important in-school factors contributing to student learning (Aaronson, Barrow, & Sander, 2007; Chetty, Friedman, & Rockoff, 2013; Rivkin, Hanushek, & Kain, 2005). To improve teacher quality, federal and local education agencies have devoted substantial resources to teacher training and development (Correnti, 2007; Darling-Hammond & Sykes, 1999). Many of these professional development (PD) programs are designed based on different theories of how teachers learn and how teachers use the acquired knowledge in practice (Kennedy, 2016; Richardson & Placier, 2001). While these theories have different conceptions of teacher learning and teaching, they all perceive teacher knowledge to be an important outcome of PD and necessary condition for improving teaching practices and student outcomes (Cochran-Smith & Lytle, 2001; Guskey, 2000; Putnam & Borko, 2000; Vermunt, 2014).
As a result of the teacher learning research and the recent development in teacher assessments, teacher knowledge measures have been increasingly used as an outcome in empirical studies that examined the effectiveness of teacher PD programs (Carney, Brendefur, Thiede, Hughes, & Sutton, 2014; Hill & Ball, 2004; Koellner & Jacobs, 2015; Phelps, Kelcey, Jones, & Liu, 2016; Polly, Neale, & Pugalee, 2014). The logic for using measures of teacher knowledge as the outcome is straightforward and based on a basic set of causal assumptions: effective PD should lead to improvements in teacher knowledge, this knowledge should then have a positive impact on instruction, and in turn, effective instruction should have an impact on student achievement (Desimone, 2009; McCutchen et al., 2002; Scher & O’Reilly, 2009; Yoon, Duncan, Lee, Scarloss, & Shapley, 2007).
However, there is one critical issue with this set of assumptions that has received limited attention in research on teacher PD programs. For gains in knowledge to have an effect on teaching, these knowledge gains must persist. This may not hold because cognitive research on knowledge acquisition shows that newly learned knowledge is malleable and often impermanent rather than stable and unchanging (Grider, 1993; Vosniadou, 2007). One important finding is that for knowledge to be maintained, it needs to be applied or used in a relevant setting. Without use, knowledge tends to move from being active and available to inactive and difficult to retrieve (Arthur, Bennett, Stanush, & McNelly, 1998; Stothard & Nicholson, 2001). Another finding is that the application of new knowledge is inconsistent and this knowledge is often modified (Farr, 2012; Spiro, Collins, Thota, & Feltovich, 2003). These two processes, which can be labeled as forgetting and adaptation, can lead to treatments on knowledge and skills having immediate desired outcomes but undesirable long-term outcomes.
These ideas from cognitive research could also be applied to the field of teacher learning and teacher education. Just as teachers could develop knowledge, they can also forget what they have learned from the training programs. This is especially the case if the programs do not have sustained support or if knowledge is not applied or used in daily classroom instruction (Abdal-Haqq, 1995; Birman, Desimone, Porter, & Garet, 2000; Guskey & Yoon, 2009). If teachers lose the knowledge needed for effective teaching, then the PD might have little to no effect on student learning, even when there are substantial initial program effects on teacher knowledge outcomes.
The study of teacher knowledge loss, or what is often called knowledge decay, is important for both designing and evaluating teacher PD programs. On one hand, studying teacher knowledge decay is essential to properly understand the effectiveness of teacher PD programs. So far, prior research evaluating teacher PD has found mixed evidence on the effect of these programs. While many studies have found significant impacts, there is still plenty of evidence showing that PD programs have null effects on teacher and student outcomes (Battey et al., 2013; Bell, Wilson, Higgins, & McCoach, 2010; Blank, De las Alas, & Smith, 2008; Garet, Heppen, Walters, Smith, & Yang, 2016; Gersten, Taylor, Keys, Rolfhus, & Newman-Gonchar, 2014; Heller, Daehler, Wong, Shinohara, & Miratrix, 2012; Kraft, Blazar, & Hogan, 2018; Santagata, Kersting, Givvin, & Stigler, 2010). Although different possibilities have been hypothesized, knowledge decay may be one important reason for the null results. Specifically, if decay occurs after teachers complete the training, the estimated effects would be sensitive to how soon after a program the teacher knowledge outcome measures are administered. On the other hand, the estimated knowledge decay could also inform the design of the teacher PD programs. For example, in organizational contexts, estimated decay rate is used to optimize the scheduling of refresher trainings (Healy et al., 1998; Rail Safety and Standards Board [RSSB], 2011; Wisher, Sabol, Ellis, & Ellis, 1999). Similarly, in the context of teacher PD, the decay rate could be useful for deciding the number of training sessions and appropriate training intervals for maintaining the knowledge and skill level required for effective teaching.
Recognizing the importance of studying teacher knowledge decay, this study represents one of the initial efforts to understand knowledge decay by estimating the decay rate of content knowledge for teaching in mathematics (CKT-M) after completing PD. There are two reasons why the CKT measures are preferred to other teacher knowledge measures in estimating teacher learning and knowledge decay. First, these assessments are designed to measure a specialized type of content expertise used in teaching, including analyzing students’ work, determining the validity of a mathematical argument, or using proper representations for teaching particular math ideas (Ball, Thames, & Phelps, 2008; Baumert et al., 2010). As a result, they are more likely to capture teacher learning and knowledge decay associated with teacher PD programs that aim to improve teaching. Second, there is mounting empirical evidence showing that assessments of CKT are associated with teaching quality as well as student learning (Baumert et al., 2010; Hill, Rowan, & Ball, 2005; Hill, Umland, Litke, & Kapitula, 2012). These findings suggest that CKT is a key mediator that is closely connected with both program treatment and student learning and thus should be included for examining the effectiveness and the underlying mechanism of the PD programs. Based on a set of PD programs for mathematics teachers that took place during 2008 to 2013, this study investigates the following two research questions:
The rest of the article is organized as follows. In the next section, we review prior studies that examine the sustainability of teacher learning from PD programs and the extent to which the sustainability varies across different conditions. Next, we describe the analytic sample and empirical strategy used to estimate teacher knowledge decay. We then present results which suggest that there is significant knowledge decay after program treatment ends. We conclude with a discussion of these results and their implications for using teacher knowledge as an outcome to evaluate and study the improvement of teacher PD programs.
Review of Prior Research on the Sustainability of PD Outcomes
Substantial research has been conducted to examine the effectiveness of mathematics PD programs (Bell et al., 2010; Hill & Ball, 2004; Koellner & Jacobs, 2015; Polly et al., 2014). However, an equally important question is to what extent the effect persists after the completion of the programs. So far, studies that examined the sustainability of PD outcomes were based primarily on small samples of teachers, and some of them used self-report outcome measures. These studies provided preliminary evidence on how much teachers sustained their knowledge or teaching practices, and discussed possible factors that were associated with the sustainability. In this section, we summarize evidence from these studies and provide a framework for understanding the conditions that may affect the sustainability of PD outcomes.
One of the early studies that examined the sustainability of the PD outcomes was conducted by Supovitz, Mayer, and Kahle (2000). The program promoted inquiry-based instruction for science and math teachers. Using self-report survey data, the study found that teachers’ attitude and use of inquiry-based teaching practices sustained 3 years after the PD. This pattern holds good for teachers of different genders, race groups, grade levels, and school types (public vs. private). In another study, Boston and Smith (2011) found that a PD program improved math teachers’ ability to select high-level instructional tasks, but not all teachers were able to sustain what they learned from the training. Those teachers who showed sustainable changes demonstrated higher inclinations to reflect on their practices and to apply strategies learned from the training to their own teaching practices. Finally, Sandholtz and Ringstaff (2016) investigated how contextual factors influenced the sustainability of learning from a PD that provided science assistance to K-2 teachers. Based on survey and interview data from 15 teachers, the study found that school-level factors, such as administrative and collegial support, were important to teachers in sustaining science instruction. High turnover of principals and other colleagues often made it challenging for teachers to maintain the instruction quality over time.
In addition to the sustainability of teaching practices, several studies provided evidence on the presence of knowledge retention, which is more close to what is examined in this study. Franke, Carpenter, Levi, and Fennema (2001) followed up with teachers 4 years after they completed a PD that focused on improving understandings on students’ mathematical thinking. In addition, Steinberg, Empson, and Carpenter (2004) conducted a case study of one teacher who participated in a similar PD. Overall, findings from both studies suggested that after the completion of the training, engagement with student thinking continued to improve teacher knowledge of students’ mathematical thinking and the quality of math instruction. Evidence on knowledge retention or decay was also available in other subject areas. Goldschmidt and Phelps (2010) estimated the impact of teacher PD and its sustainability using measures of content knowledge for teaching in the area of elementary reading instruction and found that learning gains in comprehension and word analysis were not sustained 6 months after the training was completed. Roth et al. (2011) examined the impacts of two types of science PD programs and their sustainability on teachers’ content knowledge. Their results showed declining trends in knowledge for teachers in the program that focused only on deepening science knowledge (content only program), but teachers who participated in the program with analysis of teaching practice sustained their knowledge gain afterwards.
To summarize, prior research found variation in how much teachers sustained their learning after completing PD programs. Factors related to the content and design of training, to individuals, and to the school context in which teachers worked have been shown to influence sustainability of teacher learning. In terms of training content, PD programs with sustainable impact were usually embedded in teacher work and promoted understanding of students’ thinking (Franke et al., 2001; Roth et al., 2011; Yoon et al., 2007; Zehetmeier, 2008). The sustainability may also be affected by teacher characteristics, their attitude toward the program, and their ability to internalize what they learn during the training (Boston & Smith, 2011; Steinberg et al., 2004). Finally, teacher learning was more likely to sustain for those working in a stable and well-supported school environment (Gaikhorst, Beishuizen, Zijlstra, & Volman, 2017). Principal support and teacher collaboration increased the sustainability of PD impact on teachers (Sandholtz & Ringstaff, 2016).
In this study, we expand on previous small-scale research by estimating the decay of content knowledge for teaching using a large sample of K-12 math teachers from multiple PD programs representing diverse contexts in the United States. In addition, the data set contains a set of variables on teacher characteristics and program features, which allows us to examine empirically how some of the factors identified or hypothesized in prior studies affect knowledge decay after teachers complete their PD programs. Overall, this study provides more reliable and generalizable empirical evidence on the decay of teacher knowledge after they complete PD and how much the decay rates vary by teacher characteristics and program features.
Data and Methods
Sample
The analytic sample consists of data collected in the Teacher Knowledge Assessment System (TKAS), which is a Web-based platform administering assessments of mathematical knowledge for teaching, developed through the Learning Mathematics for Teaching project at the University of Michigan (http://sitemaker.umich.edu/lmt/home). TKAS was launched in June 2008 and has been widely used by researchers, PD providers, teacher educators, and program evaluators to assess teachers’ mathematical knowledge and its development. While the system allowed for pre, pre-post, and pre-post-delayed post designs, most of the TKAS users (over 90% of programs) chose to have a pre-post or pre-post-delayed post design to evaluate the knowledge change of teachers over time. To enable these administration designs, the system contains multiple forms that are equated to generate comparable test scores across multiple test administrations. In addition, based on the training content and evaluation need, administrators could select a combination of test forms from the following elementary and middle-school assessments in mathematics: (a) Elementary number concepts and operations (ELNCOP); (b) Elementary patterns, functions, and algebra (ELPFA); (c) Grade 4 to 8 geometry (GEO); (d) Middle school number concept and operations (MSNCOP); (e) Middle school patterns, functions, and algebra (MSPFA). Along with the assessments, TKAS administers surveys at each testing administration to teachers and PD providers. These surveys collect self-report data on program characteristics, teacher background, teacher motivation for participating in the training, and teacher reflection on the training experiences.
To ensure that the sample only included inservice teachers enrolled in PD programs, we limited the sample in several ways. First, we dropped preservice teachers and preservice programs from the sample. We also excluded programs with fewer than 10 teachers, given concerns about the reliability of the estimated knowledge decay rates drawn from such small programs. Finally, we only included teachers who had both pretest and posttest scores. After these exclusions, our sample includes 3,652 inservice teachers from 161 programs across the nation.
Next, we used latency data on how long each teacher took to complete the assessment to identify teachers who did not take adequate time to consider answers to the assessment questions. We excluded teachers who completed an assessment in less time than it takes to simply read through the test form at an average speed. After removing teachers with random careless responses, the sample size dropped by approximately 10%, but the total number of programs remained unchanged. Summary statistics on teacher and program characteristics look similar across the full sample and the sample of teachers who completed the test with a good faith effort. For the remainder of the analysis, we only include the 3,340 teachers who made a good faith effort.
The final sample contains a mix of K-12 teachers who differ in education background and teaching experience (Table 1). Specifically, 27% of the teachers have a mathematics degree at either the undergraduate or graduate level. About 14% of the teachers have less than 3 years of teaching experience and 48% of them have more than 10 years of experience. In terms of program characteristics, approximately, half of the programs in the sample span both summer and school year. Summer institutes account for a slightly larger percentage compared with programs that occur during school year. In addition, the program dosage varies considerably. The number of sessions and total contact hours range from 1 to 32 sessions and from 2 to 120 hours. In terms of the uses of different tests, there are 14 test combinations in the sample. The five most frequently used combinations are ELNCOP (47 programs), MSPFA (30 programs), ELNCOP & ELPFA (16 programs), MSPFA & MSNCOP (15 programs), and ELPFA (12 programs). The complete list of test combinations used by the programs is available upon request.
Teacher and Professional Development Characteristics.
Note. The full sample has 3,652 teachers and 161 professional development programs. After excluding teachers who did not make an appropriate effort at the test, the new sample includes 3,340 teachers and 161 programs. All the categorical variables are coded as 0/1. For example, mathematics degree is coded as 1 if teachers have a math degree and coded as 0 otherwise. The mean of a categorical variable shows the percentage of teachers in a particular category.
The percentages of grade and credential do not add up to 100, as multiple options could be selected.
Empirical Strategy
The most common empirical approach to studying knowledge decay is to use pre, post, and delayed post designs to model growth trajectories as a result of an intervention (Goldschmidt & Phelps, 2010; Heyns, 1978). However, for TKAS data, and many other large-scale studies with a delayed post design, such a growth analysis often suffers from a series of issues associated with substantial missing data at the delayed posttest occasion. In TKAS, about 44 programs (27.3%) had a design with pretest, posttest, and delayed posttest occasions. In these programs, the percentage of teachers who have pretest and posttest but do not have a delayed posttest is 64%. As a result, the common approach of modeling growth would lead to significant reduction in sample size (89%) and potentially biased estimated effects.
In this analysis, we take a different approach to estimating knowledge decay that takes advantage of one of the administration features built into the TKAS system. When designing an administration plan for a PD program, it is possible to set up the TKAS system to give teachers the flexibility to choose when they would like to complete the assessment. In the study sample, 49.69% of the programs allow teachers more than 2 weeks to complete their posttest and 26.71% allow more than a month. More importantly, these wide administration windows create large variations in how long teachers wait to take the test, and therefore, it is possible to use this feature of TKAS to construct a measure for each teacher for the number of days that they wait to take the test.
Using a multilevel model with teachers nested in programs, we estimate the relationship between the number of wait days and posttest scores across teachers after accounting for program- and teacher-level differences. The model is specified below as follows:
or
where Waitij is the key explanatory variable that indicates how long teachers waited to take the test. Specifically, it is 0 if teacher i is the first test-taker in program j. We use this relative measure of wait time across teachers so that the coefficient of Waitij reflects the change in posttest score across teachers having different wait times within a program. 1 In addition, we include quadratic and cubic function forms of the Waitij to test if the decay rate estimate is relatively stable across time.
To control for individual differences that are associated with both Waitij and PostScoreij, we included PreScoreij as well as a set of teacher characteristics (Xij). The test scores are normed to the TKAS population, with a mean of zero and a standard deviation of one. Both the pretest and posttest scores used in the model are the average of these test scores from all assessments used by each program. These average scores should capture what teachers learned from the PD programs, assuming that the program administrators choose assessments aligned with their program content. PreScoreij is included to account for differences in preexisting knowledge. In addition, our data suggests that the pretest and posttest scores have a quadratic relationship. To account for this, a square term of pretest score is added into the equation. The list of teacher characteristics as denoted by Xij include demographics (race and gender dummy variables), teacher background variables (teaching experience and whether they have a math degree at either undergraduate or graduate level), and four composite variables that were built from Likert-type-scale items in teacher surveys. The composite variables include teachers’ reflections on the relevance and importance of the test questions, satisfaction with the PD program, and intrinsic and extrinsic sources of motivation for attending the programs.
This multilevel model allows the intercept and coefficient for Waitij to vary at the program level. As teachers are clustered within programs, a varying intercept (Equation 2) is required to capture the random program effect on teachers’ posttraining knowledge. 2 The random coefficient, as specified by Equation 3a, allows for program-specific estimates of teacher knowledge decay rates. In Equation 3a, π300 represents the estimated average knowledge decay rate across all programs.
Furthermore, to address the second research question, Equation 3b includes program-level covariates that are associated with decay rate. Here, π31 estimates differences in decay rate across different types of programs and is essentially the coefficient for the interaction term between wait time variable and program-level characteristics. Two kinds of program-level predictors (X3j) for the random coefficient for Waitij are examined. The first set of variables representing the design features of the programs is constructed as follows. (a) Program type, defined using the starting month and the length of the program testing window. This variable categorizes programs into three types: summer institute, programs taking place during the school year (school year program), and programs that span both summer and school year (whole year program). (b) Number of training sessions, which indicates how many times teachers are required to meet. (c) Contact hours, which is based on providers’ reports of total hours of program training. (d) Intensity, defined as contact hours divided by the number of training sessions. (e) Uses of different tests, which characterize programs by different combinations of assessments they chose to use and, therefore, also reflect the content focus of the program. The second set of program-level variables is constructed by aggregating participants’ characteristics to the program level. These variables include: (a) average initial knowledge level measured by average pretest score; (b) average years of classroom teaching experience; and (c) percentage of teachers who have a mathematics degree.
Assessing Threats to the Validity of Knowledge Decay Estimate Using Robustness Tests
In the aforementioned multilevel model, the estimated coefficient for Waitij, if positive would be attributed to evidence of continued learning after PD programs, and if negative would be attributed to knowledge decay. However, as teachers are allowed to self-select, the effects of waiting to take the test could also reflect teachers with lower posttest scores selecting into later testing dates. This selection effect is an obvious threat to the inference that the effect for waiting to take an assessment can be attributed to knowledge decay.
We tested and controlled for the potential unwanted selection in the following three ways. First, we controlled for program-level differences to avoid having selection effects that are produced due to differences in program features. For example, good PD programs, which are more effective at improving teacher knowledge, could also motivate teachers to respond more promptly. Our multilevel model compares teachers who participated in the same PD program and thus could control for this particular source of self-selection.
Second, we controlled for a series of teacher-level covariates that are related to both the number of wait days and the posttest score to reduce any confounding effects on the knowledge decay estimator. For example, teachers who do well on the pretest might be more eager to take the posttest and are also more likely to perform better on the posttest. The knowledge decay estimate might be downward biased if the difference in pretest score is neglected. The teacher-level covariates included should be able to correct for the bias caused by observed individual differences.
Despite the rich set of covariates available, the list of control variables remains inevitably incomplete. As a result, the relationship between the test score and the number of wait days might still be driven by factors other than teacher knowledge decay. To further investigate, we conducted three additional analyses to rule out alternative hypotheses regarding the relationship between wait days and test scores. In the first analysis, we ran a series of sensitivity analyses to test if the decay rate estimate is driven by a small number of extreme values that come from teachers who waited for a long time to take the test. In the full sample, a small percentage (3.4%) of teachers waited for 50 days or more to take the test. This group of teacher may not represent the average teachers attending PD programs. To find out if our decay rates estimate changes after excluding extreme groups, we fit the same multilevel model with teacher-level covariates described earlier using samples that include teachers who completed the test within 50, 60, or 70 days.
In the second analysis, we ran a counterfactual test to estimate the effect of wait time at the pretest administration. Pretest data provides a powerful counterfactual test because these data are collected before the PD and because the treatment and associated learning has not taken place. Therefore, knowledge decay should not be observed at pretest. If a similar pattern is observed for wait days at pretest, then the observed effect of wait time could be attributable to teachers self-selecting into later testing according to a variety of unobserved personal characteristics. Conversely, if under this condition, there is no effect for wait days, then the counterfactual argument would indicate that any effect observed for wait days at posttest is more likely to be associated with the knowledge decay and not these unobserved characteristics. To estimate the effect of wait time at pretest administration, we used a similar multilevel model with pretest scores as the outcome variable. In this model, the key predictor is the number of wait days at pretest administration. The covariates used include gender, race, teaching experience, mathematics degree, and the four composite variables.
Finally, we tested whether our knowledge decay estimate is clouded by reverse causality. Reverse causality occurs if the amount of teacher learning affects when teachers take the test—that is, teachers who learned the most take the test first. This is problematic because the decay estimate from our model captures the true knowledge decay as well as this pattern caused by reverse causality. To test if teacher learning has an impact on wait time, we used a self-report measure of learning. Participating teachers, in addition to completing the preknowledge and postknowledge tests were asked to self-rate their learning after PD. While self-reports may not be as accurate as tests of knowledge, we think it is less likely that personal impressions of learning will decay. Self-reports therefore provide an alternative measure of learning that does not reflect knowledge decay in the same way as test scores. Specifically, we constructed a variable for teachers’ self-report learning by averaging two Likert-type scale 3 items ranging from 1 to 6 and used this variable as the outcome measure. In addition, we used the wait time variable as the predictor and other teacher characteristics as the covariates. Results from this model could help us find out whether wait days and teacher learning are related, and a lack of a relation between wait time and self-report learning suggests that the knowledge decay rate estimate should not be biased due to reverse causality.
Results
Knowledge Decay Estimates
Table 2 presents the main results from the multilevel model. The first model specification only includes the number of days teachers wait to take the test as the independent variable and allows the intercept and the coefficient for the wait variable to vary at the program level (Column 1). Next, we add teacher-level covariates into the model to test if the estimated decay rate is driven by teacher characteristics (Column 2). The decay rate estimate does not change after adding teacher covariates into the model, indicating the estimate is robust to the inclusion of these teacher characteristics. For heterogeneous analyses, we chose to build on the model in column 2 because all model-fit statistics suggest that this model provides a better fit than the one without teacher-level covariates. In addition, the function form of the model has important implications for how the decay estimates should be interpreted. So, we tested the quadratic and cubic function form of the model, but found that the square and cubic terms are statistically insignificant. The model-fit statistics (i.e., log likelihood, Akaike information criterion [AIC], and Bayesian information criterion [BIC]) also suggest that the linear model fits the data better. Overall, the evidence suggests that the decay rate is relatively stable following the completion of the training. 4
Estimates of Teacher Knowledge Decay Rate From Multilevel Models.
Note. Teacher-level covariates include pretest score and its square term, gender, race, teaching experience, mathematics degree, and composite measure of teachers’ reflection on the relevance and importance of the test questions, satisfaction with the professional development programs, intrinsic and extrinsic motivation for program participation. We include log likelihood, AIC, and BIC to inform the model selection. The difference among these indicators is that the log likelihood could be increased by simply adding covariates. Therefore, increased log likelihood does not necessarily reflect better model fit. In contrast, AIC and BIC are not sensitive to the number of covariates included because they introduce a penalty for including more parameters in the model. All the indicators suggest that the model in column 2 is a better fit than the model in column 1. AIC = Akaike information criterion; BIC = Bayesian information criterion.
p < .05. ** p < .01. *** p < .001.
A consistent pattern in both models is that the knowledge decay estimate is negative and significant at −.003 (p value < .01), indicating that on average, waiting one day to take the test would be associated with scores that are a 0.003 standard deviations lower. 5 To understand the substantive implication of this estimate, we compare the estimated decay rate to pre–post teacher knowledge change as a result of the PD training. Specifically, we calculate the average difference of pretest and posttest score in the TKAS sample. The average observed change of .1109 would be lost in 37 days (.1109/0.003) with the decay rate being stable over time. To provide a more accurate benchmark on the magnitude of teacher learning during PD programs, we use estimates from a recent study (Phelps et al., 2016) that estimated pre–post teacher knowledge change using cross-classified hierarchical growth model and TKAS data. This study estimated the average change in scores for each of the TKAS assessments. As the test on ELNCOP is used by almost two thirds of the programs in TKAS, we use changes in teacher knowledge measured by this assessment as the proxy. In that study, the estimated average change in test scores in ELNCOP has a mean of 0.18 and a standard deviation of 0.17 at the program level. Again, assuming the decay rate is stable over time, the average knowledge gains for programs using this assessment would be lost in approximately 60 days (0.18/0.003). Overall, these findings suggest that some PD programs could have significant effects on improving teacher knowledge, but the effects do not sustain—our knowledge decay estimate shows that the average knowledge gains from the programs in our sample would be lost in just 37 days after teachers complete the training.
Robustness Tests
One piece of evidence demonstrating the robustness of the decay estimate is that the variation in estimated decay rate across samples that include different subsets of teachers is small (Table 2). Our results show that there is only a small number of teachers who waited 50 to 70 days after the completion of the trainings (43 teachers), and the estimated decay rate for teachers who completed the test within 70, 60, or 50 days are −0.0032, −0.0033, and −0.0034, respectively. This suggests that the estimated decay rate is not driven by a number of extreme cases and is relatively stable across samples.
In addition, the counterfactual test at the pretest administration generates insignificant estimates using a similar model, and the estimates for wait days are very close to zero in the model with or without teacher-level covariates (Table 3). Furthermore, the estimated program-level standard deviation for the coefficient of wait time variable is close to zero, indicating that there is no significant relationship between pretest score and the number of wait days at the pretest for all programs in the sample. These zero estimates suggest that any teacher characteristics that remain constant at pretest and posttest would not produce a relation between the number of wait days and the posttest scores. More broadly, the counterfactual test helps to rule out the possibility that the pattern captured at the posttest is attributed to any variation in observed or unobserved variables that stay constant across the two testing points. This is a good piece of evidence supporting the model in addition to the program-level and teacher-level covariates that account for possible self-selection.
Counterfactual Estimates Using Pretest Score as Outcome Variable.
Note. Teacher-level covariates include gender, race, teaching experience, mathematics degree, and variables of factor score for teachers’ reflection on the relevance and importance of the test questions, satisfaction with the professional development programs, and intrinsic and extrinsic motivation for program participation. We included log likelihood, AIC, and BIC to inform the model selection. The difference among these indicators is that the log likelihood could be increased by simply adding covariates. Therefore, increased log likelihood does not necessarily reflect better model fit. In contrast, AIC and BIC are not sensitive to the number of covariates included because they introduce a penalty for including more parameters in the model. AIC = Akaike information criterion; BIC = Bayesian information criterion.
p < .05. ** p < .01. *** p < .001.
Finally, we find that the knowledge decay estimate may not be biased by reverse causality. Specifically, we find an insignificant relationship between wait time and self-report learning (Table 4), indicating that there is no obvious pattern that teachers who self-report learning the most during the training take the test first. The result suggests that the estimates of knowledge decay rate is not influenced by this particular mechanism.
Estimates of Teacher Knowledge Decay Rate Using Self-Report Learning Measure.
Note. Teacher-level covariates include gender, race, teaching experience, mathematics degree, and composite measure of teachers’ reflection on the relevance and importance of the test questions, satisfaction with the professional development programs, and intrinsic and extrinsic motivation for program participation. We include log likelihood, AIC, and BIC to inform the model selection. The difference among these indicators is that the log likelihood could be increased by simply adding covariates. Therefore, increased log likelihood does not necessarily reflect better model fit. In contrast, AIC and BIC are not sensitive to the number of covariates included because they introduce a penalty for including more parameters in the model. All the indicators suggest that model 2 is a better fit than model 1. AIC = Akaike information criterion; BIC = Bayesian information criterion.
p < .05. ** p < .01. *** p < .001.
Knowledge Decay Estimates by Program and Teacher Characteristics
We find that there is substantial variation in the magnitude of the knowledge decay estimate across programs. The estimates range from −0.009 to 0.003 at 95% confidence level, 6 although estimates for most of the programs fall within −0.004 to −0.002 (Figure 1).

Density Plot of Estimated Knowledge Decay Rate at Program Level
In our analysis examining the heterogeneous knowledge decay rates, the variation in the decay rate is associated with program features and participants’ characteristics. The estimate of decay rate varies significantly by program types. Table 5 shows that the average knowledge decay rate for summer institute is −0.009, whereas the estimates for school year and whole year programs is −0.002. 7 When comparing the decay rate of different types of programs to their corresponding gains, we find that summer institutes have the lowest gain yet the highest decay rate. The approximate knowledge gain is 0.05 standard deviations for summer institutes and 0.13 standard deviations for the other two types of program designs.
Heterogeneous Decay Rate by Programs and Teacher Characteristics.
Note. Teacher-level covariates include pretest score and its square term, gender, race, teaching experience, mathematics degree and variables of factor score for teachers’ reflection on the relevance and importance of the test questions, satisfaction with the professional development programs, and intrinsic and extrinsic motivation for program participation. All the program and teacher characteristics variables (except for program type) are centered to the mean for the ease of interpretation.
π30 represents the estimated knowledge decay rate for the reference group. π31 is the estimate of difference in knowledge decay rate between the current group and the reference group. The estimated knowledge decay rate for whole year and school year program is −0.0024 (−0.0092 + 0.0068) and −0.002 (−0.0092 + 0.0072).
p < .05. **p < .01. ***p < .001.
In addition to the program types, we examine three other program feature variables: program intensity, number of session, and contact hour. All of these program-level variables are continuous variables and are centered to the mean to allow for examination of the heterogeneous effect at the average decay rate rather than at the extreme values. The effect of intensity on knowledge decay is negative and moderately significant, suggesting that the more compressed the training, the faster the acquired knowledge is lost. The effect of the number of sessions is only significant at .1 level (p value = .0680). The positive estimate indicates that more sessions are associated with higher knowledge retention. Finally, the results show that decay rates do not vary by the number of contact hours. Together, these findings suggest that contact hours might not reduce knowledge decay if training sessions were conducted within a short period of time. Teachers are more likely to sustain their knowledge after attending PD programs with longer duration and ongoing activities over time.
It is worth noting that the heterogeneous decay rates by program types might be partly explained by the fact that summer institutes are usually more intense with fewer sessions compared with school year programs. A group comparison using analysis of variance (ANOVA) shows that the intensity differs significantly across program types, with summer institutes having the highest intensity and whole year programs having the lowest (F value = 4.92, p value = .0091). As higher intensity is associated with higher knowledge decay rate, summer institutes with the highest intensity would be expected to have the highest decay rate among the three program types. Finally, the results also show that there is no significant variation in decay rate estimates across different test combinations. The interaction term is very small and statistically insignificant.
We next investigate how decay rate varies by participants’ characteristics. As the decay rate is estimated at the program level, we can only examine how it varies by the average characteristics of participants in a program. Results show that programs with a higher initial average score have a higher level of knowledge sustainability. A program with the lowest average level of initial knowledge has a decay rate that is twice as high as the rate for a program with an average knowledge level. However, neither the average teacher experience nor the percentage of teachers with a mathematics degree is found to be significantly associated with the decay rate.
To summarize, we find significant and substantial knowledge loss after teachers complete the PD programs. The estimated knowledge decay rate is −0.003, which indicates that the knowledge gain from pre to post for programs in this sample would be lost in 37 days, with decay rate being stable across time. Furthermore, the decay rate is affected by program features. Summer institutes have a much higher decay rate than programs that span a school year or the whole year. Programs with higher intensity or fewer number of training sessions are also associated with higher decay rate. These program features might affect the decay rate through a similar mechanism, as summer institutes are usually characterized by high intensity and small number of training sessions. Finally, programs with higher average teacher initial knowledge have less knowledge decay, but other teacher characteristics, such as average teaching experiences or having a math degree, are found to be unrelated to the decay rate in the context of teacher PD programs.
Discussion
This study makes use of a large-scale data set to empirically examine the extent to which teachers retain knowledge following PD program. The results show that teacher knowledge is not stable after PD. In fact, the loss of knowledge, or knowledge decay, does take place, and it may be quite substantial. If improving student learning is the goal of the PD programs, the focus of these programs should not only be on what teachers learn but also on whether the learning lasts long enough to make it to the classroom. In addition, the study shows that knowledge decay rate varies considerably across programs, and the decay rate is affected by teachers’ initial knowledge level as well as the design of PD programs (e.g., program type and intensity).
Implications
While there is more to learn, results from this study suggest several components of PD programs that are important for sustaining posttraining teacher learning. One important component is the duration of the program. Prior studies showed that longer duration of PD led to larger improvements in teacher knowledge or teaching practices (Copur-Gencturk & Papakonstantinou, 2016; Garet, Porter, Desimone, Birman, & Yoon, 2001). Likewise, our study found that programs with greater number of training sessions or longer duration are associated with lower knowledge decay rate, suggesting that longer program duration not only increases the impact of PD on teachers but also sustains the impact afterwards. Another component is the provision of follow-up activities after intensive training sessions. We found that summer programs have higher knowledge decay than whole year or school year programs. While alternative explanations are possible, a major difference between these programs is that unlike summer program, whole year or school year programs provide opportunities for teachers to immediately apply new knowledge in their own classrooms. School year programs also often provide follow-up support to teachers. These supports are important for teachers to internalize and apply key ideas learned from the training and, as a result, may reduce knowledge decay following the completion of PD. In designing or implementing teacher PD program, it could be important to consider these findings on knowledge decay together with those design components that improve program effects to increase teacher learning in the short and long run.
In addition to program design, the study informs the design of research studies evaluating the effectiveness of teacher PD. Policymakers might consider including the sustainability of program effects as another criterion for assessing program effectiveness. Existing research has found that teacher PD programs can have significant and substantial effects on teacher knowledge, supporting claims by policymakers and researchers that teacher PD training is an important lever toward improving instruction and student learning (Desimone, 2009). However, to date, research on the effectiveness of these programs has treated the program effects as if they are stable. This is against the viewpoint from cognitive scientists that knowledge or skills tend to be unstable and changing over time. The decay of knowledge gains is particularly salient when there is a lack of use of the knowledge acquired (Kim, Ritter, & Koubek, 2013). Those program evaluation studies that neglect the changing nature of teacher knowledge may provide incomplete or even misleading evidence on the effectiveness of the programs.
More broadly, knowledge decay has important implications not only in the context of teacher PD, but also throughout the entire continuum of teacher learning. For example, previous research suggests that teacher candidates and teachers often have gaps in their knowledge (Depaepe et al., 2015; Donna & Hick, 2017; Wilburne & Long, 2010). This occurs because candidates may not be adequately prepared by their teacher education programs and may not be able to keep deepening their knowledge after they start teaching. Other than that, teacher knowledge decay is another possible reason for the existing knowledge gaps. So far, there has been a lot of research discussing promising programs and strategies that prepare teachers and support them throughout their career (Abd-El-Khalick & Lederman, 2000; Alves et al., 2018; Christ, Arya, & Chiu, 2017; Darling-Hammond, 2012; Feiman-Nemser, 2001; Tatto & Menter, 2019). While these programs have focused more on teacher learning, our study suggests that it is worthwhile to consider knowledge decay as an equally important factor when designing, implementing, and evaluating programs that support teacher learning.
Limitations and Future Directions
This study estimates teachers’ knowledge decay rate after they complete the PD. While this approach does not require a second posttest score to estimate knowledge decay and thus does not suffer from missing data issues, it is limited in several other ways. The biggest concern is that the effects we observe might result from self-selection rather than true knowledge decay. However, as presented in the results section, our robustness test does not show evidence of self-selection. Furthermore, both the models we used to estimate the knowledge decay rate and the controls we entered in the models further guard against the possibility that the findings are due to selection. Still, we recognize that it is not possible to definitively claim that the effects of waiting to take a test are due to knowledge decay and not to other factors. To avoid this problem, future PD evaluation studies could collect data on treatment and control groups and use the standard pre, post, and follow-up design to examine the knowledge sustainability after the training. In practice, this design usually requires larger sample size and time commitments. In addition, the follow-up data collection is usually challenging with possible nonrandom sample attrition. For example, high performers are often more likely to respond to the follow-up, which leads to insignificant knowledge decay estimates. So, the choice of this design should be considered with its feasibility in the specific context as well as resources available to the study.
Second, while we know some aspects of the PD programs, it would be helpful if we have more information regarding the PD intervention, the teachers who participated, and also the school context in which these teachers work. For example, we do not know whether different decay rates for summer, school year, and whole year programs are a result of different timing, specific training activities, or other key measures of the program quality. In addition, prior research (Arthur et al., 1998) suggests that differences in decay rate might be attributed to whether there is a coordinated effort between schools and PD programs to train teachers on the most relevant knowledge or skill and allow them to apply it in schools. As it is suggested by our previous review of the literature (Roth et al., 2011; Yoon et al., 2007), details on different attributes of PD, including its context, content focus, and approach, would have been very useful to decipher how the key program features are connected to knowledge decay rates. Due to the unavailability of the information, our current analysis is unable to uncover what the key components of the programs are that lead to different knowledge decay rates.
Finally, the functional form of the model is another issue and a potential limitation of the study. In previous research, decay has typically been modeled with a power-form curve or log-linear learning model (Jaber, Kher, & Davis, 2003; Wright, 1936). However, we found that the relationship between the wait time and test score appears to be linear. For example, we tested a number of nonlinear specifications by adding polynomials of wait time, but none of these fit the data better than the simple linear model. Furthermore, the simple linear model allows for a straightforward interpretation of the average knowledge decay estimates. While these points explain the approach we took in our analyses, we acknowledge that our findings are limited by the lack of accurate prediction of knowledge decay rates at different points in the distribution of wait days. Furthermore, we do not have data on instruction quality and student outcome measures. Future studies with data on all these measures could provide a better understanding on the underlying mechanism of how decay of teacher knowledge is connected to the impact on teaching instruction and student outcomes.
Conclusion
The findings from this study strongly suggest that in studying the outcomes of PD and the effects of PD on teaching and learning, it is important to account for more than immediate teacher learning outcomes. A more comprehensive approach would consider a program’s true effect as the combination of learning and forgetting. From the way studies are currently conducted, we simply do not know to what extent learning is maintained, and consequently, we do not know the programs’ true value. Our findings suggest that it will be just as important to study both the stability of these learning effects and the conditions that support teachers in retaining knowledge. The importance of such studies becomes even more salient, given the current efforts to use teacher PD programs as one of the main policy levers for improving teacher knowledge and skill, teaching quality, and ultimately student achievement. Districts and states continue to devote substantial research funds to teacher education and PD programs; the true benefit of these investments will be difficult to ascertain without understanding both how teachers learn and forget critical knowledge and skills.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This study was supported by a grant from the National Science Foundation (award no. 0927725). The opinions expressed herein are those of the authors and not the funding agency.
