Abstract
Leadership preparation is essential in developing critical skills of effective school leaders. Educational leadership preparation programs vary greatly in terms of program features, it is important to identify which specific elements are likely associated with program quality. Using multiple years of Initiative for Systemic Program Improvement through Research in Educational Leadership (INSPIRE) data, this study analyzes the relationships between program features and graduates’ assessments of overall program quality. Our main finding is that multiple program features matter in quality preparation, including cohort model, face-to-face course delivery, field-based assessment, and strong district partnerships. Implications for policy and future research are discussed.
Keywords
Introduction
Principal leadership is recognized as instrumental in shaping teacher satisfaction, instructional effectiveness, and school improvement (Grissom & Youngs, 2016; Hitt & Tucker, 2016; Leithwood et al., 2020). The effects of principal leadership on student learning are considered largely indirect (Hallinger & Heck, 1996), operating through improving various school processes including professional capacity, school climate, parent-school ties, and instruction guidance (Bryk et al., 2010; Sebastian & Allensworth, 2012). Prior research suggests that, of all school factors, the effect of school principals on students’ learning is second only to that of classroom instruction, explaining about one quarter of all school effects (Leithwood et al., 2004). A recent Wallace Foundation sponsored literature review suggested that the impact of an effective principal on students is even greater in magnitude than previously estimated, and the scope of the impact is much broader and extends beyond students’ academic achievement (Grissom et al., 2021).
There is a growing awareness that high quality leadership preparation is essential in developing critical skills of effective school leaders who can advance organizational effectiveness and promote teaching and learning in schools (Ni et al., 2017, 2019). Educational leadership preparation programs (ELPP) have extended beyond the walls of universities, including emerging as district “home-grown” programs. Yet, they remain largely university-based in educational leadership departments (Perrone, 2019). ELPPs vary in terms of program features, including recruitment and selection of candidates, curriculum focus, instructional strategies, program delivery methods, internship practices, and program and/or candidate assessment and evaluation (Young et al., 2009). These features are shaped by both internal elements within university settings and external or contextual changes. Internal elements include such factors as greater expectations of faculty engagement with local community, [mis]alignment of faculty rewards and incentives, faculty recruitment and retention challenges, and academic norms and standards. External and contextual changes include such factors as increased state and professional policy focus on assessment and accountability, critiques of university-based programs, market competition, growing state influence, and changes in the call for what the content of preparation programs should attend to (Crow & Whiteman, 2016; LaMagdeleine et al., 2009).
Although critics have accused ELPPs, especially university-based ELPPs, of failing to adequately equip individuals with the knowledge and skills necessary to be successful in their roles in schools (e.g., Levine, 2005), a growing body of research on leadership preparation suggests that selected program features are likely to be associated with better gradate outcomes (Darling-Hammond et al., 2010; Grissom et al., 2019; Ni et al., 2019). However, the current research base on ELPP features and their effectiveness is both largely descriptive and emergent (Ni et al., 2017). This study extends our knowledge about ELPPs by examining the relationships between various elements of ELPP and program quality. Specifically, this study addresses:
How do educational leadership preparation programs vary by program features?
Are specific program designs and features associated with graduates’ assessments of the quality of their leadership preparation programs?
Leadership Preparation Programs: The Need to Be a Learning Organization
This study is grounded in the theory of organizational learning. At least rhetorically, the field of education has relied upon the theory broadly to both explain and drive the needs and outcomes associated with continuous improvement in the past 20 years. Argyris and Schon (1978) explained that organizational learning relies on the detection and correction of error. Despite the pronouncement of continuous improvement as a rationale for many changes in the field, we argue that giving attention to the roots of organizational learning, particularly the focus on single- and double-loop learning (Argyris & Schon, 1978), provides needed insight and guidance into whether university-based ELPPs remain central to leadership preparation (Young et al., 2019). Based on his long-standing work in the area of organizational learning, Argyris (1999) described single-loop learning as occurring “whenever an error is detected and corrected without questing or altering the underlying values of the system,” and contrasted it to double-loop learning that occurs “when mismatches are corrected by first examining and altering the governing variable [assumptions—values, beliefs, norms, goals] and then the actions” (p. 68).
Arguably, preparation programs have engaged in continuous improvement efforts, albeit at different levels and for different reasons (Young et al., 2009). However, often these efforts remain at the level of single-loop learning. External demands (e.g., the adoption of leadership standards for both preparation and practice; shifts in state policy; accreditation; accountability; available resources to support program improvement; increased market competition among university and non-university entities; and even a pandemic) have prompted examination and changes in faculty appointments, practitioner engagement and partnerships, and curriculum. Concurrently, internal demands such as faculty foci, growing knowledge and research-base on program quality factors, and increased commitment to equity and social justice have resulted in further and deeper programmatic changes (e.g., double-loop learning).
Situating ELPPs as a learning organization allows us to consider the long-term role that university-based ELPPs serve in preparing future generations of school, district and state educational leaders. As Garvin (1993) explains, a learning organization reflects “an organization skilled at creating, acquiring, and transferring knowledge, and at modifying its behavior to reflect new knowledge and insights (p. 80).” Learning organizations typify six practices: systemic problem solving, experimentation, learning from past experience, learning from others, transferring knowledge, and measuring learning (Garvin, 1993; Goh et al., 2012). To better prepare leaders, ELPPs are situated to advance purposeful program design and improvement for individual, organizational, and system learning to advance specific, targeted goals. Studying ELPP effectiveness from the perspective of learning organizations, as we do here, will help us understand the extent to which preparation programs are (or could be) learning organizations. It will help determine a base for analyzing whether being a learning organization is predictive of meeting the preparation needs of future leaders and of remaining viable long-term as a program.
Empirical Evidence on Effective Leadership Preparation Programs
Empirical studies focusing on ELPP program features and the relationships between program features and graduates’ learning and practices as school leaders have increased in recent decades. The Handbook of Research on the Education of School Leaders (Young et al., 2009) was the first to provide the most thorough and in-depth examination of the literature up to 2007 on various dimensions of the preparation of educational leadership and the second edition of the handbook updates the examination by reviewing research that has emerged since 2008 (Young & Crow, 2017). About 70 authors contributed to the two handbooks identifying major characteristics of exemplary leadership preparation program design including, but not limited to, rigorous student recruitment and selection, standards-based and coherent curriculum, problem-based learning strategies, field-based internships, cohorts for learning, and collaboration between universities and school districts (Young & Crow, 2017; Young et al., 2009). Here we offer an overview of the research to date on these characteristics and the areas that need to focus attention in the field.
Program Admissions
With respect to student recruitment and selection, many ELPPs, often driven by demands for enrollment, rely on minimally selective admission approaches such as the minimum institutional academic criteria (e.g., Grade Point Averages and/or Graduate Record Examination scores) (Browne-Ferrigno & Muth, 2009). However, there is increased understanding that it is important to identify individuals who aspire to a career in educational administration and leadership and who have special abilities and experience (Jacobson et al., 2015). More rigorous selection criteria should include clear commitment and intentions, academic expectations, career aspirations, and prior experiences that suggest candidates have leadership potential (Browne-Ferrigno & Muth, 2009; Fuller et al., 2017). Yet there is limited extant research that establishes the link between selection efforts and any outcome measures (Fuller et al., 2017).
Program Content and Delivery
Curriculum
ELPP accreditation can occur nationally, state, or regionally, depending on university or state level policies. In recent years, public demand for more effective schools and accountable public schools has placed an increased demand on ELPPs. Part of this demand has required ELPPs to intentionally realign their curriculum to meet national standards based on effective leadership research, such as the recent Professional Standards for Educational Leaders (PSEL) and National Educational Leadership Preparation (NELP) standards. These national standards emphasize different dimensions of leadership knowledge, skills and dispositions, including ethical and professional norms, strategic leadership, operations and management, instructional leadership, professional and organizational culture, equitable and supportive learning environment, and family and community engagement (NPBEA, 2015; NPBEA, 2018). Concurrently, there has been an increase in a curricular emphasis on social justice leadership in ELPPs (Ylimaki & Henderson, 2017). Since most studies on ELPP curriculum have been normative and descriptive, there is a lack of research on how these content areas are taught and their effectiveness (Crow & Whiteman, 2016), despite the affirmation that coherent and comprehensive standards-based curriculum matters in preparation programs (Darling-Hammond et al., 2010).
Instructional strategies
Research on instructional strategies in ELPPs is largely built upon transformational learning theories, including the expectation that ELPPs should offer opportunities for students to engage in inquiry and problem-solving as a method to better prepare candidates to adopt, or adapt, a broad and inclusive approach in addressing issues of student learning and equity in schools (Byrne-Jimenez et al., 2017). Darling-Hammond et al.’s (2010) case study of exemplary programs argues that active instructional approaches linking theory and practice are effective, including problem-based learning, action research, and field-based projects. Reviewing a large amount of literature, Byrne-Jimenez et al. (2017) identified five key instructional strategies that should exist in ELPPs: problem-based learning, simulations, case studies, critical reflection, and critical discourse. Despite this proclamation, the research base of the effectiveness of these instructional strategies remains thin.
Faculty
Faculty quality is essential to program rigor and graduates’ learning experiences in ELPPs (Ni et al., 2019). Although tenure-track faculty remain as the majority of the program personnel, many ELPPs continue to involve a critical mass of clinical and adjunct instructors (Hackmann & McCarthy, 2013). Although scholars suggest that the majority of courses should be taught by full-time faculty members (Young et al., 2012), ELPP faculty are encouraged to collaborate with field-based practitioners for instruction and other activities such as admissions, supervision, and assessment. However, there is a lack of research in terms the effectiveness of these practices.
Cohort model
Many ELPPs use cohort models for program delivery. Research has identified the strength of cohort models, including building peer relationships, trust, and peer networking, creating opportunities for collaboration, developing a sense of trust and community, increasing the likelihood of program completion, and improving leadership practices (Donmoyer et al., 2012; Greenlee & Karanxha, 2010; Salazar et al., 2013). On the other hand, there are several concerns about cohort models, including negative group dynamics, a lack of flexibility for students, and limited research on the effectiveness of cohort models in graduate outcomes (Byrne-Jimenez et al., 2017; Crow & Whiteman, 2016).
Technology use
Technology use in ELPPs refers to how to prepare school leaders to be better technology leaders and use digital technology to teach leadership content and deliver the programs (Dexter et al., 2017). However, there exist significant gaps on how ELPPs are incorporating technology instruction into their programs and on the effectiveness of different practices. Limited research showed that online delivery models created less sense of community than face-to-face models, but made no difference in graduate learning outcomes (Choi et al., 2005; Ritter et al., 2010).
Student Assessment
Despite calls to assess ELPP students, there remains a significant amount of controversy regarding how to best assess the competency of these students. Research has shown ELPPs use various types of assessments (e.g., portfolios, journal mapping, simulations, and formal multi-rater assessment instruments) at different stages of program study (formative, quarterly or mid-program, and summative) (Crow & Whiteman, 2016; Korach & Agans, 2011). Some scholars have called for valid assessments with authentic techniques featuring field-based experiences that examine not simply knowledge, but students’ abilities to perform essential duties of school administrators (Kochan & Locke, 2009). Because different ELPPs are accredited by different state or national bodies with different standards, assessments to date have varied. Similar to other areas of research on ELPPs, research on assessment tends to be descriptive rather than evaluative, and there is only limited evidence indicating positive relationships between assessment practices and student progress and degree completion (Crow & Whiteman, 2016; Orr & Hollingworth, 2018).
Internship and Local Partnership
High quality and authentic internship/residency experiences have long been recognized as a core preparation component. Positive internship experiences provide intensive, developmental opportunities to apply leadership knowledge and skills supported by high-quality mentoring and coaching (McCarthy, 2015; Pounder, 2011). They also improve candidates’ learning about their leader knowledge and skills (Ni et al., 2019), which in turn increases the likelihood of pursuing leadership careers (Crow & Whiteman, 2016; Orr & Barber, 2006). A fair amount of research has been devoted to the importance of mentoring and coaching as an important component of the internship experience, although it still lacks rigorous examinations of mentor selection and the effectiveness of mentoring (Crow & Whiteman, 2016; Reyes-Guerra & Barnett, 2017).
The Wallace Foundation (2016) identified strong university-district partnerships as essential to high-quality preparation. In authentic and active partnerships, local districts are involved in leadership preparation in multiple ways, including collaboration on selection and recruitment of students, ongoing support of program elements, and providing a clinically-rich internship experience (The Wallace Foundation, 2016). Several studies tout benefits of the partnership as bridging theory and practice, creating more delivery options, emphasizing collaborative leadership, increasing levels of commitment of candidates, and providing reciprocal benefits for districts (Darling-Hammond et al., 2010; Jacobson et al., 2015). Although the development of solid and formal university-district partnerships are greatly encouraged (Young, 2010), the knowledge base remains limited regarding the characteristics of meaningful partnerships and their effects.
In conclusion, although some evidence has emerged from the literature that ELPPs with major exemplary elements are found to generate more positive graduate outcomes, including their perceptions of learning, career intentions and job placement, leadership practices, and indirectly student learning outcomes, the amount of the empirical evidence remains thin and preliminary. In addition, many existing studies are case studies that focus on a single program and rely on limited samples and inadequately developed analytic frameworks. The few quantitative studies tend to be descriptive and rely on cross-sectional investigations (Crow & Whiteman, 2016). Taken together, these studies produce a grounded hypotheses and serve as the base of our knowledge to develop more rigorous, large-scale, quantitative evaluation studies that incorporate multiple data sources and types of data and promise richer understanding of leader preparation program elements and their effects or outcomes.
Data and Methods
The Initiative for Systemic Program Improvement through Research in Educational Leadership (INSPIRE) Research Collaborative has developed a leadership survey suite with multiple components to better understand the experiences of graduates from their ELPP completion to practice. Aligned with national standards such as the PSEL and NELP standards, the INSPIRE survey provides users an opportunity to utilize highly reliable and valid standards-based instruments focused on program improvement, leadership experience development, improved leadership practice, and stakeholder support. The INSPIRE surveys have been psychometrically tested for reliability and validity, and improved through multiple years of field-testing (Ni et al., 2019). The full INSPIRE Leadership Survey Suite includes: (a) a Preparation Program Survey (INSPIRE-PP) that describes program features and processes that comprise a program, (b) a Graduate Survey (INSPIRE-G) that includes graduates’ assessments of program quality and leadership learning outcomes, and (c) a Leaders in Practice Survey (INSPIRE-LP) and (d) a Leadership 360, which include both the graduate as school leaders and their staff and supervisor to assess the school leader’s leadership practices and resultant school conditions.
Study Sample and Data Collection Procedure
This study uses multiple years of data collected from ELPPs and their graduates through INSPIRE-G and INSPIRE-PP. All participating ELPPs were university-based, and most of them were members of University Council for Educational Administration (UCEA), a large professional organization for university-based ELPPs. The INSPIRE-PP survey was administered to programs every 2 years, asking program coordinators or faculty to describe program features and processes that comprise a program in the current or most recent year. This study focuses specifically on those programs with building-level certification programs. In total, 100 different preparation programs responded, most of which responded both in the 2015-2016 and 2017-2018 surveys and only a few responded either 1 year or the other.
INSPIRE-G data were collected every year. For this study, we used 4 years of data collected from 2015–2016 to 2018–2019. In addition to information from recent graduates about their ELPP experiences, INSPIRE-G also asks about their demographics, professional background, and professional/career positions and aspirations. While some ELPPs have participated only one round of the INSPIRE-G data collection, others participated in multiple years. The total number of graduate respondents varied by program, ranging from 4 to 136 in a single administration. In total, 2,994 individual graduates across 60 programs responded to the survey during the study period. Using unique institution identifiers, we merged de-identified individual graduate data (INSPIRE-G) with preparation program feature data (INSPIRE-PP). Since INSPIRE-PP were only available for every 2 years, we matched multiple years of graduate level INSPIRE-G data with one cycle of INSPIRE-PP data. Our final merged dataset included 2,168 graduate students from 42 ELPPs.
Measures
Program quality
Program quality serves as the outcome variable in our study and is measured at the individual graduate student level. In the INSPIRE-G survey, all graduates are asked to rate the Overall Program Quality of their program by a single item “please rate the overall quality of this program” on a Likert scale from 1-very poor to 5-very good.
Program features
Data on program features are collected at the program level through the INSPIRE-PP survey. Some program features are measured by a single question, including whether the program has a cohort structure, whether aligned to national standards (e.g., NELP, and/or PSEL), whether reviewed by national accreditation bodies (e.g., Council for the Accreditation of Educator Preparation [CAEP]), whether students take a national licensure assessment for leadership licensure and certification (e.g., School Leaders Licensure Assessment [SLLA]), whether the institution is public or private, and percentages of different types of personnel teaching in the programs (tenured/tenure track faculty, full-time clinical faculty/instructors, and other).
Some program features are measured with multiple Likert-scale items (1-not at all to 5-to a great extent), including student admission criteria, curricular emphasis, instructional strategies, technology use, student assessment, internship experiences, practitioner roles, and district-university partnership(s). Sample survey items measuring these latent program features are shown in Table 1.
Latent Program Features—Descriptive Statistics, Reliability, Factor Loadings, and Selected Items (N = 200).
Student admission criteria
Program coordinators are asked what candidate qualities their programs use to evaluate candidates for admissions. Items include previous leadership experience, communication skills, evidence of teaching effectiveness, academic transcripts, prior degrees, and standardized test score(s).
Curricular emphasis
Whether a program has a coherent standards-based curriculum is measured by asking “How much emphasis is given to the content areas below in this program’s curriculum?” The areas include organizational culture, instructional leadership, school improvement, management, and family and community relations, constructed to align with the national PSEL and NELP standards (NPBEA, 2015,2018).
Instructional strategies
Instructional strategies refer to what extent the instructional practices of the programs emphasize applied problem-solving and interactive learning. Items include field-based projects, problem-based learning, action research or inquiry projects, case studies, collaborative activities or assignments, lecture, small group activities, and in-class/online discussion.
Technology use
This program feature measures the extent technologies are used in the programs, including electronic learning management system, computer labs with internet access, video cameras and recording equipment, online research databases, social networking technology, and assessment management systems.
Student assessment
These include a broad range of questions to measure how candidates’ knowledge, skills, and performances are evaluated in formative and summative assessments. For example, one set of questions ask “to what extent are the following types of candidate formative assessments used to evaluate candidate learning?” The items include filed-based assignment, research paper or essay, and clinical and field-based projects or performance evaluation. Another set of questions asks “Which of the following summative assessment strategies does your program use to evaluate students to be recommended for program completion?” The items include a portfolio of student professional preparation work, projects and accomplishments, completion of a capstone or culminating project, a final comprehensive exam, master thesis or research paper, state and/or national assessment, and evaluation feedback from internship/field supervisors.
Internship experiences
Program coordinators are asked to describe the internship experience, including purpose, design, components, and mentoring support. It asks the extent a program’s internship has authentic clinical/field-based work, course-embedded field work, in-depth clinical/field-based work, sustained clinical/field-based work, and clinical/field work regularly evaluated by program faculty. It also asks whether the experiences provide students opportunities to synthesize and apply program content knowledge, develop perspectives on school improvement, engage colleagues in shared problem solving and collaboration, and work in schools serving students with a variety of student populations. In addition, there are multiple items around how mentors are selected, whether training is provided to mentors, and whether internship mentors have demonstrated successful experience as educational leader.
Practitioner roles
A set of questions asks the extent practitioners serve in the program in curriculum development, instruction, guest speaking, supervising field work, serving on an advisory board, and assessing students for program admissions and/or for program completion/graduation.
Local and district partnership(s)
These are questions around whether the program has affiliations with local district partners and to what extent these partners are included in a formal advisory committee, shared decision making on student selection, curriculum and/or program design, and program teaching. It also asks whether district partners provide support on internship arrangements and/or supervision, or give priority to hiring program graduates into leadership positions.
Graduate characteristics
We included individual graduates’ demographic and professional backgrounds (e.g., gender, race and ethnicity, age, whether currently in a leadership position, and total years of different professional experience). In addition, for programs with partial or de facto cohort models, we supplement the data with graduates’ own experiences in the program (whether they were part of a full cohort or not).
Analytic Strategies
Confirmatory factor analysis (CFA) was used to generate latent program features in different domains using 2 years of program data collected through INSPIRE-PP. We then used two different analytical strategies, bivariate analyses (i.e., t-tests and chi-squared tests) and multilevel mixed-effects models, to examine how various program features were associated with graduates’ assessment of program quality.
Comparing program features between high and low quality programs
To detect any significant differences in specific features among programs, we first aggregated graduates’ ratings on Overall Program Quality (OPQ) to the program level and identified two groups of programs whose mean OPQs were either very high or very low. The rationale for focusing only on programs whose qualities were rated by graduates in the more extreme ends of the rating continuum is that the respondents of INSPIRE-G and INSPIRE-PP surveys were overwhelming from research or doctoral-granting institutions and the variances on program OPQ were more limited than may have been the case with a less homogeneous sample. As shown in Table 1, the mean of program OPQ is 4.4 on a 5-point scale indicating that graduates generally rated their ELPPs’ quality high. It has a standard deviation (SD) of 0.75, suggesting strong agreement among graduate respondents. To tackle the issue of small variation, we created the sub-sample of two groups of programs with “high” versus “low” program quality, whose average OPQ were more than 0.5 SDs from the mean of the original sample. A series of t-tests and chi-squared tests were then run, depending on whether a specific program feature variable was continuous or categorical, to detect any significant difference on each of the program features between the high and low quality ELPPs.
Examining the associations between program features and program quality
Since the bivariate analyses (t-tests and chi-squared tests) were performed at the program level and only looked at the relationships between OPQ and one program feature at a time, we turned to multilevel mixed effects models that examine relationships among multiple program features and program quality, taking into account the nesting nature of the multilevel data of graduate students within programs. To do this, we first merged all program features, including both single-item features and latent features generated from the CFA analyses (at the program level), with the INSPIRE-G data (at the individual graduate level), which included individual graduates’ assessment of OPQ and their background characteristics. Multi-level mixed effects models were then performed, where the standardized OPQ at the individual graduate level was the dependent variable, and various program features at the program level were predictors. In addition, individual graduate characteristics were included as covariates. A series of year dummies were also included in the models to account for aggregate factors that change each year and are common to all programs for a given year.
Results
Descriptive Statistics on Program Features
Table 2 shows descriptive statistics of graduate characteristics. Among all graduates who responded to the INSPIRE-G, a majority were female (67%). Three quarters of the graduates were white (76%), 10% were Black, and 7% were Hispanic. The average age of respondents was 38 years old. On average, graduates had almost 13 years of professional education experience and over 5 years working in their current school. On average, they had approximately 10 years as classroom teachers, 3 years as teacher leaders (e.g., department chair, instructional coach, etc.), just under one and a half years as K-12 administrators (e.g., principal, assistant principal, and central office administrator), and just over one and a half years (1.6) in other K-12 professional educator positions (e.g., school counselor, psychologist, librarian, or in other types of educational agency). Approximately 20% of graduate respondents were working as a school principal or assistant principal, and another 42% were a classroom teacher and/or in some school-level leadership positions (e.g., instructional coach, curriculum specialist, or department chair).
Descriptive Statistics on Graduate Demographic and Professional Characteristics, 2015–2016 to 2018–2019 (N = 2,994).
Table 1 shows the CFA results on latent ELPP program feature variables, including sample survey items measuring different program features, scale reliability (Cronbach’s α), and the factor loading ranges. Several program features, such as standard-based curricular emphasis, problem-based instructional strategies, technology use, local partnership, and practitioner roles in the program are represented by one latent variable. Other program features are represented by multiple latent variables. For example, CFA on admissions criteria for candidates yielded two factors: leadership potential and academic competency. Internship experiences also generated two sub-scales: internship design and mentoring. Finally, assessment practices included evaluating both field-based work and research-based work of students. All factor scores on program features were then standardized to have a mean of 0 and SD of 1 for subsequent analyses. All CFAs showed great model fit, with satisfactory chi-square statistic, comparative fit index (CFI), and root mean square error of approximation (RMSEA).
Table 1 also shows the mean and SD of all the items included in each of the latent program features. On average, ELPPs emphasize standard-based curricula, program solving and interactive instructional pedagogies, internship programs that provide students with authentic and in-depth clinical and field-based experiences, and assessment on field-based performance. All these features have high means (all of which were over 4 on a 5-point Likert scale) and small variations (mostly with SD < 0.7). Lower means and higher variations were found in other program aspects. For example, on average, district partnerships has a mean of 2.88 and SD of 0.99, indicating the practices of building partnerships with districts and involving them in decision making and support roles were not universal and varied greatly among different ELPPs. Additionally, practitioner involvement was moderately emphasized by ELPPs on average and the variation of the level of involvement among different programs was larger than all other features (mean = 2.53, SD = 1.07).
Table 3 presents the descriptive statistics of program features measured by single items. While the vast majority of the sampled programs had a cohort model, 37.7% of which were partial or de facto models. Among all the courses, 46% were taught entirely online or partially online. About half (50.6%) courses were taught by tenure-track or tenured faculty, while 29.6% were taught by full-time clinical faculty, and around 25.4% were taught or co-taught by part-time adjunct instructors and/or other practitioners. Among all the sampled ELPPs, 82% were located in public institutions. Around 70% were reviewed or accredited by national accreditation bodies, 72% were aligned to national standards, and 33.3% assessed students on national SLLA tests for leadership licensure and certification.
Program Features Measured by Single Item (N = 200).
The percentage of courses taught by different types of personnel may not add to be 100% because of measurement error caused by the fact that the respondents were asked to identify the percentages by dragging the sliders on the scale bar.
Differences Between High and Low Quality Programs
We performed a series of bivariate analysis (t-tests for continuous variables and chi-square tests for categorical variables) to detect differences of program features between two groups of programs whose mean OPQ were rated either very high or low by graduates. The target sample resulted in 23 programs, 14 of which had high (0.5 SDs or more) mean standardized OPQ and nine had low (−0.5 SDs or less) mean standardized OPQ, based on the ratings of graduates.
Table 4 shows the means of the all program features and their differences between highly- and lowly- rated programs. The latent features were measured by the standardized factor scores generated from the CFAs. The features measured by single items have values of either 0 or 1, or ranging from 0 to 1. Descriptively, compared to low quality programs, highly rated programs seemed to focus more on leadership potential in admissions criteria, standards-based curricula, technology use, field-based assessment, and were more likely to offer internship programs that provide students with authentic and in-depth clinical and field-based experiences. Highly rated programs also tended to have more courses taught by tenure-track faculty, stronger district partnerships, and greater roles for practitioners in the program. Finally, they were more likely to be in private institutions, cohort-based and nationally accredited, more aligned to national standards, and more likely to rely on national tests for leadership licensure and certification. By contrast, lower quality programs emphasized academic performance of candidates in program admissions, offered more online courses, and focused on applied and interactive instructional pedagogies, and research-based assessment.
Differences of Program Features Between Programs with High and Low Quality Ratings, t-Test and Chi-Square Test Results.
p < .05. **p < .01. in two-tailed t-tests or chi-square tests.
However, with the small sample sizes of the programs, even with substantial differences between “high” and “low” overall rated programs, only a few features were statistically significant, including admission criteria focused on leadership potential, building and maintaining close local partnerships, and whether or not nationally accredited.
Multi-level Mixed Effect Model Results
The descriptive statistics and the significance tests provide information on the general direction and magnitude of the associations between individual program features and perceived program quality by graduates, and set the stage for the multi-level mixed effects models, which further investigate the associations of multiple program features and overall program quality. We first estimated a null model with no independent variables and obtained the intraclass correlation coefficients (ICC) of 10.2%. This indicates that around 10% of the variation in the overall program quality is at the program-level and the multi-level analysis is warranted.
The results of all the multi-level mixed effect models are presented in Table 5. The single mixed effect models show the results where each of the program feature variables was entered separately, while in the full model, all program feature variables were entered simultaneously. In all the single and full models, the graduate characteristics and the year dummies were included as control variables.
Program Features and Overall Program Quality, Results from Mixed Effects Models.
Note. Single models are estimated with only one program feature variable. The full model includes all program feature variables. Graduate characteristics and year dummies are included as control variables in all the single and full models. Sample sizes of all models are 2,168 graduates within 42 ELPPs.
p < .05. **p < .01. ***p < .10.
In the single models, when ELPPs emphasized students’ leadership potentials in admissions criteria, graduates tended to rate their programs’ quality as high (coefficient = 0.08, p < .05). Emphasis on academic achievement in admissions, on the other hand, had a negative association with perceived overall program quality, although the association is only marginally statistically significant (coefficient = −0.07, p < .10). Emphasizing standard-based curriculum contents were unrelated to quality, while emphasizing applied problem-solving instructional strategies was negatively related to program quality (coefficient = −0.10, p < .05). The internship practices, either in terms of design or mentoring, seemed to be unrelated to program quality. In terms of assessment practices, graduates tended to perceive higher overall program quality if assessments were field-based work (coefficient = 0.17, p < .01) but lower quality if assessments were research-based (coefficient = −0.11, p < .05). In addition, strong local district partnerships in the program were positively related to program quality (coefficient = 0.13, p < .01), whereas whether practitioners played a significant role in the program seemed to be unrelated to graduates’ assessments of overall quality.
Students who were in a program with a full cohort tended to rate their program as high quality (coefficient = 0.16, p < .01). In terms of course teaching, students tended to rate their program quality lower if a large proportion of their courses were taught online (coefficient = −0.22, p < .01) or by full-time clinical faculty (coefficient = −0.46, p < .01). On the other hand, the percentages of courses taught by tenure-track faculty, adjunct, or other practitioners had no significant association with program quality.
A program that was nationally accredited (coefficient = 0.29, p < .01) and relied on national tests for leadership licensure and certification (coefficient = 0.22, p < .05) was more likely to be rated as high quality, while the alignment to national standards (vs. state standards or other standards) seemed to be unrelated to program quality. Institution type (private vs public) had no association with program quality. Although not reported in Table 5, most personal characteristics, including age, race, and leadership experience, were statistically insignificant, while female graduates tended to rate their programs higher than their male peers.
In the full model, all program feature variables were entered simultaneously. Several program features remained statistically significant after other program features and graduate background information were controlled. For example, graduates tended to rate their program quality higher if the program emphasized field-based work assessment (coefficient = 0.22, p < .01), had a full cohort model (coefficient = 0.16, p < .05), had strong district partnerships with school districts (coefficient = 0.21, p < .05), were nationally accredited (coefficient = 0.29, p < .05), and relied on national tests for licensure and certification (coefficient = 0.27, p < .05). On the other hand, applied instructional strategies (coefficient = −0.13, p < .01), online courses (coefficient = −0.18, p < .05) and research-based assessment (coefficient = −0.17, p < .01) remained negatively related to program quality. Several program features, previously statistically significant in the single models, became insignificant after controlling for other features. This includes whether a program emphasizes academic achievement or leadership potential of candidates in program admissions and the percentages of courses taught by different personnel. The ICC became 5% in the full model, indicating half of the variance at the program level were explained by all these program features.
Discussion
Given the critiques of university-based ELPPs and limited research on their effectiveness, it is important to understand what specific program features are associated with preparation quality. Using multiple years of survey data from both program coordinators and program graduates, our study describes ELPP program features and analyzes the relationships between program features and perceived program quality.
ELPPs in our sample shared many common program features and emphasize them to a good extent (greater than 4 on a 5-point Likert scale), including offering standard-based curriculum, utilizing program solving and interactive instructional pedagogies, providing authentic clinical and field-based internship opportunities, and emphasizing field-based performance in student learning assessment. In addition, the vast majority of the ELPPs use some forms of cohort structure. More than two-thirds of the programs are nationally accredited and aligned to national standards. All these features are indicative of the alignment of professional standards and quality preparation in the previous literature (Byrne-Jimenez et al., 2017; Crow & Whiteman, 2016; Darling-Hammond et al., 2010). Strong representation of these features may reflect the fact that most ELPP programs in our sample were members of UCEA, which promoted effective ELPP practices through communication and dissemination of research and scholarship among its member institutions. On the other hand, several program features are less common in these ELPPs, including building local district partnerships, involving practitioners in the program, and relying on research projects for student performance assessment (with mean values less than 3 on a 5-point Likert scale and SDs close to 1). Although district partnerships are promoted in both preparation program standards and best practices, their development and maintenance require dedicated time and staff. It may be that these factors are often outside of the control of the ELPP, making it challenging to build and sustain local partnerships or to engage practitioners in the program (The Wallace Foundation, 2016).
When comparing programs with “high” versus “low” overall quality ratings, we find many differences in program features between the two groups, although only a few differences were statistically significant due to small sample sizes. The multilevel mixed effects models, relying on both graduate and program data, taking into account the nesting nature of the multilevel data, and controlling for graduates’ demographic information, reveal several significant relationships that are similar to the results from the comparisons between high and low quality programs. Consistent with previous research, programs with a full cohort model, assessments focusing on field-based projects, and strong local partnerships were perceived by graduates as high quality (Donmoyer et al., 2012; Greenlee & Karanxha, 2010; Kochan & Locke, 2009; Salazar et al., 2013; The Wallace Foundation, 2016). On the other hand, large proportions of online courses seem to be an indicator of poor program quality perceived by graduates (Choi et al., 2005; Ritter et al., 2010).
Surprisingly, despite previous evidence on the importance of coherent curriculum and authentic field-based internship experience with high-quality mentors (Darling-Hammond et al., 2010; McCarthy, 2015), our analysis shows no significant relationships between these features and graduates’ assessments of program quality. This could be caused by several research conditions. First, our measure of curriculum has a narrow focus of on standards-based curriculum content areas, which fails to consider other dimensions of curriculum, including the development of values, beliefs, and identities (Scribner & Crow, 2012). In addition, field-based internship design and mentoring may be not be experienced by students the way intended or conceptualized by program faulty or coordinators, as students’ quality of field experiences may be influenced substantially by school-specific factors unrelated to the internship design and mentoring.
Contrary to the previous case studies suggesting that applied instructional strategies are effective program features (Byrne-Jimenez et al., 2017; Darling-Hammond et al., 2010), our analysis shows that emphasizing field-based projects, problem-based learning, and action research projects were negatively associated with graduates’ assessments of program quality. It might suggest that students equally value theory-based instructional approaches linking theory and practice, or that applied instructional strategies and action research often require a significant amount of additional time and effort beyond the traditional program time commitments (e.g., course time or additional layers of access and reciprocity with the schools and districts involved). Moreover, this may indicate that our measures of applied instruction need to be refined since the internal consistency is relatively low (α = .68).
In terms of program admissions criteria, although emphasizing candidates’ leadership potential and academic potential each showed some relationships with graduates’ assessment of program quality, the relationships disappeared in the full model where other program features were controlled. This might suggest that the relationship between admissions criteria and program quality is mediated through program features. It also suggests by the time graduates complete their program of study and respond to the survey, they may perceive program experiences as more important indicators of quality than initial candidate selection processes.
While ELPPs can receive accreditation from different organizations at national, state, or regional level, our analysis shows that programs that are nationally accredited and relying on national assessments (e.g., SLLA) tended to receive higher quality ratings by graduates. This is an interesting finding. Perhaps, there is a perception among candidates that successful completion of a national assessment while acquiring a license and/or a degree is additional legitimacy and/or portability of the degree or license.
Conclusions
Over the past three decades, the focus on ELPP evaluation has expanded both topically and methodologically. Small but consistent evidence indicates what program features are necessary to develop leadership-ready graduates. This study explores relationships between multiple program features and graduate assessed overall program quality. Our main finding is that cohort model, face-to-face course delivery, field-based assessment, and strong district partnerships matter in quality preparation, providing more concrete support of previous literature (Young & Crow, 2017; Young et al., 2009). On the other hand, the associations of program quality with standards-based curriculum, applied instruction strategies, and authentic field-based internship experience are less consistent as suggested in previous literature. Taken together, the current study deepens our understanding of preparation program delivery and the pivotal role of intentionally improving ELPP quality. In particular, our collective attention to organizational learning, especially double-loop learning, raises the important question of “what will ELPPs do to be more effective, more agile, and more impactful” and creates an opportunity for purposeful changes within the ELPP field that are responsive to (and less reactive to) the internal and external demands for relevance, viability, and legitimacy (Young et al., 2019). Ultimately, our study highlights the value of organizational learning as a framework for program improvement, including attention to the values and beliefs that drive our decision about which program characteristics to focus upon in light of demands for increased enrollment, faculty line retention and productivity, state and national standards, accreditation, and accountability.
In addition, as the COVID-19 pandemic continues to prevail, many university programs, including ELPPs, have been forced to switch to online course-offering intermittently. New evidence suggested that the shift to online education in the pandemic had negative results for college student engagement and learning, peer relations, and access to learning supports (Kofoed et al., 2021). Although it is too early to fully assess its impact on programs, graduates, or school districts hiring these graduates, this shift elevates our finding that large proportions of online courses was an indicator of poor program quality perceived by graduates to a heightened level of concern. Given the fact that universities will continue to rely on online or hybrid learning post-pandemic, often as a strategy to increase enrollment, it is imperative for researchers and ELPP faculty to intentionally engage in organizational learning for program improvement, including having more nuanced conversation of impact of changes in our environment, carefully assessing various online or hybrid models, and seizing the opportunities to learn from programs with successful models.
Several implications for future research can be derived from this study. First, given the status of university-based preparation programs as the primary mechanism for creating pathways to the principalship and similar positions, additional data collection from a broader range of ELPP institutions, particularly those that serve more diverse populations is needed. A majority of the ELPPs in our sample were from a subset of research or doctoral-granting institutions, which potentially led to limited variations in the program features and quality ratings. It would be interesting to see whether the relationships between features and program quality in this study exist among different types of ELPPs such as programs in Masters Comprehensive institutions or the vast majority of leadership preparation institutions nationally. Also, collecting similar data across a variety of programs and institutions within entire states would also provide institutional program variation while controlling for state-level extraneous factors.
Second, this study has indicated that further instrument revisions are needed. Our measures on several program features have lower internal consistency (e.g., applied instructional strategies and research-based assessment) and several other scales are defined too narrowly. For example, technology use is defined as ELPPs providing technology-based learning environment, and instructors using digital technologies to deliver content and improve the learning experience. However, ELPPs need to provide candidates with the necessary learning and organizational support and prepare them for their future roles as technology leaders to employ technology as a lever of teaching, learning, and leading. This aspect of technology use is equally important but not measured. In addition, ongoing changes in the field provide an opportunity for instrument revision to be attentive to next generation of ELPPs (e.g., curriculum focus on competency, culturally responsive leadership, and crisis leadership).
Third, it is important to note potential limitations of this study because all data are self-reported by respondents. Future research should expand to include additional objective measures of graduate outcomes, such as employment placement, and more distal outcomes, such as their leadership practices in schools and school outcomes.
In conclusion, by identifying program components that are associated with graduates’ assessments of quality, this study not only contributes to the knowledge base of ELPP quality and effectiveness, but also informs policies in terms of preparing effective, innovative change agents for schools. From this study, considerably more research can be developed in which researchers can expand and meet the challenges and limitations of conducting leader preparation evaluation research (American Institutes for Research, 2016; Ni et al., 2017). This also offers an opportunity for the field to take stock and reflect from what we know seems to matter, from which ELPPs as learning organizations can grow and determine what deep changes need to be adjusted and what can simply be tweaked in promoting responsive and relevant educational leadership preparation.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The author(s) received financial support from the Research Incentive Seed Grant Program at the University of Utah for the research of this article.
