Abstract
We conducted a study of predictive analytics (PA) applied to state data on post-school outcomes (PSO) of exited high-school students with disabilities (SWD). Data analyses with machine learning Random Forest algorithm and multilevel Bayesian ordered logistic regression produced two key findings. One, Random Forest models were accurate in predicting PSO. Two, Bayesian models found high-school graduation was the strongest predictor of higher education and reliably predicted the specific type of outcome relative to other outcomes. Limitations of this study are the data source and small number of predictors. Implications of the study for researchers and educators are discussed in conclusion.
Keywords
Introduction
In the United States, special-education research has historically had an in-school focus, in areas such as science and reading instruction. In the most recent decade, another area has garnered increased attention, the post-school outcomes (PSO) of students with disabilities (SWD) (see Lipscomb et al., 2017). At the same time, newer technology has produced new tools, such as predictive modeling with machine learning (see Nassif et al., 2016), to analyze large administrative data, such as PSO, to support decision making (see Larkan-Skinner & Shedd, 2021). Utilizing these data and analytic tools is especially important for educators in closely examining the PSO of SWD (Mandinach, 2012), who continue to experience significant barriers to and challenges in successfully transitioning from high school to young-adult life (Liu et al., 2018). Thus, the focus of this study was to utilize these new tools to analyze the PSO of exited high-school SWD, to assess the modeling performance of these tools, and examine the implications of their utilization for state and local educators and administrators.
Literature Review
Post-School Outcomes
Socioeconomic disparities between people with disabilities and their peers without disabilities have been well documented. In April 2022, the U.S. Department of Labor Office of Disability Employment Policy reported that the unemployment rates of people with and without disabilities aged 16 and older were 8.3% and 3.1%, respectively. Putting those rates into a pre-pandemic historical context, in November 2019 the unemployment rates of people with and without disabilities aged 16 and older were 6.9% and 3.2%, respectively (U.S. Department of Labor, 2019). Even after accounting for the level of educational attainment, people with disabilities had higher unemployment rates (U.S. Department of Labor—Office of Disability Employment Policy, 2022). Moreover, among all people with lower rates of educational attainment and employment regardless of age group, the poverty rate remains higher for people with disabilities (Yin et al., 2014). These employment disparities can also be seen among the younger groups of workers when the data are disaggregated. The unemployment rates of people with and without disabilities aged 16–19 years were 21.1% and 11.4%, respectively. The unemployment rates of people with and without disabilities aged 20–24 were 16.5% and 8.7%, respectively.
The Individuals with Disabilities Education Act, 2004; P.L. 108–446) was enacted to address the educational needs of SWD. Since 2008, one of the requirements of this law is that states have to collect annual data of exited high school SWD, Indicator B14—otherwise known as PSO. This measure refers to the number of SWD in the state who exited high school, had an Individualized Education Program (IEP) when they exited, and within a year of exit were (a) enrolled in higher education, (b) employed in a competitive integrated setting, (c) enrolled in some other post-secondary education or training program, or (d) employed in some other (not competitive) employment (see Individuals with Disabilities Education Act, 2004, 20 U.S.C. 1416[a][3][B]). The exiters who meet criteria for one of these outcomes are counted as being “engaged” in PSO; otherwise they are counted as “not engaged.”
Over the last decade, research in PSO has notably advanced since the seminal work of Test et al. (2009), which identified 16 specific predictors of PSO. Eaves et al. (2012) advanced that work, creating and validating a survey to empirically connect high-school transition services to PSO. Haber et al. (2016) further advanced that work, conducting a meta-analysis on the 16 predictors; and their findings indicated some variance in employment, education, and independent living outcomes based on different in-school predictors and demographic characteristics. Burnes et al. (2018) found that high-school SWD being involved in their IEP, interacting with others (i.e., social skills development), and receiving support from their community were also associated with positive PSO.
Much of the peer-reviewed studies in the PSO research literature have utilized data from the National Longitudinal Transition Study (i.e., NLTS-1, NLTS-2, NLTS-2012) (see Blackorby & Wagner, 1996; Newman et al., 2011), which have been conducted since the 1990s. Over those many years, studies utilizing the NLTS data have consistently reported gender and graduation from high school as significant predictors of post-high school competitive employment and education (e.g., Aud et al., 2010). Carter et al. (2013) also reported that among those with severe disabilities, having paid employment experience during high school, parent expectations, and social skills were significant predictors of post-high school employment.
There are fewer peer-reviewed PSO studies in the research literature that do not utilize NLTS or other nationally-representative samples. A small number of these studies include state-level PSO. Across several states and school districts, McConnell et al. (2015) reported that “student GPA and the percentage of time instruction received in the general education setting cannot indicate a student has all skills needed after high school for employment and further education…more is needed for college and career success” (p. 334). In a Wisconsin pilot study, interventions involving school-led coaching led to increased work experiences in high school, which predicted post-high school paid employment for young adults with intellectual and developmental disabilities (Molfenter et al., 2017). In a study of educators in Indiana, Sprunger et al. (2018) recommended that in order to improve PSO, special-education teachers have to incorporate both academic and work components into the high-school transition plans of SWD, and to receive regular professional development that is focused on transition programming.
The latest and potentially most important development in the PSO literature has been research regarding high-school STEM for SWD and their post-high school outcomes in STEM education and career (see Hwang & Taylor, 2016). This research is also occurring amidst the passage of the STEM Education Act (2015), which places federal priority on STEM education and recent government projections of STEM occupations growing at a faster rate than non-STEM occupations and paying higher wages/salaries over the coming decades (see Bureau of Labor Statistics, 2020). Also, some research indicates that career-technical education and STEM enrichment efforts (e.g., summer, after school programs) in high school may lead to positive PSO for SWD in STEM (Falkenheim et al., 2017; Plasman & Gottfried, 2018). Despite these efforts, however, disparities in STEM access and outcomes between students with and without disabilities remain a persistent problem (National Science Foundation, 2021). More importantly, what is missing from the PSO literature is any federal or state-level study of different types of outcomes (e.g., college, career training, full-time job) analyzed concurrently, rather than as several single-level or single-outcome analyses.
Utilizing Data and New Analytics
Beyond meeting a federal requirement, states can utilize their own annual PSO data to drive and/or support policies, programs, and practices across their local schools and districts (see Gingerich & Crane, 2021). As part of best-practices in data utilization in education decision making, some researchers (e.g., Schildkamp, 2019) have recommended that educators access and utilize more than just assessment data, such as state administrative (e.g., PSO) data. Other researchers have asserted that “decision makers must interpret the data to make informed decisions about how to effectively support students” (Wilcox et al., 2021, p. 2) and called for improving data literacy among educators.
As an increasingly popular approach to data analysis, predictive analytics (PA) is generally described as “statistical or machine learning methods to make predictions about future or unknown outcomes” (Abbasi et al., 2015, p. 35). Machine learning is generally referred to or described as a set of algorithms that use statistics and probabilities in a process of “learning” from the data in order to find patterns and relationships among variables (Meserole, 2018). In recent years, PA have become popularized in mainstream culture, through television shows such as “Numbers” and movies such as “Moneyball.” (see Banerjee, 2018). In research, PA have become more widely used as advances in computing power and data storage have enabled most home computers to run complex algorithms to analyze data (Zhang, 2020). Much of the research application of PA occurs in public health and medical research, for example, to know who may be at greater risk for certain health or medical conditions (see Yperman et al., 2020) and/or to drive targeted interventions and effective treatments (e.g., Desai et al., 2020; Panesar et al., 2019). In these studies, PA application also often includes the use of historical data (see Brown et al., 2015).
Compared to other fields, the application of PA is comparatively newer to education. These studies have focused on assessing the risk of students falling behind academically or dropping out of college (e.g., Herodotou et al., 2019; Wagner & Longanacker, 2016) and K-12 schools (e.g., Porter & Balu, 2016). These studies, however, have not focused on SWD. Studies applying PA in special education and disability/rehabilitation research have centered on the assessment or diagnosis of disability conditions such as autism (e.g., Maenner et al., 2016), multiple sclerosis (e.g., Flauzino et al., 2019), and dyslexia (e.g., Rello et al., 2020). What is missing from the literature are studies regarding PSO of SWD with data national or state level data utilizing PA, specifically machine learning and Bayesian multilevel modeling. These are powerful new analytic tools that are widely used in other fields to support decision makers.
The Current Study
This study seeks to fill two gaps in the PSO research literature. One, educators at the state and local levels need to know whether the predictors of PSO identified through previous studies are also observed when analyzing state-level PSO data. While the studies in the literature have utilized nationally-representative data (i.e., NLTS), local educators and state administrators need real-time and/or annual data of their own populations of SWD and post-high school engagement in further education/training and/or employment after exiting high school. This study utilized state-level, annual PSO data from recent years, taking into account specific school and district contexts in the data analyses. Two, while studies in the PSO research literature have identified variables with statistically significant associations to outcomes, they are missing important aspects of PA—utilizing machine learning and assessing predictive accuracy. In order for local educators and state administrators to reliably use PA, they need to know how accurately these models predict which SWD would and would not become engaged in PSO after exiting from high school. Filling these two gaps in research knowledge is critical for states’ PSO data utilization and effective PA model development to support policies, programs, and practices for SWD. Thus, based on these two knowledge gaps in the literature, this study was framed around two a priori research questions: (1) How did the models of PSO engagement differ in terms of predictive performance across time? (2) Were there differences in the significant predictors of the PSO categories across time?
Method
Data Collection
We received 5 years of de-identified PSO data from a northwest state. The data years represented the exit year of SWD from high school; the outcome data are collected a year after exit. Because the state agreed to provide only de-identified data, other information that the authors would have liked to utilize in their data analyses could not be included.
The frequency counts of the exited high-school SWD and the number of schools and districts in the analyses were tabulated for each year. In 2014, there were N = 2196 exiters from 311 schools and 148 districts. In 2015, there were N = 2937 exiters from 336 schools and 158 districts. In 2016, there were N = 2950 exiters from 318 schools and 150 districts. In 2017, there were N = 3257 exiters from 312 schools and 141 districts. In 2018, there were N = 3102 exiters from 309 schools and 140 districts. The frequency counts differed each year due to annual differences in the number of exiters and the annual PSO response rates. The response rates of the exited high-school SWD for these years ranged from 76% to 79%, which is higher than what is typically observed for this population (i.e., young adults with disabilities) using this method of data collection (i.e., statewide census) (see Dillman et al., 2014).
The state PSO data included one binary outcome variable, “post-school engagement” (i.e., engaged, not engaged in PSO). This outcome is defined by law (Individuals with Disabilities Education Act, 2004) as SWD who, during their first year after high-school exit, had been: (a) enrolled in higher education, (b) competitively employed, (c) enrolled in some other post-secondary education or training program, or (d) employed in some other noncompetitive employment (20 U.S.C. 1416[a][3][B]). Exiters are counted as “engaged” in PSO if they meet any of the four criteria above; otherwise they are counted as “not engaged.” For those who meet criteria for more than one outcome, they are still only counted once in the highest ordered category, higher education, followed by competitive employment, some other postsecondary education or training, and some other employment (Individuals with Disabilities Education Act, 2004). In this study, the state coded 1 if exiters had engaged in PSO (i.e., reference), and 0 (comparator) if they had not engaged in PSO.
The PSO data included five predictors: gender, ethnicity, disability category, least restrictive environment (LRE), and high-school exit status. Research has shown these variables predict PSO (e.g., Eaves et al., 2012; Haber et al., 2016). In this study, the gender predictor was coded with male as the reference (coded as 1) and female as the comparator (coded as 0). Ethnicity was coded with Caucasian/white as the reference and Students of Color as the comparator. Disability was coded with specific learning disability (SLD) as the reference and the other disability categories as the comparator. LRE was coded with the amount of time spent in a general education classroom settings 80% or more per instructional day as the reference and 79% or less per instructional day as the comparator. High-school exit status was coded with graduating from high school as the reference and not graduating as the comparator.
The state data did not contain missing or impossible values. For a couple of predictors, disability category (10 categories) and ethnicity (7 categories), there were little to no data for most of their categories in most schools and districts. These very low counts were concerning, and in order to ensure statistical modeling would produce reliable estimates and an admissible solution (see Kline, 2005), these two predictors were recoded into binary variables. This meant that we coded students with a SLD as the reference and the other disability groups as the comparator. Students with SLD were not only the largest group of SWD in this state but they are also the largest group of high-school SWD in the U.S. served each year in special education (National Center for Education Statistics, 2020). For the ethnicity predictor, Caucasian/white students were coded as the reference and Students of Color as the comparator to reduce problems in model/parameter identifiability (see Gelman & Hill, 2007). Moreover, given the complexity of PA models, our end goal was to be as clear as possible in the interpretations of results for educational stakeholders in describing how these models could be subsequently utilized and/or further refined (see DeCoster et al., 2009).
Data Analysis
All data analyses were conducted using the Stata 17 statistics software (StataCorp, 2021). The decision rule for judging statistical significance of model testing outcomes was set a priori at p < 0.05. Data analyses were conducted in three connected sequential steps: (a) analyses of PSO engagement with machine learning Random Forest algorithm, (b) analyses of Random Forest model predictive performance with the Receiver Operating Characteristic (ROC) curve, and (c) analyses of the outcome types (described in “Data Collection” above) with Bayesian ordered multilevel logistic regression.
The three PA steps taken in the data analysis were necessary to address the knowledge gaps we identified in the PSO literature and to answer the two research questions in this study. The steps were also intentionally sequenced. First, we analyzed whether exited high-school SWD were engaged in PSO, which would represent the states’ main legal reporting requirement for PSO each year. Second, we assessed model predictive performance, which could empirically establish the potential usefulness and usability of PA engagement models for state and local educators. Third, in addition to knowing whether an exited SWD was engaged in PSO, we also analyzed the relationship of the same PSO engagement predictors and each outcome type, which could provide an understanding of why any specific predictor may have been strongly associated with one outcome type but not with another. The three PA steps are each described further.
Machine Learning
The type of machine learning algorithm utilized in this study, “Random Forest” (see Breiman, 2001), is based on an ensemble decision-tree approach and is particularly useful for predicting binary outcomes (Cutler et al., 2011). The algorithm randomly separates data into training and validation samples, refines its predictive modeling on the former and finished modeling on the latter through thousands of iterations of random samples. Thus, the first main result of interest in Radom Forest models is the “out of bag” (OOB) error (i.e., abbreviating “bootstrap aggregating” as “bagging”), indicating the degree of model predictive error out of the training sample iterations. The second main result of interest is the final error, indicating the degree of predictive error out of the validation sample. The difference between the OOB error rate and the final error rate indicates the degree of “learning” by the algorithm (Breiman, 1996). The general expectation is that the predictive algorithm improves or “learns” through this process, and improves model predictive performance (e.g., accuracy in binary-outcome predictions) from the training to validation modeling stages (Couronné et al., 2018).
The specific reason for utilizing the Random Forest model in this study was to determine whether a machine learning algorithm could perform well in model predictive accuracy and be stable across several recent years of moderately-sized state PSO data and predictors. That knowledge would fill a key gap in the PSO literature, as this study would be the first to use machine learning to analyze the socioeconomic outcomes of young adults with disabilities who had recently exited high school. The Random Forest analysis also produced the first part of the answer to the first research question in this study.
Receiver Operating Characteristic Curve
This next PA step in the data analysis provided the other part of the answer to the first research question. We assessed the predictive performances of the Random Forest models for PSO engagement using the ROC Curve analysis. This is a commonly used measure with two key related indices, Sensitivity and Specificity (Metz, 1978). Sensitivity refers to the proportion of people predicted to experience the desired outcome (e.g., engaged in PSO), out of the total number who did in fact experience that outcome. Specificity refers to the proportion of people predicted not to experience the desired outcome (e.g., not engaged in PSO), out of the total number who did not in fact experience that outcome. These two indices also provide important information about the rates of true/false positives, and true/false negatives in these binary predictive models.
The Sensitivity and Specificity indices constitute integral parts of the “area under the curve” or AUC (i.e., two-dimensional area under the ROC Curve), a measure of a binary model’s overall predictive performance (Hosmer et al., 2013; Mason & Graham, 2002). AUC values range from 0 to 1, with 0 indicating 100% inaccurate model prediction and 1 indicating 100% accurate prediction. Although the importance of AUC varies by field (e.g., medicine, business), an AUC value of “0.7 to 0.8 is considered acceptable, 0.8 to 0.9 is considered excellent, and more than 0.9 is considered outstanding” (Mandrekar, 2010, p. 1316).
Bayesian Multilevel Ordered Logistic Regression
The answer to the second research question was produced in this third and final PA step, Bayesian multilevel ordered logistic regression. All Bayesian models are based on the Bayes Theorem of conditional probabilities (Bolstad, 2007; Gelman et al., 2004). In hypothesis testing, Bayesian statistics models the probability of a hypothesis being “true” given the observed data (Stone, 2013). The fundamental and unique aspect of Bayesian statistics is the use of prior information (i.e., “priors”), such as the mean of intercepts or SD of slope parameters (see Van de Schoot et al., 2021). This information (e.g., changes in disease progression) can be derived from previous experiments or the research literature, and then combined with the probability distribution of the newly collected data to produce a new posterior probability distribution, which updates our knowledge of an outcome (e.g., likely prognosis for healthy recovery). The general recommendation is to avoid selecting completely uninformative or completely informative priors (Gelman & Hill, 2007), an especially important consideration for multilevel models given their complexity. Therefore, we chose hierarchical priors, which are between those two types of priors and appropriate given the study design and type of data.
Bayesian statistics also utilizes “Markov Chain Monte Carlo” (MCMC), a large class of algorithms that estimate the posterior distribution of parameters (see Hastings, 1970; Kruschke, 2011). Monte Carlo is a process by which a large number of random samples are drawn from a particular distribution; and these simulations—numbering in the thousands—enable an efficient way to approximate parameters. Markov Chain is a process by which those random samples are sequentially generated, whereby a sample draw depends on the one preceding it and not any other (i.e., Markov property) (Van Ravenzwaaij et al., 2018). A key feature of MCMC is that “the approximate distributions are improved at each step in the simulation, in the sense of converging to the target distribution” (Gelman & Hill, 2007, p. 409). In this study, we used a type of MCMC algorithm called Random-Walk Metropolis-Hastings.
The “ordered” part of the Bayesian modeling begins with an outcome variable (i.e., PSO) containing several categories in a meaningful order, such as lowest to highest or least important to most important (Hedeker, 2015). While similar to a multinomial (i.e., unordered) logistic regression (see Agresti, 2007), the key distinction is that the ordered regression does not assume equal distances, either conceptually or in measurement between the outcome categories (Bauer & Sterba, 2011). The ordered regression is also more parsimonious than fitting the data to several separate models for each individual outcome category and then synthesizing model results post-hoc. Thus, in this study, based on the lowest to highest order of outcome categories (see Post School Outcomes section above), the outcome variable was coded from 0 to 4: (0) no PSO engagement, (1) employed in some other non-competitive employment, (2) enrolled in some other post-secondary education or training program, (3) employed in a competitive integrated setting, and (4) enrolled in higher education (see Individuals with Disabilities Education Act, 2004, 20 U.S.C. 1416[a][3][B]). The analysis was specified as a three-level model with mixed effects (see Austin & Merlo, 2017; Hox, 2002). The fixed effects were each of the five predictors at level-1 (i.e., student), gender, ethnicity, LRE, disability category, and high-school exit status. The interaction effect was gender-by-disability, which was specified in the model based on the research literature. No predictors were available in the state’s data at level-2 (i.e., school) or level-3 (i.e., district).
`The main outputs for our Bayesian multilevel ordered logistic regression included the odds ratios, estimated posterior mean, 95% Credible Interval (95% C.I.), Monte-Carlo standard error (MCSE), and the mean cut points. The odds ratio is an indication, expressed in terms of likelihood, of the magnitude of each predictor’s unique (not shared) effect on the outcome. The 95% Credible Interval assumes that the interval (i.e., Lower − and Upper +) is fixed and the estimated posterior (i.e., mean) parameter is random. This is not to be confused with the more familiar statistical output “95% Confidence Interval” (although it is also abbreviated as “95% C.I.”), which assumes the interval is random and the estimated posterior parameter is fixed. The interpretation of the Credible Interval is also different. It indicates that there is a 95% probability the “true” parameter lies within that interval given the observed data (e.g., PSO state data from 2014 to 2018). The MCSE is an indicator of the degree of imprecision in the Monte-Carlo sampling and estimation of the posterior mean (Gong & Flegal, 2015).
The mean cut points are the mean values representing the estimated threshold at or below which individuals would be classified at the corresponding categories of the outcome variable. In the specified Bayesian model, there were four ordered cut points. Cut Point 1 differentiated “no PSO engagement” (i.e., the lowest ordered outcome) from the other four outcomes. Cut Point 2 differentiated the “no PSO” and “some other employment” outcomes from the other three. Cut Point 3 differentiated “no PSO,” “some other employment,” and “other education or training” outcomes from the other two. Cut Point 4 differentiated the “no PSO,” “some other employment,” “other education or training,” and “competitive employment” outcomes from “higher education,” the highest ordered outcome. With the posterior distribution of parameter estimates, any of the exiters in the 5 years (2014–18) of PSO data could be located within one of these thresholds indicating their most likely outcome. In other words, a model of engagement with a particular outcome could be derived for all exiters, providing information about how likely these SWD are to become engaged in any particular outcome (or not engaged) 1 year after high school, and which predictors (described above) were the most strongly associated with that outcome.
Results
The analyses of data across the 5 years produced results for (a) demographic frequencies (%) of the exited high-school SWD and their PSO, (b) statistical modeling outcomes to answer Research Question 1, and (c) statistical modeling outcomes to answer Research Question 2. Each of these results are described below in turn.
Demographic Characteristics
Demographic Characteristics (%) of Exited High-School Students With Disabilities.
aAsian/Pacific Islander category was disaggregated starting in 2015–16.
Research Question 1
Assessment of Machine Learning Random Forest Model Predictive Performance.
Random Forest Algorithm: Actual Versus Predicted PSO Engagement Frequencies (n) in 2014.
Research Question 2
Bayesian Multilevel Ordered Logistic Regression Model 2014.
Bayesian Multilevel Ordered Logistic Regression Model 2015.
Bayesian Multilevel Ordered Logistic Regression Model 2016.
Bayesian Multilevel Ordered Logistic Regression Model 2017.
Bayesian Multilevel Ordered Logistic Regression Model 2018.
Random Effects for School and District Intercepts 2014 to 2018.
The second part of the answer to Research Question 2 produced, perhaps, the most notable finding in this study—using Bayesian statistical modeling to predict the type of outcome for the exited high-school SWD. Among these model results, there were two notable patterns. One, the largest distance (i.e., numerical difference) were between Cut Point 3 and Cut Point 4: 1.98 in 2014, 1.94 in 2015, 2.08 in 2016, 2.13 in 2017, and 1.68 in 2018. This indicated how different “higher education” was as one type of PSO compared to the other types. Two, the smallest distance occurred between Cut Point 2 and Cut Point 3: 0.27 in 2014, 0.22 in 2015, 0.37 in 2016, 0.25 in 2017, and 0.31 in 2018. This indicated that Cut Point 2 (no PSO and some other employment) and Cut Point 3 (no PSO, some other employment, and other education/training) were the least dissimilar among outcomes, and especially compared to the highest ordered outcomes, competitive employment and higher education.
In addition to the distance-differences across outcomes, the threshold values themselves differed from year to year (see Tables 4–8). In 2014, the values ranged from 0.50 to 3.29. In reference to Cut Point 1, SWD with values of 0.50 or lower would be predicted as not engaged in PSO 1 year after exiting high school. Likewise, for SWD with values of 3.29 or higher would be predicted as engaged in higher education 1 year after exiting high school. SWD with values in between 0.50 and 3.29 would be predicted as engaged in one of the other three outcomes. In 2015, these threshold values ranged from 0.67 to 3.40. In 2016, these values ranged from 0.50 to 3.41. In 2017, these values ranged from 0.17 to 2.97. In 2018, these values ranged from 0.31 to 2.96. These values can be thought of as scores on a latent trait variable. In the output (Tables 4–9), these values are identified as “Mean”, to specify (a) all of the predictors coded as the comparator group, coded as 0 (e.g., less than 80% per instructional day in general education classrooms, female SWD), (b) fixed effects, and (c) average of random effects (i.e., school- and district-level intercepts) in the model.
Finally, in assessing overall model fit and specification by inspecting the Deviance Information Criterion (DIC), we observed that the best model fit was in 2014, with a DIC of 5707, and the worst model fit was in 2018, with a DIC of 8490. This information criterion does represent a tradeoff between model fit and complexity (see Gelman & Hill, 2007). In addition, across 5 years the Monte Carlo Standard Error (MCSE) for the mean/intercept estimates were also small, indicating a sufficient level of reliability in their posterior estimation.
Discussion
The results of the data analyses not only provide answers to the two research questions but also contribute to the PSO literature and offer potential new tools or approaches to state and local educators for their decision making. In this section, we further elaborate on the results of PA with machine learning and Bayesian modeling, and we also describe the study’s limitations. We conclude this study by describing its implications for education research and practice.
Machine Learning
The first research question focused on the use of machine learning with Random Forest modeling of several recent years of a state’s PSO data. Those results contribute new knowledge to the extant research literature that differs from the current studies in the research literature using data from NLTS, which provide a broadly descriptive and national picture of outcomes for SWD after exiting high-school among student-disability groups (see Blackorby & Wagner, 1996; Liu et al., 2018). This current study takes an important next step. In the U.S., education remains decentralized and mostly managed at the state level, with some autonomy given to local (i.e., district, school) levels. As such, the issue for educators regarding PSO is that much of the important practical decisions for programs, practices, and resourcing have to be made with data of their own students, not with national studies. That is, this study presents, for the first time, an empirical data-analytic approach that researchers and educators could take to examine PSO for an individual state, its local-level Random Forest models and assessing predictive accuracy across several years. With these PA models, states can take the PSO data they collect each year and utilize a set of variables to reliably estimate (e.g., AUC of 80%–90% in this study) exiting high-school SWD that are most likely to be engaged in PSO 1 year after high school. That knowledge could also serve as a valuable resource in educating and supporting current high-school SWD in transition planning for college or career.
As it relates to this study and PA modeling, we acknowledge that there may be misunderstandings or misgivings about its broad application to socioeconomic data. The fact is, PA is not deterministic, and its results should never be considered the “final word” in decision-making for any policies, programs, or practices. It is a tool that needs to be applied carefully and appropriately. There are healthy ongoing debates in research about PA and its appropriate applications. We are also keenly aware of the public’s concerns regarding the over-reliance on technology-heavy and algorithm-driven methods in decision making. There is general agreement across research fields that human judgment remains central, and that these analytic tools should only be used in support of that judgment, not in its place (Chin-Yee & Upshur, 2017). As technology continues to advance, ethical and equitable applications of PA to data (see Palmer & Carpenter-Hubin, 2020) must always be ensured to move these fields forward for everyone.
Bayesian Logistic Regression
The one result in this study that was consistent with findings from other studies in the PSO literature is that high-school graduation was the strongest predictor of SWD engagement in higher education. What this study newly contributes to the literature, however, is that the Bayesian models are also able to predict which type of outcome each SWD is most likely to engage 1 year after high school, relative to the other outcomes. For example, if an exited SWD in 2014 had an overall mean (i.e., “score”) of 3.45, by looking at the corresponding thresholds in Table 4, this clearly met the threshold of the higher education outcome category. Conversely, if that exited SWD in 2014 had a mean of 0.32, this lies within the threshold for the “no PSO” category. On the other hand, if exited SWD have scores that are much closer to the threshold boundary, the interpretations of outcomes would be more nuanced. This would be especially true in the middle categories, such as Cut Point 2 and Cut Point 3, where the differences between threshold values were the smallest in our model results. Overall, when educators at state and local levels have detailed information not only about whether their exited SWD are engaged in any PSO, but also the type of PSO and which set of predictors seem to be driving the outcome, they have more usable information. Their own PSO data (which they collect yearly) and PA can become useful tools in their decision making.
Limitations
There were two limitations in this study. First, the state data we received contained sufficient total numbers of exited high-school SWD across all 5 years, but many schools and districts in the state had small numbers of exiters in their data, or more exiters in one category than in others. In addition, smaller samples (e.g., small numbers of SWD) within many level-2 and level-3 units (i.e., schools and districts) can pose challenges to producing multilevel models that converge to an admissible solution, or models that produce reliable parameter estimates. Grilli and Rampichini (2012) assert that the sample size required for producing reliable estimates in multilevel ordinal models depends on a number of factors, such as the estimation method, cluster variance values, and model complexity. These authors recommend, “In the random intercept case, the estimates are reasonably good with most estimation methods even with 10–15 clusters as long as the average cluster size is at least 10. If the clusters are smaller, more clusters are needed. In the random slope case, the requirement is considerably higher, say 30 clusters of size 30” (p. 6). As described earlier in the Method section, a unique feature of Bayesian statistics is the use of prior information. Thus, some of the challenges with regard to conducting complex multilevel statistical modeling with smaller sample sizes across the school and district levels could be addressed in subsequent studies by utilizing parameter and other distributional information garnered from this study.
The second limitation of this study is the statistical modeling; and it is related to the first limitation. We acknowledge that any statistical model is only as “good” as the limits of the available data being analyzed. This state’s data contained a very small number of predictors due to the state’s need to reduce the identifiability of individual SWD in the PSO data; there certainly could have been some degree of model misspecification as a result. There could have been other variables (not present in these data) that could contribute to exiters becoming engaged in PSO 1 year after high school, for example, individual motivation, access to transportation, and family circumstances. We also recognize that, as this was the first PSO study in the literature to use Bayesian modeling, our selection of “hierarchical priors” in the data analyses could therefore not have been based on previously established studies to identify and select the “right” priors to use. We took a cautious approach, but still could have erred. As Gelman (2007) noted, “All models are wrong, and the purpose of model checking (as we understand it) is not to reject a model but rather to understand the ways in which it does not fit the data” (p. 349).
Implications for Stakeholders
Educators
With an increasing emphasis in U.S. education on using data in decision making for programs, practices, and policies, the application of PA to administrative data, such as PSO of SWD, will become more necessary. Educators can delve into more granular analyses of their PSO data by applying PA modeling approaches from this study to better understand how individual, school, and community factors are connected to PSO across time. Such information could also be shared and utilized at the district and state levels, for example, to identify areas where high-school SWD can obtain additional necessary supports either through their IEP or in community services, mental health, continuing education and job training from local community colleges, leading to better outcomes after exiting high school.
In addition, as state data systems become more complex, the need for real-time actionable information takes on increased significance. Thus the models tested in this study could be considered a template for future model testing, as a starting point for subsequently adding more predictors and addressing this study’s limitations. One possible way to increase the number and types of variables in future analyses is to link a state’s education-agency databases to the relational databases at other state agencies (e.g., health services, workforce/labor) via interagency collaboration. Moreover, the results from Bayesian modeling in this study could be utilized as “priors” (i.e., prior information) in those analyses. Such work could be important for educators to understand how different sets of predictors are more strongly associated with one type of PSO versus another type, and continue to refine their understanding of these analyses with each year’s PSO data. As the economy continues to evolve, the needs of educators, SWD, and families will also have to evolve. This places even greater emphasis on PSO and the need for effective, data-driven transition planning for SWD during high school. For young adults with disabilities already facing obstacles and challenges to post-high school adult life, this need is urgent. For policymakers, supporting educators in the acquisition and application of PA tools is important. They can help to embed these tools and best practices in the everyday work of educators at the district and local levels.
Researchers
The implications of this study for researchers come from the specific modeling techniques we utilized—and within the context of the limitations we noted above. This study was the first in the literature in which PSO data were analyzed using PA with machine learning and ROC for predictive accuracy, and using Bayesian multilevel ordered logistic regression for the different types or categories of outcomes. These models can and should be further refined and tested using state administrative data not only from education but also from state health/human services, labor and employment, and other state agencies. Future studies could also further examine reliability and stability of these models over time, combining state administrative PSO data with other state administrative data. As more states implement longitudinal data systems that link k-12 education data to higher education, employment, and other state data systems, opportunities will expand for researchers to conduct more comprehensive PA modeling of many different predictors over many years. Examining outcomes in the context of students’ in-school experiences, such as participation and completion of career and technical education (CTE), could also further the development of these predictive models. Another important area is to utilize fuller detailed information from students’ IEPs. This would aid researchers in developing models that include the types of special education and related services students are receiving, to the degree they are appropriate, and/or whether the IEP is being implemented as written, or is not being written in such a way as to be able to be effectively implemented.
As described in the implications for educators (above), the Bayesian models developed in this study could also serve as informative “priors” in subsequent state-level PSO research studies. That information could assist researchers in their technical understanding of how well or poorly the predictors within these models perform over time (i.e., empirical testing of reliability and stability of models). This study’s model results could also be used to help researchers develop ways to refine and re-specify models to elaborate where such models were misspecified and how to address such issues. This information could also help researchers further clarify how threshold values near the boundaries of cut points should be understood in terms of probabilities of any of the several PSO categories.
Lastly, although the data analyses in this study were completed using Stata 17 (StataCorp, 2021), it can also be completed using other well-known applications. Commercial options include MPlus, SAS, and SPSS, with varying costs depending on the type of license and features. Open-source options include R, R-Studio, and Python programming platforms (i.e., integrated development environment) such as Jupyter Notebook and Spyder. The authors of this study are not advocating for or recommending the use of any particular program for conducting PA modeling. These programs all provide extensive user support and online resources (e.g., technical specifications and requirements, and programming suggestions) for a range of users, from novice to expert. There are also numerous user-group platforms that offer an array of well-documented, programming advice and step-by step instructions for conducting different PA modeling.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
