Abstract
Identifying incarcerated individuals with poor adaptive functioning (AF) is critical to ensuring their safety and preserving their legal rights, particularly when a diagnosis of intellectual disability (ID) is considered. This study examined the utility of the Problems in Everyday Living Test (PEDL) to identify incarcerated individuals with deficits in AF who may warrant further assessment for ID. The sample consisted of 154 incarcerated adults housed in mental health units in a large urban jail. Latent class analysis supported a three-class model consistent with Impaired, Borderline, and Intact groups, which differed in the level of AF and other indicators of ID. A cutoff score of 13 in the PEDL was optimal to identify incarcerated individuals with deficits in AF, possible intellectual impairment, and a history of special education. Study findings provided preliminary support for using a 12-item modified version of the PEDL as a screening tool in correctional settings.
Keywords
Since the decision in Ruiz v. Estelle (1980), correctional institutions are required to screen for mental health issues, including intellectual disability (ID), among incarcerated individuals. Legal decisions in the context of the death penalty (e.g., Hall v. Florida, 2014; Moore v. Texas, 2016) have further increased awareness of the challenges faced by defendants with ID and the need to improve screening practices in correctional institutions to preserve their rights and ensure their safety while incarcerated. Although guidelines for mental health screening in correctional settings have been developed, they focus primarily on the assessment of intellectual functioning and little is known about how adaptive functioning (AF) is assessed among incarcerated individuals in the United States (American Psychiatric Association [APA], 2016; International Association for Correctional and Forensic Psychology, 2010; National Commission on Correctional Health Care, 2018). To date, there are no well-established ID screening guidelines in use in U.S. correctional institutions (Wijetunga, 2020). Moreover, research suggests that AF is often overlooked in the screening process, despite being essential to a diagnosis of ID (McBrien, 2003).
The conceptualization of ID has moved away from focusing solely on deficits in intellectual functioning to placing equal emphasis on deficits in the adaptive capacities necessary for independent living that emerge during development (APA, 2013). Specifically, for a diagnosis of ID, an individual must present with significant deficits in both intellectual functioning (most commonly identified using an intelligence quotient [IQ] score below 70–75) and deficits in at least one adaptive skill domain: conceptual skills (communication, academics, and self-direction), social skills (interpersonal, social, and leisure), and practical skills (home living, self-care, work, health and safety, and community use). Deficits in AF are directly related to impairments in general mental abilities (e.g., reasoning, problem solving, abstract thinking, and learning from experience) and hinder an individual’s capacity to adequately perform some activities of daily living. The level of severity of ID (i.e., mild, moderate, severe, profound) is determined based on the individual’s level of impairment in AF. An individual’s AF also helps inform the type of support needed, as well as the intensity of services the individual requires.
The estimated prevalence of ID among adults incarcerated in the U.S. ranges between 4% and 10% (Petersilia, 2000), while the estimated prevalence in the U.S. general population is roughly 1% (Zablotsky et al., 2017). There are a number of possible factors contributing to the elevated prevalence of ID in incarcerated samples such as vulnerability to confessing during interrogations, greater willingness to accept a plea bargain, or more difficulty assisting an attorney in preparing a defense (e.g., Ryba & Zapf, 2011; Schatz, 2018). Making a diagnosis of ID among incarcerated individuals is especially challenging because as many as 25% to 30% of those with ID have intellectual and AF levels that lie around two standard deviations below the population mean, also referred to as “mild” ID (Verheijen et al., 2022).
Despite limitations in cognitive functioning, incarcerated individuals with mild ID may not present obvious signs of impairment and largely resemble incarcerated individuals without ID. They often learn ways to mask their limitations to avoid stigma and victimization (Salekin et al., 2010; Schalock et al., 2010). These individuals can sometimes maintain employment and live independently, but like individuals with more severe ID, they typically have difficulty adjusting to institutional life and are at a higher risk of victimization and misconduct while incarcerated compared to individuals without ID (Nixon & Trounson, 2017; Scheyett et al., 2009). In addition, individuals with ID may have difficulty understanding their rights and legal proceedings and are vulnerable to leading questions, acquiescence, and involuntary confessions (Drizin & Leo, 2004; O’Connell et al., 2005). Consequently, early identification of incarcerated individuals with poor cognitive and adaptive functioning is key to provide individualized services, facilitate institutional management, reduce disciplinary misconduct, and preserve individuals’ legal rights and safety. Conducting a valid and reliable assessment of AF is crucial to enable incarcerated individuals with ID to receive the support they need.
Although many tools exist to measure intelligence, relatively few exist to gauge AF and there is no gold standard for its assessment (Salekin et al., 2018). Assessing AF in correctional settings is particularly challenging, making most existing measures of AF unsuitable for use in correctional settings (Tassé, 2009; Young et al., 2007). First, measures of AF often fail to assess for deficits common among individuals with mild ID, such as social competence and gullibility (Greenspan, 2008). As a relatively high proportion of incarcerated individuals with ID have AF deficits that fall under the mild ID category, it is important that measures of AF assess relevant vulnerabilities including more subtle ones. Second, existing AF measures typically require collateral information about a person’s developmental history and day-to-day functioning. Contact with collaterals may be difficult when someone is incarcerated, and the accuracy of information provided may be outdated or of questionable validity, especially when incarcerated individuals come from economically and educationally disadvantaged contexts, have had episodes of homelessness, or do not have identifiable or reliable third-party informants (e.g., family members, close friends, teachers) available who can speak about their functioning in the community, which may be many years in the past (Fisher, 2013; Salekin & Doane, 2009). The retrospective assessment of AF also seems to be problematic, especially when it relies on the potentially inaccurate (and perhaps distant) memory of relatives (Salekin et al., 2018) or correctional staff reports of incarcerated individuals’ functioning in jail/prison (Boccaccini et al., 2016). Third, because of the limited freedom and responsibility permitted in correctional settings, AF cannot be accurately assessed based on functioning while incarcerated (Bonnie & Gustafson, 2007; Everington & Olley, 2008). Finally, existing measures are time-intensive, often taking 30 to 60 min to administer (Sparrow et al., 2016). Consequently, it is not feasible to comprehensively evaluate AF (or intelligence) in all incarcerated individuals who enter an overcrowded criminal legal system.
Given the challenges to using existing measures of AF in correctional settings, a growing emphasis has been placed on the development of objective measures that rely on skills and abilities that are directly observable. For example, the Adaptive Functioning Assessment Tool (AFAT; Smith, 2014) was developed to assess AF inside U.K. prisons. It relies on behaviors that are present in correctional settings rather than in the community, based on the observations and judgment of prison staff (Boccaccini et al., 2016). Although the AFAT may be a useful measure of daily life functioning in correctional settings (Ross et al., 2020), the skills needed to function in a correctional setting, where liberty is highly restricted, are far more limited than those needed to function in the community, where a considerable degree of independence is necessary. Hence, the assessment of AF when an ID diagnosis is suspected should focus on the individual’s ability to function independently in the community, not just their ability to adapt to institutional life (Bonnie & Gustafson, 2007; Everington & Olley, 2008). Particularly when time constraints require a rapid assessment of AF as it existed in the community, there is a clear need for brief, performance-based measures of AF that can be administered without specialized training. A brief measure based on everyday scenarios and decisions may help assess one’s adaptive skills without needing to directly observe their prior behavior in the community (i.e., through collateral informants).
One such measure that was developed specifically to assess AF using a self-report format is the Problems in Everyday Living Test (PEDL; Beatty et al., 1998). The PEDL is a 14-item problem-solving measure intended to assess an individual’s responses to a range of different hypothetical situations. Existing research examining the psychometric properties of the PEDL has found excellent inter-rater reliability, good convergent validity, and utility in differentiating individuals with and without cognitive deficits. In Beatty et al.’s (1998) initial validation study, the PEDL total score displayed excellent inter-rater reliability based on an intraclass correlation coefficient (ICC = .94). They found that the performance of 43 adults with multiple sclerosis on the PEDL was significantly and positively correlated with problem-focused coping (r = .46), emotion-focused coping (r = .32), and verbal abstract reasoning abilities (r = .32), providing some support for the convergent validity of the PEDL.
In a later study by Leckey and Beatty (2002), 22 adults with cognitive impairment secondary to Alzheimer’s Disease scored significantly lower on the PEDL than 18 adults without cognitive or neurological impairment, generating a very large effect size (d = 2.06). Performance on the PEDL was also significantly correlated with patients’ performance of instrumental (r = .71) and basic (r = .58) activities of daily living as rated by caregivers, further supporting its convergent validity. Leckey and Beatty indicated that the PEDL could be used as a screening tool, with “suspiciously low scores” prompting additional examination of particular skills (p. 52). However, they did not provide a cutoff score for identifying significant deficits in AF.
The authors of the PEDL reported that the PEDL structure and scoring system was based on the Wechsler Adult Intelligence Scale–Revised (WAIS-R) Comprehension Test (Wechsler, 1981). Specifically, the PEDL includes three items from the WAIS-R Comprehension test as well as 11 items created by the authors of the PEDL (Beatty et al., 1998). Although the internal consistency and factor structure of the English language version of the PEDL have not been previously described, a Chinese language version (C-PEDL) displayed fair internal consistency, Cronbach’s alpha of .69 (Law et al., 2014). Consistent with research using the English PEDL, Law and colleagues (2014) found significant differences in C-PEDL scores between patients with mild cognitive impairment (MCI) and cognitively healthy controls in both literate and illiterate groups. They found a cutoff score of 21 was optimal to identify participants with MCI (Area Under the Curve [AUC] = .74, SE = 0.09, p = .02; sensitivity and specificity were not provided).
Although potentially useful, the PEDL has rarely been systematically evaluated and studies are needed to evaluate its use in a correctional setting. Particularly, research is needed to determine the utility of the PEDL with incarcerated individuals in the U.S., as well as to identify a possible cut-score that would correspond to significant deficits in AF in this population. This study examined the utility of the PEDL as a screening tool for AF in correctional settings because AF is often overlooked in the screening process of ID, despite being a component of an ID diagnosis. Specifically, this study explored whether the PEDL could be useful as a brief measure of AF that could be integrated into a routine jail/prison intake and administered without specialized training to detect incarcerated individuals with AF deficits. If useful, the PEDL could be used as a first step, in combination with a cognitive screener, or following a positive result on a screening tool for IQ scores in the below average range or lower (Guilmette et al., 2020) to identify incarcerated individuals who may need a more comprehensive ID assessment.
Furthermore, the assessment of AF is critical to rehabilitation efforts in prison. Deficits in AF are likely to interfere with an individual’s capacity to adjust to institutional life and benefit from prison- and community-based programs that address both clinical and criminogenic needs. Thus, if we identify ID accurately when it is present we will be better able to adapt any form of treatment/rehabilitation to the person’s abilities. Unlike other measures such as the Dangerousness, Understanding, Recovery and Urgency Manual (DUNDRUM) (Kennedy et al., 2013), the PEDL is not intended to identify rehabilitation needs or monitor changes (i.e., response to treatment), but it could assist in identifying responsivity factors or possible obstacles to rehabilitation, such as ID, to inform treatment planning.
Thus, this exploratory study focused on a screening measurement of AF to assist in the identification of incarcerated individuals with ID. The primary objective was to examine the psychometric properties of the PEDL in a correctional mental health setting. We conducted latent class analysis (LCA) for PEDL items given the expectation that individuals with ID represent a distinct subgroup of incarcerated individuals. Hence, distinguishing this particularly vulnerable subgroup from incarcerated individuals with average or below average (e.g., borderline) intellectual functioning, who reflect the “typical” offenders in carceral settings, is critical. We also examined whether the unobserved groups differed on demographic characteristics and ID-related variables (e.g., estimated overall IQ, history of special education). Finally, this study aimed to identify a cutoff score to detect incarcerated individuals with significant deficits in AF.
Method
Participants
Participants were recruited from mental health housing units in a large urban jail. These mental health units housed adults who had or were suspected of having a psychiatric disorder, including ID, and required observation (e.g., suicide watch), medication management, and/or intensive clinical services. All fluent English-speaking adults (21 and over) housed in those units during the study period were eligible to participate. Individuals were excluded from participation if they demonstrated difficulties understanding or speaking English or exhibited severe psychiatric symptoms or behavioral problems that interfered with their ability to consent to participate and/or engage in the study.
The initial sample consisted of 201 incarcerated individuals who volunteered to complete an assessment battery designed to evaluate the validity and classification accuracy of ID screening tools (Wijetunga, 2020). Of those participants, 150 (97.40%) men and four (2.60%) women completed all the items of the PEDL, demonstrated adequate effort based on the Test of Memory Malingering (TOMM; Tombaugh, 1996) and the Reliable Digit Span (RDS; Greiffenstein et al., 1994), and were included in the present analyses (further detailed below). As presented in Table 1, in the final sample (n = 154), approximately two fifths (n = 67, 43.51%) of the individuals identified themselves as Black and the average age was 38.03 years old (SD = 11.35, range = 21–73). The racial/ethnic and age composition of the sample resembled the larger Department of Corrections population (New York City Department of Corrections, 2019).
Sample Characteristics (N = 154)
Note. ESL = English as a second language; FSIQ-2 = full-scale IQ assessed by the two-subtest WASI-II.
Roughly one fourth of the sample (n = 38, 24.70%) indicated English was their second language (ESL) but were capable of completing all testing in English. Researchers used their judgment as native or proficient English speakers to assess participants’ English language proficiency during the consent process and throughout the interview. The average full scale IQ based on the two-subtest version of the Wechsler Abbreviated Scale of Intelligence–Second Edition (WASI-II, Wechsler, 2011), referred to as the FSIQ-2, fell in the low average range (80.83, SD = 14.43, range = 54–128), with 40.26% (n = 62) of the sample having a FSIQ-2 equal to or below 75 (i.e., FSIQ scores within the standard error of measurement for the traditional FSIQ threshold for ID, FSIQ ≤ 70).
According to intake evaluations conducted upon admission to the jail, most participants had been diagnosed with a schizophrenia spectrum disorder (n = 88, 57.14%), while a substantial minority were diagnosed with a mood/anxiety disorder (n = 34, 22.08%) and only two individuals were diagnosed with ID at the time of intake (n = 2, 1.30%). Most participants self-reported a history of psychiatric hospitalizations (n = 114, 74.03%), and about half were supported by disability benefits from the government (n = 76, 49.35%).
Measures
As part of a larger assessment battery, participants were administered a demographic and mental health questionnaire, the TOMM, followed by the two-subtest version of the WASI-II, the RDS, and the PEDL. The order of test administration was consistent across participants. Both the TOMM (Tombaugh, 1996) and the RDS (Greiffenstein et al., 1994) were used to assess the possibility of insufficient effort during testing or deliberate symptom exaggeration (see Wijetunga, 2020). The two-subtest version of the WASI-II (Wechsler, 2011), which is comprised of the Vocabulary and Matrix Reasoning subtests, was administered to minimize burden while generating an estimate of current intellectual functioning (i.e., FSIQ-2). In this study, the WASI-II demonstrated excellent inter-rater reliability, ICC (k = 2, N = 23) = 1.00, .99, .99 for the Vocabulary subtest raw score, Matrix Reasoning subtest raw score, and FSIQ-2, respectively. In addition to the FSIQ-2, four other variables were considered as indicators of possible ID: (a) educational attainment, (b) self-reported history of special education due to learning problems, (c) history of employment (i.e., not having maintained employment for more than 6 months at any time in the past), and (d) self-reported history of independent living (i.e., not having lived on their own at any time in the past or having required considerable support from others).
In addition, participants were administered a modified version of the PEDL. The PEDL (Beatty et al., 1998) is a 14-question, interviewer administered measure intended to assess practical problem-solving skills. Each question presented a scenario or problem (e.g., what should the person do if they spot smoke and fire while in a movie theater) and the interviewer wrote the examinee’s responses verbatim. Responses were scored on a 0 to 2 scale following Leckey and Beatty’s (2002) scoring guidelines, with higher scores representing more adaptive responses. In this study, some additions to the original scoring guidelines were made to facilitate consistent scoring of vague or short answers frequently provided by study participants. For example, we added additional possible answers for several PEDL items and specified the scoring value they merited and gave a score of 1 when participants’ responses were too vague or incomplete to justify a score of 0 or 2 (our scoring guidelines are available from the corresponding author upon request). Scores were summed to generate a total score that ranged between 0 and 28. The administration of the PEDL took approximately 5 to 10 min.
For this study, eight PEDL questions were adapted to make their content more relevant to urban life and our participants. For example, a question asking how the person would find their way if lost in a forest was changed to “If you were lost in a part of the city you are not familiar with and your cell phone is dead, how would you go about finding your way home?” A question about what the person would do if their new coffee maker stopped working was changed to “Last month you purchased a new cell phone; it worked well for about three weeks but now it does not turn on. What do you do?” (all adapted PEDL items used in this study are available from the corresponding author upon request). The scoring criteria for the adapted questions were consistent with the original PEDL. Furthermore, an optional prompt was added for each question asking the individual “What should you do?” as an attempt to differentiate between those who were unable to provide a correct answer (i.e., impaired AF) and those who may have been unwilling to do so (Greenspan & Switzky, 2006). This question prompt was an attempt to control for the potential effect of antisocial or psychopathic traits on AF scores (Young-Lundquist et al., 2012) by providing a socially unacceptable response (e.g., “beat up the person who sold me the broken phone”). When the optional prompt was used, the total PEDL score was calculated using participant’s answer to that prompt.
Procedure
Data were collected on five mental health units within a large jail complex, all of which had private office space available where the assessments were conducted, helping to protect participants’ confidentiality and test validity. Graduate-level research assistants approached individuals residing in those units and gave them a brief explanation of the purpose of the study. If the individual expressed interest in learning more about the study, they were invited to the office space, where the research assistant explained the purpose of the study and provided a written consent form. Participants were informed that no incentive was offered and that confidentiality would be breached if they were in imminent risk of harming themselves or others or disclosed that they were planning an escape from the facility. Participants could ask questions as well as discontinue their involvement at any time during the study process. Participants also provided consent to having their electronic jail records reviewed. Upon obtaining written informed consent, research personnel conducted a semi-structured interview to obtain background information (i.e., academic, mental health, medical, and vocational history) and administered the assessment battery. For 28 interview and testing sessions, a second research assistant was present to allow assessment of inter-rater reliability.
To ensure the PEDL was scored reliably and assess its inter-rater reliability, two researchers independently double-coded 103 (66.88%) participants’ responses to the PEDL. Twenty-five of these cases were used to train researcher staff to reliably score the instrument. First, the researchers coded five cases together and discussed the coding guidelines and ratings. Second, the coders independently coded 10 cases and compared their coding to identify and discuss any differences in scoring and find resolutions. This process was repeated until both coders agreed on most codes and were able to score five consecutive cases without inconsistencies. Based on the training cases, the PEDL coding guidelines were clarified to more reliably score common responses that were not originally captured in the scoring (Leckey & Beatty, 2002). Once the training process was completed, two researchers independently coded the remaining 78 (50%) PEDL cases. They continued to discuss any coding uncertainties or discrepancies, which were resolved by consensus. Independent ratings were used to calculate inter-rater reliability, whereas consensus ratings were used in statistical analyses. To protect confidentiality and integrity of the data, each participant was assigned a numeric code and the researchers who scored the PEDL were not aware of participants’ performance on other tests. This study received institutional review board (IRB) approval from the correctional institution and Fordham University.
Data Analyses
Descriptive statistics were calculated for the participant demographic characteristics (Table 1), the PEDL, WASI-II, and other indicators of ID. Inter-rater reliability of the PEDL item ratings was assessed using Cohen’s Kappa, which corrects for rater agreement due to chance. Kappa coefficients are typically interpreted with the following guidelines: poor ≤ .20, fair ≤ .40, moderate ≤ .60, good ≤ .80, and excellent > .80 (Cohen, 1960). Inter-rater reliability of the PEDL total score and WASI-II were assessed using two-way, mixed effects, absolute agreement ICC, based on the following interpretive guidelines: poor < .50, moderate = .50 to .74, good = .75 to .90, and excellent > .90 (Koo & Li, 2016).
There were no missing data for any of the PEDL items. Missing data analysis was conducted before the LCA. Educational attainment, self-reported history of special education, employment, and independent living had missing data rates ranging from 0.00% to 16.20%. Missing completely at random (MCAR) tests were conducted for these variables (Enders, 2022). The correlation of the missingness between these variables and other analyzed variables was investigated. Results supported that the missingness indicators (0 = observed, 1 = missing) of educational attainment (M = 2.47, range = 1–6), self-reported history of special schooling (M = 0.69, range = 0–1), employment (M = 0.64, range = 0–1), and independent living (M = 0.76, range = 0–1) were not associated with other analyzed variables (median absolute r = .055, range = .004–.187), supporting that they were MCAR and listwise deletion was appropriate and used.
For the PEDL items, LCA was conducted using robust maximum likelihood (MLR) estimation in Mplus 8.8 (Muthén & Muthén, 1998–2017). We used 5,000 sets of random start values and all analyses had replicated best log-likelihood values indicating model convergence (Collins & Lanza, 2010). The one- to three-class models were estimated using all 14 items; models with four or more classes did not converge and were therefore uninterpretable. Multiple criteria were used to evaluate model fit and identify the appropriate number of latent classes (Nylund et al., 2007; Nylund-Gibson & Choi, 2018): (a) Bayesian information criterion (BIC) with a lower value indicating better model fit, (b) Sample-size adjusted Bayesian information criterion (SABIC) with a lower value indicating better model fit, (c) Akaike information criterion (AIC) with a lower value indicating better model fit, (d) the Vuong–Lo–Mendell–Rubin adjusted likelihood ratio test (VLMR-LRT), where significant results supported a k-class model over the (k − 1)-class model, and (e) Bootstrapped likelihood ratio test (BLRT), where significant results also supported a k-class model over the (k − 1)-class model. Although it was not used as a model selection criterion, entropy was also reported as an index of the classification accuracy of participants of different latent classes (Masyn, 2013; Nylund-Gibson & Choi, 2018). Model selection was also based on the substantive interpretation of the latent classes, based on item response probabilities and class proportions. Items that did not contribute to identifying latent classes (i.e., no class differences on the item) were removed. After determining the best-fitting model, sensitivity analysis was conducted to examine its stability and reliability. In the sensitivity analyses, LCA was conducted with other sets of items being removed.
In addition, the latent class differences in age, FSIQ-2, educational attainment (11th grade or lower/no general educational diploma [GED] = 1, GED = 2, high school diploma = 3, some college coursework = 4, undergraduate degree = 5, graduate degree = 6), self-reported history of special education (yes = 0, no = 1), employment (yes = 1, no = 0), and independent living (yes = 1, no = 0) were tested using the three-step procedures in Mplus. Exploring these correlates enabled us to understand how they distinguished latent classes, providing insights into the associations between ID-related variables and latent classes based on performance on the PEDL. Different from other clustering methods like k-means clustering, LCA does not assign a participant to a class; it produces latent class posterior probability for each participant belonging to each class. To properly account for these probabilities when testing latent class differences, we followed the suggestions from previous studies (Asparouhov & Muthén, 2014a, 2014b) and used the BCH procedure for continuous variables (Bolck et al., 2004) and the DCATEGORICAL procedure for categorical variables (Lanza et al., 2013).
Finally, the classification accuracy of the PEDL total score was examined using the Receiver Operating Characteristics (ROC) curve analysis. An AUC statistic of .80 or greater was deemed to reflect good discrimination between the latent class with the lowest mean PEDL total score (i.e., the class with the most impairment in AF based on PEDL responses, more suggestive of a diagnosis of ID) and the other two latent classes combined (Mossman, 1994). Sensitivity and specificity statistics for different cutoff scores of the PEDL were also calculated.
Results
PEDL Descriptive Statistics
The mean PEDL total score was 17.46 (SD = 3.54, range = 7–27, skew = −0.23, kurtosis = −0.07). Individual item means ranged from 0.71 to 1.83 (Mdn = 1.21; see Table 2). The correlations among the PEDL items ranged from r = .00 to .32 (Supplemental Table S1, available in the online version of this article). While some items were significantly correlated with one another, the magnitude of the correlations was generally small, indicating that the PEDL items were largely independent from each other. The PEDL item scoring generated excellent inter-rater reliability, with κ ranging from .83 to 1.00 (Mdn = .95; Table 2). The inter-rater reliability of the PEDL total score was also excellent, with an ICC (k = 2, N = 78) of .98.
PEDL Item Characteristics and Inter-Rater Reliability
Note. κ = Cohen’s Kappa; ICC = Intraclass correlation coefficient.
The optional prompt, “What should you do?,” was used in at least one question with only 22 participants and only impacted the total PEDL score of 13 participants by increasing it by 1 (n = 3) or 2 (n = 10) points. When asked the optional prompt, the remaining nine participants did not provide an answer that granted a higher score than the answer provided to the original prompt and, therefore, it did not impact their total PEDL score. As the optional prompt was rarely used, we could not assess whether adding the optional prompt had any impact on the PEDL.
LCA
Table 3 presents the model fit and classification diagnostics of the one- to three-class latent class models using all 14 PEDL items (the PEDL items had no missing values). The BIC suggested a one-class model, but SABIC and AIC suggested a three-class model, and BLRT and VLMR supported a two-class model. It is common for these criteria to disagree on the number of latent classes (Nylund-Gibson & Choi, 2018), so the group proportions and item response probabilities were examined in each of the three models (one-, two-, and three-classes). The proportions of group membership were 100% for the one-class solution, while the two-class solution evenly divided the sample into higher and lower performing groups (50.80% and 49.20%); the three-class solution generated group sizes of 53.30%, 38.10%, and 8.60%. Moreover, the size of the smallest group (8.60%) in the three-class solution was roughly consistent with the prevalence rate of ID found in the U.S. correctional system (Petersilia, 2000), which further supported interpretation of the three-class model. Supplemental Figures S1 to S3 (available in the online version of this article) present the item response probabilities and the class proportions for the one-class, two-class, and three-class models, respectively.
Model Fit of Latent Class Models
Note. LL = Log-likelihood; BIC = Bayesian information criterion; SABIC = Sample-size adjusted BIC; AIC = Akaike information criterion; BLRT = Bootstrapped likelihood ratio test; VLMR-LRT = Vuong–Lo–Mendell–Rubin adjusted likelihood ratio test. Bold indicated the selected model based on the criterion.
In the selected three-class model, the item response probabilities of Item 7 (what the person would do if they realized their friend left their backpack on a park bench) and Item 14 (how the person would find their way if lost in the city) were about the same between the three classes, meaning that these two items lacked power in identifying latent classes (i.e., lack of class separation). Thus, these two items were removed. The revised 12-item three-class model was supported by the SABIC, AIC, BLRT, and VLMR-LRT (Table 3). The entropy of this model was .87, supporting satisfactory and high classification accuracy. The item means and the class proportions of each latent class are presented in Figure 1. Class 1 (n = 81, 52.60%), Class 2 (n = 58, 37.66%), and Class 3 (n = 15, 9.74%) were labeled for the purposes of this study as “Intact,” “Borderline,” and “Impaired,” respectively, as the Intact latent class obtained adequate scores across most PEDL items (M total score = 17.37, SD = 2.13, range = 13–24), while the Impaired group performed poorly on most items (M total score = 10.76, SD = 2.16, range = 6–13) and the Borderline group showed variable performance, with marginal (i.e., 1-point) responses to many items (M total score = 13.48, SD = 1.70, range = 9–17).

Item Class Profile Plot for the PEDL for the Three-Class Model
As a sensitivity analysis, LCA with fewer items were conducted. Items that did not discriminate between classes relatively well were removed, and models with 8, 9, and 11 items were examined (see Supplemental Table S2 and Figures S4 to S7, available in the online version of this article). The three models yielded similar results to those based on 12 and 14 items, with model indices supporting the three-class model. Furthermore, we examined whether participants remained in the same class when using 14 items and after removing the two items that lacked power in identifying classes (i.e., Items 7 and 14). Only the class membership of six (3.90%) participants changed. Four participants moved from the Intact to Borderline group, and two participants moved from the Borderline to Intact group. Overall, these results supported the 12-item three-class model was stable and reliable.
Latent Class Differences on Age and Indicators of Possible ID
Table 4 shows the results of latent class differences in age, FSIQ-2, and educational attainment. There was a significant class difference on age (χ2 [df = 2] = 12.78, p = .002). The Intact group (M = 40.90) was significantly older than the Borderline group (M = 35.76, χ2 [df = 1] = 5.52, p = .02) and the Impaired group (M = 31.38, χ2 [df = 1] = 11.33, p = .001). The Borderline and Impaired groups did not differ significantly in age (χ2 [df = 1] = 2.33, p = .13). The Intact group (M = 84.97) obtained a significantly higher mean FSIQ-2 than the Borderline group (M = 76.84, χ2 [df = 1] = 8.23, p = .004) and Impaired groups (M = 73.92, χ2 [df = 1] = 9.23, p = .002), but the Borderline and Impaired groups were not significantly different from one another (χ2 [df = 1] = 0.58, p = .45). The three classes did not differ significantly on educational attainment (χ2 [df = 2] = 1.73, p = .42).
Latent Class Differences on Age and ID-Related Variables
Note. Age: years old; FSIQ-2: full-scale IQ based on two-subtest version of the WASI-II; Education: 11th grade or lower/no GED = 1, GED = 2, high school diploma = 3, some college coursework = 4, undergraduate degree = 5, and graduate degree = 6.
Different superscripts indicate statistically significant differences (p < .05) between groups.
p < .05. **p < .01.
Table 4 presents the results of the latent class differences in self-reported history of special education, employment, and independent living. While the three classes did not significantly differ on their history of employment (lasting more than 6 months) (χ2 [df = 2] = 3.25, p = .20) or living independently (χ2 [df = 2] = 2.01, p = .37), they significantly differed on special education (χ2 [df = 2] = 6.07, p = .048). There was a significantly higher proportion of incarcerated individuals with a history of special education in the Impaired group than in the Intact (χ2 [df = 1] = 6.06, p = .01) or Borderline (χ2 [df = 1] = 4.15, p = .04) groups.
Potential Cutoff Scores for Identifying Impaired Functioning
The distribution of PEDL scores across the three latent classes is presented in Table 5. ROC curve analysis demonstrated strong discrimination between the Impaired group and the combined Intact and Borderline groups, with an AUC of .94, 95% confidence interval [CI] = [.90, .98], p < .001, for the 12-item modified version of the PEDL. A cutoff score of 13 or less on the PEDL maximized both sensitivity and specificity; 100% of the participants in the Impaired group were correctly classified by the PEDL (i.e., sensitivity = 1.00), whereas only 21.58% of the participants (n = 30 of 139) in the Intact and Borderline groups obtained scores of 13 or lower (i.e., 78% specificity). Only three of 81 participants (3.70%) from the Intact group obtained PEDL scores of 13 or lower along with 27 of 58 participants (46.55%) from the Borderline class. Participants who scored at or below the cutoff score of 13 had an average FSIQ-2 = 73.47 (SD = 12.51, range = 54–105) and those who scored above had an average FSIQ-2 = 83.87 (SD = 14.13, range = 55–128). A cutoff score of 12 on the PEDL improved specificity (.87) but had poorer sensitivity (.80), failing to identify 3 of the 15 (20.00%) participants in the Impaired group but misclassifying 18 of 58 participants (12.90%) from the Borderline group.
Distribution of 12-Item PEDL Scores Across Classes
Discussion
Correctional institutions in the United States are required to screen incarcerated individuals for ID (Ruiz v. Estelle, 1980). However, guidelines for mental health screening in correctional settings and research on this area suggest the assessment of AF is often overlooked, despite being essential to a diagnosis of ID. Brief screening tools of AF, as it existed in the community, that can be administered directly to the incarcerated individual without specialized training could help in accurately and efficiently screening the high volume and fast turnover of people seen in prisons and jails (Catalano et al., 2020). This is the first study exploring the psychometric properties and utility of the PEDL, a brief performance-based measure of AF, to identify incarcerated individuals with deficits in everyday problem-solving abilities who may warrant further assessment for ID.
We found excellent inter-rater reliability for the PEDL item and total scores (κ = .83–1.00, ICC = .98), which were comparable with those found in the PEDL initial validation study (Beatty et al., 1998) and a subsequent Chinese translation (Law et al., 2014). It is, however, important to note that in this study, the scoring guidelines were expanded to facilitate consistent scoring of the frequent vague or short answers. This likely increased the inter-rater reliability of the PEDL items. When using the PEDL with correctional populations, it may be helpful to ask follow-up questions to allow for a more accurate assessment, to reduce the potential for rater’s bias and improve inter-rater reliability.
Our analysis also found that two items (Items 7 and 14), which were modified from the original to accommodate an urban setting, did not discriminate between the latent classes. Therefore, these items may not be particularly useful among incarcerated individuals. When the LCA was conducted on the remaining 12 PEDL items, the results supported a three-class model that was consistent with and labeled for the purposes of this study as Impaired, Borderline, and Intact groups. These three classes differed primarily in the level of AF, which supports the assumption that AF occurs along a continuum, but is not necessarily normally distributed (i.e., the Impaired class was substantially larger than would be expected if AF was normally distributed, but it is roughly consistent with the proportion of individuals with ID found in correctional settings). Of note, other characteristics of the sample, such as the high prevalence of schizophrenia spectrum disorder, might have also contributed to the proportions of the classes (this is further described in the Limitations section below). The three classes were also differentially associated with age and ID correlates. Members of the largest group, the Intact class, were more likely to be older, which is consistent with prior research and theories suggesting that everyday problem-solving skills are derived from accumulated knowledge and experience and, therefore, increase with age (e.g., Mienaltowski, 2011). However, Law et al. (2014) did not find age-related effects on C-PEDL performance. This difference in results may be due to the broader range of ages in this study compared to Law et al.’s sample.
Consistent with previous research examining the association between adaptive and intellectual functioning, the Intact group had a significantly higher mean FSIQ-2 (M = 84.97, SD = 1.68) compared with the other two classes (Hayes, 2005; Law et al., 2014). While individuals in the Impaired and Borderline groups shared many characteristics, they differed in their history of special education services (i.e., members of the Impaired class were significantly more likely to have a history of special education services than the Borderline and Intact classes). These findings provide some preliminary support for the validity of the emerging classes, despite the lack of significant differences in educational attainment, history of employment, or independent living. This discrepancy likely reflects the operationalization of these historical variables, as a history of 6 months of employment or living independently (for any length of time) does not preclude a diagnosis of ID, while the absence of this history could be due to many other factors (e.g., substance abuse, serious mental illness).
Finally, this study’s findings suggested that a cutoff score of 13 on the 12-item modified version of the PEDL was optimal in identifying individuals in the Impaired class. This cutoff score correctly identified all individuals from the Impaired class, but also included several individuals from the Borderline class (nearly half of that subgroup) and a handful of the Intact group (3.70%). While these may indeed be “false positives,” some may also be individuals who have ID. Regardless, given that the goal of a screening tool is to cast a somewhat wider net so that subsequent, more rigorous testing can identify precisely which individuals require accommodations, these findings are reassuring, albeit preliminary. Indeed, a cutoff score of 13 identified a significant number of incarcerated individuals with deficits in AF, possible intellectual impairment (MFSIQ = 73.92, SD = 3.18), and a history of special education.
Clinical Implications
Although further research supporting this study’s findings is needed, the preliminary findings suggest the PEDL could assist in the legally required screening for the presence of ID in U.S. correctional institutions. It is first important to note the high prevalence of deficits in intellectual functioning (FSIQ ≤ 75) among study participants (40.26%). Despite that, only 17.53% of the sample had both an FSIQ ≤75 and a PEDL total score ≤13. Although neither the 2-subscale WASI-II nor the PEDL is sufficient for a diagnosis of ID, based on study findings, the PEDL may allow identifying incarcerated individuals with AF deficits who might meet the criteria for ID in an effective yet cost and time-efficient manner. While the PEDL item content, self-report nature, and prior research showing correlations between PEDL scores and measures of AF skills (Beatty et al., 1998; Leckey & Beatty, 2002) suggest the PEDL assesses adaptive skills across the three domains (conceptual, social, practical) noted in the Diagnostic and Statistical Manual of Mental Disorders (5th ed.; DSM-5; APA, 2013), performance on the PEDL does not provide information about the adaptive skill domain in which the person has deficits. Additional examination of particular skills would be needed to identify in which area a person is having difficulties, whether they are consistent with an ID diagnosis, and to inform effective management and treatment of individuals with ID in correctional settings.
The PEDL could be used as a first step, in combination with an IQ screener, or following a positive result on a screening tool for IQ scores in the below average range or lower to identify individuals with AF deficits who might meet the criteria for ID. The PEDL could be integrated into a routine admission jail/prison intake as it could be administered relatively fast by nonclinical staff, helping them triage the need for a more comprehensive evaluation to assess for ID. Nonetheless, prior to its implementation, it would be crucial to investigate whether nonclinical staff could reliably score the PEDL when provided with clear scoring guidelines.
While the tentative cutoff score of 13 on the 12-item modified version of the PEDL may identify some individuals without ID, it maximizes the identification of individuals with deficits in AF who would require a more thorough evaluation to determine whether an ID diagnosis is warranted. Nonetheless, it is important to note that this tentative, proposed cutoff score warrants additional examination across multiple samples before its adoption in clinical practice. Institutions must also decide for themselves whether the balance between sensitivity and specificity rates is acceptable or whether a different cutoff score is preferred to maximize sensitivity or specificity.
Strengths, Limitations, and Future Directions
This study included a relatively large, racially, and ethnically diverse correctional sample, particularly compared to other studies examining ID in correctional settings. The sample was also diverse with respect to age and educational attainment. Unfortunately, the very small number of women participants in this study did not allow for the examination of the impact of gender on the PEDL. Although gender has not been found to impact the assessment of ID or AF, most studies have had predominantly male samples. Hence, the use of the PEDL with women and nonbinary gender samples should be evaluated.
This study was conducted in a large urban jail in the U.S. and, therefore, the results may not generalize to other areas of the U.S. or other countries. In addition, study participants had current or suspected mental health issues that may have put them at a greater risk for presenting deficits in intellectual and adaptive functioning regardless of ID (Sheffield et al., 2018). Research has shown cognitive impairments and decreased functional abilities are common among individuals with severe mental disorders (Bouras et al., 2004; Woodberry et al., 2008). Recent genomic studies suggest schizophrenia, bipolar disorder, and neurodevelopmental disorders lie on a gradient of severity, differing to some extent quantitatively and qualitatively (Owen & O’Donovan, 2017). For instance, individuals with mental illnesses that typically have an onset later in the life span (late teens and early twenties) will have had opportunities to acquire skills in a way that those who have had impairment in their function from birth may not have. Thus, the high prevalence of mental illness/substance abuse may have confounded the assessment of AF and IQ in this study, compromising the generalizability of the findings to incarcerated individuals in general population. To control, or at least minimize, the potential confounding impact of mental illness on AF assessment, future research could examine the utility of the PEDL in the general prison population. In practice, nonetheless, the comprehensive assessment expected to follow a positive screening with the PEDL could aid in distinguishing the etiology of AF deficits and clarify whether an ID diagnosis is warranted.
Perhaps the most notable limitation of this study is that no other measure of AF was administered alongside the PEDL. The association between the PEDL and other established measures of AF designed specifically for ID assessment needs to be evaluated to establish the construct validity of the PEDL. However, this was not possible in this study for several reasons, including resource and time constraints and the challenges in assessing AF in correctional settings described earlier. The PEDL was administered as part of a larger assessment battery and adding a comprehensive measure of AF, even if reliable collateral sources were available, would have significantly extended the time required to complete the battery, likely increasing attrition or decreasing effort. If feasible, future research should further examine the construct validity of the PEDL using a comprehensive AF measure. Relatedly, given the lack of an external measure of AF, the use of the labels “Impaired,” “Borderline,” and “Intact” for the three classes identified in this study was tentative and was not intended for clinical purposes (i.e., to establish norms to classify individuals). Although those labels appeared to fit our identified classes, further research is needed to determine whether those are the most appropriate labels.
Because the wording and scoring of many PEDL items were modified, it is unclear whether comparisons between this study findings and (limited) past research are meaningful or appropriate. The wording of eight PEDL items was adapted to make their content relevant to urban life and our participants, while maintaining the skill being assessed. For example, the last item of the PEDL was modified from what the person would do if they were lost in the forest to if they were lost in the city. While we modified the context (forest vs. city), both items assessed how the person would go about finding their way home. However, given that the original items were not administered in this study, we could not examine whether wording changes had any unintended impacts on measurement of the skill. We also added an optional prompt for each PEDL item (what should you do?) as an attempt to control for the potential effect of antisocial or psychopathic traits on AF scores; however, its rare use and minimal changes on participants’ scores suggest this prompt did not have any significant impact on the results of this study. In addition, to accurately score the PEDL in this sample, it was necessary to adjust the scoring guidelines to better classify frequent responses elicited from this sample. It is unknown whether raters in previous studies ran into similar issues in their coding or if this result was unique to the current sample. The vague and simplistic answers often elicited from the participants could reflect a lack of effort, and differentiating lack of effort from impaired AF is challenging. Thus, it would be necessary to evaluate whether prompting participants to expand on answers (i.e., permitting follow-up questions) would facilitate a more accurate assessment, reducing the potential for rater’s bias and improving inter-rater reliability.
Taken together, our findings are promising, albeit preliminary, highlighting the need for further research to provide additional support for the utility of the PEDL as an AF screening tool in U.S. correctional settings. The use of LCA rather than factor analysis would help identify possible distinct groups and examine whether the three-class model and proposed labels replicate across multiple samples. The power of Items 7 and 14 in differentiating classes must also be investigated before determining whether their removal is justified in correctional settings. Relatedly, the proposed cutoff warrants external confirmation before its adoption.
Conclusion
The present exploratory study was the first attempt to study the utility of the PEDL in identifying incarcerated individuals with deficits in AF who may need a comprehensive ID evaluation. Study findings provided preliminary support for the use of a 12-item modified version of the PEDL as a screening tool in correctional settings. Although a brief measure like the PEDL should not substitute for a comprehensive assessment of AF, in an overcrowded setting where resources and time are limited, a screening tool such as the PEDL may be critical to help identify incarcerated individuals with ID who are at an increased risk of having difficulties adapting to institutional life and to respond to their needs to preserve their legal rights and safety. To date, no instrument seems suitable for assessing AF in a short period of time and without relying on informants’ interactions with and observations of the incarcerated individuals being evaluated. Therefore, further investigating the utility of the PEDL as a potential screening tool for AF is crucial to support its clinical use as part of the assessment of ID in correctional settings.
Supplemental Material
sj-docx-1-cjb-10.1177_00938548241268135 – Supplemental material for Measuring Adaptive Functioning in a Correctional Setting: An Analysis of the Problems in Everyday Living Test (PEDL)
Supplemental material, sj-docx-1-cjb-10.1177_00938548241268135 for Measuring Adaptive Functioning in a Correctional Setting: An Analysis of the Problems in Everyday Living Test (PEDL) by Maria Aparcero, Hyunjung Lee, Charity Wijetunga, Erin M. Conley, Barry Rosenfeld and Heining Cham in Criminal Justice and Behavior
Footnotes
Authors’ Note:
We have no known conflict of interest to disclose. Maria Aparcero is now at Patton State Hospital.
Supplemental Material
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
