Abstract
Score reliability and validity of parent responses concerning their 10- to 17-year-old students were analyzed using the Screening Test for Emotional Problems–Parent Report (STEP-P), which assesses a variety of emotional problems classified under the Individuals with Disabilities Education Improvement Act. Score reliability, convergent, and discriminant validity estimates were adequate.
Children with emotional difficulties make up one of the fastest growing segments of the public school system in the United States (Lane et al., 2009). Prevalence estimates of emotional disturbance in school-age children ranged from 2% to 20% (Walker, Ramsey, & Gresham, 2004), although less than 1% of students actually received special education services under the provisions of the Individuals with Disabilities Education Improvement Act (IDEA; U.S. Department of Education, 2004). Detection of early warning signs lead to early intervention in support of students with possible emotional disturbances (Lane et al., 2009). This detection is hindered by a lack of screening tools that accurately identify emotional disturbances according to IDEA standards (Erford, Balcom, & Moore-Thomas, 2007), as well as confusion and disagreement over the definition of emotional disturbance (Epstein, Nordness, Cullinan, & Hertzog, 2002). The IDEA definition of emotional disturbance is any condition exhibiting one or more of five general emotional characteristics for a prolonged period and to a marked degree: social problems, behavioral problems, depression, anxiety, or some other emotionally based difficulty in learning (U.S. Department of Education, 2004).
Many assessment tools exist today for useful screening of childhood emotional and behavioral difficulties, including the Emotional Disturbance Decision Tree (Euler, 2007), the Behavioral and Emotional Rating Scale–Second Edition (Epstein, 2004), the Achenbach System of Empirically Based Assessment (Achenbach & Rescorla, 2001), the Beck Depression Inventory–Second Edition (Beck, Steer, & Brown, 1996), the Behavior Assessment System for Children–Second Edition (Reynolds & Kamphaus, 2006), and the Conners 3 (Conners, 2008). However helpful, none of these instruments identifies students in accord with all five of the IDEA categories, and most do not have parent, teacher, and self-report versions. Therefore, a truly comprehensive assessment requires educators to use a combination of tools in order to identify children in need of special education services, and this process of administering and interpreting multiple instruments can be very costly and time intensive.
The absence of a screening measure that encompasses all IDEA (U.S. Department of Education, 2004) emotional disturbance facets may be because of obstacles such as high cost, time intensity, lack of training, and legal liability (Erford et al., 2007). Time and financial resources for identifying emotional problems in public school systems often are limited and confined to quick, inexpensive, screening level tools. In addition, mental health professionals in the educational system may not have the appropriate training to use more sophisticated instruments. Another area of concern involves potential liability. School systems are apprehensive about the legality of diagnosing or misdiagnosing a student with mental illness. This combination of factors underscores the need for inexpensive and quick tools that are easy to administer and interpret by educators and school-based mental health practitioners, but which address screening needs without the potential legal ramifications of misdiagnosis or misidentification.
Uhing, Mooney, and Ryser (2005) identified that best practices for assessing emotional disturbance are comprehensive and multifaceted and that the individual’s behavior should be directly observed across various settings. In addition, many individuals with direct knowledge of the targeted client should be interviewed or administered standardized rating scales or informal checklists. Wrobel and Lachar (1998) stated that while teacher reports are the most frequent referral source in the school setting and are very useful in identifying hyperactivity and inattentiveness, parent reports are considered just as useful in the observation of behaviors over long periods of time and in different contexts. Child self-reports, on the other hand, are considered the most useful in the areas of social maladjustments, delinquent behaviors, and internalizing symptoms such as depression. Here, triangulation of data becomes especially important as different respondents may interpret the same questions differently. For example, some studies showed that the presence of anxiety or depression in the caregiver leads to increased reporting of maladjusted behaviors in their children (Briggs-Gowan, Carter, & Schwab-Stone, 1996; Youngstrom, Loeber, & Stouthamer-Loeber, 2000). Therefore, considering information from multiple respondents may lead to clearer diagnostics and a more focused intervention.
Purpose of This Study
Considering the underidentification of students with emotional disturbance (U.S. Department of Education, 2004; Walker et al., 2004) and the lack of efficient multidimensional screening tools (U.S. Department of Education, 2004, Erford et al., 2007), the Screening Test for Emotional Problems–Parent Report (STEP-P) was developed as a cost-effective, quick, and comprehensive instrument that fills the need. This article provides results of preliminary reliability and validity studies for the STEP-P—a new instrument that screens students for emotional disturbances categorized by IDEA. Although self-report and teacher-report versions of the STEP also exist and were reported on elsewhere (Erford et al., 2007, Erford, Short, & Freeman, 2011) the studies below report scores for the parent report only.
Score reliability of the STEP-P was established through studies of internal consistency and test–retest stability using Pearson coefficients. Confirmatory factor analysis (CFA) was used to establish the fit of scores to the hypothesized 5-factor model based on IDEA categories. Convergent evidence was demonstrated by calculating Pearson correlations between subscale scores of the STEP-P, Disruptive Behavior Rating Scale–Second Edition–Parent Report Version (DBRS-II-P; Erford & McCarter, 2012), and Conners-3 parent-report short version (Conners, 2008). Finally, discriminant evidence was established through correlations with the Self-Evaluation Scale–Parent Report (SES-P; Erford, 2012) and Self-Efficacy Parent Report Scale (SEPRS; Erford & Gavin, 2012).
Method
The following four studies consisted of two independent samples of convenience. Some of the children composing Samples 1 and 3 and Samples 2 and 4 had the same mother and father. Norms and samples were differentiated because of norm group differences on several of the subscales. Participants were volunteers who were parents or guardians of children in public and private schools. All guidelines for human subjects study were followed and informed consent was obtained.
Participants
Study 1
Participants were the mothers of 234 students (114 boys, 120 girls) aged 7 to 17 years (M = 11.70; SD = 2.93) from a dozen schools in the mid-Atlantic region of the United States. Of the participants, about 69% were White, 18% African American, 8% Hispanic American, 3% Asian American, and 2% Other. About 16% were from urban settings, 74% from suburban settings, and 10% from rural settings. Of the parents of the participants, 8% of the fathers did not complete high school, 44% had only a high school diploma, 22% completed some college, and 26% were college graduates. Of the mothers, 5% did not complete high school, 41% had only a high school diploma, 28% completed some college, and 26% were college graduates.
Study 2
Participants were the mothers of 71 students (40 boys, 31 girls) aged 7 to 17 years (M = 11.55; SD = 3.08) from four schools in the Mid-Atlantic region of the United States. About 74% of the students were White, 12% African American, 8% Hispanic American, 3% Asian American, and 3% Other. About 91% were from urban or suburban settings (communities larger than 2,500 people) and the remaining 9% were from rural settings. Of the parents of the participants, 7% of the fathers did not complete high school, 48% had only a high school diploma, 20% had completed some college, and 25% were college graduates. About 4% of the mothers did not complete high school, 48% had only a high school diploma, 28% completed some college, and 20% were college graduates.
Study 3
Participants were the fathers of 187 students (92 boys, 95 girls) aged 7 to 17 years (M = 12.08; SD = 3.09) from a dozen schools in the Mid-Atlantic region of the United States. Of the participants, about 66% were White, 17% African American, 11% Hispanic American, 3% Asian American, and 3% Other. About 20% were from urban settings, 70% from suburban settings, and 10% were from rural settings. Of the parents of the participants, 9% of the fathers did not complete high school, 44% had only a high school diploma, 20% completed some college, and 27% were college graduates. Of the mothers, 5% did not complete high school, 41% had only a high school diploma, 26% completed some college, and 28% were college graduates.
Study 4
Participants were the fathers of 68 students (36 boys, 32 girls) aged 7 to 17 years (M = 12.12; SD = 3.15) from four schools in the Mid-Atlantic region of the United States. About 75% of the students were White, 10% African American, 8% Hispanic American, 3% Asian American, and 4% Other. About 88% were from urban or suburban settings (communities more than 2,500 people) and the remaining 12% were from rural settings. Of the parents of the participants, 9% of the fathers did not complete high school, 44% had only a high school diploma, 23% had completed some college, and 24% were college graduates. About 4% of the mothers did not complete high school, 41% had only a high school diploma, 31% completed some college, and 24% were college graduates.
Instruments
The Screening Test for Emotional Problems–Parent Report
The STEP-P (Erford et al., 2007; see Figure 1) can be administered to individual or groups of parents and scored in less than 15 minutes. The test consists of 40 items, with 8 items logically assigned to each of 5 subscales designed in accordance with IDEA categories for identifying emotional disturbances: (a) Academic Problems, (b) Social Problems, (c) Behavior Problems, (d) Depression, and (e) Anxiety. These 8 items on each subscale were selected from a pool of 12 potential items through procedures and each correlated .30 or higher with each subscale score during a pilot study (n = 60; Erford et al., 2007). Teacher and self-report versions of the STEP were also developed with parallel items. Simple sum of score procedures were used to determine raw scores for each subscale, which can be converted into T scores (M = 50; SD = 10) and percentile ranks. Psychometric characteristics of the STEP-P are reported in the Results section. Erford et al. (2007) reported the five factors of mother responses to the STEP-P accounted for 46.2% of variance among items. STEP-P subscale alphas ranged from .76 to .88 and test–retest reliabilities ranged from .74 to .81. Convergent validity yielded moderate to strong validity coefficients, and decision accuracy indicated an overall diagnostic accuracy rate of 91%.

The Screening Test for Emotional Problems–Parent Report version
Conners 3rd Edition–Parent–Short (Conners 3 P-S)
The Conners 3 P-S is a parent report instrument used to assess problem behavior in students aged 8 to 18 years (Conners, 2008). The Conners 3 is a multi-instrument system composed of long and short versions for teachers, parents, and self-report. This study used the short form version (Connors 3 P-S) of the parent-report instrument, which contains 45 items assessing problematic behaviors across 5 factors. A 4-point scale was used to rate each item as follows: 0 = Not true at all (Never, Seldom); 1 = Just a little true (Occasionally); 2 = Pretty much true (Often, Quite a bit); 3 = Very much true (Very often, Very frequent). The normative sample for the Conners 3 P-S included parent or guardian ratings of 1,200 children aged 6 to 18 years (600 males and 600 females). The race and ethnicity of participants were generally representative of the U.S. population. Reliability studies indicated that the Connors 3 P-S demonstrated high levels of internal consistency with a Cronbach’s alpha of .90 (range = .85-.92). Test–retest reliability also showed adequate temporal stability with a correlation of .86 (range = .73-.97; Connors, 2008). Exploratory factor analysis (EFA) of items yielded 5 factors each ranging from 6 to 14 items: Learning Problems, Aggression, Hyperactivity/Impulsivity, Peer Relations, and Executive Functioning. Through confirmatory factor analysis (CFA), a variety of indicators revealed adequate fit to the data: normed fit index (NFI) = .92, nonnormed fit index (NNFI) = .93, comparative fit index (CFI) = .93, and root mean square error of approximation (RMSEA) = .06. Minimum indications of adequate fit were described as NFI > .90, NNFI > .90, CFI > .90, and RMSEA < .10. A well-documented and lengthy discussion of factorial and convergent evidence can be found in the Connors 3 manual (Conners, 2008).
The Disruptive Behavior Rating Scale–2nd Edition–Parent Report version
The DBRS-II-P (Erford & McCarter, 2012) is administered to individuals or groups of parents or guardians and scored in less than 10 minutes. The second edition of the DBRS contains 35 items, with 7 items designated to each of 5 subscales devised to evaluate childhood behavioral difficulties: Distractible, Hyperactive/Impulsive, Oppositional, Antisocial Conduct, and Anxiety. A 4-point scale is used to rate each item: 0 = Never, rarely, or hardly ever; 1 = Occasionally; 2 = Frequently; and 3 = Most of the time. Simple sum of score procedures were used to determine raw scores for each subscale that can be converted into T scores (M = 50; SD = 10) and percentile ranks. Coefficients alpha for a sample of parents on the subscales were the following: Distractible = .90; Hyperactive/Impulsive = .90; Oppositional = .82; Antisocial Conduct = .72; and Anxiety = .75. CFA results for the DBRS-II-P showed a marginal fit of data to the 5-factor model. Correlations between the DBRS-II-P subscales and similar subscales from the STEP-P and Conners 3 P-S reported significant convergent validity correlations. Subscales correlations with the SEPRS and SES-P were used to demonstrate adequate discriminant evidence.
The Self-Evaluation Scale–Parent Report
The SES-P (Erford, 2012) can be administered to individuals or groups of parents and scored in less than 15 minutes. The SES-P consists of 28 items each rated on a 3-point scale (i.e., [U] = Usually; [S] = Sometimes; or [R] = Rarely; the numerical equivalents were 2, 1, and 0, respectively) and interpreted as a total through simple sum of score procedures, which can be converted into T scores (M = 50; SD = 10) and percentile ranks. Coefficient alphas resulting from analysis of standardization sample participant responses indicated the total scale α = .88; test–retest reliability yielded a coefficient of r = .82. Factorial, convergent, and discriminant evidence were demonstrated adequate for a screening-level test.
The Self-Efficacy Parent Report Scale
The SEPRS (Erford & Gavin, 2012) is administered to groups or individuals of parents or guardians and scored in less than 10 minutes. The test contains 19 items rated on a 3-point scale ([U] = Usually; [S] = Sometimes; or [R] = Rarely; the numerical equivalents were 2, 1, and 0, respectively). Simple sum of score procedures were used to determine raw scores for the total score and each subscale, which can be converted into T scores (M = 50; SD = 10) and percentile ranks. Coefficient alpha for a sample of mothers’ and fathers’ scores was .92, indicating a high level of internal consistency. A 14-day test–retest reliability of .87 and .86 for mothers and fathers, respectively, revealed a high degree of temporal stability as well. Factorial, convergent, and discriminant evidence were demonstrated adequate for a screening level test.
Procedures
Institutional review board approval was obtained for conducting the data collection. The STEP-P, DBRS-II-P, SES-P, and SEPRS were administered to mother and father participants from Samples 1 and 3, respectively, in a counterbalanced order and according to standardized procedures. Internal consistency coefficients (Coefficient α) were computed for each STEP-P scale, and the Pearson r was used to conduct interscale correlations. CFA of STEP-P items was conducted to determine the scores’ fit to the hypothesized 5-factor model using Mplus version 5.21 (Muthén & Muthén, 2008). Mothers and fathers in Samples 2 and 4, respectively, were administered the STEP-P and Conners 3 P-S. After exactly 14 days, the STEP-P was administered again. Test–retest Pearson correlation coefficients were computed using subscale standard scores for each administration. A correlation of r = .45 was used as the cutoff point for convergence and divergence because it was used successfully in previous studies (Erford et al., 2007) to distinguish between modest and moderate relationships. Some scale overlap would be expected due to common methods variance and because some students are likely to have problems in more than one area.
Results
Four independent samples of parents were used to explore external and structural aspects of validity (Dimitrov, 2012; Messick, 1989, 1995) and coefficient alpha and test–retest reliability.
Factorial Validity
CFA was conducted on scores of the mothers in Sample 1 (n = 234) using Mplus version 5.21 (Muthén & Muthén, 2008). CFA procedures were used to determine how well the scores fit the hypothesized 5-factor model. Mplus was used because the STEP-P yields ordinal scores. The output statistics generated by Mplus for the mother sample indicated a chi-square to degrees of freedom ratio (χχ2:df) of 2.00, CFI of .95, Tucker–Lewis index (TLI) of .91, RMSEA of .091 (90% confidence interval [CI] = .081-.096), and weighted root mean square residual (WRMR) of 1.36. Dimitrov (2008a, 2008b, 2012) suggested that for adequate model fit, the χ2:df ratio should be <2.00, CFI and TLI should be ≥.90, RMSEA should be ≤.06, and WRMR should be ≤2.00. Thus, the results of this CFA indicate an adequate fit of the data to the 5-factor model; χ2:df, WRMR, TLI, and CFI indicate a good fit, whereas RMSEA indicates a marginal fit. For mother responses, the five factors accounted for approximately 46% of total variance.
CFA was also conducted on scores from the fathers in Study 2 (n = 187) also using Mplus version 5.21 (Muthén & Muthén, 2008). The χ2:df ratio was 2.52, CFI of .96, TLI of .89, RMSEA of .089 (90% CI = .084-.094), and WRMR of 1.14. Thus, the results of this CFA indicate an adequate fit of the data to the 5-factor model; WRMR and CFI indicate a good fit, whereas χ2:df, RMSEA, and TLI indicate a marginal fit. For father responses, the five factors accounted for approximately 45% of the total variance.
Reliability
Internal consistency coefficients of rationally derived subscales of the STEP-P item responses for the mother samples used in Studies 1 and 2 are presented in Table 1 (range = .69-.90). The 14-day test–retest reliability (Pearson) coefficients of rationally derived subscales of the STEP-P mother responses in Study 2 also are presented in Table 1 (range = .67-.84). Likewise, internal consistency coefficients for father responses to the STEP-P items used in Studies 3 and 4 are presented in Table 2 (range = .71-.91). The 14-day test–retest reliability (Pearson) coefficients of rationally derived subscales of the STEP-P father responses in Study 4 also are presented in Table 2 (range = .66-.84). Overall, mothers’ and fathers’ responses to the STEP-P items reached adequate levels of response reliability both in terms of internal consistency and test–retest stability for screening-level purposes.
Convergent and Discriminant Validity and Reliability Coefficients for Mother Responses
Note: N = 234, except for Conners 3 (n = 71). Boldface coefficients indicate expected convergent (i.e., theoretically related) correlations (r ≥.45); italicized coefficients indicate expected discriminant (theoretically unrelated) correlations (r < .45); α1 = coefficient alpha for Study 1 (n = 234); α2 = coefficient alpha for Study 2 (n = 71); rtt = test–retest reliability for Study 2 (n = 71).
Convergent and Discriminant Validity and Reliability Coefficients for Father Responses
Note: N = 234, except for Conners 3 (n = 71). Boldface coefficients indicate expected convergent (i.e., theoretically related) correlations (r ≥ .45); italicized coefficients indicate expected discriminant (theoretically unrelated) correlations (r < .45), α1 = coefficient alpha for Study 3 (n = 234); α2 = coefficient alpha for Study 2 (n = 71); rtt = test–retest reliability for Study 4 (n = 71).
Convergent and Discriminant Validity
Convergent validation of the STEP-P responses involved the calculation of Pearson correlations between the standard scores of theoretically similar subscales of the STEP-P, DBRS-II-P, and Conners 3 P-S, whereas discriminant validation involved correlations with subscales of the SEPRS and SES-P. These results are presented in Tables 1 and 2 for mother and father responses, respectively. Regarding mother responses, generally, the STEP-P displayed moderate to high correlations (i.e., r ≥ .45) with 15 of the 16 scales measuring similar constructs (94%) and low correlations (i.e., r < .45) with 53 of 64 theoretically unrelated criterion measures (83%), indicating acceptable convergent and discriminant evidence, respectively. Father responses yielded a bit more favorable results. Regarding father responses, generally, the STEP-P displayed moderate to high correlations (i.e., r ≥ .45) with 15 of the 16 scales measuring similar constructs (94%) and low correlations (i.e., r < .45) with 60 of 64 theoretically unrelated criterion measures (94%), also indicating acceptable convergent and discriminant evidence, respectively.
Discussion
This investigation of item reliability and validity indicated the STEP-P is an efficient screening instrument for students with a broad range of emotional disturbances as defined by IDEA. In terms of internal consistency and test–retest reliability, the mothers’ and fathers’ responses to the STEP-P items generally reached adequate levels of response reliability for screening-level purposes (α ≥ .80). Regarding factorial analysis, both the mothers’ and fathers’ results for CFA indicate an adequate fit of data to the 5-factor model. Although convergent and discriminant validity correlations for mothers’ and fathers’ responses were both in the acceptable range, fathers’ response correlations were slightly more favorable on measures of theoretically unrelated criteria.
Although promising, these initial results were accompanied by a number of limitations deserving consideration. First, while the Conners 3 (Conners, 2008) and DBRS-II (Erford & McCarter, 2012) were previously published and have good associated psychometric data, the SES-T and SEPRS are instruments developed concurrent with this study. Future studies should use a diverse array of criteria against which STEP subscale scores can be compared.
The 3-point response scale is another possible limitation of the current study. For low-frequency items, a 3-point scale may decrease reliability and validity estimates. Future studies should assess whether a 4- or 5-point scale would lead to more heterogeneous scores, especially for low-frequency items. If so, the psychometric integrity of the STEP could be enhanced by replacing the current 3-point scale with the 4- or 5-point response scale. In addition, associated cross-validation studies should be conducted on independent samples to evaluate the underlying factor structure. These findings confirm that the factor structure proposed in the original five IDEA criteria is supported for theoretical, rational, and practical reasons. With greater understanding of the factor structure underlying the STEP items, and some minor item and scale modifications, scale validation could help researchers develop stronger CFA and cross-validation evidence with new samples.
Including preadolescent with adolescent participants may inflate the correlations and factor distinctions because older students’ problems may be more developed and therefore more detectable. Similarly, combining male and female participants into a single sample may have also inflated correlations since the percentages of male and female students in the problem areas likely differ.
Finally, future studies should address sampling issues regarding selection and diversity. The majority of these studies were conducted in urban and suburban areas. In addition, students whose parents graduated from college and had a higher socioeconomic status were overrepresented in this study. Future research should include members of various residential, socioeconomic, and educational levels. Moreover, this study was influenced by self-selection in that only willing parents completed the response forms. If those who chose not to respond are significantly different than those who did, this could influence the findings of the present study. These possible differences, however, are unknown to the researchers. Future studies must take this potential bias into consideration as it is an ongoing concern within the context of social science research.
Despite these limitations, preliminary studies suggest that the STEP-P is a psychometrically sound screening tool useful for assessment of wide-ranging emotional disturbances. Replication studies could provide the multifaceted, consistent, and useful evidence necessary for the establishment of effective screening instruments, a critical occurrence if stakeholders are to provide timely and consistent interventions for those students identified with emotional disturbances (Cullinan & Sabornie, 2004).
The STEP-P may have important practical applications for professional counselors and other educational stakeholders. The instrument provides an effective, simple, and quick method for identifying students with possible emotional disturbances through teacher, parent, and self-report. In the school setting, increasing instructional time is essential, so an instrument that can be administered and scored in less than 15 minutes could be a valuable tool. The STEP could also help determine whether further diagnostic testing is needed. Such information could help multidisciplinary student support teams make better decisions regarding special education services and to develop targeted early response to intervention strategies. Early intervention programs are time effective and cost effective and could prevent children from developing more severe emotional difficulties.
Footnotes
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
The author(s) received no financial support for the research, authorship, and/or publication of this article.
