Abstract
The Strengths and Difficulties Questionnaire (SDQ) is a screening measure commonly used to assess behavioral and emotional symptoms and strengths among children and adolescents. However, despite its frequent use, its underlying factor structure remains an important area of inquiry. Whereas the original five-factor structure has often been supported through exploratory factor analysis, results from confirmatory analyses continue to yield mixed results. We analyzed data from youth in Grades K through 12 from a large epidemiologic study in the Southeastern United States. Teacher-report SDQ data were used to test three confirmatory factor models by school level (i.e., elementary [Grades K–5] and secondary [Grades 6–12]): The original five-factor model, a three-factor model, and a bifactor model. Model fit indices and reliability measures supported the original five-factor model as the preferred model when using the teacher-reported SDQ with both elementary and secondary school children. Implications for using the SDQ in applied research and predictive modeling are discussed.
Keywords
The Strengths and Difficulties Questionnaire (SDQ; R. Goodman, 1997) is a brief, 25-item screening measure used to assess behavioral and emotional symptoms and strengths (e.g., conduct problems, hyperactivity, emotional symptoms, peer problems, and prosocial behaviors) among children and adolescents. The SDQ is a free open-source measure that has been translated into many different languages. Self- and other-informant (i.e., parent and teacher) forms are available. The SDQ is widely used in clinical and research assessments and in large community epidemiological studies, including the National Health Interview Survey.
Despite its popularity as a screening tool, the underlying factor structure of the SDQ remains a subject of debate (e.g., Caci et al., 2015; A. Goodman et al., 2010). The original five first-order factor structure of the SDQ (R. Goodman, 1997) has not been consistently supported in confirmatory analyses (Caci et al., 2015; Kóbor et al., 2013; Niclasen et al., 2013; Tobia et al., 2013). Thus, further investigations into alternate factor structures are warranted. The aim of the current study was to compare alternative measurement models, including a bifactor model, of the teacher-report SDQ using data from a large population-based epidemiological study in the Southeastern United States.
Development of the SDQ
The SDQ was originally developed to capture information similar to the well-validated Rutter Questionnaire (Rutter, 1967) while simultaneously improving on the utility of the measure. Although the Rutter Questionnaire was widely used and demonstrated strong psychometric properties, Goodman (1997) determined a need for an updated behavioral screening measure for youth to better match conceptualizations of psychopathology and functioning that aligned with Diagnostic and Statistical Manual of Mental Disorders criteria (4th ed.; DSM-IV; American Psychiatric Association, 1994). The SDQ was designed with the following five primary specifications: (a) to fit easily on one sheet of paper; (b) to be applicable to children between the ages of 4 and 17; (c) to have similar self-report and informant-report versions that demonstrate high inter-rater reliability; (d) to represent behavioral strengths in addition to psychological and behavioral difficulties; and (e) to include equal numbers of items for each of the constructs represented by the measure. The original exploratory factor analysis (EFA) of the SDQ revealed a five-factor structure in which each construct (Conduct Problems, Emotional Symptoms, Attention/Hyperactivity, Peer Problems, and Prosocial Behavior) was represented by five items (R. Goodman, 1997).
The SDQ has been shown to effectively discriminate between clinical and nonclinical populations among both sexes, while also demonstrating strong predictive utility (R. Goodman, 2001; R. Goodman et al., 1998), adequate ratings of internal consistency for exploratory and basic research (Caci et al., 2015; R. Goodman, 2001; A. Goodman et al., 2010), and reliability across different informants (i.e., self-, parent-, and teacher-report; R. Goodman, 2001). Reliability among different informant ratings is critical, as it allows for a more complete perception of youth’s behavioral and emotional functioning (Achenbach et al., 1987), particularly because symptoms of psychopathology may manifest exclusively in certain contexts and not others (e.g., at school but not at home).
SDQ Factor Structure
Multiple examinations of the teacher-report SDQ employing EFAs demonstrate support for the original five-factor structure (e.g., Becker et al., 2004; R. Goodman, 2001). However, given that EFAs are investigative in nature, researchers have stressed the importance of utilizing confirmatory factor analyses (CFAs) to determine the best-fitting factor structure (Caci et al., 2015). Although one study (Hill & Hughes, 2007) found a marginally adequate fit for the five-factor structure for teacher- and parent-report in a U.S. sample, another study (Dickey & Blumberg, 2004) in the United States showed more support for a three-factor structure consisting of internalizing problems (a combination of the Emotional Symptoms and Peer Problems subscales), externalizing problems (a combination of the Conduct Problems and Attention/Hyperactivity subscales), and prosocial behaviors (Prosocial subscale). However, this study (Dickey & Blumberg, 2004) only utilized parent-report SDQ data for their CFA.
Given this inconsistency in U.S. samples and other studies showing support for factor structures other than the original five-factor structure in European countries (e.g., Mellor & Stokes, 2007; Van Leeuwen et al., 2006), A. Goodman and colleagues (2010) conducted a CFA to test factor structure and construct validity for all versions of the SDQ (i.e., teacher-, parent-, and self-report) using a large British sample. Findings from this study revealed that, although the original five-factor structure of the SDQ was the most parsimonious model, the five factors demonstrated poorer cross-informant discriminant validity than did the two broader factors of externalizing and internalizing problems. However, all five subscales on the parent and teacher SDQs did show good convergent and discriminant validity when predicting clinical disorder. These findings led A. Goodman et al. (2010) to suggest that different utilizations of the SDQ may be warranted for different purposes. They suggested that the original five subscales should be utilized when the SDQ is used to screen for disorder or when studying high-risk children; however, if the SDQ is used as a predictor or criterion variable in low-risk epidemiological research, they suggested using the broader internalizing and externalizing subscales that were identified in the Dickey and Blumberg (2004) study. More recent examinations of the SDQ employing CFA have continued to yield mixed results (Caci et al., 2015; Español-Martín et al., 2021; Kóbor et al., 2013; Smid et al., 2020; Stone et al., 2015; Tobia et al., 2013.
Specifically, Tobia and colleagues (2013) tested the factor structure of the SDQ by employing a CFA using a large Italian sample (n = 1,000); results indicated that the three-factor structure identified by both A. Goodman and colleagues (2010) and Dickey and Blumberg (2004) was a superior fit to the original five-factor structure. Conversely, in a sample of Dutch children ages 4 to 7, Stone and colleagues (2015) found strong support for the original five-factor model of the SDQ teacher version (n = 2,238). Likewise, more recently, in a sample of Spanish children and adolescents, aged 5 to 17, the original five factor model had the best model fit for the teacher SDQ (n = 6,775; Español-Martín et al., 2021). Another study (Caci et al., 2015) used a French sample (n = 889) and identified the best-fitting model as a bifactor model that grouped the five a priori factors into two general psychological symptoms factors (internalizing and externalizing problems) in addition to one strengths/prosocial factor that correlated negatively with the two general psychological symptom factors.
Utility of Bifactor Models
Although originally described over 80 years ago (Holzinger & Swineford, 1937), bifactor modeling has recently appeared again in the psychometric literature (Reise, 2012), especially among constructs that are known to be multidimensional (Reise et al., 2013). Bifactor models are advantageous in that they include a general or primary factor measured by the instrument, as well as specific or secondary factors, usually conceptualized as subscales (Rodriguez et al., 2016a). As with all measurement models, bifactor models are estimated based on the variance of the scale items. The general factor is a function of the common variance among all scale items; secondary factors are based on additional common variance among clusters of items, typically with highly similar content (Rodriguez et al., 2016b). In addition, bifactor modeling assumes that the primary and secondary factors all are orthogonal (Holzinger & Swineford, 1937); thus, when bifactor models are supported, clinicians and researchers can use the overall primary factor as well as secondary factors in predictive models. Likewise, by explaining additional variance of each of the individual items used in the subscales, bifactor models can yield more accurate estimates of the traits being examined (Rodriguez et al., 2016b).
Although bifactor models are being used more frequently to model various psychological constructs (Hammer & Toland, 2017; McDermott et al., 2017; Mészáros et al., 2014; Thomas, 2012), only two studies (Caci et al., 2015; Kóbor et al., 2013) have examined bifactor models of the SDQ, and these studies either utilized only a subset of the SDQ items or examined the model in a small international sample. Moreover, the bifactor models estimated in these studies do not conform to the traditional bifactor model requirements. Specifically, bifactor models are supposed to include orthogonal general and secondary factors (Holzinger & Swineford, 1937), though the bifactor models presented by both Caci and colleagues (2015) and Kóbor and colleagues (2013) contained oblique factors. Ultimately, researchers have indicated that there is a need for further validation of the SDQ factor structure to substantiate the clinical validity of the SDQ in screening for clinical disorders (Caci et al., 2015; A. Goodman et al., 2010; Stone et al., 2010; Tobia et al., 2013) and to inform how the SDQ is best utilized in predictive modeling and other applied research. Thus, further examining a bifactor model is an important next step in SDQ validation studies.
Because the SDQ can produce a total score using all items, or individual subscales, it is important to understand if the subscales proposed by R. Goodman and colleagues (1998) and A. Goodman and colleagues (2010) are indeed the latent constructs that should be utilized. Thus, to better understand the factor structure of the teacher SDQ, as well as further investigate the utility and validity of the teacher SDQ in screening for childhood clinical disorders and in applied research, this study employed three separate CFAs, including a bifactor model, in a large sample of youth enrolled in Grades K–12 in the southeastern United States (n = 7,199). After determining the best-fitting model, we also examined it for measurement invariance across boys and girls.
Method
Participants
The sample for this study, 7,199 children in regular and special education classrooms in Grades K–12 in a rural school district in the Southeastern United States, is from the Project to Learn about Youth – Mental Health I (PLAY – MH I), a large epidemiological study funded by the U.S. Centers for Disease Control and Prevention (CDC). The distribution of children between elementary school (i.e., Grades K–5) and secondary school (i.e., Grades 6–12) was nearly equal (51% and 49%, respectively). Using teacher-reported data, among the elementary school children, 52% were male, 65% were White, 27% were African American, 5% were Hispanic/Latino, and the remainder (~3%) were from other racial/ethnic groups. Three percent of participants were indicated as English language learners and 53% were on free or reduced lunch.
Among the secondary school participants, 52% were male, 63% were White, 27% were African American, 5% were Hispanic/Latino, and the remainder (~5%) were from other racial/ethnic groups. Two percent of participants were indicated as English language learners and 62% were on free or reduced lunch. Overall, the sample was representative of the geographic area in terms of racial and ethnic diversity and socioeconomic status.
Procedure
Data were drawn from the first stage of a population-based study of prevalence of emotional and behavioral health disorders in children and adolescents. Figure 1 presents the data collection process. All 10,567 children in regular and special education in the school district were eligible for participation, which involved one teacher for each child completing a brief online screener for emotional and behavioral health concerns. Parents were informed of the procedures through two informational letters sent to the home, one via U.S. Mail and the other sent home from school with the child. Each letter contained a prepaid postcard that parents could return if they did not want their child’s teacher to complete a screener on their child. Parents were also informed of the study procedures through a “robocall” from the school district, press releases to local media outlets, and a detailed study website. Ten percent of parents returned the postcard, declining participation for their child.

Flowchart of the Data Collection Process.
For those children whose parents did not return an opt-out postcard, one teacher was contacted to complete the online screener. For elementary students, this was the main classroom teacher, and for secondary school students, the first period or first block teacher was contacted. Teachers received a detailed letter describing the procedures and more detailed information was also provided on the study website, including the screener items. Teachers also received a list of students in their classroom whose parents had not opted out. Shortly after receiving this letter, teachers were sent an email with further instructions and a weblink to the secure online screener on Qualtrics. Teachers used a unique identification number assigned to each child when completing the survey, rather than child name or other identifying information, to ensure confidentiality/anonymity of responses. The screener took approximately 5 minutes to complete per child and teachers were compensated US$4 per survey completed. The Institutional Review Board approved these study procedures.
Teachers started a screener for 7,207 students. The total number of students with completed screeners was 7,199 students; this was 76% of eligible students (i.e., students whose parents did not opt out) and represented 69% of all students enrolled in the district during the 2014–2015 academic year. Most teachers (85%) reported knowing students for 2 months or less before the screener was completed in October to December 2014.
Measures
Demographics
Questions at the beginning of the screener asked the teacher to indicate child school, grade, English as a Second Language (ESL) status, and length of time the teacher had known the student. Each of these items had several response options to choose from, and teachers could not move on to the next item until they responded. Response options for the length of time the teacher had known the student included less than 1 month, 1–2 months, 3–5 months, 6–11 months, and 12 months or more. Free/reduced lunch status, biological sex, race, and ethnicity data were obtained directly from the school district.
SDQ
The SDQ (R. Goodman, 1997) is a 25-item brief behavioral screening questionnaire with parent, teacher, and self-report versions available. The teacher-report version was included in the screener for the current study. The teacher SDQ for children ages 4 to 10 was used for participants in elementary school, whereas the teacher SDQ for children ages 11 to 17 was utilized for participants in secondary school. Each child in the sample had only one teacher complete the screener. Teachers were asked to rate the child’s behavior over the last 6 months or since the beginning of the school year using a 3-point Likert-type scale ranging from “not true” to “certainly true.” The SDQ was developed with five scales, each with five items: Emotional Symptoms, Conduct Problems, Hyperactivity/Inattention, Peer Relationship Problems, and Prosocial Behavior. More information about the SDQ can be found at sdqinfo.org.
Statistical Analyses
Using Mplus Version 8, we examined three different measurement models. The first CFA (Model 1) tested the original five-factor model whereby each subscale is represented by its own factor and all factors are correlated with one another (R. Goodman, 1997). The second CFA (Model 2) tested a model including three factors based on the proposed factor structure generated from a large British sample (A. Goodman et al., 2010), which was further validated by Tobia and colleagues (2013). The third CFA (Model 3) tested a bifactor model based on findings from Kóbor and colleagues (2015). Figure 2 shows the schematic illustration of the three models examined in this study.

Schematic Illustration of Three Alternative Measurement Models of the Strengths and Difficulties Questionnaire for Teachers. Model 1: Original Five-Factor First-Order Model; Model 2: Three-Factor First-Order Model; and Model 3: Five-Factor Bifactor Model.
We estimated each of the three models separately by school level (i.e., elementary school vs. secondary school) because of inherent differences in students’ amount of time spent in one classroom and the nature of the teacher-student relationship. Specifically, elementary school teachers are likely more reliable reporters of their students’ behavior, as compared with the first period teacher of a secondary school student, because they spend most of their day in a single classroom with a single group of students. We used mean- and variance-adjusted diagonally weighted least squares (WLSMV) estimation given the categorical responses for each SDQ item and all models used Mplus’ default methods for handling item-level missingness. We examined model fit using multiple fit indices according to usual practice: comparative fit index (CFI), Tucker–Lewis index (TLI), and the root mean squared error of approximation (RMSEA). Model fit is deemed acceptable when CFI and TLI are close to .95 and RMSEA is at least .08, with smaller RMSEA values representing stronger fit (e.g., RMSEA values closer to .05 are representative of stronger fit than RMSEA values of .08). We also examined closeness of model fit (CFIT of RMSEA) which is a statistical test of the null hypothesis that RMSEA = .05. After identifying the best-fitting model for both samples (i.e., elementary school and secondary school), we examined factor reliabilities and dimensionality. Finally, we examined the best-fitting model for both samples for measurement invariance across biological sexes.
Results
Results for the three measurement models are in Tables 1 through 4. As shown in Table 1, the three-factor model (Model 2) had the poorest model fit for both elementary and secondary school. The original five-factor model (Model 1) and the bifactor model (Model 3) both demonstrated acceptable fit for both school levels, with the bifactor model having slightly better fit than the original five-factor model. Factor loadings for each of the three models that we examined, by school level, are presented in Table 2 (original five-factor model), Table 3 (three-factor model), and Table 4 (bifactor model). Because the model fit between the original five-factor model and the bifactor model were close, as measured by CFI, TLI, and RMSEA, and because several items did not load on their main factors in the bifactor model, we examined factor reliabilities and dimensionality for both the bifactor model and the original five-factor model (Table 5). Likewise, given the poor model fit for the three-factor model, although the detailed model results are provided on Table 3, we focus our written results on the original five-factor model and the bifactor model.
Model Fit Indices for Three Competing Measurement Models of the Strengths and Difficulties Questionnaire for Teachers.
Note. Elementary = Grades K–5; Secondary = Grades 6–12; CFIT of RMSEA = statistical test of null hypothesis that RMSEA = .05. CFI = comparative fit index; TLI = Tucker–Lewis index; RMSEA = root mean squared error of approximation; CI = confidence interval.
Confirmatory Factor Analysis Standardized Factor Loadings of the Strengths and Difficulties Questionnaire for Teachers Original Five-Factor Model by School Level (n = 3,509 for Elementary School and n = 3,690 for Secondary School).
Note. Elem = elementary school, Grades K–5. Sec = secondary school, Grades 6–12.
Confirmatory Factor Analysis Standardized Factor Loadings of the Strengths and Difficulties Questionnaire for Teachers Three-Factor Model by School Level (n = 3,509 for Elementary School and n = 3,690 for Secondary School).
Note. Elem = elementary school, Grades K–5. Sec = secondary school, Grades 6–12.
Confirmatory Factor Analysis Standardized Factor Loadings of the Strengths and Difficulties Questionnaire for Teachers Bifactor Model by School Level (n = 3,509 for Elementary School and n = 3,690 for Secondary School).
Note. Elem = elementary school, Grades K–5. Sec = secondary school, Grades 6–12.
Score Reliabilities for the Strengths and Difficulties Questionnaire for Teachers Bifactor and Five-Factor Models by School Level.
Note. Elementary School = Grades K–5; Secondary School = Grades 6–12.
Reliability data from bifactor model. b Reliability data from five-factor model.
As shown in Table 4, none of the Prosocial Behavior items loaded on the General Problems factor in the bifactor model. In addition, among secondary school youth, five additional items did not load on the General Problems factor: somatic symptoms, worries, many fears, picked on/bullied, and better with adults. The item, better with adults, also did not load on the General Problems factor for the elementary school model. Thus, a total of ten (secondary school) and six (elementary school) items, respectively, did not have adequate factor loadings on the General Problems factor. These results, in conjunction with the low explained common variance (ECV = .61 for both school levels), suggest that the teacher-reported SDQ does not represent a unidimensional construct (Quinn, 2014). Moreover, the low reliability scores for most of the subscales in the bifactor model (omega hierarchical values from 0.15 to 0.40 for elementary school and 0.15 to 0.31 for secondary school) also suggest that, although the measurement model fit is slightly better for the bifactor model, conceptually, it is not a superior model compared with the original five-factor model in either elementary or secondary school students.
Results from the original five-factor model (Table 2) are more theoretically aligned and more clinically relevant. Specifically, most items had strong factor loadings on their original factors (i.e., factor loadings > .70). Only one item, better with adults, exhibited weak factor loadings in both the elementary and secondary school models. From a research perspective, almost all of the subscales yielded acceptable levels of reliability (i.e., McDonald’s ω >.80); the Peer Relationship Problems [ω = .68 (elementary students) and .74 (secondary students)] and Conduct Problems [ω = .77 (elementary students) and .79 (secondary students)] subscales were on the lower end of acceptable levels of reliability.
After identifying the original five-factor model as the best-fitting model, we examined the model for measurement invariance (MI) between girls and boys, for each of the samples (elementary school and secondary school). When examining MI, there are three types of invariances that are examined: configural, which assumes that the number of factors and pattern of loadings are the same for each group, metric, which assumes that the magnitude of loadings are the same for each, and scalar, which assumes that configural and metric invariance exist and that thresholds are also equal across groups. We first attempted to examine configural invariance using the exact five-factor model used in the CFA process—this approach generated a nonpositive definite latent variable covariance matrix for both girls and boys, in both the elementary school and secondary school samples. To simplify the model, we decided to exclude the correlations among the latent factors and reexamined the data for configural invariance, with item residual variances all fixed to 1 and factor mean equal to 0 and variance equal to 1, for identification. We once again ran into estimation troubles—model nonconvergence. The item “unhappy” which is from the SDQ item “often unhappy, depressed, or tearful” was the cause of the models not converging, for both the elementary and secondary school samples. We reexamined configural invariance again, after removing the item “unhappy” from the models for both girls and boys. The models converged; however, configural invariance between girls and boys was not supported for elementary school or secondary school.
For the elementary school sample, model fit data include RMSEA = .230, CFI = .203, TLI = .214, and SRMR = .352. Likewise, for the secondary school sample there was also poor model fit: RMSEA = .234, CFI = .175, TLI = .187, and SRMR = .343. The latent factors that exhibited the greatest differences between girls and boys differed for the elementary school students and the secondary school students. As shown in Table 6, among elementary school students, the items in the Hyperactive/Inattentive factor perform differently for girls and boys. Yet, in the secondary school sample, the greatest differences are found with Emotional Symptoms (see Table 7 for more details). The one item that had drastically different factor loadings between girls and boys in both of our samples was the item labeled “considerate” which is the SDQ item “considerate of other people’s feelings.” Because configural invariance must be met before testing for metric and scalar invariance, the lack of configural invariance indicates that although the original five-factor model was the best-fitting model compared with the three-factor model and a bifactor model, it is not invariant across girls and boys.
Standardized Factor Loadings for Configural Measurement Invariance Test Between Girls and Boys on the Strengths and Difficulties Questionnaire for Teachers Original Five-Factor Model for Elementary School Students (n = 3,509).
Had to remove item for the models to estimate. Secondary school includes Grades 6–12. Invariance testing was conducted without correlations between latent factors; when testing was conducted on models that included the correlations between latent factors, the latent variable covariance matrix, psi, was not positive definite, which prevented complete model estimation.
Standardized Factor Loadings for Configural Measurement Invariance Test Between Girls and Boys on the Strengths and Difficulties Questionnaire for Teachers Original Five-Factor Model for Secondary School Students (n = 3,690).
Had to remove item for the models to estimate. Elementary school includes Grades K–5. Invariance testing was conducted without correlations between latent factors; when testing was conducted on models that included the correlations between latent factors, the latent variable covariance matrix, psi, was not positive definite, which prevented complete model estimation.
Discussion
This study is the first to explore the global and specific factor structure of the SDQ using a large U.S.-based sample of children and adolescents across Grades K–12. The results of our study do not support the three-factor SDQ model reported by Dickey and Blumberg (2004) and A. Goodman and colleagues (2010). Instead, the bifactor model yielded the best fit to our data in both elementary and secondary school students. However, most of the subscales included in the bifactor models for both elementary and secondary school students demonstrated very low score reliabilities (.40 and below). Therefore, whereas the bifactor model demonstrated slightly better model fit than the original five-factor model, results from our study do not support using the subscales given these low score reliabilities. We thus accepted the original five-factor model as the best-fitting model.
Despite the overall acceptable fit to our data, the subscales in the five-factor SDQ model had varying levels of score reliabilities, with the Conduct Problems and Peer Relationship Problems subscales for both elementary and secondary school youth exhibiting lower reliabilities than the preferred level of .80 or higher for basic research (Lance et al., 2006). This pattern was also observed in van den Heuvel et al.’s (2017) study. Overall, the number of subscales with lower reliability in our study is more than those reported by Caci and colleagues (2015) and less than those reported by A. Goodman and colleagues (2010). Specifically, in the Caci et al. (2015) study, Peer Relationship Problems (ω = .76) was the only subscale with reliability less than .80. In the A. Goodman et al. (2010) study, three of the five subscales had reliability levels less than .80: Emotional Symptoms = .78, Peer Relationship Problems = .69, and Conduct Problems = .75.
Reliability levels for the Prosocial Behavior and Hyperactivity/Inattention subscales were not only greater than .80 but they were also close in value. Thus, the 10 items in these two subscales seem to function better than the 15 items in the other three subscales. One possible explanation could be that the items for these two subscales are easily observed by teachers (e.g., restless, fidgety, easily distracted, volunteers, shares, and kind to younger children). Items used to measure Peer Relationship Problems, Conduct Problems, and Emotional Symptoms might be harder for teachers to assess, either because of not knowing the students very well or for very long or because the behaviors might not manifest in the classroom.
For example, in our study, over 75% of the teachers reported knowing students for one to two months before the screener was completed. This is an inherent limitation of our study, and the possibility exists that these teachers did not function as reliable informants for these domains of functioning. Yet, given the lack of information provided in other published research that has also examined the structure of the teacher-report SDQ (i.e., only one study in addition to ours reports the length of time teachers knew the students before completely their ratings), it is difficult to determine the impact that this factor might have on the reliability of the teacher-report data in our study. Moreover, van den Heuvel et al. (2017), the only study to report length of time teachers knew students, also report low reliability for conduct problems and peer problems and in their study, teachers reported knowing children for at least 2 months when they completed the SDQ and results. Future research should be conducted to determine if the length of time that teachers know the students impacts the reliability of teacher-report SDQ data.
Finally, although the reliability scores for some of the teacher SDQ subscales are sufficient for use in basic research endeavors, these results call into question the clinical utility of the teacher-report SDQ. Our findings suggest that this measure does not demonstrate reliability scores sufficient for making diagnostic or treatment decisions in clinical settings in the absence of other screening measures for adolescents (Lance et al., 2006). Moreover, the continued inconsistency with the factor structure of the teacher-report SDQ across samples of various sizes and from multiple countries suggest that when the SDQ is used in predictive models, researchers need to first estimate a CFA to determine how to best operationalize the latent constructs.
The findings regarding measurement invariance between girls and boys are another area that researchers should take into consideration when determining how to operationalize these latent constructs. For example, although no other U.S.-based study has examined measurement invariance with the teacher SDQ, results from other studies conducting with Dutch children (Stone et al., 2015) and Spanish children (Español-Martín et al., 2021) have revealed mixed findings. Similar to our results, Stone and colleagues (2015) did not find measurement invariance by student sex; however, the data did support measurement invariance based on student age and race/ethnicity. Español-Martín et al. (2021) also reported measurement invariance by student sex and age. Given these mixed findings, until more research is conducted in this area, when using the teacher-report SDQ in predictive models where student sex is part of the research focus, researchers should run CFAs, by student sex, to determine how to operationalize the latent constructs.
Another limitation of our study is that we only collected data using the teacher-report SDQ; thus, we are unable to compare these results to parent-report or child-report SDQ factor structures. Because of this, it is unclear if the inconsistency in model fit observed in our study is specific to the teacher-report SDQ, or if perhaps there are larger factor structure issues with the items in the SDQ. Future research should aim to extend our study by examining the factor structure of all three SDQ versions from a single sample. Doing so would help determine how items function similarly or differently based on respondent—teacher, parent, or child. Given the age of the SDQ, results from this type of research can help determine if perhaps the items in the SDQ might need to be revised.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the U.S. Centers for Disease Control and Prevention (grant number 5U01DD001007).
