Abstract
This study used the correlated trait–correlated method minus one model to examine the convergent and discriminant validity of the scales of the Strengths and Difficulties Questionnaire (SDQ). The SDQ scales are emotional symptoms (ES), conduct problems (CP), hyperactivity (HY), peer problems (PP), and prosocial behaviors (PS). A total of 202 adolescents provided self-ratings and were also rated by their mothers and teachers. The findings indicated support for convergent validity for all five SDQ scales for all three respondents. Generally there was more convergence between mother–adolescent ratings than mother–teacher and adolescent–teacher ratings, especially for ES and PP. There was support for the discriminant validity between the traits in all scales, except between CP and HY. The findings are discussed in relation to the construct validity and clinical use of the SDQ.
Keywords
The Strengths and Difficulties Questionnaire (SDQ) is a rating scale for screening emotional and behavioral problems of children and adolescents, aged 4 to 16 years (R. Goodman, 1997; R. Goodman, Meltzer, & Bailey, 1998). Identical versions exist for parent and teacher completion. There is also a near identical version for self-completion by adolescents between 11 and 16 years of age (R. Goodman, Meltzer, & Bailey, 2003). The present study used a multiple-traits by multiple-methods (MTMM) approach called the correlated trait–correlated method minus one (CT-C(M − 1; Eid, 2000; Eid, Lischetzke, Nussbeck, & Trierweiler, 2003) to examine the convergent (the extent to which the same trait is measured by different methods) and discriminant (the extent to which different traits in a measure are distinct) validity of the SDQ, based on ratings from mothers, adolescents, and teachers.
The SDQ contains five scales, with five items in each scale: emotional symptoms (ES), conduct problems (CP), hyperactivity-inattention (HY), peer problems (PP), and prosocial behavior (PS; R. Goodman, 1997; R. Goodman et al., 1998). The proposed factor structure of the SDQ is an oblique five-factor model, with the ES, CP, HY, PP, and PS items loading on their own respective factors. The selection of the SDQ items and their organization into the different scales concur with current nosological models. The SDQ has been offered as a useful screening instrument for clinical research and practice, including epidemiological and cross-cultural research (R. Goodman, 2001). In light of the roles advocated for the SDQ, it is critical that we have a good understanding of its construct validity, such as its convergent and discriminant validity.
A number of studies involving parent, teacher, and adolescent ratings of the SDQ have reported adequate support for the proposed five-factor oblique model (e.g., Sanne, Torsheim, Heiervang, & Stormark, 2009; Van Leeuwen, van Meerschaert, Bosmans, Medts, & Braet, 2006; Van Roy, Veenstra, & Clench-Aas, 2006). However, there are also studies that have either found only marginal support (e.g., Hill & Hughes, 2007; Rønning, Handegaard, Sourander, & Mørch, 2004) or no support (e.g., Dickey & Blumberg, 2004; Mellor & Stokes, 2007) for this factor structure. Studies have generally shown moderate to high correlations between the SDQ scales (Sanne et al., 2009; Van Leeuwen et al., 2006; Van Roy et al., 2006), thereby indicating some, albeit mixed, support for their discriminant validity. Existing data also show significant, but low to moderate levels of cross-informant agreement (correlations of around .25 to .35) for the same SDQ scales across parent, teacher, and adolescent ratings (Becker, Hagenberg, Roessner, Woerner, & Rothenberger, 2004; R. Goodman, 2001; R. Goodman et al., 2003; Van Leeuwen et al., 2006; Sanne et al., 2009). These findings can be interpreted as providing some, albeit modest, support for the convergent validity of the SDQ scales. Studies also show more agreement and therefore convergence for parent–adolescent and parent–teacher ratings than teacher–adolescent ratings (R. Goodman, 2001; R. Goodman et al., 2003), with relatively more convergence for CP and HY than ES (Beaker et al., 2004; R. Goodman, 2001; Van Leeuwen et al., 2006; Sanne et al., 2009).
It is worth noting that the low to moderate cross-informant agreement is consistent with findings involving other child and adolescent measures (Choudhury, Pimentel, & Kendall, 2003; Grills & Ollendick, 2003). The meta-analysis by Achenbach, McConaughy, and Howell (1987), involving 119 studies, also found low to moderate agreement between different informants, including parents, children, and teachers for ratings of social, emotional, and behavioral problems in children and adolescents. It also found more agreement for parent–adolescent and parent–teacher than teacher–adolescent information and more agreement for externalizing behavior problems (inattention, hyperactivity, and conduct problems) than internalizing behavior problems (anxiety and depression).
Relative to simple correlations, such as the ones mentioned earlier, the convergent and discriminant validity of a measure can be more robustly evaluated using MTMM models (Campbell & Fiske, 1959; Lance, Noble, & Scullen, 2002). In general, this approach involves data for two or more traits measured by two or more methods. The original Campbell and Fiske approach evaluates convergent and discriminant validity through inspection of correlation matrices involving the scores obtained through the different methods. For a multidimensional rating scale, evidence of convergent validity of a subscale is inferred when there are relatively high monotrait–heteromethod correlations (same subscale across different methods). Evidence of discriminant validity of the subscales is inferred when there are relatively low hetrotrait–monomethod correlations (different subscales within methods). Method effects are inferred when the correlations for the subscales within a method are larger than correlations across methods but within traits. MTMM analysis can also be conducted within a confirmatory factor analysis (CFA) framework (Lance et al., 2002). Two commonly used models are the correlated trait–correlated method (CT-CM) and the correlated trait–correlated uniqueness (CT-CU) models. These models allow evaluation of the convergence of a measure after taking into account the method and error variance in it. In the CT-CM model, the same subscales from the different methods load on their own trait factor. Also, all the subscales from the same method load on their own method factor, with the different trait factors correlating with each other and the different method factors correlating with each other. Trait and method factors are not correlated. In the CT-CU model, the trait factors are similar to that for CT-CM model. However, the method effect is modeled by within method correlated error variances. For both models, support for convergent validity is inferred if there are relatively high trait variances for the subscales. In the CT-CM model, high method effect is inferred if there are relatively high variances for the different subscales within methods. In the CT-CU model, high method effect is inferred if there are relatively high correlations for the different error variances within methods. In both models, low correlations between the latent trait factors are taken as evidence for their discriminant validity (Lance et al., 2002).
To date, at least three studies have examined the convergent and discriminant validity of the SDQ using MTMM models (A. Goodman, Lamping, & Ploubidis, 2010; Hill & Hughes, 2007; Van Roy et al., 2006). A. Goodman et al. and Hill and Hughes used the original Campbell and Fiske (1959) approach. A. Goodman et al. examined ratings provided by parents, teachers, and adolescents, whereas Hill and Hughes examined parent, teacher, and peer ratings. Both studies found moderate levels of monotrait–heteromethod correlations for the SDQ scales, thereby supporting their convergent validity. However, in both studies there was evidence of moderate method effect. This weakens but does not eliminate the support for the convergent validity of the SDQ scales in these studies. A. Goodman et al. found more convergence between parent–adolescent than parent–teacher and teacher–adolescent and also more convergence between parent–teacher than teacher–adolescent. Hill and Hughes found more convergence for parent–teacher and teacher–peer than parent–peer. Both studies found some support for discriminant validity of the SDQ scales. Van Roy et al. (2008) used the CT-CM model and examined parent and preadolescent self-ratings. The Hill and Hughes (2007) study, cited earlier, also examined their data with the CT-CM model, which however did not converge, and the CT-CU model. Van Roy et al. found support for convergent validity, with the support being higher for adolescent ratings of ES and PP and parent ratings of HY. For the CT-CU model, Hill and Hughes also found support for the convergent validity of the SDQ scales. However, for both studies there was also considerable method effect. Again, this weakens, but does not eliminate, the support for the convergent validity of the SDQ scales in these studies. Both studies found some support for the discriminant validity of the SDQ scales, and Van Roy et al. found support for the discriminant validity of their methods (parent and preadolescent). Overall, therefore, like the correlation studies involving the SDQ, the findings from past MTMM studies do provide some support for the convergent validity of the SDQ. Also, they have shown more convergence between parent–adolescent than teacher–adolescent (and also parent–teacher) and some support for the discriminant validity of the SDQ scales. Unlike the correlation studies that have generally shown more convergence for CP and HY, the MTMM studies have shown more convergence for adolescent self-ratings for ES and PP.
Despite useful existing MTMM convergent and discriminant findings for the SDQ, these data are limited. First, the only study (A. Goodman et al., 2010) that has examined the convergent and discriminant validity of parent, adolescent, and teacher ratings of the SDQ simultaneously used the Campbell and Fiske (1959) approach. As this approach uses observed scores, which comprise error variances, the findings may be confounded, as demonstrated in studies of construct validation of other measures (Green, Goldman, & Salovey, 1993). Second, there are problems in the CT-CM and the CT-CU methods that have been used in past studies in this area (Eid, et al., 2008; Geiser, Eid, & Nussbeck, 2008; Höfling, Schermelleh-Engel, & Moosbrugger, 2009; Nussbeck, Eid, Geiser, Courvoisier, & Lischetzke, 2009). Eid and his associates have suggested that a key issue that has to be considered when selecting a CFA model for MTMM analysis is the types of methods in the model. Methods can be either interchangeable (all respondents have the same access to the target and, therefore, rate the target from the same perspective) or structurally different (all raters have different access to the target, and they would respond from different perspectives). According to these terms, when the SDQ is completed by adolescents, mothers, and teachers, the data constitute structurally different methods. According to measurement experts, the CT-CU is not appropriate for structurally different methods. Although the CT-CM can be used, there are a number of substantive and psychometric problems with this approach (Eid, 2000; Eid et al., 2003). Among them are that if there is a general trait factor across the different traits for a method, then the variance for this trait component will be incorporated as part of the variance of the method factor, thereby confounding results; and its application often leads to inadmissible solutions, as found for the SDQ analysis by Hill and Hughes (2007).
According to Eid and others, the CT-C(M −1) model, which is derived from the CT-CM model, is the preferred model for structurally different methods (Eid et al., 2008; Geiser et al., 2008; Nussbeck et al., 2009). In this model, one of the methods is selected as the reference method (Eid, 2000; Eid et al., 2003). The true scores (consistency coefficients) of the reference method indicators are used to predict the true scores (consistency coefficients) of indicators of the other methods or nonreference methods. The consistency coefficient of the reference method is its true-score variance, whereas the consistency coefficients of the nonreference methods are the amount of variance in them that are predicted by the true score of the reference method. For any scale, the consistency coefficients of the nonreference methods indicate the convergence of the nonreference methods with the reference methods. The proportion of the true score variance in the nonreference methods that are not predicted by the true score of the reference method are their method-specific variance or method-specific coefficients.
The aim of the current study was to examine the convergent and discriminant validity of the SDQ scales for ratings provided by adolescents, mothers, and teachers using the CT-C(M − 1) model. Each trait for each method unit was represented by its total scale score (single indicator). As the method selected as the reference method contributes substantively to the meaning of the trait and method factors (Geiser et al., 2008), we conducted two different CT-C(M − 1) analyses. In the first analysis (MTMM Analysis 1), adolescent self-rating was used as the reference method. In the second analysis (MTMM Analysis 2), mother rating was used as the reference method. Conducting these complementary analyses enabled examination of the convergence of adolescent ratings with mother ratings (in MTMM Analysis 1 or 2), adolescent ratings with teacher ratings (in MTMM Analysis 1), and mother ratings with teacher ratings (in MTMM Analysis 2). Given the wide age-group, the equivalencies of both the CT-C(M − 1) models were also examined. Based on exiting data, the expectation was to find some support for convergent validity, with more convergence between mother–adolescent ratings than mother–teacher ratings. Based on the findings from past MTMM studies, the expectation was more convergence for ES and PP. Some support for discriminant validity of the SDQ traits was expected. Discriminant validity for the different method factors was also expected.
Method
Participants
There were 202 adolescents, of whom 98 were males and 104 were females. Their age ranged from 12 to 17 years. The mean age of all participants together was 14.03 years (SD = 1.39). The mean age of male (M = 14.04, SD = 1.39) and female (M = 14.02, SD = 1.39) adolescents did not differ significantly, t(200) = 0.11, ns. The adolescents were from 14 secondary schools. These schools were selected randomly, and 50% of the randomly selected schools agreed to participate in the study, as described in the Procedure section below.
In terms of background relating to socioeconomic status, occupational status of father’s (or mother’s, when father’s was not available) was coded according to the Australian Standard Classification of Occupations (Australian Bureau of Statistics, 1997). This has nine major hierarchically organized occupational categories defined in terms of occupational levels related to skills and specializations. In decreasing order, they are managers and administrators (coded 1); professionals (coded 2); associate professionals (coded 3); tradespersons (coded 4); advanced clerical and service workers (coded 5); intermediate clerical, sales, and service workers (coded 6); intermediate production and transport workers (coded 7); elementary clerical, sales, and service workers (coded 8); and laborers (coded 9). Those not employed were also coded 9 in this study. The overall mean occupational level was 4.58 (SD = 3.17), which is about “middle-class.”
For the adolescents involved in the study, approximately 96% were of European background, 3% Asian, and 1% others (including indigenous Australian). These figures compare to around 90% European, 7% Asian, and 3% others (including indigenous Australian) in the general Australian population (Australian Bureau of Statistics, 2007). There was close match in ethnicity between the Australian general population and the group involved in the study, χ2(df = 2) = 2.79, p = ns.
Measures
The study used the adolescent self-rating, parent, and teacher versions of the SDQ (R. Goodman, 1997). All three versions have 25 items, with five scales (EP, CP, HY, PP, and PS). Each scale has five items, and each item is rated by the informant as either “not true” (scored 0) or “somewhat true” (scored 1) or “certainly true” (scored 2). The internal consistency (Cronbach’s alpha) values for adolescent ratings for the ES, CP, HY, PP, and PS scales in this study were .67, .68, .70, .60, and .67, respectively. These values for mother ratings were .74, .71, .85, .57, and .85, respectively. For teacher ratings, they were .81, .69, .85, .59, and .86, respectively.
Procedure
Stratified random sampling was used to select schools to approach for participation in the study. The population was divided into nine groups, corresponding to the nine regions of the State of Victoria, Australia. A total of 28 secondary schools from the nine regions were contacted. Within each region, a random number table was used to determine the schools to be contacted. Of the schools contacted, 14 consented to participate.
Following consent from directors of education and school principals, classroom teachers were issued the appropriate numbers of large, sealed envelopes to be forwarded to mothers, through their students. Each envelope contained two sets of documents plus questionnaires and a return envelope. One set was for the mother and the other set was for the adolescent targeted for participation in the study. Each set comprised a letter describing the study, the parent version of the SDQ (for mothers) and the self-rating version of the SDQ (for adolescents), and a number of other questionnaires. It also included a consent form for permission to have the adolescent’s teacher complete the same set of questionnaires. The letters describing the study stressed the importance of completing one’s own questionnaires independently and mentioned that this was essential for the success of the study. To minimize bias in ratings, the letters to mothers and adolescents indicated that the study was about adolescent behaviors and that the questionnaires were not identified by name. Only the SDQ ratings are of interest in this study.
In all, about 56%, or a total of 342, of the questionnaires distributed to mothers and adolescents were returned with completed scores for the SDQ and consent for teachers to also complete the questionnaires. Of these, teachers rated about 59% of the adolescents, resulting in 202 ratings with complete (usable) sets of adolescent, mother, and teacher ratings. Only these 202 ratings were included in the analyses. Mother, adolescent, and teacher ratings were obtained within 1 month of each other, and all ratings were obtained between April and September 2002. Teachers had been interacting with the adolescents for at least a minimum of 3 months prior to the ratings.
For the 140 adolescents who were not rated by their teachers, there were 67 males and 73 females, and their overall mean age was 13.90 years (SD = 1.24). Their fathers’ occupation levels had a mean score of 4.80 (SD = 3.08), and approximately 95% were of European background, 4% Asian, and 1% others (including indigenous Australian). There was no difference between these 140 adolescents and the 202 adolescents who were included in this study for sex distribution, χ2(df = 1) = 0.01, p = ns; age, t(340) = 0.88, ns; father occupation level, t(340) = 0.63, ns; and ethnicity, χ2(df = 2) = 0.15, ns.
Statistical Procedures
Figure 1 shows a conceptual representation of the CT-C(M − 1) model tested in this study. In line with model specifications, all the indicators (ES, CP, HY, PP, and PS) of the reference method (adolescent in MTMM Analysis 1 and mother in MTMM Analysis 2) were linked to their appropriate traits factors and not to any method factor. Also, these indicators for the nonreference methods (mother and teacher in MTMM Analysis 1 and adolescent and teacher in MTMM Analysis 2) were linked to the appropriate traits factors and to their method factors. The trait factors correlated with each other, and the method factors correlated with each other. This CT-C(M − 1) models estimated 66 parameters. This means that with a sample size of 202, there were about three participants for every parameter estimated (ratio of 3:1).

Conceptual path diagram of the correlated trait–correlated method minus one model, used in the study.
In a CT-C(M − 1) model, the percentages of consistency coefficients and method-specific coefficients for the indicators can be examined for both observed and true scores (see Eid et al., 2003, Appendix A, for appropriate formulas for computation of these). The square roots of the consistency coefficients or latent correlations of the indicators reflect the latent correlations between the true scores of reference and nonreference methods, with higher values indicating more convergence. Convergent validity of an indicator is inferred if it has large latent correlation and significant consistency coefficient. If in this instance the indicator has a larger consistency coefficient than method-specificity coefficient, then it means that there is good support for its convergent validity. If the opposite is the case, then the support for convergent validity is diminished but not eliminated. The degree of correlations of the trait factors indicate the discriminant validity of the traits (as reflected in the reference method), whereas the degree of correlations between the method factors indicate the discriminant validity of the methods. In both cases, low values support their discriminant validity. The strength of all correlations (including latent correlations) was interpreted using the guideline proposed by Hemphill (2003) for correlation effect sizes: <.2 = small; .2 to .3 = medium or moderate; and >.30 = large. The guideline suggested by Brown (2006) for CFA MTMM was used to assess discriminant validity between factors (i.e., correlation <.85 being supportive of discriminant validity).
Invariance of both the CT-C(M − 1) models across age was tested using the multiple indicators multiple causes approach (MIMIC; Joreskog & Goldberger, 1975). Based on the mean age of 14.03 years, participants were allocated into two groups: younger adolescents (from 11 to 13 years, n = 93) and older adolescents (from 14 to 17 years, n = 109). A stepwise forward model building procedure was used to identify noninvariance indicators (Muthén, 1988; see also Brown, 2006). This involved an initial model in which all the trait and method factors were regressed on age-group (the predictor), with all paths between age-group and the indicators constrained to zero. Using the modification indices (MI) values of these paths, the path with the highest MI was freely estimated in a subsequent model. This process was repeated until all significant path coefficients between age-group and the SDQ indicators were identified. These indicators were considered to have noninvariance across the age-groups.
The CFA model in the study was analyzed using Mplus (Version 6) software (Muthén & Muthén, 2010). Robust maximum likelihood was used for estimation. The robust scaled chi-square statistic (called Satorra–Bentler or S-Bχ2), the comparative fit index (CFI), the root mean squared error of approximation (RMSEA), and the standardized root mean square residual (SRMR) were used to ascertain the model fit. The guidelines suggested by Hu and Bentler (1998) are that CFI values close to .95 or more and RMSEA values close to .06 or less be taken as good fit. For the RMSEA, values between .07 and .08 are considered acceptable fit. For the SRMR, values close to .08 or lower are taken as good fit.
Results
Missing Data
Out of a total of 3,030 scores (5 scales × 3 methods × 202 participants), there were 18 scores missing (i.e., around 0.6%). Maximum likelihood (direct ML) was used to handle missing data (Brown, 2006).
MTMM Analysis 1 (Adolescentas the Reference Method)
The fit indices for the CT-C(M − 1) model with adolescent as the reference method (MTMM Analysis 1) were S-Bχ2(df = 69) = 121.46, p < .001, CFI = .95, RMSEA = .06, and SRMR = .058. Thus, this model showed good fit.
Table 1 provides the variance components and latent correlations of the ES, CP, HY, PP, and PS indicators for MTMM Analysis 1. It shows that for all scales completed by mothers and teachers there were large latent correlations. Also, the values for all scales completed by mother were much higher that the corresponding scales rated by teacher. For all three respondents, all consistency coefficients were significant. Taken together, these findings indicate support for their convergent validity, with more support for adolescent–mother ratings than adolescent–teacher ratings. For both mother and adolescent ratings, ES, HY, PP, and PS had more consistency coefficients than method specificity coefficients, whereas CP had more method specificity coefficient than consistency coefficient. These findings indicate good support for the convergence of adolescent–mother and adolescent–teacher ratings for the ES, HY, PP, and PS scales compared with the CP scale. For both mother and adolescent ratings, there was much stronger support for convergence for ES and PP.
Variance Components in the CT-C(M − 1) Model With Adolescent as the Reference Method.
p < .05. **p < .01. ***p < .001.
Table 2 shows the correlations between the trait factors and between the method factors. As shown, all the trait correlations were large. However, with the exception of the correlations between CP and HY, all correlations were less than .85. These findings can be taken as providing support for the discriminant validity between all trait factors, except between CP and HY. There was medium significant correlation (.24 or <.85) between mother and teacher method factors, thereby supporting their discriminant validity.
Correlations of the Trait and Method Factors in the CT-C(M − 1) Model With Adolescent as the Reference Method.
Note. Values in parentheses are the amount of shared variances between variables.
p < .05. **p < .001.
MTMM Analysis 2 (Motheras the Reference Method)
The fit indices for the CT-C(M − 1) model with mother as the reference method (MTMM Analysis 2) were S-Bχ2(df = 69) = 128.01, p < .001, CFI = .94, RMSEA = .064, and SRMR = .060. Thus, for this model, the SRMR showed good fit, whereas both the CFI and RSMEA indicate values that were very close to good fit. Examination of localized areas of strain indicated that the standardized residual covariance ranged from −.76 to .70, thereby suggesting no large residual covariance. There was also no large MI involving correlated residuals, with the highest value being 10.40 (standardized expected parameter = 0.23). Given this, the initial model was retained for the CT-C(M − 1) analyses.
Table 3 provides the variance components and latent correlations of the ES, CP, HY, PP, and PS indicators. It shows that for all scales completed by adolescent and teacher there were large latent correlations. Also, the values for all scales completed by adolescent were much higher than the corresponding scales rated by teacher. For all three respondents, all constancy coefficients were significant. Taken together, these findings indicate support for their convergent validity, with more support for mother–adolescent than mother–teacher ratings. For adolescent ratings, ES, HY, PP, and PS had more consistency coefficients than method specificity coefficients, whereas CP had more method specificity coefficient than consistency coefficient. These findings indicate more support for the convergence of mother–adolescent ratings for the ES, HY, PP, and PS scales compared with the CP scale. The consistency coefficients for ES and PP were especially high. For teacher ratings, ES and PP had more consistency coefficients than method specificity coefficients, whereas for CP, HY, and PS there were more method specificity coefficients than consistency coefficients. These findings indicate more support for the convergence of mother–teacher ratings for ES and PP compared with CP, HY, and PS.
Variance Components in the CT-C(M − 1) Model With Mother as the Reference Method.
p < .05. **p < .01. ***p < .001.
Table 4 shows the correlations between the trait factors and between the method factors. As shown, all the trait correlations were large. However, with the exception of the correlations between CP and HY, all other correlations were less than .85. These findings can be taken as providing support for the discriminant validity between all trait factors, except between CP and HY. There was no significant correlation between mother and teacher method factors, thereby supporting their discriminant validity.
Correlations of the Trait and Method Factors in the CT-C(M − 1) Model With Mother as the Reference Method.
Note. Values in parentheses are the amount of shared variances between variables.
p < .001.
Examination of the latent correlations for adolescent with teacher (Table 1) and mother with teacher (Table 3) suggests comparable levels.
MIMIC Analyses of Noninvarianceof the SDQ Scales Across Age
For the model with adolescent as the reference method, the goodness-of-fit values of the initial MIMIC model were S-Bχ2(df = 77) = 133.86, p < .001, CFI = .95, RMSEA = .060, and SRMR = .053. Thus, this model showed goof fit. This model flagged teacher completed ES scale as having the highest MI (MI = 5.04, standardized EPC = −0.15) for the paths between age and the SDQ scales and methods indicators. The revised MIMIC model that included this path indicated that it was not significant, thereby indicating that none of the SDQ scales had noninvariance across the age-groups. Although details are not present, the findings were the same for the MIMIC model that examined the CT-C(M − 1) model with mother as the reference method.
Discussion
The results indicated support for the convergent validity of all the SDQ scales for adolescent–mother ratings (MTMM Analyses 1 and 2) and also adolescent–teacher ratings (MTMM Analysis 1), with more convergence between adolescent–mother ratings (MTMM Analysis 1). The results also indicated support for the convergence validity of mother–teacher ratings, with more convergence between mother–adolescent than mother–teacher ratings (MTMM Analysis 2). There was also little difference in the convergence of mother–teacher and adolescent–teacher. Despite the support for convergent validity of the SDQ scales, the level of agreement was moderate, especially for mother–teacher and adolescent–teacher ratings. Across the two analyses together (MTMM Analyses 1 and 2), the findings indicated support for the discriminant validity between all trait factors, except between CP and HY. There was also support for the discriminant validity of the mother and teacher methods (MTMM Analysis 1) and adolescent and teacher methods (MTMM Analysis 2). There was also evidence that these findings were applicable to younger and older adolescent groups (MIMIC analysis).
The support for the convergent and discriminant validity for the SDQ found in the current study is consistent with existing data (A. Goodman et al., 2010; Hill & Hughes, 2007; Sanne et al., 2009; Van Leeuwen et al., 2006; Van Roy et al., 2006). The moderate level of agreement between the respondents found in the current study is also consistent with existing data involving the SDQ (Becker et al., 2004; R. Goodman, 2001; R. Goodman et al., 2003; Van Leeuwen et al., 2006; Sanne et al., 2009), as well as other child and adolescent measures (see Achenbach et al., 1987, for a meta-analysis). Also similar to the current study, the only other study that involved parent, adolescent, and teacher ratings generally found more convergence between parent–adolescent than parent–teacher and teacher–adolescent ratings and no difference between parent–teacher than adolescent–teacher ratings.
Although the convergent and discriminant findings for the SDQ found in the current study are generally consistent with existing data, they provide data from a new perspective. First, this is the first study to examine simultaneously the convergent and discriminant validity of the parent, adolescent, and teacher versions of the SDQ using a CFA MTMM approach. Thus, as is typical of CFA, latent scores that are free of error variances were analyzed. The study by A. Goodman et al. (2010) that examined simultaneously the convergent and discriminant validity of the parent, adolescent, and teacher versions of the SDQ used the Campbell and Fiske (1959) approach, and therefore used observed scores that have error variances in them. This methodological difference is an important one as observed scores can distort construct validation findings (Green et al., 1993). Second, as the current study used the CT-C(M − 1), it applied a CFA methodology that is well suited for structurally different methods that are reflected in SDQ ratings provided by mother, teacher, and adolescent ratings.
To date, a number of explanations have been proposed for the low to moderate cross-informant agreement for child and adolescent measures. In general, these explanations relate to either real differences in children’s and adolescents’ behaviors at home, school, and other settings (situation specificity hypothesis) or differences in respondents’ perceptions of children’s and adolescents’ behaviors (bias hypothesis). Existing data have provided more support for the situation specificity hypothesis (Achenbach et al., 1987). In the current study, it can be assumed that adolescents’ self-ratings of the SDQ would most likely be cross-situational. If so, the findings indicating substantial more convergence for adolescent–mother than adolescent–teacher suggests that relative to teachers, mothers are more able to provide a more broader judgment of adolescents’ behaviors and also that they are more able to indentify adolescents’ own perceptions of their behaviors.
The findings in this study also showed more support for the convergence of ES and PP. These findings were predicted and are consistent with existing MTMM findings (Van Roy et al., 2006). Also, R. Goodman et al. (2003) found higher agreement for ES than for HYP and CP. Generally, cross-informant differences have been explained in terms of higher informant agreement for more observable behaviors (De Los Reyes & Kazdin, 2005). If so, the findings here that ES and PP have more convergence suggest that these behaviors are more observable than the other behaviors. Also, the finding that CP had the least convergence suggests that this behavior is less observable than the other behaviors. One reason for this may be that at least two of the five items in the CP scale (lies/cheats and steals) are covert behaviors that have low access to parent and teacher observations.
The findings in this study have clinical implications. First, since the amount of convergence for corresponding SDQ scales across the different respondents was only moderate, it follows that the different respondents are not providing identical measures and that they each have different information to offer. This highlights the need for researchers and clinicians to be cautious when interpreting scores derived from a single source and conversely the need for obtaining SDQ ratings from multiple sources. Indeed, R. Goodman (2001) showed that combining the information from parent, teacher, and adolescent self-rating of the SDQ resulted in the highest sensitivity in predicting psychiatric disorders compared with using information from single sources or combinations of two sources. Second, as this study found no support for the discriminant validity between CP and HY, it can be argued that these scales are confounded with each other. This also means that together these SDQ scales may constitute a single higher order factor. Thus, they may be better viewed as an overall measure for externalizing problem behaviors rather than separate measures for CP and HY. Third, the findings here suggest that the ES and PP scales provide more consistent information across mother, teacher, and adolescent ratings than the other scales. CP provided low agreement across these respondents. This underscores the need to be cautious when using the CP scores.
It is worth noting that the findings and interpretations made in this study have to be viewed with some limitations in mind. First, as the ratio of participants for every estimated parameters was low (3:1), the number of participants in the study can be consider low for stable estimates. A recent simulation study showed that a minimal sample size of 50 is sufficient to provide stable estimates for the CFI (Iacobucci, 2010). Given that the CFI values for all the models tested here indicated either good or very close to good fit, it could be argued that the findings obtained here are reliable. Second, approximately 44% of parents and adolescents invited to participate in the study did not participate. As ethics approval for this study did not permit collection of information about individuals prior to inviting them to participate, there is no information about those who did not participate and, therefore, how this affected the results. Also, 41% of adolescents with self-ratings and mother ratings were not rated by their teachers. However, there was no difference between those rated and not rated by teachers for sex distribution, age, father occupation level, and ethnicity. Third, the ratings for teachers had a hierarchical structure, as each teacher rated a number of adolescents. Although it is possible with MPlus to account the hierarchical structure using the TYPE=COMPLEX option, this was not possible in the current study as there was no information on which teachers completed which ratings. Thus, it is possible that the findings may be biased (Muthén & Muthén, 2010). Despite these limitations, it is argued that the collective results in this and previous studies involving the SDQ and the comparability of the findings for the SDQ with other child and adolescent measures do provide a strong psychometric and empirical basis for the continued use of the SDQ in applied and research work.
Footnotes
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
