Abstract
Anecdotal evidence has suggested that rater-based measures (e.g., parent report) may have strong across-trait/within-individual covariance that detracts from trait-specific measurement precision; rater measurement-related bias may help explain poor correlation within Autism Spectrum Disorder (ASD) samples between rater-based and performance-based measures of the same trait. We used a multi-trait, multi-method approach to examine method-associated bias within an ASD sample (n = 83). We examined performance/rater-instrument pairs for attention, inhibition, working memory, motor coordination, and core ASD features. Rater-based scores showed an overall greater methodology bias (57% of variance in score explained by method), while performance-based scores showed a weaker methodology bias (22%). The degree of inter-individual variance explained by method alone substantiates an anecdotal concern associated with the use of rater measures in ASD.
Autism Spectrum Disorder (ASD) diagnosis and assessment depend on the validity of those scales intended to capture and quantify the behavioral features that are essential to the diagnosis. The Autism Diagnostic Observation Schedule (ADOS) (Lord et al., 2000) and the ADOS, second edition (Lord et al., 2012) are clinician-administered, performance-based assessments of ASD symptoms. Other widely used assessments of ASD symptom magnitude are the Social Responsiveness Scale (SRS) (Constantino & Gruber, 2005) and the SRS second edition (SRS-2) (Constantino & Gruber, 2012), which are questionnaires based on parent report. Although both ADOS and SRS aim to quantify ASD-specific symptomology, the two measures have been found to fail to correlate as two measures of the same construct should (e.g., Leung et al., 2016 showed a nonsignificant correlation, with r = .06). Such a lack of correlation, at its most severe interpretation, calls into question the fundamental coherence of the ASD construct (Waterhouse et al., 2016).
However, one alternative possibility is that there exists substantial measurement-type bias in one or both measures. Campbell and Fiske (1959) described the output of assessments as encompassing a “trait-method unit,” whereby a score is a function of two things: the quantitative response to content (i.e., a particular psychological trait, domain, or construct) and the quantitative response to features of the method of assessment. The content representing a valid construct/trait should be identifiable by multiple different methods. Optimally, measurements should be insensitive to methodology, yet this method-independence does not hold in practice. Within ASD samples, Wodka and colleagues (Wodka et al., 2016) reported that a performance measure of tactile perception and a performance measure of attention individually correlated more strongly with one another than with the parent-rating-based instrument corresponding to the respective trait. Further, parent ratings of attention and parent ratings of somatosensory processing were more correlated with one another than with their performance-based counterparts. In the case of ASD, at least, there was a stronger alignment of scores based on method than trait.
Our goal in the current research was to examine method-based variance among measures of differing domains in both parent-report measures and performance measures across multiple traits in an ASD convenience sample with restricted IQ range. Our prediction was based on anecdotal concerns within the neuropsychological community, viz., that parent-report instruments suffer from substantial correlation across cognitive domains (traits), moreso than is the case with performance-based instruments.
Methods
We report how we determined our sample size, all data exclusions, all manipulations, and all measures in the study.
Participants
Data came from multiple case-control studies of children diagnosed with ASD (e.g., McAuliffe et al., 2017, 2020). Participants were recruited from local schools, from those who receive services at the Kennedy Krieger Institute and express interest in research participation, and from community clubs/organizations. The collection and analysis of data were approved by the Johns Hopkins Medicine Institutional Review Board. All data were collected by psychology associates supervised by doctoral-level psychologists. Inclusion for ASD criteria consisted of age 8.0 to 12.9 years, meeting criteria based on the ADOS-G or ADOS-2, module 3 (Lord et al., 2000, 2012) and final diagnosis by a single child neurologist specialized in the diagnosis of ASD. Some studies additionally required meeting criteria via the Social Communication Questionnaire (Rutter et al., 2003), Autism Spectrum Screening Questionnaire (Ehlers et al., 1999), and/or the Autism Diagnostic Interview-Revised (ADI-R) (Lord et al., 1994). In addition, participants were required to have had a Full-Scale IQ (FSIQ)
Demographic Data of Study Participants.
Note. Values represent means and (SD, ranges), unless otherwise noted. ADHD = attention-deficit/hyperactivity disorder; WISC-IV = Wechsler Intelligence Scales-IV; FSIQ = Full-Scale IQ; ADOS = Autism Diagnostic Observation Schedule; SRS = Social Responsiveness Scale; WMI = Working Memory Index; BRIEF = Behavior Rating Inventory of Executive Function; DCDQ = Developmental Coordination Disorder Questionnaire; PANESS = physical and neurological examination of subtle signs; DSF = digit span forward.
Measures
Autism Diagnostic Observation Schedule, First and Second Editions, Module 3 (ADOS-G and ADOS-2, Respectively)
The ADOS-G (Lord et al., 2000) and ADOS-2 (Lord et al., 2012) are behavioral observation protocols designed as a semi-structured interview and assessment. They aim to elicit behaviors characteristic of ASD. There are five modules of the ADOS, of which one is selected for the individual based on language level. For the present study, module 3 was administered to all participants, based on the recommendation for verbally fluent children and adolescents younger than age 16. Calibrated severity scores for ADOS-G were used to compare scores between ADOS-G and ADOS-2 (Wiggins et al., 2019). Administration time for the module 3 is 40-60 minutes. For module 3 of the ADOS-2, inter-rater reliability is well-established between 0.91-0.94 and test–retest reliability between 0.81-0.87 (which is comparable to the ADOS-G). For the current analyses, we used the ADOS Total score as the performance assessment of the Core ASD Feature trait.
Social Responsiveness Scale, First and Second Editions (SRS; SRS-2)
The SRS (Constantino & Gruber, 2005) and SRS-2 (Constantino & Gruber, 2012) are parent-rating measures of the constellation of qualities that are posited to differentiate autism from other conditions. Sixty-five items are rated on a four-point Likert-type scale, which make up the five domains of interest (Social Communication, Social Motivation, Social Awareness, Social Cognition, and Restricted and Repetitive Interests and Behaviors); completion time is estimated at 15-20 minutes. Internal consistency reliability is well-established at Cronbach’s α > 0.95 (based on a standardization sample of 1,963). In our model, we used the SRS(-2) Total score as the rater assessment of the Core ASD Feature trait.
Digit Span Forward (DSF)
Digit Span is a subtest in the Wechsler Intelligence Scales (WISC) (Wechsler, 2003, 2014) where participants are asked to repeat digit spans of increasing length that are spoken by the examiner. Participants much reach a basal level of two correct items in a digit set and a ceiling of two incorrect items in a digit set. The DSF raw score is calculated by adding all correct items before ceiling and converted to an age-corrected scaled score, which was used in analyses as the performance measure of attention. Reliability of DSF is well-established at r = .81 vial the split-half method. DSF was used as the performance instrument measuring the Attention trait.
Conners-3
The Conners-3 (Connors, 2008) is a 99-item parent-rating instrument that measures symptoms of ADHD and co-occurring behaviors in children and adolescents; completion time is estimated at 20 minutes. There are six content scales and three validity scales, as well as several DSM-IV TR symptom scales. Reliability is well-established at r = .91 for internal reliability, r = .85 for test–retest reliability, and r = .81 for inter-rater reliability. For this analysis, the Inattention T-score was used as the rater measure of the Attention trait.
Working Memory Index (WMI)
The WMI is one of the index scores that contributes to the FSIQ of the WISC. It is comprised of the Digit Span and Letter-Number Sequences subtests of the WISC-IV and the Digit Span and Picture Span subtests of the WISC-V. Internal reliability is well-established at 0.92 for the WISC-V, comparable to the WISC-IV. For the present analysis, WISC-IV and WISC-V WMI Standard Score was used as the performance measure of the Working Memory trait.
NEPSY Statue
The NEPSY is a battery of assessments measuring a variety of neuropsychological constructs (Korkman et al., 2007). The Statue subscale is aimed at assessing motor inhibition, asking the child to stand still for 75 seconds while environmental distractions are presented. The child’s score on the measure reflects their ability to stay still, and errors are noted for eye opening, vocalization, and movement. The raw total errors scores were used as the performance measure of the Inhibition trait, as scaled scores are not available for our age range. NEPSY Statue raw scores have been used to measure inhibition in other clinical populations in this age range, with stronger performance than other instruments (Soltani et al., 2023).
Behavior Rating Inventory of Executive Function (BRIEF)
The BRIEF is an 88-item parent-rating form measuring executive functioning (Gioia et al., 2002); completion time is 15 to 25 minutes. There are eight subscales as well as three index scores that are norm-referenced. Mean internal consistency ratings reported for clinical populations using the BRIEF Parent Form range from 0.82 to 0.98. Three-week test–retest correlations for clinical populations on the Parent Form range from 0.72 to 0.84. For this analysis we used the Inhibition subscale T-score as the rater measure of the Inhibition trait and the Working Memory subscale T-score as the rater measure of the Working Memory trait.
Physical and Neurological Examination of Subtle Signs (PANESS)
The PANESS is an examination of motor coordination, ranging from qualities of gait to more subtle hand and finger movements (Denckla, 1985). Inter-rater reliability is estimated at 91% agreement. The PANESS Total Score was used as the performance measure of the Motor Coordination trait.
Developmental Coordination Disorder Questionnaire (DCDQ)
The DCDQ is a parent-rated questionnaire targeting motor problems in children (Wilson et al., 2009). The measure contains 15 items that make up three subscales (Control During Movement, General Coordination, and Fine Motor/Handwriting). Higher scores indicate better motor skills; cut-offs by age are offered such that scores 15 to 46 (ages 5–7 years), 15 to 55 (ages 8–9 years), and 15 to 57 (ages 10–15 years) suggest/indicate Developmental Coordination Disorder. Reliability and validity of this measure have been established in a clinical context. We used the DCDQ Grand Total as the rating measure of the Motor Coordination trait.
Statistical Analysis
We had both performance-method and report-method data in support of a total of five psychological domains (traits) (Table 2). We modeled each measurement according to its trait and method; for participant i, trait j, and method k, we have the following statistical model:
where μjk was a fixed mean effect of the jth trait and kth method,
Seven Psychological Domains (Traits, Constructs) Examined, and the Report- and Performance-Method Instruments Used to Assess Each.
Note. BRIEF = Behavior Rating Inventory of Executive Function; ASD = autism spectrum disorder; SRS = Social Responsiveness Scale; ADOS = Autism Diagnostic Observation Schedule.
Data were normalized to account for directionality of severity and the use of raw and scaled scores. To estimate the parameters in the model designed above, we estimated
Results
Sixty-eight participants had the WISC-IV, and 15 had the WISC-V. Thirty-five participants had the SRS, and 48 had the SRS-2. Thirty-three participants had the ADOS-G, and 50 had the ADOS-2. A few participants had a computed FSIQ < 80, in line with the caveat relating to index-factor discrepancy, as described above.
Parent-rating instruments were susceptible to a greater amount of method-attributed variance in each psychological domain (trait) compared to that of the corresponding performance measures (Table 3). The aggregate estimate of shared variance attributable to method among parent-rating instruments was 57%, whereas the aggregate estimate of shared variance attributable to the performance method was 22%.
Percentage of Shared Factor Variance Explained by Method for Each Trait of Interest.
Table 3 reflects the percentage of the variability in each instrument score that is attributable solely to the method. Note that percentages do not add to 1 since they are percentages of factor variability. Usual care should be taken when comparing normalized percentages across traits. The key conclusion, however, is that the percentage of score variability attributable to method is greater for rater instruments than for performance instruments (rather than across-trait comparisons). This pattern is true not only across all domains (traits) but also within each and every domain measured.
Discussion
Our results indicate that, among autistic children with a restricted IQ range, a substantial percentage of between-individual variation seen in parent-report scores can be attributed to a bias that is inherent to rater methodology. That is to say that an individual’s score on a particular trait-specific rating scale is based to a large degree on how the individual is rated (by a parent, in this analysis) in a nonspecific fashion rather than the individual’s true level of expression of the targeted psychological trait. This large rater-method bias could explain, to a degree, the lack of correlation in ASD samples between ADOS and SRS (Leung et al., 2016), despite the fact that they were both designed to measure core ASD features, or between performance and rater measures of sensory reactivity (Wodka et al., 2016), as the rater bias effect may overwhelm true correlations between the performance and rater instruments targeting the same trait.
We know that this across-trait correlation of rater-instrument scores is not solely a function of true covariance of inter-individual abilities across traits (e.g., due to biological variation that affects performance in multiple psychological domains), because performance measures showed a substantially smaller degree of covariance across traits. (The mathematical alternative, that rater instruments truly capture real between-trait covariance and that the performance instruments show greater within-individual, between-trait differences due to enhanced susceptibility to random noise, seems less likely on its face.) Future work using samples from multiple raters per participant (e.g., parents and teachers) could better triangulate the specific role of the rater in the across-trait covariance; the convenience sample used did not include information from multiple raters. Moreover, future work may also evaluate the psychometric properties associated with different types of raters (self, parent, and teacher).
Although our results were primarily dependent on within-method covariation across traits, we are also conscious that our model also takes into account the within-trait correlation between instrument method. Here, there is a question about the degree of alignment of targeted traits within performance and rater pairs targeting the same trait. In choosing a performance instrument and a rater instrument for the same trait, our goal was to select instruments that would be viewed as representing the designated trait in a way that matches clinical practice. That said, we are also conscious that perfect alignment around the same trait is challenging. For one thing, the empirical delineation of the “primitive,”“most fundamental,” or “irreducible” traits/domains within Neuropsychology or Clinical Psychometrics is often limited (Poldrack & Yarkoni, 2016). Although the psychometric field has empirical approaches for defining the boundaries of a construct (Cronbach & Meehl, 1955; Nunnally & Bernstein, 2010; Strauss & Smith, 2009), and highly rigorous methods for the empirical validation of constructs have been proposed in research contexts (e.g., Shallice & Cooper, 2011, §3.3), actual clinical and research practice has often generated novel constructs and distinctions “ad lib and ad hoc” (Poldrack & Yarkoni, 2016, p. 588, citing W. R. Uttal). As one example, what level of granularity represents true and natural distinctions, for example, attention versus sustained attention versus auditory sustained attention?
The second consideration is that all instruments’ scores are sensitive not only to the traits they are purported to measure but are also sensitive to forms of artifact (such as bias within rater instruments, or the quality of the prior night’s sleep within performance instruments) as well as to other level of function of other psychological traits/domains (such as general intelligence tested via verbal IQ scales being sensitive to language ability/disability). We have generally spoken about the rater-instrument bias term observed within this analysis as being a product of rater bias, but it is equally plausible that rater instruments could be sensitive to a broader range of nontargeted psychological traits than performance instruments. Again, having rater data with ≥2 raters per participant would help disentangle these possibilities, as would multiple rater and performance instruments per trait (under the assumption that each instrument is sensitive to different confounds, which will average each other out when combining across instruments). Moreover, the use of instruments representing an even more diverse range of methods (e.g., interview and ethological observation) would help better characterize the testing artifacts and sensitivity to nontarget psychological domains than the current work. Doing so would involve a larger sample as well as other analysis methods, such as structural equation modeling. Such work may be particularly important within the study of core autism symptoms, where our conceptual models (compared with other psychological domains) are still under evolution. Optimally, the field would also subject each instrument used to be subject to a comprehensive task analysis (Friedrich & Rader, 1997) based on an empirically defined ontology of psychological traits (Poldrack & Yarkoni, 2016).
In our sample of autistic children, a condition that is especially heterogeneous in its presentation, the bias identified in rating forms further motivates the use of in-person testing for nuanced neuropsychological assessment in children with ASD, whether for research or clinical purposes. We could also envision future work to develop statistical adjustments against rater bias within a battery of rater measures, to render them more effective at making performance distinctions across traits.
The limited range of IQ, verbal skill, and core symptom presentation mean these results are not generalizable to the entirety of the autism spectrum. The visibility and impact of psychological variation in attention, working memory, inhibitory, motor coordination, and core autistic features may be different in autism groups with IQs different from the ones reported here. In fact, the evidence that exists suggests that the discriminatory ability of tests intended to measure autistic core features may perform less well in populations exhibiting IQs lower than in our current sample (Havdahl et al., 2016), whether or not the cause is method-related bias. The male-to-female ratio in this sample was higher than the typically reported 4:1 ratio for autistic individuals, and sex differences were not analyzed. However, more heavily male-dominated ratios in ASD samples have been reported among higher-IQ-range samples (de Giambattista et al., 2021; Werling & Geschwind, 2013), as was the case here.
While the applicability of the current project is toward understanding how best to characterize the psychological profile of individuals already diagnosed with ASD, future work may leverage transdiagnostic samples (see Tang et al., 2023) to understand how psychometric instruments function at the borderlands of differential diagnosis and with regard to co-occurring conditions (for the instruments used here, most notably Attention-Deficit Hyperactivity Disorder and Intellectual Disability).
In conclusion, the method by which psychometric data are gathered in ASD has an impact on interpretation. Data collected by surveying parent observations of their child’s behavior share much more variance (over 50%) than data collected by directly measuring the child’s performance on a task. As such, parent-rated scales may be no more specifically descriptive of inter-individual differences on the construct they are reportedly measuring than they are with the mélange of factors that goes into parent report across traits.
Footnotes
Acknowledgements
The authors wish to acknowledge the Center for Neurodevelopmental and Imaging Research for sharing data in support of this work.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: Funding for this work came from the NIH grant R01MH113652 to J.B.E. and P50HD103538. Funding sources were not involved in the analysis or interpretation of data, writing of the report, or the decision to submit for publication.
