Abstract
Students with complex communication needs may require augmentative and alternative communication systems to supplement or replace their speech abilities. To effectively identify a communication system, a feature-matching process should be implemented as it considers the student’s present levels of performance. Due to the unique communication characteristics of students with complex communication needs, informal assessment tools are often used to help determine the students’ skills in natural contexts. A challenge often faced with informal assessment tools is the reliability among evaluators. As such, this pilot study attempted to evaluate the reliability of a feature-matching screening checklist and its corresponding matrices among potential professionals who would be part of the educational team. Results indicated that (a) pre-service and in-service special education teachers were the most reliable combination when completing the screening checklist, (b) exposure to the screening checklist was an influencing factor on reliability, and (c) in-service speech-language pathologists made the most errors while completing the screening checklist. Implications for practices and future research directions are discussed.
Students with complex communication needs often communicate using a variety of conventional or non-conventional, linguistic or non-linguistic, and aided or unaided modalities (Brady et al., 2016). To meet the needs of these students, the Individuals with Disabilities Education Act (IDEA) states that “the IEP [Individualized Education Program] team shall consider whether the child needs assistive technology devices and services” (Individuals with Disabilities Education Act, 2004, §1414). Assistive technology, including augmentative and alternative communication (AAC) such as communication systems, can be used to supplement or replace the students’ spoken or written language (Beukelman & Light, 2020; Da Fonte & Boesch, 2019a).
Despite the benefits communication systems may have for students with complex communication needs, professionals often recommend communication systems based on familiarity, past experiences, or discussions with colleagues, rather than through evidence-based practices (Schlosser & Raghvendra, 2004; Sievers, Trembath, & Westerveld, 2020). For members of the IEP team to make data-driven decisions, an assessment should be conducted to identify the student’s strengths and needs (Sievers et al., 2020). In particular, the assessment protocols and procedures should ensure that the student’s communication modalities are amply considered, rather than hindering the assessment process (DiStefano, Sadhwani, & Wheeler, 2020). While the speech-language pathologist will likely take the primary role in the AAC assessment process (e.g., Lund et al., 2017; Moorcroft et al., 2019), the special education teacher should be actively involved as they spend the most time with the student (Peckham-Hardin, Hanreddy, & Ogletree, 2018). These collaborative efforts should be the grounding framework during the AAC assessment process, and, above all, professionals should ensure that the process is centered on the student.
The process of conducting a capability assessment (Beukelman & Light, 2020) or feature matching (Da Fonte et al., 2019b) emphasizes the importance of identifying a communication system that “best fits” the needs of the student while simultaneously maximizing their strengths (Da Fonte et al., 2019b; Da Fonte & Boesch, 2019a). To effectively and accurately identify a communication system, the student’s cognitive functioning, communication skills, adaptive behavior, and motor skills should be considered during the assessment (Beukelman & Light, 2020; Da Fonte & Boesch, 2019a; DiStefano et al., 2020). A key element of the evaluation should be to determine whether the student is a non-symbolic or symbolic communicator. A non-symbolic communicator uses behaviors that do not clearly demonstrate an intent in their communicative interaction, often resulting in the partner adding meaning to the attempt (Peckham-Hardin et al., 2018). A symbolic communicator, on the other hand, uses and understands concrete and abstract symbols (Peckham-Hardin et al., 2018). Equally, it is important to assess the intensity of supports the student may need (DiStefano et al., 2020). For example, deficits in adaptive behaviors such as conceptual, social, and practical skills may impact the student’s independence in daily life, and therefore, may require the identification of additional supports (Brady et al., 2016). Similarly, motor abilities may influence the student’s access to the communication system; as such, motor skills are an essential domain to assess (Beukelman & Light, 2020).
It is important to consider using various sources instead of relying exclusively on formal standardized assessments to determine a student’s adaptive behavior skills (Olgetree & Price, 2017). While formal assessments provide standardized results, they may have little validity for students with high support needs (Brady et al., 2016). The limited validity may stem from challenges in determining between a skill limitation and a performance deficit (Larriba-Quest, Byiers, Beisang, Merbler, & Symons, 2020). Thus, relying on formal assessment alone could lead to an inaccurate representation of the student’s abilities (DiStefano et al., 2020). By supplementing formal assessments with informal tools, the educational team may gain a broader understanding of the student’s abilities (Brady et al., 2016; Lund et al., 2017). A main advantage of informal assessments in the ability to evaluate the student in natural settings and across contexts (Boesch & Da Fonte, 2014; Olgetree & Price, 2017), and the flexibility of these tools allow for the evaluation of students who may be challenging to assess (Olgetree & Price, 2017). However, due to the characteristics of informal assessments, there will be subjectivity of the observers (Olgetree & Price, 2017), as personal experiences may influence the observation (Gronlund & James, 2013). Without systematic instructions or guidelines, two evaluators may interpret the same event differently (Olgetree & Price, 2017). To attempt to mitigate some of the biases inherent in observations, the data collection process should involve factual and descriptive information to increase the objectivity of the findings (Gronlund & James, 2013).
Informal Feature-Matching Assessment Tools
Informal feature-matching assessment tools with empirical support are important, particularly those that consider the student as a whole, support feature matching, and can be completed by the educational team in a feasible amount of time (e.g., Marfilius & Fonner, 2012; Ohio Center for Autism and Low Incidence, 2012). Yet, a systematic review of the literature conducted by the research team revealed a scarcity in these assessment tools (see Da Fonte et al., 2019b), uncovering three feature-matching checklists: Brady (n.d.), Marfilius and Fonner (2012), and Senner (n.d.). The three checklists were designed to evaluate the features of both low and high technology communication systems. Although these checklists may be helpful, they do not gather information on the student’s present levels of performance. Two additional tools were also found by the research team, the Plan for Evaluation of Effectiveness of AT Use, created by the Quality Indicators for Assistive Technology Services Leadership Team (QIAT, 2008) and the Student Inventory for Technology Supports (SIFTS), created by The Ohio Center for Autism and Low Incidence (Ohio Center for Autism and Low Incidence, 2012). Both tools assist the educational team during the decision-making process by recommending specific assistive technologies that match the student’s needs. Although these tools may be useful, the validity and reliability of these tools is unclear, and they require a large time commitment. Based on gaps in the literature and given the limitations of the existing evaluation tools, the aim of this pilot study was to evaluate the reliability of the research team’s screening checklist and its corresponding feature-matching matrices. Specifically, this study evaluated the reliability of pre-service special education teachers and in-service professionals (special education teachers and speech-language pathologists) in using the screening checklist. The research questions included (1) were there reliability differences among dyads and triads when completing the screening checklist for a target student; if so, what were the influencing factors? and (2) were there common errors made by participants while completing the screening checklist; if so, were these based on the type of participant?
Methods
Participants
The target participants were pre-service special education teachers and in-service professionals (special education teachers and speech-language pathologists). The rationale for targeting such participants was based on their current or future role and responsibilities in assessing, identifying, and selecting a communication system for students. To be included, pre-service special education teachers had to (a) be an undergraduate or master’s student in the teacher preparation program at a southern university and (b) be seeking an initial teaching licensure or a teaching endorsement in severe disabilities. For in-service special education teachers, participants had to be (a) serving students in a comprehensive development classroom in which the classroom was separate from general education, and special education teachers offered holistic education to students with severe disabilities across subject areas throughout the day (Baroody, 2017); (b) serving at least one student with complex communication needs (for the purpose of this study, known as the target student); and (c) working in a surrounding school district. For in-service speech-language pathologists to be included, they had to (a) provide services in the same classroom as that of the in-service special education teacher and (b) serve at least one student with complex communication needs.
Recruitment
After receiving Intuitional Review Board approval, the research team compiled 132 emails. The email list consisted of 30 emails of pre-service undergraduate and graduate special education teachers, and 102 emails of in-service professionals (70 special education teachers and 32 speech-language pathologists). The field placement coordinator provided a list of emails of candidates completing a practicum placement, and contacted two surrounding school districts for emails of in-service professionals who met the inclusion criteria. A recruitment email was then sent to all potential participants describing the research purpose, time commitment, and confidentiality information pertaining to their participation and of the target student. Two additional recruitment reminder emails were sent, each being sent 15 days apart.
Research Design
Usability testing was used, which is defined as “a process that employs people as testing participants who are representatives of the target audience to evaluate the degree to which a product meets specific usability criteria” (Rubin & Chisnell, 2008, p. 21). Both quantitative and qualitative data were collected to determine the reliability of the feature matching checklist and its corresponding matrices among participants in dyads and triads.
Materials
Several materials were created to (a) gather participants’ demographics (rater profile), (b) record target student demographics (student profile), (c) collect data on the target student’s areas of strengths and needs (screening checklist), and (d) guide participants during the identification of a communication system (feature matching matrices).
Rater and Student Profiles
The research team created a rater and a student profile form to systematically gather demographic information on the participants and target students. The rater profile consisted of five questions that were completed by each participant, with an additional question for pre-service special education teachers. The questions pertained to their (a) age, (b) gender, (c) ethnicity, (d) school setting (i.e., elementary, middle, and high), and (e) profession (i.e., pre-service special education teacher, in-service special education teacher, or in-service speech-language pathologist). Pre-service special education teachers were also asked if they were seeking an undergraduate or master’s degree. The student profile consisted of six questions and was completed by the in-service special education teacher on the target student. The questions related to the target student’s age, grade, gender, ethnicity, primary disability, and communication system use (previous or current).
Communication System Identification Screening Checklist
The content of the Communication System Identification Screening Checklist (i.e., screening checklist) was based on a systematic review of the literature conducted by Da Fonte et al., 2019b. The screening checklist consisted of two sections. Section 1 included three questions about the target student’s present levels of performance, including (a) cognitive abilities such as critical thinking and communication skills, including expressive and receptive; (b) adaptive behaviors, which were categorized into three subdomains: independent, referring to target students who did not need any support related to adaptive behavior; behavior challenges, defined as a target student who required supports related to challenging behaviors to successfully use a communication system (Da Fonte et al., 2019b); and motor challenges, which referred to target students who required supports related to fine and gross motor skills, and who may require specific access methods or mounting systems to successfully use a communication system (Da Fonte et al., 2019b); and (c) sensory needs, such as visual or hearing capabilities. Based on information from the target student’s present level of performance (Section 1), the student was identified as either non-symbolic communicators or symbolic communicators. Section 2 comprised Section 2a with relevant communication systems for non-symbolic communicators, and Section 2b with communication systems for symbolic communicators.
Based on the student’s present levels of performance, participants were given options to select (a) low technology communication systems, such as objects, tangibles, photographs, line drawings, and 1- or 2-message speech-generating switches; (b) dedicated high technology communication system, which included speech-generating devices that used either digitized or synthesized speech; or (c) non-dedicated high technology communication system, which included tablets or other mobile technologies. For target students who had motor challenges, an additional step was required to determine supports needed based on the student’s fine motor skills, gross motor skills, access method, and need for a mounting system.
Feature-Matching Matrices
Like the screening checklist, the Feature Matching Matrices were based on Da Fonte et al., 2019b systematic literature review. The feature matching matrices were designed to guide the users on how to use the compiled data on the student’s strengths and needs based on their present level of performance (Section 1) to identify a potential communication system (Section 2). The feature-matching matrices mirrored the screening checklist and outlined elements that should be considered when recommending a communication system. Based on the student’s present levels of performance, communication skills, and need for supports, the participant had access to six matrices to determine which matrix was more suitable for the target student. The matrices consisted of (a) non-symbolic communicators with independent functioning (based on their adaptive behavior skills), (b) non-symbolic communicators with behavior challenges, (c) non-symbolic communicators with motor challenges, (d) symbolic communicators with independent functioning (based on their adaptive behavior skills), (e) symbolic communicators with behavior challenges, and (f) symbolic communicators with motor challenges (see Figure 1). Decision-making flowchart for matrix identification.
Procedures
The research team matched pre-service special education teachers with in-service professionals to create dyads and triads. Dyads consisted of one pre-service special education teacher and one in-service special education teacher (n = 26), while triads comprised one pre-service special education teacher, one in-service special education teacher, and one in-service speech-language pathologist (n = 12). The groups were created to assess the reliability between different group combinations, and to evaluate the similarities and differences among them. For data analysis, pre-service special education teachers were categorized as having
Materials were disseminated to the pre-service participants in an envelope that included the rater profile, student profile, the screening checklist, and the feature-matching matrices. For pre-NEF participants, instructions were given to deliver all materials to the in-service special education teacher for distribution. For pre-EF, the in-service special education teacher was asked to meet with the pre-service special education teacher prior to completing the screening checklist to identify and provide background information on the target student. Members of the dyad or triad were instructed to independently complete the screening checklist and then contact the research team to return the completed materials.
Data Analysis
To determine differences in reliability, analyses were conducted by dyad and triad groups and by participants’ characteristics. To further evaluate the triads, they were deconstructed into all possible dyad combinations to assess the reliability among other types of participant pairings to determine if there were any influencing factors. Reliability was assessed globally (across sections of the screening checklist) and locally (within sections of the screening checklist). Using SPSS (version 27.0), percent agreement, descriptive statistics, Cohen’s kappa, and Fleiss’ kappa were used to calculate the reliability of the screening checklist. Percent agreements were calculated by dividing the number of agreements by the total number of agreements plus disagreements and multiplying by 100. The criterion set for this study was
The number of errors participants made on the screening checklist was also calculated within each section. Errors were determined (a) globally, by evaluating the number of errors across the screening checklist; (b) locally, by calculating the number of errors within each section; and (c) based on the type of participant. The research team determined that an error occurred when the matrices were not followed, and: (1) An assessment question was not answered in the target student’s present level of performance (Section 1), which determined if the student was functioning independently or required sensory supports. (2) The flowchart designed to select the correct section to fill out was not followed based on the performance levels identified in the target student’s present level of performance (Section 1). For example, an error occurred if a communication system option for symbolic communicator (Section 2b) was completed, yet the target student was a non-symbolic communicator (Section 2a). (3) The communication system selected did not correspond to the student’s present levels of performance in Section 2. For example, if the participant selected a communication system within the independent domain even though they had previously identified the target student as requiring sensory supports. (4) More than one low technology communication system was selected in the communication system option for non-symbolic communicators (Section 2a). For example, rather than only considering the highest level of symbol abstraction for the target student, the participants selected two communication systems. (5) More than one speech-generating device was chosen in the communication system option for non-symbolic communicators (Section 2a). An error occurred if the participant selected both 1- and 2-message speech-generating switches for the same target student, rather than considering the student’s highest potential performance level. (6) More than one high technology communication system was selected in the communication system option for symbolic communicators (Section 2b). For example, both dedicated and non-dedicated high technology communication systems were selected for the same target student.
Treatment Integrity
Specific procedures were outlined to ensure the instructions were delivered similarly across participants. If questions arose by a participant, general responses were given to guide participants back to the instructions on the screening checklist or corresponding feature-matching matrices. Data were collected on the participants’ independent completion of the screening checklist. Approximately 46.67% of the participants discussed the target students prior to completing the screening checklist. From these, 26.67% reviewed the feature-matching matrices and screening checklist together, while none conversed during the completion of the screening checklist. Anecdotal notes indicated that the participants’ most common question was on the differences between non-symbolic and symbolic communicators. In these cases, terms were defined by a research team member, but the screening checklist was independently completed.
Results
Participants
Participants’ and Target Students’ Demographic Information.
Note. †more than 46 participants because some worked in or collected data in multiple classrooms; % = percentage of participants; n = number of participants; SLP = speech-language pathologist; SPED = special education.
Target Students
As represented in Table 1, a total of 38 target students were assessed across dyads (n = 26) and triads (n = 12). Many of the target students were classified as having intellectual disability (57.90%) as their primary disability, with the least number of students having deaf-blindness (2.63%) and traumatic brain injury (2.63%). Most target students were Caucasian (57.89%), male (63.68%), and between the ages of 6 and 10 (44.47%). Most of the target students used low technology communication systems (52.63%), followed by high technology communication systems (26.32%) and no communication system (21.05%).
Reliability and Influencing Factors of Dyads and Triads
Dyads and triads were analyzed separately to determine reliability differences. Missing kappa suggest that dyads or triads followed the matrices and agreed 100% that a specific communication system was not appropriate given the target student’s level of performance. Therefore, this selection was not completed by participants on the screening checklists.
Dyads
Dyads consisted of one pre-service special education teacher and one in-service special education teacher (n = 26). Across all sections of the screening checklist (globally), the percent of agreements ranged from 76.92% to 100% for dyads, with means varying from 24.67 (SD = 1.528) to 25.33 (SD = 0.817) and kappa ranging from −0.013 to 1.000. Possible agreements ranged from 0 to 26, with the highest mean for the communication system option for non-symbolic communicators (Section 2a; M = 25.33; SD= .817), followed by the target student’s present levels of performance (Section 1) and potential communication systems for symbolic communicators (Section 2b) with similar mean agreements (M = 24.67; SD = 1.528 and M = 24.46; SD = 1.914, respectively). Overall, dyads reached 80% agreement across all but one item of the screening checklist. The item in question pertained to the match between the presence of challenging behaviors and a high technologydedicated communication system.
Comparison of Dyads’ Reliability Across Sections of the Screening Checklist.
Note.—kappa and percent agreements could not be calculated as no target student had characteristics for the specific component; %I = percentage of communication systems identified by dyads for target student use; %NI = percentage of communication systems not identified by dyads for target student use; C-k = Cohen’s kappa; (H) D-CS = high technology dedicated communication systems; (H) ND-CS = high technology non-dedicated communication systems; Ind = independent; (L) D-CS= low technology dedicated communication systems; n = number of dyads; SGD = speech-generating device; TS = target student.
Influencing Factors
Pre-service special education teachers were separated into two groups: (1) Those who had no previous exposure to the screening checklist but exposure to the target students and classroom (pre-NEF) and (2) participants who had previous exposure to the screening checklist, but not to the target students or to the classroom (pre-EF). Overall, findings from both groups indicated similar kappas (ranging from 0.000 to 1.000) and percentage of agreements between pre-NEF (ranging from 71.43% to 100%) and pre-EF (ranging from 78.95% to 100%). Possible agreements for pre-NEF could be from 0 to 7. The highest mean was in the communication system option for non-symbolic communicators (Section 2a; M = 6.67; SD= 0.516), followed by the communication system option for symbolic communicators (Section 2b; M = 6.54; SD = 0.577), and the lowest mean was in the target student’s present level of performance (Section 1; M = 6.0; SD = 1.000). For pre-EF, the possible agreement could be between 0 and 19. The target student’s present level of performance (Section 1) and the communication system option for non-symbolic communicators (Section 2a) had the same mean agreement (M = 18.33; SD = 0.577 and M = 18.33; SD = 0.516, respectively), with the lowest agreement in the communication system option for symbolic communicators (Section 2b; M = 17.77; SD = 1.301). Overall, results indicate that pre-EF was slightly more reliable within the dyad than pre-NEF, suggesting that exposure to the form may have been an influencing factor contributing to higher reliability among dyads (see Table 2 for a summary of dyad reliability).
Triads
Triads comprised one pre-service special education teacher, one in-service special education teacher, and one in-service speech-language pathologist (n = 12). The percent of agreements among triads ranged from 58.33% to 100%. The possible mean agreement ranged from 0 to 12, with the highest mean agreement obtained in the communication system option for symbolic communicators (Section 2b; M = 11.08; SD = 1.382), followed by the communication system option for non-symbolic communicators (Section 2a; M = 10.83; SD = 0.753), and the lowest occurred in the target student’s present level of performance (Section 1; M = 8.67; SD = 1.528). Kappa ranged from −0.029 to 0.768. Yet, it is noteworthy to highlight that kappa for 11 of the 22 items were unable to be calculated given that there was a lack of target students who had the characteristics who could benefit from using certain communication systems. Data show triads had 80% or higher agreement on only some of the communication systems.
Local data indicated that triads did not reach the set criteria of 80%, with an average of 72.22% agreement (ranging from 58.33% to 83.33%) across the three assessment domains. The target student’s present level of performance in the adaptive behaviors domain (Section 1), had the lowest percent agreement (58.33%); and the lowest reliability was in the identification of the target student’s communicative skills (k = 0.273). Even with these differences, performance finding had the highest percentage of agreements (83.33%).
Due to the inconsistency across reliability measures, the raw data were analyzed, showing that two target students’ performance levels were outlined differently between the pre-service special education teachers and in-service speech-language pathologists. Disagreement among these participants in the target student’s present level of performance (Section 1) led to additional disagreements that influenced the other sections (2a or 2b). Within the communication system option for non-symbolic communicators (Section 2a), the highest percent agreements occurred across objects, tangibles, and line drawings (100%); however, the inability to calculate kappa resulted in having no variance among target students.
Comparison of Triads’ Reliability Across Sections of the Screening Checklist.
Note. – kappa and percent agreements could not be calculated as no target student had characteristics for the specific component; %I = percentage of communication systems identified by triads for target student use; %NI = percentage of communication systems not identified by triads for target student use; F-k = Fleiss’ kappa; (H) D-CS = high technology dedicated communication systems; (H) ND-CS = high technology non-dedicated communication systems; Ind = independent; (L) D-CS= low technology dedicated communication systems; n = number of triads; SGD = speech-generating device; TS = target student.
Influencing Factors
Triads including pre-NEF had percent agreements ranging from 66.67% to 100% and kappa ranging from .080 to .755, while triads with pre-EF had percent agreements ranging from 33.33% to 100%, and kappa from −0.125 to 1.000. The mean score for triads involving pre-NEFs ranged from 6.33 (SD = 0.577) to 8.50 (SD = 0.548) with a possible range from 0 to 9. For triads with pre-EFs, the highest mean score was for the communication system option for non-symbolic communicators (Section 2a; M = 3.00; SD = 0.000), followed by the communication system option for symbolic communicators (Section 2b; M = 2.60; SD = 0.630), and the lowest for the target student’s present level of performance (Section 1; M = 2.33; SD = 1.155) with the possible agreements could range from 0 to 3.
Within the triads, both groups of pre-service special education teachers had the highest mean agreements for the communication system option for non-symbolic communicators (Section 2a) and lowest for the target student’s present level of performance (Section 1). However, the small number of pre-EF within triads made it difficult to compare with pre-NEF. Overall, findings suggest that pre-service special education teachers with exposure to the form was not an influencing factor that increased reliability among triads (see Table 3 for a complete summary of the reliability among triads).
Triad Deconstruction
Comparison of Deconstructed Triads’ Reliability Across Sections of the Screening Checklist.
Note.—kappa and percent agreements could not be calculated as no target student had characteristics for the specific component; %I = percentage of communication systems identified by deconstructed dyads for target student use; %NI = percentage of communication systems not identified by deconstructed dyads for target student use; C-k = Cohen’s kappa; (H) D-CS = high technology dedicated communication systems; (H) ND-CS = high technology non-dedicated communication systems; Ind = independent; IS = in-service; (L) D-CS= low technology dedicated communication systems; n = number of deconstructed dyads; PS = pre-service; SGD = speech-generating device; SLP = speech-language pathologist; SPED = special education teacher; TS = target student.
Evaluation of Errors
The most errors occurred in Section 1 (n = 25), while the fewest errors occurred in the communication system option for symbolic communicators (Section 2b; n = 9). Within the target student’s present level of performance (Section 1), no errors were made when filling out the performance levels of the student. Rather, errors were made in the adaptive behavior (13.64%; n = 12) and sensory needs domains (13.64%; n = 12). Within the communication system option for non-symbolic communicators (Section 2a), most errors occurred when selecting more than one communication system (7.95%; n = 7). In the communication system option for symbolic communicators (Section 2b), the most errors occurred while following the recommendations within the feature matching matrices (5.68%; n = 5), and these 5 errors were made for symbolic communicators with behavior challenges. Overall, pre-service special education teachers had the lowest percentage of errors (21.05%) when using the screening checklists, while in-service speech-language pathologists had the highest percentage of errors (41.67%). When considering the type of participant, the most errors (n = 20) were made by in-service speech-language pathologist, followed by in-service special education teachers (n = 16) and pre-service special education teachers (n = 14). Approximately 41.67% of the screening checklists completed by in-service speech-language pathologist had 2+ errors, with the lowest percentage among pre- and in-service special education teachers (15.79% with 2+ errors).
Discussion
The aim of this study was to determine the reliability of an informal assessment tool grounded in the current available literature on feature matching. Overall, results indicated that (a) participants within dyads were more reliable than triads, (b) the deconstructed triad involving an in-service speech-language pathologist had consistently lower reliability than other pairings, (c) exposure to the screening checklist was an influencing factor on reliability, (d) the most errors occurred in target student present level of performance (Section 1), and (e) the most errors were made by in-service speech-language pathologists.
Foremost, despite differences in the reliability across dyads and triads, the percent of agreement suggests that the screening checklist was reliable overall, as there was agreement above 80% with most communication systems across dyads and triads. However, when considering the global and local results, findings indicated that pre- and in-service special education teachers (dyads) were more reliable than a triad that included an in-service speech-language pathologist in completing the screening checklist. Differences among dyads and triads may be related to the participants’ roles and their educational backgrounds. For example, pre- and in-service special education teachers almost always had higher reliability across sections of the screening checklist, as compared to other combinations that included an in-service speech-language pathologist. A possible explanation may be that the group size may have influenced the reliability among evaluators. Yet, findings of the pilot study seem to be aligned to assertions suggesting that professionals who have similar roles and observe students in the same context may have higher correlated observations (Achenbach, McConaughy, & Howell, 1987; Achenbach, 2018). The notion that the differences in reliability may be attributed to context differences (Peckham-Hardin et al., 2018) may have played a role in this study. Specifically, given that the pre- and in-service special education teachers likely interacted and observed the students in a similar manner, while the in-service speech-language pathologists viewed the students’ abilities in a different light.
Furthermore, although informal non-standardized assessments tools are valid methods to collect data to inform communication interventions, limitations exist given the subjectivity of observers (Olgetree & Price, 2017). While personal experience may influence a rater’s ability to objectively determine the student’s ability (Olgetree & Price, 2017), practice in using these assessment tools may decrease bias and increase reliability (Gronlund & James, 2013). Because results from the current study indicated that pre-service special education teachers with exposure to the screening checklist form had slightly higher reliability compared to dyads involving those with no exposure to the form, this may suggest that a way to increase effectiveness and reliability of observations is for data collectors to receive additional exposure through training (Olgetree & Price, 2017). Similarly, results also suggest that it may be beneficial to have a team member with ample knowledge of the student (e.g., in-service special education teacher). By having a team member with experience with the assessment tool, reliability may increase.
Findings also indicated that most of the errors occurred when participants had to determine the target student’s present levels of performance. Although it is unclear why these errors occurred, assumptions can be made based on the type of errors. It reasonable to conclude that professionals can become reliable observers (Gronlund & James, 2013), conducting effective assessments, and making accurate conclusions when exposure and active engagement occurs (Grainger & Adie, 2014). Furthermore, based on anecdotal data, most of the questions asked by participants pertained to terminology about the differences between non-symbolic and symbolic communicators. This lack of knowledge about the terminology used in the screening checklist may explain some of the errors participants made, which may have impacted the reliability. It is important to have a good grasp of the terminology used in the operational definitions and instructions of the informal assessment tool to become reliable observers (Gronlund & James, 2013).
Implications for Practice
The fact that students with complex communication needs present heterogeneous characteristics (Light & McNaughton, 2012) implies that professionals will need to stay current in their practices to ensure effective instruction for all students (Lund et al., 2017). The combination of limited research on the decision-making during AAC assessment (Lund et al., 2017; Schlosser & Raghavendra, 2004), the lack of teacher training (Andzik et al., 2019; Costigan & Light, 2010; Da Fonte et al., 2022 accepted), and the rapid changes in assistive technology (Abbot & Bride, 2014) highlights the need for having (a) a reliable, systematic, and comprehensive approach in identifying communication systems; (b) practitioner-friendly assessment tools that are grounded in the literature; (c) a collaborative team approach for systematically identifying and supporting the needs of students with complex communication; and (d) AAC training across all disciplines. As such, it is critical to conduct comprehensive assessments in a collaborative manner when identifying a communication system for students with complex communication needs (Brady et al., 2016; Lund et al., 2017) to maximize the strengths of each team member. Additionally, it is important to use tools that can assist in gathering data more systematically and effectively within natural contexts (Boesch & Da Fonte, 2014; Olgetree & Price, 2017). Using these approaches can be instrumental in guiding the feature-matching process. Therefore, school districts are urged to consider providing interdisciplinary training opportunities so that professionals across disciplines including special education teachers and speech-language pathologists can gain the skills required to become reliable assessors and help identify the supports students need to be successful communicators.
Although the reliability of triads was low, the notion that all stakeholders should be involved in the AAC assessment process has been documented (Beukelman & Light, 2020; Brady et al., 2016; Da Fonte & Boesch, 2019a; Lund et al., 2017; Sievers et al., 2020). Each member brings their own expertise and knowledge of the student across multiple contexts to the assessment process. Collaborative efforts in a comprehensive assessment will set the stage for open discussions, which in turn will increase the reliability between assessors (Achenbach et al., 1987). In fact, Hodapp et al. (2019) highlight that disagreements between professionals from varying disciplines can reflect contextual differences, rather than a disagreement between raters. This belief supports findings about differences among the reliability of dyads and triads. Therefore, even though results from this study indicate a lower reliability when more team members were included (triads), the concept of collaboration in the assessment process among all those involved with the student maintains; and the assertion that special education teachers should be key professionals in this process stands, due to their extensive interactions with the student (Peckham-Hardin et al., 2018).
Limitations and Future Research Directions
Several limitations exist in this pilot study. First, there was limited representation of a large, diverse group of participants. The pre-service special education teachers in this study were all from the same program, and the in-service teachers and speech-language pathologists were all from school districts, which decrease the possibility of a diverse sample. Future research should recruit a larger and more diverse group of participants that allow for analysis across preparation programs, communities, years of experiences, and other demographic variables. Moreover, future research could consider exposing all participants within the dyads and triads to the study materials through a training phase prior to completing the screening checklist. This information would allow for further examination of the exposure to the screening checklist as an influencing factor. Another future research possibility is to expand the participants’ inclusion criteria, among dyad and triads. For example, future studies could consider evaluating the reliability between school professionals, such as special education teachers and speech-language pathologist, and the responses of caregivers or paraeducators. Findings of such studies may help expand the number of potential stakeholders who can partake in the feature-matching process.
Although the results of this pilot study suggest that the screening checklist was overall reliable, the reliability varied within the screening checklist, which may suggest potential changes needed to improve the tool. For instance, anecdotal data indicated that many participants were unfamiliar with the terminology of non-symbolic and symbolic communicators or did not use the corresponding matrices to complete the screening checklist. Potential revisions could include (a) operational definitions, with examples and non-examples that supplement the screening checklist and (b) outlining the need to use the corresponding matrices when completing the screening checklist. Both of these revisions could potentially increase the clarity and reliability of the screening checklist.
Another limitation was the lack of variance, which resulted in being unable to calculate kappa across different communication systems. For example, there was a lack of target students with motor challenges, and therefore, the tool was unable to be validated for this subgroup. Recruiting a larger pool of participants will inherently increase the number of target students being assessed, and as a result, future research could consider an even distribution of non-symbolic and symbolic communicators. Although AAC knowledge was accounted for with pre-service special education teachers, a limitation of this study was that the same was not done with in-service professionals. Future research should consider gathering information on in-service professionals’ background in AAC. Future research should also consider the collecting data on social validity to guide the improvement of the screening checklist and its corresponding matrices. Recommendations may help increase the reliability of the screening checklist and decrease the number of errors made by participants. Similarly, future research should consider a deeper dive into the differences among dyads and triads and determine if group size is an influencing factor on the reliability of the screening checklist.
Conclusion
Comprehensive assessments are essential to gaining a holistic understanding of the student’s strengths and areas of needs, prior to recommending a communication system. Findings of this pilot study suggest reliability differences exist among dyads and triads when using the screening checklist. Future research is warranted to determine why these differences may exist. By increasing the reliability among and across potential evaluators, results may allow for more fine-grain recommendations for an effective communication system.
Footnotes
Acknowledgments
We would like to thank Dr. Robert Hodapp for all his assistance with this project.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Author’s Note
Emily R. DeLuca, is a former master's student from Vanderbilt University, and is now a Board Certified Behavior Analyst in Franklin Tennessee. The findings of this study are part of the first author’s master’s degree thesis.
