Abstract
The purpose of this experimental study is to investigate the effects of using content acquisition podcasts (CAPs), an example of instructional technology, to provide vocabulary instruction to adolescents with and without learning disabilities (LD). A total of 279 urban high school students, including 30 with LD in an area related to reading, were randomly assigned to one of four experimental conditions with instruction occurring at individual computer terminals over a 3-week period. Each of the four conditions contained different configurations of multimedia-based instruction and evidence-based vocabulary instruction. Dependent measures of vocabulary knowledge indicated that students with LD who received vocabulary instruction using CAPs through an explicit instructional methodology and the keyword mnemonic strategy significantly outperformed other students with LD who were taught using the same content, but with multimedia instruction that did not adhere to a specific theoretical design framework. Results for general education students mirrored those for students with LD. Students also completed a satisfaction measure following instruction with multimedia and expressed overall agreement that CAPs are useful for learning vocabulary terms.
Given pervasive negative outcomes for large numbers of students with specific learning disabilities (LD; see Newman et al., 2011), it is important for researchers to continue to develop, test, and disseminate theories and interventions that support the cognitive and academic learning needs of adolescents with LD (Scammacca et al., 2007; Scruggs, Mastropieri, Berkeley, & Graetz, 2011; Swanson, 2009). Specifically, interventions should be grounded in theory and empirical findings related to the cognitive processes of adolescents with LD when interacting with academic tasks (Compton, Fuchs, Fuchs, Lambert, & Hamlett, 2011; Johnson, Humphrey, Mellard, Woods, & Swanson, 2010). In addition, designers of new interventions should carefully note the demands of specific content areas and identify the component skills that underlie these learning tasks given perceived and known needs of students with LD (Archer & Hughes, 2011; Deshler & Shumaker, 2006; Pressley & Harris, 2006). Multimedia technology offers an avenue for researchers and practitioners to purposefully build instructional materials that meaningfully convey subject-specific content while simultaneously supporting students’ learning needs and motivation to learn (Kennedy & Deshler, 2010).
The purpose of special education is to provide specially designed instruction to help students with exceptionalities access and make progress in the general education curriculum (King-Sears & Bowman-Kruhm, 2010, 2011; Zigmond, 2006). Special and general education teachers frequently face limitations with respect to their own content knowledge, pedagogical repertoire, and cognitive resources when juggling realities of teaching (e.g., large class sizes, limited planning time, accountability pressures, behavioral challenges) with the demands of designing and delivering individualized instruction (Feldon, 2007; Kennedy & Ihle, 2012; McKenzie, 2009). Thus, educators may benefit from interventions that are designed specifically to address the unique learning needs of students with LD and that are packaged to deliver evidence-based instruction in a wide variety of content areas without adding significantly to the already heavy load that teachers carry (Berkeley, Mastropieri, & Scruggs, 2011). The purpose of this article is to describe the results of the first experimental test of an emerging, multimedia-based instructional intervention that may help teachers of students with exceptionalities provide specially designed instruction that supports vocabulary learning within the general curriculum.
Identifying Academic Demands to Inform Instruction for Students With LD
In this article we define students with LD using the definition of the National Joint Committee on Learning Disabilities (1991). Specifically, students with LD who have significant difficulty acquiring and using skills and knowledge related to reading are the subgroup of interest. The rationale for adopting this definition is, in content-area classrooms, most content demands are tied to reading processes and the accompanying cognitive processing tasks (Shanahan & Shanahan, 2008). Students who have inherent difficulty learning from texts are at a substantial disadvantage for finding success on typical measures of academic proficiency (Roberts, Torgesen, Boardman, & Scammacca, 2008). In addition, given traditional general education instructional settings at the secondary level, it is unlikely a student with LD who struggles with reading will receive the type and amount of evidence-based reading instruction needed to improve reading skill and make progress within the content’s standards (Kennedy & Ihle, 2012; Mastropieri et al., 2005). It is for this reason that we explore multimedia as a tool for packaging and delivering evidence-based vocabulary instruction for students with LD. Multimedia provides educators with an opportunity to shape instruction in any needed configuration to meet the demands of the content (Clark, 2009) and also the needs of the learner (Kennedy & Deshler, 2010; Mayer, 2011).
Effective Vocabulary Instruction
Vocabulary knowledge is an example of a high-leverage academic skill (Ebbers & Denton, 2008) and is correlated with comprehension (RAND Reading Study Group, 2002) and other processing tasks required within secondary-level coursework (Roberts et al., 2008). Helping students understand the context-based and often multiple meanings of words so they can comprehend and use words when reading or during other language-rich situations is the primary goal of vocabulary instruction at the secondary level (Baumann, Kame’enui, & Ash, 2003). Thus, for students with LD, significant attention is needed to determine ways to gain sustained engagement with new vocabulary words and concepts embedded within general education courses (Ebbers & Denton, 2008; Jitendra, Edwards, Sacks, & Jacobson, 2004).
Key interventions for vocabulary instruction include (a) helping students become aware of the semantic parts of words (Bos & Anders, 1990), (b) dedicating instructional time to teaching word parts and meanings (Bryant, Goodwin, Bryant, & Higgins, 2003; Nagy, Berninger, & Abbott, 2006; Reed, 2008), and (c) explicitly teaching strategies for forming connections between semantically related terms (Baumann et al., 2003; Graves, 2006). Two categories of vocabulary instruction are necessary to help translate these themes into practice: (a) teaching the definitions of terms and (b) teaching students the skills and strategies needed to decipher word meanings (Graves, 2006; S. A. Stahl & Kapinus, 2001). These two types of instruction are referred to as nongenerative and generative teaching strategies, respectively (Harris, Shumaker, & Deshler, 2011).
For students with LD, both direct instruction in word meanings (nongenerative) and building capacity through the use of strategies (generative) are generally needed for successful learning (Bryant et al., 2003; Ebbers & Denton, 2008; Harris et al., 2011; Jitendra et al., 2004). Unfortunately, in practice, the primary methods used to teach vocabulary terms and concepts to students include mentioning the definition during lecture, expecting they will pick up words’ meanings through reading, and having students look up terms in the dictionary (Kennedy & Wexler, 2013; Lesaux, Harris, & Sloane, 2012). Although reasons why teachers do not use evidence-based practices for vocabulary instruction are not well understood (see Jitendra et al., 2004), we hypothesize the creation and introduction of instructional technology that contains embedded evidence-based instruction may provide one section of the bridge needed to span the divide between research and practice.
Use of Multimedia in Vocabulary Instruction for Students With LD
Given the cognitive needs of adolescents with LD during content-area learning tasks, instructional materials should not be left to chance with respect to their audio and visual makeup. For example, few videos that instructors purchase or download to use during instruction are likely to adhere to a validated instructional model and be appropriate for students with LD (Kennedy, 2011; Kennedy & Wexler, 2013). Multimedia provides educators with a unique opportunity to control every second of instruction by way of planning and selecting the audio and visuals that a learner receives (Mayer, 2011). However, this control alone is never sufficient to guarantee effective instruction (Clark, 2009). Instead, instruction should reflect and incorporate existing evidence-based instruction for a given content demand, but also scaffold cognitive processes so the learner’s cognitive load capacity is not overwhelmed (Kennedy & Deshler, 2010). With that said, the largest limitation of the existing research on instructional technology is the scarcity of detail describing the “looks and sounds” of instruction (Clark, 2009; Mayer, 2011).
To illustrate our point regarding the need to carefully script and shape the “looks and sounds” of multimedia instruction, we offer Richard Mayer’s (2009) cognitive theory of multimedia learning (CTML) and accompanying instructional design principles (Mayer, 2008) as a specific and applied framework to guide practice. This framework consists of 12 evidence-based practices for instructional design (Mayer, 2008) but should not be considered an evidence-based practice on its own for teaching students with disabilities. Instead, this framework is a logical starting point for anchoring instructional practices that have evidence for delivering a specific type of instruction (e.g., vocabulary instruction).
In the following discussion, we use Mayer’s model to deconstruct a sample video from the Khan Academy (www.khanacademy.org) to illustrate how validated instructional design principles, such as those from Mayer, are commonly violated by multimedia acquired via the web. In addition, we criticize Mr. Khan’s video for its lack of adherence to evidence-based practices that are appropriate for individuals with LD. To be true and fair, these videos are not in any way intended for teaching students with LD; however, as teachers use them in that vein (see Fulton, 2012), the criticism is justified, and needed. Following this deconstruction, we detail how Mayer’s model is jointly leveraged along with evidence-based practices for vocabulary instruction to create the intervention tested in this study, content acquisition podcasts (Kennedy, Lloyd, Cole, & Ely, 2012; Kennedy & Wexler, 2013).
Designing Quality Instruction Using Multimedia
Mayer’s CTML
CTML is a learner-oriented instructional theory and empirically validated design process intended to guide the process of creating effective multimedia-based instruction (Mayer, 2009). Mayer’s CTML is an applied theory in that he intends instructional designers and other educators to leverage understandings of how people use visual and auditory inputs when creating new multimedia for various learning tasks. Specifically, his model is operationalized through 12 principles of instructional design that function together as a framework to limit cognitive load and maximize human capacity to learn (DeLeeuw & Mayer, 2008). Each of Mayer’s principles is supported by at least 3 experimental studies of its effectiveness and as many as 17 (Mayer, 2008, 2011). Figure 1 reports the number of studies and mean effect size for each principle based on Mayer’s program of research in this field. For additional detail on each principle and the empirical research supporting this framework, please see Mayer (2009).

Mayer’s Instructional Design Principles as Rubric for Evaluating Multimedia Instructional Materials.
Although Mayer’s theory is well respected, some researchers have questioned its comprehensiveness. To illustrate, Astleitner and Wiesner (2004) point out that the CTML does not specifically address the impact of motivation on a learner’s cognition (p. 11), instead holding to a highly bounded and process-based description of learning and cognition. This is an important limitation, as multimedia can affect learners’ cognitive and emotional interests in unpredictable ways depending on their prior experiences and knowledge and other variables, such as mood (Harp & Mayer, 1997). In other words, any motivation that occurs during multimedia instruction will function to either consume or enhance a viewer’s limited cognitive processes. This is currently unaccounted for in Mayer’s model. Given that researchers in special education have long reported the impact of motivation on the performance of students with disabilities (e.g., Guthrie & Davis, 2003), this is an important aspect of multimedia instruction to be explored in future research.
In the following discussion we use Mayer’s applied theory to illustrate our call for caution when adopting multimedia for use in teaching students with LD and demonstrate how the intervention tested in this article reflects cognitive science and relevant evidence-based practice. In addition, we utilize a simple satisfaction measure as a proxy for motivation to gain preliminary perspective on whether students with LD report being motivated to learn using this new instructional tool.
Deconstruction of Freely Available Internet Videos Used to Teach Students With LD
As noted, it is important to deconstruct any example of multimedia that a general or special education teacher might opt to use with students with (and without) LD. Although deconstruction of videos on the Internet (e.g., the Khan Academy or YouTube) for this purpose is to some extent unfair because of their purpose (or lack thereof), as special educators, it is our job to help shed light on potential mismatches between student learning needs and instructional materials potentially being used to teach them (Kennedy & Wexler, 2013). In addition, we have the responsibility to ensure students with disabilities receive evidence-based, individualized instruction that addresses the goals and objectives noted within their individualized education program (IEP; Zigmond, 2006). Multimedia that does not meet this standard should not be used.
Figure 2 lists Mayer’s (2008) 12 instructional design principles and a side-by-side comparison of the extent to which a sample video from the Khan Academy (http://www.khanacademy.org/humanities/history/euro-hist/v/french-revolution--part-1) adheres to these principles. Figure 1 is the rubric used in this study to evaluate the extent to which multimedia adhere to Mayer’s instructional design principles.

Mayer’s 12 Instructional Design Principles and Corresponding Evaluation of Sample Video from the Khan Academy.
After review, to a large extent, the sample video from the Khan Academy does not adhere to Mayer’s principles for creating multimedia. Given its long length (17:05), significant number and use of history-specific vocabulary terms, lack of explicit instructional cues, redundant use of on-screen text, and constant presence of visuals (pictures and text) that do not correspond to the narration, according to Mayer’s principles, this video will contribute extraneous cognitive load—especially to a viewer unfamiliar with the content being presented.
In addition, this video does not make use of evidence-based practices known to help students with LD learn content. For one, the video contains substantial content about the French Revolution. To illustrate, the narrator repeatedly uses history-specific language conventions and vocabulary terms without explicitly defining them. This content may be difficult to organize and internalize for many learners given the broad scope of this multifaceted topic. The narrator uses a pedagogy of telling, which can be explicit, but falls short of recommendations for explicit instruction noted by Archer and Hughes (2011) and is not strategic, as noted by Deshler and Shumaker (2006) and others (see Pressley & Harris, 2006). There is a substantial assumption of prior knowledge without explicit references to other videos that may contain the information. The video lacks any advance organization that follows a hierarchical ordering system that might help facilitate a learner’s engagement with the content (see Dexter, Park, & Hughes, 2011). In summary, this video does not provide individualized instruction as required by students’ IEPs.
Innovation for Multimedia-Based Vocabulary Instruction
In the present study, an intervention called content acquisition podcasts (CAPs) is tested for the purpose of evaluating its effect on vocabulary performance of adolescents with and without LD. CAPs are multimedia-based instructional modules created using Mayer’s (2008) instructional design principles. For example, Figure 3 presents Mayer’s model, with a description of how CAPs reflect each of these 12 instructional design principles. In contrast to the sample video from the Khan Academy, CAPs ensure all images (text and pictures) correspond to the content being presented, narration is limited to essential content, and a formula is used with respect to helping students develop a routine for recognizing the structure of each video (see http://tecplus.org/files/download/19).

Mayer’s 12 Instructional Design Principles and Corresponding Evaluation of Sample content acquisition podcast (CAP).
Another contrast with the Khan Academy video is that each CAP contains evidence-based vocabulary instructional elements for one critical vocabulary term and lasts approximately 2 minutes. An example CAP can be viewed at www.vimeo.com/39293791. Six specific instructional practices, grounded in the empirical literature on vocabulary instruction (e.g., Bryant et al., 2003; Ebbers & Denton, 2008; Jitendra et al., 2004), constitute a menu of practices embedded into the instructional routine used within CAPs. These include (a) promoting word consciousness (e.g., pronunciation, spelling, syllables, prefix, suffix, root words; Reed, 2008), (b) providing direct instruction of word meanings (Archer & Hughes, 2011), (c) providing guided practice and scaffolding (Dexter et al., 2011), (d) providing instruction that promotes awareness of closely related terms (Graves, 2006), (e) using the keyword mnemonic strategy (Mastropieri, Scruggs, & Levin, 1987), and (f) providing a statement of purpose/rationale for why the student needs to learn a given term or concept (Deshler & Shumaker, 2006). These six elements of effective vocabulary instruction constitute a checklist used to create CAPs (see Kennedy, Lloyd, et al., 2012). When combined with Mayer’s instructional design principles, the use of these evidence-based practices for vocabulary instruction may form the base for a multimedia-based practice that can be tested for use with students.
The keyword mnemonic strategy is utilized within CAPs given its strong empirical record for improving vocabulary performance in various content areas (see Mastropieri, Scruggs, Bakken, & Whedon, 1997; Mastropieri, Sweda, & Scruggs, 2000; Scruggs et al., 2011; Scruggs, Mastropieri, Berkeley, & Marshak, 2010) and use of imagery, which lends itself to multimedia-based instruction (Austin, 2009). This is critical because empirical evidence in the field of vocabulary instruction shows that teachers should use a blend of methods, including explicit and strategic instruction, to achieve maximum learning effects (Bryant et al., 2003; Ebbers & Denton, 2008; Jitendra et al., 2004). One limitation of CAPs that use the keyword mnemonic strategy is that there is not an automated rehearsal mechanism built in, which is an important element of this strategy (Mastropieri et al., 1987). Although not used in this way during this study, practitioners who use CAPs in the future should follow up with students to ensure the rehearsal of the keyword and definition is provided.
In this study, CAPs were produced using Microsoft’s PowerPoint (see Kennedy, Hart, & Kellems, 2011, and Kennedy & Thomas, 2012, for specific production steps). Although CAPs differ slightly from the textbook definition of podcasts (audio recording or synced audiovisual recording uploaded to the Internet), our rationale for referring to this learning tool as podcasts is marketing to general education teachers. In other words, because most practitioners know what podcasts are, and possibly use generic podcasts in their teaching, we argue that including podcast in the name of this new tool may facilitate adoption and usage. In addition, although not specifically used in this way during this study given our research questions, we intend for CAPs to be uploaded to the web, and downloaded by students for use on their iPad, iPhone, or other portable device.
Thus, the purpose of this experimental study is to explore the use of CAPs as a tool to improve vocabulary learning for adolescents with LD enrolled in rigorous secondary-level coursework. Four configurations of CAPs with embedded evidence-based vocabulary instruction were tested to determine the most efficient and effective combination of multimedia and vocabulary instruction: (a) explicit instruction alone, (b) the keyword mnemonic strategy alone, (c) combination of explicit and strategy (keyword mnemonic strategy) instruction, and (d) explicit instruction without adherence to Mayer’s model. In summary, CAPs were used to answer these research questions: To what extent does multimedia instruction that adheres to various instructional design principles and is combined with evidence-based vocabulary instruction promote vocabulary learning for adolescents with and without LD? Do students who learn using various forms of CAPs report satisfaction with this learning method?
Method
Setting and Participants
The University Human Subjects Committee, the participating school district’s research review board, the principal of the school, parents of all students, and the students gave permission to conduct this research. The school district is located in an urban, Midwestern community of 146,867 residents. The researchers recruited two world history teachers responsible for teaching 12 total sections of world history to approximately 300 students to participate in the study. A total of 278 urban high school students (9th to 12th graders) enrolled in a world history course participated in the study. African American students represented the largest ethnic group (67.5%), Caucasian students were the next largest group at 21.9%, and Hispanic students constituted 8.1%. Of the 278 participants, 52% were female, 48% were male, and 91% were in 10th grade. The mean age of participants was 16.7 years. At the time of the study, the selected high school had a student enrollment of 987, 78% of whom received free and/or reduced-price lunch. Permission to collect individual socioeconomic status could not be obtained from the school district’s human subjects review board. However, given that nearly every 10th grader in the school is enrolled in one of the 12 sections of world history participating in this project and 78% of students at this school received free or reduced-price lunch, we assumed that approximately three quarters of students received free or reduced-price lunch.
Two groups of students participated: (a) students with LD in a specific area related to reading (n = 30) and (b) students without disabilities and students who receive special education services for a reason other than a reading disability (n = 248). All students in the LD group had an IEP stemming from a diagnosis of specific learning disability related to reading, which manifests as difficulty conducting cognitive processes necessary for reading. These 30 students were in the 10th grade at the time of the study and had a mean age of 16.9 years. In terms of demographic characteristics, 80% were male and 20% were female; 63.3% were African American, 26.7% were Hispanic, and 10.0% were Caucasian. The mean Wechsler Intelligence Scale for Children–IV IQ score for the 30 students with LD was 93.4 (SD = 9.3). Furthermore, the students scored homogeneously on the reading subtests (Word Reading, Reading Comprehension, and Pseudoword Decoding) of the Wechsler Individual Achievement Test–Second Edition. Using grade-based norms, the students’ mean standard score for the Word Reading subtest was 68 (SD = 4.3), which is the second percentile. The mean standard score for the Reading Comprehension subtest was 63 (SD = 9.1), which is the first percentile; and the mean standard score for the Pseudoword Decoding subtest was 66 (SD = 4.8), which is also the first percentile. Using grade-based norms, the mean reading composite score for these 30 students was 176 (SD = 18.2), which translates into a standard score of 51 and less than the first percentile.
All students received daily special education services embedded within their core academic content classes taught by a general education teacher (e.g., social studies, science, mathematics, and language arts) and also had a study skills course taught by a special educator. Although some students without LD also had an IEP (primarily students with an emotional/behavioral disorder diagnosis), they were grouped with the students without disabilities for the purpose of this study’s activities and analyses. The rationale for this decision is the desire to study the effects of the intervention on students who have documented disabilities specific to the demands of reading.
Using a table of random numbers, researchers randomly assigned students into one of four experimental conditions. Two stratification variables proportionately sorted students into the four groups: disability status (LD or not LD) and achievement status (high, typical, low). The researchers used students’ first-semester GPA in world history to sort students into one of three levels of achievement status: (a) high achiever—85% or above, (b) typical achiever—84% to 70%, and (c) low achiever—69% and below. The researchers stratified students by achievement status given the expectation that high achievers would likely learn more vocabulary content than typical or low achievers regardless of how they are taught. Therefore, it was important to proportionately distribute these students across the four experimental conditions.
Methods for Creating CAPs
CAP adherence to Mayer’s CTML
Two reviewers each with 20 hours of experience in producing and evaluating CAPs for a previous study (e.g., Kennedy & Thomas, 2012) used a production rubric to independently score the CAPs for the present study (see Figure 1). Reviewers completed one rubric per CAP (n = 30) to gauge adherence to Mayer’s instructional design principles. Therefore, reviewers used understanding of these principles and their judgment to complete this task. Feedback from the reviewers informed revisions prior to use in the study. Interscorer reliability across the reviewers was 95%. The researcher made revisions based on the feedback and then invited reviewers to review the CAPs a second time. CAPs were not used in the study until all concerns were satisfied.
Differences among the four experimental conditions
Grounded in the empirical literature for teaching vocabulary to students with LD (e.g., Bryant et al., 2003; Ebbers & Denton, 2008; Jitendra et al., 2004), researchers made the decision to test effects on student learning when explicit instruction and the keyword mnemonic strategy are provided in isolation and then combined. Although most researchers agree that combinations of explicit and strategy instruction are the most effective for teaching students with LD, for the purpose of this experiment this assumption is tested by the use of multimedia, and thus these two types of vocabulary instruction are tested together and in isolation.
Researchers created four types of multimedia-based vignettes that contain various combinations of evidence-based practices for the purpose of determining the most powerful combination of theoretically valid instructional design features and practices for learning vocabulary terms. The four experimental conditions are (a) CAPs, containing explicit instruction only (EI); (b) CAPs with the keyword mnemonic strategy only (KMS); (c) CAPs containing explicit instruction and the keyword mnemonic strategy (EI + KMS); and (d) instructional videos with the same audio narration as students in the EI group, but did not adhere visually to Mayer’s design principles (NM). The CAPs for the first three conditions feature judicious use of on-screen text, frequent use of vivid images, and meticulously scripted words delivered during narration. The vignettes created for the fourth condition are generic enhanced podcasts (audio synced with text-based slides), and do not adhere to Mayer’s principles, but make use of the same audio narration used in the EI condition’s videos.
Students assigned to watch CAPs that incorporate EI viewed vignettes that provide (a) rationales for why learning the given term is important, (b) direct instruction of word meanings, (c) awareness of closely related terms, (d) guided practice and scaffolding, and (e) word consciousness (e.g., identify morphemes in terms). Students in the KMS condition also received (a) rationale for why learning that term is important, (b) direct instruction in word meaning, (c) rationales for why the KMS is a good tool for remembering vocabulary terms, and (d) an acoustically similar remembering word (keyword) along with an image of the keyword interacting with the definition of the term. Students in the combined EI and KMS condition saw CAPs with all of the aforementioned instructional practices and the same narration. Students in the final, NM condition saw multimedia-based vignettes that contain the same content and narration as the EI condition, but the content was presented as text only, instead of images and occasional text. The same two independent reviewers noted above scored each CAP to ensure each practice was embedded within each CAP. Interscorer reliability across reviewers was 98%.
With respect to length, students in the EI and NM conditions watched vignettes that contain the exact same audio track and thus have the same running time. Students in the KMS condition watched vignettes that, on average, are approximately 10 seconds shorter than the EI CAPs. Students in the EI + KMS watched the longest vignettes, which, on average, are approximately 30 seconds longer than the EI CAPs. Although students in the EI + KMS viewed the longest CAPs and thus, received the most instruction, there is reason to hypothesize that “bigger is not necessarily better” for students with LD when it comes to duration of instruction (Deshler & Shumaker, 2006; Swanson, 2001, 2009). Therefore, an important question undertaken by this study is related to the duration of multimedia-based instruction for teaching vocabulary terms and concepts. With that said, it is not trivial that the students in the EI + KMS group received more instruction for each term. Any significant differences between the groups would at least in part be attributable to this difference in duration of instruction.
Selection of vocabulary terms/concepts
Researchers and two teachers recruited to participate in this study selected 30 vocabulary terms and/or concepts to be turned into CAPs. This group reviewed a list of all relevant vocabulary terms/concepts for the World War I unit based on a review of the course textbook, district curriculum, and state standards. The selection of terms was guided by two principles. First, researchers selected vocabulary terms that are similar in terms of complexity and conceptual density. Table 1 contains the list of terms. The participating teachers used their professional judgment and prior experience teaching this unit to judge a term’s complexity and density and ease of student acquisition. The second selection criterion was that terms needed to be represented clearly with visuals and translated into a keyword for use within the KMS. The teachers provided scripts for definitions and related content that students needed to learn based on the curriculum and standards. Finally, terms were selected based on the teachers’ course calendars to ensure the first in-class exposure to the terms/concepts would come through this study’s research activities.
List of Vocabulary Terms/Concepts Used to Create Content Acquisition Podcasts.
Measures Used to Evaluate Student Vocabulary Performance and Satisfaction
Multiple-choice instrument
The researchers created a multiple-choice (MC) instrument with 30 items to measure students’ ability to use their knowledge to identify correct definitions for important vocabulary terms and concepts from a world history unit. The score range is 0 to 30. The MC instrument includes items that correspond to the 30 vocabulary terms/concepts selected by the two history teachers. The stem for each item simply includes the term and the appropriate verb (e.g., “Imperialism is . . .”). The answer choices and distractors for each item are definitions from the textbook glossary. The decision to use definitions as answer choices and distractors corresponds to a common reading requirement within most high school history courses (VanSledright, 2008) and research on curriculum-based measures (e.g., Espin, Busch, Shin, & Kruschwitz, 2001). Answer choices were selected based on length (number of words), relevance to the correct answer (as distractors), and language density (ease of reading).
The construction of this instrument reflects best practice for MC item construction as detailed by Haladyna, Downing, and Rodriguez (2002). Three experts in world history reviewed each of the MC items for difficulty, clarity, and errors in content or grammar and provided comments for revision. The two teachers in the study also reviewed the items and provided comments for revision. Cronbach’s alpha for the MC items at posttest was .87, thereby demonstrating adequate internal consistency for use in this study.
Open-ended instrument
The second instrument was open ended (OE). Its purpose is to evaluate students’ ability to produce a definition for the term in writing and also to probe deeper knowledge of terms (e.g., synonyms, antonyms) and any contextual understanding based on knowledge provided within each CAP. Specifically, the OE instrument asked students to “write what you know” about each of the 30 terms/concepts. The score range for the OE instrument is 0 to 60. To add structure, the test form provided space for students to write (a) the definition, (b) a synonym for the term, (c) an antonym, and (d) any additional information they know about the concept. Thus, this instrument required much more than simple matching, a form of vocabulary assessment that has been widely criticized (K. A. Stahl & Bravo, 2010). The alpha level for the OE instrument was .95. An answer key of acceptable responses for each component was developed and used in scoring. The researcher and each of the two teachers scored student responses using a rubric of acceptable answers for the purpose of establishing interscorer reliability. Following discussion, 100% agreement was achieved.
Satisfaction survey
A satisfaction survey was given to all participating students following completion of the maintenance probe. The survey took approximately 5 minutes to complete and contains eight Likert-type items. The scale for the Likert-type items was 1 to 10, with a score of 1 designating a response of strongly disagree and a score of 10 designating that the respondent strongly agreed with the statement or question. Students were free to select any score along the range from 1 to 10. The use of a 10-point scale allows larger differentiation and interpretation of responses than is possible with typical 5- or 7-choice Likert-type items (Fowler, 2009).
The items in the survey reflect two constructs relevant to this study’s research questions: (a) ease and function of technology within CAPs and (b) usefulness of CAPs for learning new vocabulary terms/concepts in world history. Because students remained in the same experimental group throughout the study, they were asked to note which of the four groups they were in for the experiment. This allowed survey responses to be sorted by group and analyses of responses based on the version of the CAPs that were watched. The survey was reviewed by three doctoral students, who were asked to provide feedback on wording of questions, order of questions, and overall quality of the instrument. Feedback was used to make updates to the survey. The reliability alpha for the survey was .73.
Procedure for Watching CAPs and Taking Assessments
The first author and the two teachers used a checklist to ensure fidelity during research activities. The checklist contains procedures for the research activities for each day, including the script for directions that was read to students. To begin research activities, students completed a pretest during their regularly scheduled class, which is the same instrument as the posttest and maintenance probes. Students were informed that the pretest would not count toward their grade, but would be used as a class work/participation grade.
One week later, students watched an orientation CAP that introduced this format of instruction and cued them to its various features. Students watched CAPs at individual laptop terminals while wearing headphones. CAPs were loaded onto the school’s intranet and sorted into password-protected folders for each experimental group. Students watched a total of 10 CAPs during one class period. The first author and teacher circulated the room to ensure students were watching the CAPs and not navigating to other websites or programs.
After finishing the fifth video, students were to raise their hand and close their laptop. The teacher handed students the OE instrument, which asked students to (a) write the definition for each term/concept, (b) provide a synonym/antonym, and (c) provide any other related information for each term/concept. When students completed this instrument, they raised their hands and the teacher collected the paper and handed the students the MC assessment. The decision to give students the OE instrument first was purposeful to eliminate the possibility of being cued to the answer with the MC instrument. After completing the MC items, students again raised their hands, the paper was collected, and students watched the next five CAPs, and the process repeated. In sum, students watched a total of 30 CAPs, in the order of 10 per day across 3 days of the experiment.
Maintenance Procedure
The first author returned to the school approximately 3 weeks after concluding the posttest for all 30 vocabulary terms to administer a maintenance probe to measure durability of learning using CAPs. Following the students’ initial learning of 30 terms using CAPs, the two teachers retaught Terms 1 to 20 as a part of their required course curriculum. The researcher did not monitor how the terms were retaught to students. As a result, Terms 1 to 20 were assessed only at posttest. However, Terms 21 to 30 were not retaught or assigned as part of assignment during this period of time. Thus, the maintenance probe consisted of an MC test to evaluate student knowledge of Terms 21 to 30; the probe is the same as the posttest measure.
Research Design
Researchers conducted main and secondary analyses for this study. There are two parts of the main analysis. In Part 1, the between-subjects factors for students with LD are four approaches to delivering vocabulary instruction with multimedia (EI, n = 7; KMS, n = 8; EI + KMS, n = 7; NM, n = 8). The four groups are evaluated on a dependent measure with two levels: performance on the pretest and posttest for Items 1 to 30 (evaluated using a 4 × 2 split-plot, fixed-factor repeated measures ANOVA). To create the one dependent variable score, researchers respectively standardized and then averaged student scores on the MC and OE instruments. In the second part of the main analysis, the between-subjects variables are the same as above; the dependent measure has three levels: performance on the pretest, posttest, and maintenance probes for Items 21 to 30 (evaluated using a 4 × 3 split-plot, fixed-factor repeated measures ANOVA).
In the secondary analysis, the above analyses are repeated, but for students without LD (between- and within-subjects variables are the same; EI, n = 60; KMS, n = 62; EI + KMS, n = 63; NM, n = 63). The reason for including this secondary analysis here is twofold. First, because of the small sample size and thus limited statistical power in the four groups among students with LD, it is important to confirm observed results using a larger group (albeit with students without LD). Second, if there were any effects for any of the types of multimedia instruction when comparing students without LD, that intervention would potentially be more attractive to a general education teacher.
Finally, results from the satisfaction survey were analyzed using quantitative data from the Likert-type items. Responses are organized by group assignment and compared to one another using a series of one-way ANOVAs for the purpose of evaluating differences and/or emerging trends that may be attributed to the different versions of the CAPs.
Results
Main Analysis, Students With LD: Part 1
In this analysis, researchers conducted a 4 × 2 split-plot, fixed-factor repeated measures ANOVA. The between-subjects factor was group assignment (group), and the two levels of the within-subject factor (time) were performance on the pretest and posttest for Terms 1 to 30. Levene’s test for equality of variances was not significant for this evaluation. Given the random assignment of students to experimental groups, there was no further need to statistically control for between-group variance and differences. The raw score means, standard deviations, p values, and effect sizes (Cohen’s d) for the pretest, posttest, and maintenance probe are listed in Tables 2 and 3. There was not a significant main effect for group, F(3, 26) = 1.7, p = .187, ω2 = .03, or time, F(1, 26) = > 1, p = .516, ω2 = .00; however, the interaction between group and time, F(3, 26) = 13.0, p < .000, ω2 = .26, was significant. Therefore, 27% of the variance in this model is explained by the sum of effects from the predictor variable (EI, KMS, EI + KMS, NM) in the experiment.
Raw Mean Scores and Standard Deviations for Students With Learning Disabilities (LD) and Without LD (Not LD) on the Multiple Choice and Open-Ended Pretest and Posttest (Terms 1–30) and Maintenance Instruments (Terms 21–30).
Note. The pretest and posttest MC has a score range of 0–30; the pretest and posttest OE has a score range of 0–60. The maintenance instrument has a score range of 0–10 for MC and 0–20 for OE. LD = students with an LD in an area related to reading; MC = multiple-choice instrument; Not LD = students without a learning disability in an area related to reading; OE = open-ended instrument.
Maintenance Mean Scores and Standard Deviations for Multiple-Choice and Open-Ended Instruments.
Note. The MC instrument has a score range of 0–10; the OE instrument has a score range of 0–20. MC = multiple-choice instrument; NSWD = students without a learning disability (LD); OE = open-ended instrument; SWD = students with LD.
To examine the direction and location(s) of the significant interaction, researchers conducted post hoc pairwise comparisons using Tukey’s honestly significant difference (HSD) test. A Bonferroni correction was used to evaluate the individual group differences at the two time points (.05/2 = .025). These results demonstrate that there was not a significant difference in scores between any group of students with LD on the pretest; however, students with LD in the EI +KMS group (MPost = 1.45, SD = 2.7) had significantly higher scores on the posttest than students with LD in the NM group (MPost = −2.8, SD = 2.3; d = 1.97). No other post hoc comparisons were statistically significant at the .025 level; however, the mean scores show the EI + KMS group (MPost = 1.45, SD = 2.7) scored higher on the posttest than the students with LD in the EI (MPost = −0.92, SD = 1.5; d = 1.09) and KMS (MPost = −1.6, SD = 1.6; d = 1.40) groups. These large effect sizes should be interpreted with caution given the limited statistical power and use of a researcher-created instrument; however, they do provide preliminary evidence that for students with LD learning these vocabulary terms using the full CAP model had an impact on their learning.
Main Analysis, Students With LD: Part 2
In this analysis, researchers conducted a 4 × 3 split-plot, fixed-factor repeated measures ANOVA. The between-subjects factor was group assignment (group), and the three levels of the within-subject factor (time) were performance on the pretest, posttest, and maintenance probe for Terms 21 to 30. Levene’s test for equality of variances was not significant for this evaluation. However, Mauchly’s test of sphericity was significant, Mauchly’s W = .551, p = .001; thus, Greenhouse–Geisser corrections are used to evaluate and interpret results. There was a significant main effect for group, F(3, 26) = 6.3, p < .002, ω2 = .21, and the interaction between group and time, F(4.1, 52) = 8.9, p < .000, ω2 = .18. However, the effect for time, F(2, 52) = > 1, p = .460, ω2 = .00, was not significant. Therefore, 39% of the variance in this model is explained by the sum of effects from the predictor variable (EI, KMS, EI + KMS, NM) in the experiment.
To examine the direction and location(s) of the significant interaction at maintenance, researchers conducted post hoc pairwise comparisons using Tukey’s HSD test. A Bonferroni correction was used to evaluate the individual group differences at the three time points (.05/3 = .017). Student performance on the posttest was evaluated for the 30 vocabulary terms in the main analysis, and so for the sake of brevity the results are not reported here for Items 21 to 30.
At maintenance, students with LD in the EI +KMS group (MMaint = 1.1, SD = 1.2) had significantly higher scores on the maintenance probe than students with LD in the NM group (MMaint = −1.4, SD = 1.3; d = 2.40). No other post hoc comparisons were statistically significant at the .017 level; however, mean scores show the EI + KMS group (MMaint = 1.1, SD = 1.2) scored higher on the posttest than the students with LD in the EI (MMaint = −0.87, SD = 0.75; d = 1.96) and KMS (MMaint = −0.34, SD = 0.71; d = 1.45) groups. These large effect sizes also should be interpreted with caution given the limited statistical power and use of a researcher-created instrument; however, there is preliminary evidence that students with LD who learned using the full CAP model made gains in vocabulary performance.
Secondary Analysis, Students Without LD: Part 1
To mirror the main analysis but with students without LD (to help confirm findings given limited statistical power), researchers again conducted a 4 × 2 split-plot, fixed-factor repeated measures ANOVA. The between-subjects factor was group assignment (group), and the two levels of the within-subject factor (time) were performance on the pretest and posttest for Terms 1 to 30. Levene’s test for equality of variances was not significant for this evaluation. Given the random assignment of students to experimental groups, there was no further need to statistically control for between-group variance and differences. The raw score means, standard deviations, p values, and effect sizes (Cohen’s d) for the pretest, posttest and maintenance probe are listed in Tables 2 and 3. There was a significant main effect for group, F(3, 243) = 6.7, p < .001, ω2 = .05, and the interaction between group and time, F(3, 243) = 25.3, p < .001, ω2 = .05. However, the evaluation for time, F(1, 243) = > 1, p = .827, ω2 = .00 was not significant. Therefore, 10% of the variance in this model is explained by the sum of effects from the predictor variable (EI, KMS, EI + KMS, NM) in the experiment.
To examine the direction and location(s) of the significant interaction, researchers conducted post hoc pairwise comparisons using Tukey’s HSD test. A Bonferroni correction was used to evaluate the individual group differences at the two time points (.05/2 = .025). In a repetition of findings from the main analysis, these results demonstrate that there was not a significant difference in scores between any group on the pretest; however, students without LD in the EI +KMS group (MPost = 1.8, SD = 2.4) had significantly higher scores on the posttest than students with LD in the NM group (MPost = −1.7, SD = 2.6; d = 1.41). No other post hoc comparisons were statistically significant at the .025 level; however, the mean scores show the EI + KMS group (MPost = 1.8, SD = 2.4) scored higher on the posttest than the students with LD in the EI (MPost = 0.30, SD = 2.6; d = 0.61) and KMS (MPost = 0.10, SD = 2.3; d = 0.74) groups. These medium to large effect sizes should be interpreted with caution given the use of a researcher-created instrument, however, they do provide corroborating evidence that students who learned these vocabulary terms using the full CAP model (EI + KMS) made gains in performance relative to other students who did not receive the full intervention or a different intervention.
Secondary Analysis, Students Without LD: Part 2
Again to mirror the main analysis for students with LD, researchers conducted a 4 × 3 split-plot, fixed-factor repeated measures ANOVA. The between-subjects factor was group assignment (group), and the three levels of the within-subject factor (time) were performance on the pretest, posttest, and maintenance probe for Terms 21 to 30. Levene’s test for equality of variances was not significant for this evaluation. However, Mauchly’s test of sphericity was significant, Mauchly’s W = .705, p = .001; thus, Greenhouse–Geisser corrections are used to evaluate and interpret results. There was a significant main effect for group, F(3, 243) = 14.672, p < .001, ω2 = .10, and the interaction between group and time, F(4.6, 375) = 19.2, p < .000, ω2 = .06. However, the effect for time, F(1.5, 375) = > 1, p = .914, ω2 = .00 was not significant. Therefore, 16% of the variance in this model is explained by the sum of effects from the predictor variable (EI, KMS, EI + KMS, NM) in the experiment.
To examine the direction and location(s) of the significant interaction at maintenance, researchers conducted post hoc pairwise comparisons using Tukey’s HSD test. A Bonferroni correction was used to evaluate the individual group differences at the three time points (.05/3 = .017). Student performance on the posttest was evaluated for the 30 vocabulary terms in the main analysis and so for the sake of brevity will not be reported here for Items 21 to 30. At maintenance, students with LD in the EI +KMS group (MMaint = 1.1, SD = 1.6) had significantly higher scores on the maintenance probe than students without LD in the EI (MMaint = 0.10, SD = 1.3; d = 0.66), KMS (MMaint = −0.08, SD = 1.3; d = 0.80), and NM (MMaint = −0.86, SD = 0.99; d = 1.47) groups. In addition, students without LD in the EI and KMS groups, respectively, significantly outperformed students without LD in the NM group (d = 0.81; 0.69) at maintenance. These medium to large effect sizes should also be interpreted with caution given the use of a researcher-created instrument; however, this finding corroborates the observed impact of the CAP intervention for students with LD noted above.
Student Satisfaction Survey
Table 4 reports the data from the satisfaction survey. Because students completed the survey at the conclusion of the maintenance probe, and students with LD were embedded within the general education classroom for the experiment, it was not possible to isolate those students’ survey responses. Thus, their responses are included within the larger group disaggregated by treatment condition (EI + KMS, EI only, KMS only, NM).
Student Satisfaction Survey Results.
Note. Item scores range from 1 to 10. EI = explicit instruction; KMS = keyword mnemonic strategy; NM = non-Mayer.
Does not include Item 8, which was answered only by students in Groups 1 and 3.
Based on the results of the survey, few students reported having technical problems when watching CAPs (M = 0.08, SD = 0.276). On average, students reported that the narrator was easy to understand (M = 8.59, SD = 1.78) and the CAPs looked good with respect to pictures and on-screen text (M = 8.2, SD = 1.95). Mean scores demonstrate that students tentatively agreed that the content of the CAPs was interesting (M = 6.32, SD = 2.16). On average, students reported that they do not have a difficult time learning new vocabulary terms in history (M = 4.97, SD = 2.56); however, students generally agreed that the CAPs helped them learn the meaning of the terms (M = 7.38, SD = 2.06) and that the CAPs prepared them to do well on the posttest and maintenance probes (M = 6.91, SD = 2.16). Finally, students in the EI + KMS and KMS only groups, on average, agreed that the KMS is helpful in learning the meaning of vocabulary terms (M = 7.58, SD = 2.19).
Post hoc analyses of student survey responses were completed using one-way ANOVAs. Tukey post hoc comparisons of the four groups’ survey responses for Item 3 (The podcasts looked good) indicated that the students in the EI + KMS group (n = 70, M = 8.53, 95% CI = 8.11, 8.95), EI only group (n = 68, M = 8.41, 95% CI = 7.96, 8.86), and KMS only group (n = 70, M = 8.66, 95% CI = 8.31, 9.00) all expressed significantly higher satisfaction with the visual elements of the CAPs than did students in the NM group (n = 70, M = 7.26, 95% CI = 6.70, 7.81), p < .000, p < .002, and p < .000, respectively. In addition, students in the EI + KMS group (n = 69, M = 7.91, 95% CI = 7.42, 8.41) expressed significantly higher satisfaction on Item 6 (The podcasts helped me learn the meanings of the vocabulary terms) than students in the NM group (n = 69, M = 2.13, 95% CI = 6.26, 7.28, p < .006). Finally, students in the EI + KMS group (n = 70, M = 7.69, 95% CI = 7.26, 8.11) expressed significantly higher satisfaction on Item 7 (After watching the podcasts I was ready to do well on the quizzes) than did students in the NM group (n = 70, M = 6.17, 95% CI = 5.64, 6.69, p < .000).
Discussion
General and special education teachers should seek opportunities to use various instructional techniques and strategies to support individual learning needs of students with LD (Swanson, 2001), especially during rigorous secondary-level content-area coursework (Deshler & Shumaker, 2006). Although it can be difficult to fully understand the complex learning needs of students with LD and then translate that understanding into individualized instruction given competing demands (e.g., curriculum pacing, planning), this is the meta-space educators must occupy to support the learning of children with the significant learning needs (Kennedy & Deshler, 2010; King-Sears & Bowman-Kruhm, 2010; Zigmond, 2006). The results of this study confirm the need to leverage evidence-based teaching practices with delivery mechanisms that do not solely present content in a novel way. Instructional technology, such as CAPs, may provide a functional mechanism for designing and delivering this type of individualized instruction. This does not preclude the use of technology in an assistive role (see King-Sears, 2009, for an excellent discussion) but does provide compelling rationale to invest resources into ways instructional technology can deliver individualized instruction for a variety of critical content areas. Teachers should use curriculum-based measures to evaluate the ongoing progress of students to determine the extent to which current approaches to instruction are working (Espin et al., 2001). Using a progress monitoring system like that described by Espin and her colleagues (2001) makes particular sense for evaluating the impact of vocabulary instruction delivered using multimedia.
The results of this study are consistent with those of previous studies in terms of (a) positive effects of multimedia-based instruction based on Mayer’s CTML and instructional design principles (Kennedy, Ely, et al., 2012; Kennedy et al., 2011; Kennedy & Thomas, 2012), (b) successful use of multimedia to teach vocabulary terms and concepts to adolescents (Xin & Rieth, 2001), and (c) augmented performance on measures of vocabulary learning for adolescents with LD following use of a blend of explicit and strategic approaches (Bryant et al., 2003). In addition, this study provides preliminary evidence that extends existing theories of multimedia learning and evidence-based practices for vocabulary instruction into new space in the name of augmenting academic skills and outcomes for all students. With that said, some researchers have criticized Mayer’s theory (e.g., Astleitner & Wiesner, 2004) for its lack of explanation regarding the impact of motivation on a learner’s cognition. Important next steps for this line of research are to further explore the extent to which the motivation to learn using multimedia of students with LD correlates and/or predicts achievement.
Limitations
This study has several limitations. First, although an experimental design was used, only 279 students participated. In addition, only 30 (9.3%) students with LD participated. Although these are not small numbers in social science research, the students were enrolled in one high school, thus potentially representing an overly homogeneous group.
Second, the researcher created all of the CAPs used in the study. Although the CAPs were created using a production rubric based on Mayer’s CTML and a checklist for effective elements of vocabulary instruction and were reviewed by experienced colleagues, important questions remain about the ability of other teachers or researchers to create CAPs. This is an important question to be answered by future research. Furthermore, instruction was provided using individually issued laptops and headphones to all students. The availability of laptops on a 1:1 ratio is unlikely in many schools; therefore, the controlled nature of this experiment to some extent limits its external validity. One possible way to address this limitation is the template and process used in this study could be captured in an application that would allow teachers and students to easily create their own CAPs (or other multimedia).
Third, the researchers created the measures used in the study. Standardized measures of vocabulary knowledge for specific content areas (e.g., world history) do not exist, and other standardized measures were not appropriate for use in the study given the research questions. Furthermore, given the limited scope of this experiment with respect to terms that were taught as well as the duration of the study (approximately 3 weeks), growth on a standardized measure would likely be unachievable. An important question to be addressed by future research is the extent to which CAPs used across a semester or entire year may affect achievement.
Implications for Future Research and Practice
Despite limitations and the preliminary nature of its findings, this study has important implications for research and practice. With respect to research, future studies should leverage Mayer’s CTML and instructional design principles to create and then test multimedia-based instructional materials for a wide assortment of content areas. For example, this study is restricted to vocabulary terms and students enrolled in a world history course. Future explorations should be expanded to other courses within social studies (e.g., U.S. history) as well as other subject areas that require substantial vocabulary knowledge (e.g., science, mathematics, foreign languages). In addition, in this study the primary objective was to promote vocabulary learning among students with LD; however, students with other types of disabilities may also benefit from multimedia instruction designed using validated instructional design principles.
For reasons of experimental control, researchers used a clinical approach in this first study of the utility of CAPs to augment vocabulary knowledge. Future research to be conducted in partnership with practitioners should explore socially valid methods for using CAPs to deliver vocabulary instruction. This may include teachers using CAPs during large-group lectures or assigning students to watch CAPs at home or during other study times in and out of school. CAPs may then be utilized in response to intervention frameworks as a Tier 1, 2, or 3 intervention. Related to this, studies in which other researchers and teachers create CAPs and evaluate their impact on student learning should be conducted. Another interesting opportunity to extend the validity of the intervention is for students to participate in the production of CAPs. Comparison of student performance on various dependent measures following viewing CAPs created by teachers, researchers, and students is important for future research to address.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
