Abstract
Among the variety of selected response formats used in L2 reading assessment, multiple-choice (MC) is the most commonly adopted, primarily due to its efficiency and objectiveness. Given the impact of assessment results on teaching and learning, it is necessary to investigate the degree to which the MC format reliably measures learners’ L2 reading comprehension in the classroom context. While researchers have claimed that the longer the reading test (i.e., more test items and passages), the higher its overall reliability, few studies have investigated the optimal number of items and passages required for reliable classroom-based L2 reading assessment.
To address this research gap, I adopted generalizability (G) theory to investigate the score reliability of the MC format in classroom-based L2 reading tests. A total of 108 ESL students at an American college completed an English reading test that included four passages, each of which was accompanied by five MC comprehension questions. The results showed that the score reliability of the L2 reading test was critically influenced by the number of items and passages, inasmuch as a different combination of the number of passages and items altered the degree of reliability. Implications for practitioners and educational researchers are discussed.
Keywords
Introduction
In L2 reading assessment, the selected response format is commonly adopted in standardized and classroom assessment due to its capacity to cover a large domain of test content and effectively represent the knowledge, skills, and abilities of the construct being assessed (Downing, 2006; Haladyna & Rodriguez, 2013). Because selected response items are scored objectively, inter-rater reliability is not a concern; as such, the threat of scoring subjectivity is negligible (Green, 2020; Plakans & Gebril, 2015). Moreover, technology (i.e., scoring machines and computer algorithms) can be used to ensure the scoring efficiency of selected response items, thereby lowering the cost of hiring human raters (Linn, 2006).
This does not mean that selected response test items are always preferred; indeed, researchers have criticized various aspects of selected response item formats. First, even though the scoring of these items is objective, humans must decide what selected response test items to ask. Item writers predetermine what the correct and incorrect answers will be, meaning that decisions regarding what to test and what to accept as a correct answer are subjective. Moreover, the item writers tend to shape their items into questions that test factual or procedural knowledge, which is not effective for assessing higher order knowledge (Hughes, 2003). More importantly, examinees who are tasked with selected response items may not be actively engaged in the test content because they are more focused on choosing the correct option than on applying their knowledge (Cohen & Upton, 2007). For this reason, great care must be taken when developing selected responses tests to ensure that the tests fairly and reliably measure the examinees’ higher and lower order knowledge and performance abilities (see FairTest, 2007). Assuming that the selected response items are carefully and appropriately developed (and that the use of test scores from selected response tests are mostly formative and have low to mid stakes), selected response testing can generate meaningful and positive impact, especially in terms of teacher time management, test scoring, and feedback efficiency.
Test developers use a variety of selected response formats (e.g., true-false, multiple-choice, and matching) to assess students’ mastery of the target content area and potentially minimize the occurrence of random guessing. Among these question formats, multiple-choice (MC), which requires examinees to select the answer from a list of prescribed options, is the most popular format (Haladyna & Rodriguez, 2013). Despite its popularity, however, there are certain challenges that arise when creating reading items in the MC format.
This is especially the case for classroom teachers because constructing items often requires extensive training and expertise to ensure quality. In this vein, researchers in the field identified a number of factors that must be taken into account when developing test items, including students’ reading proficiency, the number of prescribed options per item, the quality of distractors, and test length. In more specific terms, the test items should be well-written and comprehensible to students based on their L2 reading proficiency levels. Poorly written items can confuse examinees, thereby contributing to the construct irrelevant variance and impacting assessment outcomes (Alderson, 2000; Downing, 2006). In terms of the optimal number of prescribed options, researchers generally agree that test items that offer three prescribed options are most effective in assessing target knowledge (e.g., Grier, 1975, 1976; Haladyna & Downing, 1993; Haladyna et al., 2019; Lord, 1944, 1977; Rodriguez, 2005; Tversky, 1964). Another common challenge occurs when attempting to create plausible distractors since if distractors for a test item are not plausible, students who have only a limited understanding of the tested knowledge can easily guess the correct answer, confounding results (Haladyna & Rodriguez, 2013). In order to generate plausible distractors for a test item, teachers must carefully map the thinking patterns that high- and low-achieving students might use to respond to each item (Haladyna & Rodriguez, 2013).
Another important factor to consider when creating reading items in the MC format is test length, as it can affect the reliability of L2 reading test scores (Sinharay et al., 2007). Indeed, the longer the test is, the more effectively it serves as a representative indicator of the ability that is being measured (Bachman, 1990). Haladyna and Rodriguez (2013) indicated that a reading test should consist of at least three passages, with 3–12 items per passage. Although there is no direct evidence to support Haladyna and Rodriguez’s (2013) suggestions, research studies that examined the values subscores added over total scores revealed that each subsection of a content or skill area test (e.g., reading, science, and social studies) should contain a large number of items (at least 20 or even 30) to generate results that examiners and examinees can use to interpret test scores meaningfully (e.g., Haberman & Sinharay, 2010, 2013; Sinharay, 2010). Regardless, these studies did not focus on the score reliability of reading tests, meaning that they did not consider how the number of items and number of passages affect score reliability. In addition, it is unknown whether these suggestions are applicable to L1 or L2 reading assessments.
With limited research studies available, practitioners are left to wonder what test length (i.e., the number of passages and test items) is most appropriate for gauging students’ L2 reading proficiency. This lacuna is especially problematic when one considers how frequently assessment results inform teaching and learning. Even so, the corollary is that assessments that produce reliable and valid indicators of student achievement can help teachers monitor learning progress and ensure their instruction meets the students’ L2 learning needs. Moreover, reliable assessment results can provide teachers with the additional information they need to make critical decisions about whether students pass or fail a course (Green, 2020). To provide research-based implications, I conducted a series of generalizability (G) theory studies which investigated the number of test items and passages that a reading test should include for a particular classroom assessment.
Literature review
Reliability in L2 reading assessment
Developing a language test for L2 learners is a complex task that involves a number of factors. Researchers indicate that the most essential of these factors is whether the test is useful for its intended purposes (e.g., Bachman & Palmer, 1996; Buck, 2001; Douglas, 2010). According to Bachman and Palmer (1996), test usefulness is “the essential basis for quality control throughout the entire test development process” (p. 17). This factor considers six elements, namely, reliability, (construct) validity, authenticity, interactiveness, impact, and practicality (see Bachman and Palmer, 1996). This study focuses primarily on reliability, which Bachman (1990) defines as the consistency of measured results under various test situations. Bachman (1990) further clarified that “if we wish to estimate how reliable our test scores are, we must begin with a set of definitions of the abilities we want to measure, and of the other factors that we expect to affect test scores” (p. 163).
Researchers argue that it is essential to identify these other factors since they are potential sources of measurement error that could affect an individual’s test performance (Bachman, 1990; Brennan, 2001). Indeed, identifying these factors will help L2 test developers ensure that an examinee’s test scores best represent his or her language abilities. Identifying and reducing the effects of construct irrelevant factors can contribute to minimizing measurement error and maximizing the reliability of language test scores (Bachman, 1990). While certain factors that affect scores are unsystematic and unpredictable “noise” (for instance, the examinees’ health conditions or how much sleep they received the night before their test), researchers have identified certain systematic factors related to test items and relevant texts that may affect the score reliability of L2 reading tests.
The systematic factors that impact test reliability the most include test methods and the authenticity of content and language use. In the context of reliability, “test methods” are often referred to as “test conditions.” To promote reliability, test developers may try to fix or control these test conditions by employing, as an example, a reading test with passages that are similar in terms of topic, text length, text difficulty, and comprehension question type. Although standardizing test conditions may minimize threats to reliability, it does pose a challenge for practitioners wishing to extrapolate their examinees’ test performance to real-world abilities, especially when the test conditions are not appropriately standardized (Bachman, 1990).
Test authenticity is another factor that determines the degree to which examinees’ performance on a language test represents their language competency. According to Bachman and Palmer (1996), language tests should generate scores that are predictive of the examinees’ real-life use of the target language. In other words, L2 tests must replicate scenarios that mimic real-life language use in order to reliably measure the examinees’ language proficiency. To achieve this goal, L2 tests must include authentic language and content that is situated in a way that authentically stimulates the examinees’ target language abilities (Bachman, 1990; Bachman and Palmer, 1996; Buck, 2001; Douglas, 2010; Weir, 2005). Factors that should be considered in terms of test authenticity include (a) determining the purpose of the test, (b) specifying the examinees’ target language proficiency, (c) assessing whether the test content has an appropriate difficulty level and can differentiate examinees’ ability levels, and (d) defining proficiency scales for the test scores that correspond to various components of the tested ability (e.g., Bachman & Palmer, 1982; Chapelle, 2008; Clark, 1978; Lane et al., 2015; McNamara, 2000). For example, if a reading test is meant to primarily assess examinees’ L2 academic reading abilities at the Common European Framework for Reference for Languages (CEFR) B2 level, it is inappropriate to include passages that only reference general topics and conversational texts with a text difficulty at the CEFR A1 level. Moreover, even if the topics and test content of a L2 reading test are appropriate for the examinees, the test score scales may not be well-developed, in which case the test developers cannot reliably predict the specific tasks that the examinees are able to realistically perform using their current L2 reading abilities.
Other important considerations when developing L2 reading test items include the language difficulty of the questions, item independence, and question type. If the language that is used to construct the questions is confusing or excessively difficult to understand, the examinees’ test performances will not reflect the targeted reading abilities (Green, 2020; Hughes, 2003). As an example, if the comprehension questions are written in a language of higher difficulty than the passages themselves, the test developers may not be able to deduce whether the examinees’ poor performance is due to the content difficulty of the passages or to the language difficulty of the questions. Second, item independence requires that test developers ensure all items in a subtest (i.e., a set of questions that are all related to the same text) have no interdependency. If the examinees’ response to one item affects or helps determine their response to other items, their L2 reading ability cannot be reliably measured (Hughes, 2003).
Finally, question types may vary in terms of their difficulty because they require different kinds of strategies for comprehension processing. Research has revealed that question types significantly influence the use of reading strategies; specifically, more difficult items tend to require higher level cognitive reading skills (e.g., Anderson et al., 1991; Davey, 1988; Liao, 2021; Liu, 2021; Rupp et al., 2006). Davey (1988), for instance, found that question items that require inferential processing (e.g., inferential and vocabulary questions) are more difficult than those that require the examinees to locate explicitly stated information in the text (e.g., factual questions), thereby motivating them to use higher order reading strategies. Making inferences about academic texts is a cognitively demanding skill because it requires varying levels of inferencing, including (a) integrating new information with background knowledge, (b) synthesizing information from the text, and (c) evaluating information that is given in the text (Grabe, 2009). As such, in order to elicit a variety of reading skills from the examinee pool, it is important that reading tests include question types that demand varying degrees of cognition. If, however, a reading test only includes one question type, then the resulting limited scope will hinder any attempt to capture a rounded assessment of the candidate (Alderson, 2000).
Text content and text length should also be carefully considered when developing L2 reading texts. Although test developers often include a variety of reading topics, they must do so with the understanding that when the text content is too specialized, it runs the risk of assessing subject matter knowledge over reading ability (Alderson, 2000; Bachman, 1990; Green, 2020). Regarding text length, research suggests that longer texts (for adult learners at more advanced levels of proficiency preparing for college study in the target language) more fully reflect their real-life reading situations, namely, as students being required to read and study extensive academic texts in the target language (Alderson, 2000). Lengthy texts also allow such examinees to demonstrate different kinds of reading skills, such as scanning, skimming, and summarizing (Grabe, 2009). Conversely, shorter texts allow test developers to include a greater number of passages in a reading test and a wider range of topics (Alderson, 2000). Overall, the intended audience and purpose of a test determines whether the use of longer texts to promote authenticity or of shorter texts to reduce potential content bias is more suitable. To ensure the reliability of test scores, test developers must therefore consider and minimize the effects of those factors that may prevent test scores from fulfilling their purpose.
Generalizability theory
Identifying and minimizing measurement error is essential if the score reliability of tests is to be increased (Brennan, 2001). Classical test theory (CTT) is one of the most common approaches to the investigation of test score reliability, but it provides only a single undifferentiated error term that is itself incapable of differentiating the sources of measurement error associated with the observed test scores (Shavelson & Webb, 1981). Another approach that is commonly used to investigate score reliability is G theory (American Educational Research Association, American Psychological Association, & National Council on Measurement in Education, 2014). G theory has been applied in various academic areas to estimate measurement error and collect reliability evidence for a variety of tests, including medical credential tests, standardized tests for American college admissions, cognitive ability tests, and K-12 subject matter tests (e.g., Brennan et al., 1995; Clauser et al., 2006; Lakin & Lai, 2012; Solano–Flores & Li, 2006; Webb et al., 2000). Unlike CTT, G theory helps identify and estimate multiple sources of measurement error (called facets) that are associated with the observed test scores (Brennan, 2001, 2021). In addition, G theory allows researchers and L2 assessment users (e.g., test developers, teachers, and English as a Second Language (ESL) program directors) to design various measurement scenarios and estimate score reliability based on their available resources and needs (Webb & Shavelson, 2008). G theory uses some aspects of analysis of variance (ANOVA) procedures. However, G theory does not involve hypothesis testing nor statistical assumptions (Brennan, 2001). G theory focuses solely on the magnitude of the estimated variance components, not statistical significance (Brennan, 1983, 2001).
Two types of studies are involved in a G theory analysis, namely generalizability (the G study, which comes first) and decision (the D study, which comes second). The G study is used to estimate variance components that contribute to the total score variance, which helps researchers pinpoint where the major sources of measurement error (i.e., systematic bias and error) originate from. Once the results of the G study are established, the D study can be employed to replicate a series of measurement procedures in order to obtain reliability estimates (i.e., G and phi coefficients) and error variances (i.e., relative and absolute error variances) (Brennan,2001). Both G and phi coefficients 1 represent an estimated reliability of the test scores, with the former being used for relative decisions (i.e., norm-referenced score interpretation) and the latter is used for absolute decisions (i.e., criterion-referenced score interpretation). Error variances inform the degree of measurement error that is involved when estimating students’ true proficiency based on their test scores (Brennan, 2001; Webb & Shavelson, 2008).
A growing number of language assessment studies have adopted G theory to examine the reliability of L2 assessment results (e.g., Bachman et al., 1995; Bolus et al., 1982; Brown, 1990, 1993; Brown & Bailey, 1984; Gebril, 2009, 2010; Kunnan, 1992; Lee & Kantor, 2005; Ohta et al., 2018). Most of these studies, however, are related to performance assessment, such as speaking and writing. Moreover, these studies primarily focus on examining the impact of task, rater, occasion, and/or scoring criteria, rather than the items and/or passages themselves. Although the application of G theory in L2 assessment has increased in recent years, only a handful of studies have examined the score reliability of reading assessment (Brown, 1984, 1999; Sawaki & Sinharay, 2013; Zhang, 2006). It is worth noting that Brown (1999) and Sawaki and Sinharay (2013) did not address the impact of items and passages on reading score reliability. Brown’s (1999) study investigated the contribution of examinees’ L1 language backgrounds to the reliability of the Test of English as a Foreign Language (TOEFL) scores, while Sawaki and Sinharay (2013) examined the score generalizability of the TOEFL section scores (i.e., reading, writing, listening, and speaking) for relative and absolute decisions.
The most relevant studies that have explored the effects of items and/or passages are Brown (1984) and Zhang (2006). Brown (1984) investigated the impact of items and passages on the score reliability of engineering English tests that were provided to 116 students from an American university. His findings indicated that a reading test should include at least three passages and a total of 60 items in order to generate a satisfactory level of reliability. The purpose of the reading task that was used in Brown’s (1984) study, however, was to measure content-area knowledge (i.e., engineering); also, the participants of this study included both non-native and native speakers. Conversely, Zhang (2006) assessed the test scores of 91,223 examinees in order to investigate how the test items contributed to the score reliability of the test of english for international communication (TOEIC). Zhang concluded that L2 reading assessments should offer a total of 90 test items to produce reliable scores; nonetheless, these findings are based on standardized assessment (i.e., TOEIC) rather than classroom-based assessment. Moreover, Zhang did not consider the effects of passages on the reliability of test scores. The number of passages may have impact on determining the number of items and, subsequently, the generalizability of the reading score.
As previously indicated, few research studies have focused on the score reliability of L2 reading assessment—particularly in the classroom setting—or on the possible impact of items and passages. More studies are thus needed to provide insight into the reliability of classroom-based L2 reading assessment. For this reason, the purpose of this study is to enhance the understanding of how MC-test items, passages, and their interaction contribute to L2 reading score reliability in a classroom context.
Research question
To investigate the effects of items and passages on L2 reading score reliability, I attempted to answer the following question concerning the application of L2 reading tests in the classroom context:
RQ. How does changing the number of items and passages affect the reliability of the L2 reading score?
Methodology
Participants
Participants were 108 students from ESL reading classes of an American college. Among these participants, 49 were male and 59 were female. In terms of L1 language, 45% of the participants spoke Arabic, 43% spoke French, and the rest (12%) spoke other languages, including Chinese, Spanish, Thai, Vietnamese, Swahili, and Russian. English reading proficiency was measured using the Accuplacer test that was developed by the College Board (https://accuplacer.collegeboard.org). According to the results, the participants’ reading scores ranged from 20 to 80 in a total of 120 points (low-intermediate to high-intermediate), 2 with a mean of 53.86 and a standard deviation of 19.77.
Reading instrument
The English reading comprehension test in this study served as a formative assessment to monitor students’ learning progression and improve teachers’ instruction. The reading test consisted of four reading passages, with five comprehension items per passage for a total of 20 MC items. The reading passages and comprehension items were adopted from the following reading test-booklets published by Pearson Education: Get Ready to Read: A Skills-Based Reader (Blanchard, 2004, pp. 24–26 & pp. 34–37) and Ready to Read Now: A Skills-Based Reader (Blanchard, 2005, pp. 5–7 & pp. 43–45). Several factors were taken into account when choosing the passages for the reading test, including range of reading topics, text difficulty, and types of reading questions. To ensure content coverage, the reading test covered the following four reading topics: science, a historical figure, a historical event, and business. Together with two experienced English teachers, we reviewed the reading test and identified the science and historical figure passages as narrative texts, and the historical event and business passages as informational texts. Moreover, we reviewed the comprehension items (a total of 20) and identified seven vocabulary questions (each passage had two, except for the historical figure, which only had one), eight factual questions (the historical event passage had three, the science and business had two, and the historical event only had one), and five inferential questions (the historical figure passage had two, and other passages only had one).
On average, each reading passage was about 300 words long. I employed the Flesch Reading Ease 3 and The Lexile Analyzer®4 to calculate the difficulty level of the four texts. The Flesch Reading Ease scores ranged from 40 to 65, meaning that the difficulty of these four texts fell between intermediate (standard) and high-intermediate (difficult to read). The Lexile Analyzer® showed the Lexile scores ranging from 610 L to 1200 L, which are considered A2 to B2 CEFR levels (Smith & Turner, 2016). The internal consistency coefficient for the reading test was α = .72, and the item-total correlation average was .30.
Data collection and analysis
Five teachers administered the English reading test to their ESL reading class 1 week after the mid-term exam in order to gauge their students’ current English learning progress. The students had 2 hours to complete the test, though most were able to do so in only 1.5 hours. After the participants completed the test, the teachers scored their responses dichotomously using a 0–1 scale (0 being incorrect and 1 being correct). I then analyzed the collected data using G theory.
Regarding the research question, I conducted a univariate random facet p × i G theory design (indicating that students completed the same question items) to investigate the effect of changing the number of items and a univariate random facet p × h design (indicating that students read the same passages) to examine the effect of changing the number of reading passages. Moreover, I used a univariate random facet p × (i : h) design (indicating that students took the same reading test in which the question items were nested within the passages) to investigate how the relationship between the number of items and the number of passages influences reading score reliability. Finally, in responding to the research question, I used the software program GENOVA (Crick & Brennan, 1983) to conduct univariate G and D studies to estimate variance components, relative and absolute error variances, and G and phi coefficients. Depending on the G theory designs, I entered different types of codes into the control card, which were then read and analyzed by GENOVA.
Results
Descriptive statistics showed that students’ test mean score was 31.69 out of 40, with a standard deviation of 6.23. The results of each G theory design are reported in the following sections.
The impact of the number of items
Table 1 details three variance components for the p × i G study: person (p), item (i), and person-by-item (pi) plus undifferentiated error (i.e., random error). The largest variance component (80.52% of the total score variance) was attributed to the person-by-item interaction and undifferentiated error, suggesting either that the students inconsistently performed on different items or that some undifferentiated error occurred.
Estimated variance components for p × i design.
In order to understand how the number of question items used in a reading test might affect score reliability, I conducted the p × I D study to derive generalizability (Ep2) and phi (Φ) coefficients as well as relative and absolute error variances (see Table 2). Obtaining G and phi coefficients is helpful for estimating score reliability across different measurement scenarios. For example, in an L2 reading class, a teacher may wish to divide the students in a way that ensures that each group contains students of varying proficiency levels; the information that is provided by G coefficients would no doubt be helpful in this scenario. The information that is obtained from phi coefficients would also prove useful if the teacher simply wished to make a pass or fail decision based on the students’ performance on an L2 reading test. As the study is intended for a wider range of practitioners, both G and phi coefficients are provided as a reference for estimating reading score reliability. In this study, I focused more heavily on G coefficients when reporting reliability estimates because the reading test in this study was for a relative decision (i.e., monitoring and ranking students’ learning progress). With this established, the score reliability coefficient should be .80 (Shavelson & Webb, 1991).
Estimated G and phi coefficients for p × i design.
As shown in Table 2, increasing the number of items had a positive influence on the G coefficient. Indeed, the G coefficient increased dramatically when the number of items increased from 5 to 10, 10 to 15, and 15 to 20 (e.g., an 0.17 increase from 5 to 10). Even so, when the number of items increased from 20 to 25, 25 to 30, and so on, the G coefficient only slightly increased (e.g., an 0.04 increase from 20 to 25). These results reveal a limited return when increasing the number of items beyond 20. A reading test that contains about 30 items is necessary, however, to reach the acceptable G coefficient level of 0.80.
The impact of the number of passages
Table 3 details three variance components for the p × h G study: person (p), passage (h), and person-by-passage (ph) plus undifferentiated error. The largest variance was attributed to the person-by-passage interaction and undifferentiated error (51.67%), indicating either that the students inconsistently performed on different reading passages or that some undifferentiated error was involved.
Estimated variance components for p × h design.
I carried out the p × H D study in an attempt to understand how the number of reading passages affected score reliability. Similar to the D study for the item impact, Table 4 shows that increasing the number of reading passages is helpful for promoting the G coefficient. When increasing the number of passages from one to two, two to three, and three to four, the G coefficient increased sharply (e.g., an 0.17 increase from one to two). Still, the increase of the G coefficient decreased when the number of passages was increased beyond four. As such, a reading test that contains five or six passages is needed to attain the G coefficient of 0.80.
Estimated G and phi coefficients for p × h design.
The number of items in relation to the number of passages
Table 5 details five variance components for the p × (i : h) G study: person (p), passage (h), item-nested-in-passage (i : h), person-by-passage (ph), and person-by-item-nested-in-passage (pi : h) plus undifferentiated error. Person-by-item-nested-in-passage plus undifferentiated error was the largest variance component, accounting for 79.26% of the total score variance. This suggests that the students performed inconsistently on different items that were associated with different reading passages, or that some undifferentiated error occurred.
Estimated variance components for p × (i : h) design.
A series of D studies was conducted for the p × (i : h) design to account for the impact of the interaction between the number of items and passages on score reliability. The results indicated that the number of items should be at least 30 across all of the passages combined to achieve the G coefficient of 0.80 (as shown in Table 6 and Figure 1). Also, to reach the G coefficient of 0.90, the total number of items should be at least 70. For example, if the desirable G coefficient is 0.80 for a reading test that contains only two passages, then the number of items per passage should be approximately 15.
Estimated G and phi coefficients for p × (i : h) design.

G coefficients for different number of items in relation to different number of passages.
Discussion
The purpose of this study was to investigate reading test score reliability, with a particular emphasis on the impact of the number of items and passages. G studies presented estimated variance components for p × i, p × h, and p × (i : h) designs, while D studies indicated the effect of the number of items and/or the number of passages on the score reliability.
The relationship between the number of items and the number of passages
As previously stated, the longer an assessment task is, the higher reliability it will produce (Alderson, 2000; Gebril, 2009, 2010; Lee & Kantor, 2005; Sinharay et al., 2007). Having said this, however, it is unclear what the optimum length of an L2 reading test is in order to reliably measure learners’ reading comprehension. Moreover, because the reading items are nested within the passages themselves, it is essential that both the number of items and the number of passages are considered when developing L2 reading tests.
The results of this study correspond with Brown’s (1984) and Zhang’s (2006) findings, in that they showed that the total number of test items and reading passages plays an important role in determining the reliability of students’ test performance. Just as prior studies revealed that between 20 or 30 items per content or skill area test are needed to generate meaningful test interpretations (e.g., Harberman & Sinharay, 2010, 2013; Sinharay, 2010), this study found that a reading test being implemented for the purposes of formative assessment should include at least 30 items in order to provide teachers with reliable information about their students’ current English learning progress. The results further indicated that, regardless of the number of reading passages, a total of 30 items is necessary to ensure the reliability of test results.
The setting of this study was college ESL reading classrooms. For learners whose level of English reading proficiency was proximal to the proficiency level of those in this study, the appropriate number of passages in a reading test should be between three and six, thereby giving students enough time to complete the test during a single class session. In addition, the total number of items in the L2 reading test should be structured accordingly: 10 items for each of the three passages (30 items total), eight items for each of the four passages (32 items total), six items for each of the five passages (30 items total), or five items for each of the six passages (30 items total; see Table 6 and Figure 1). These results support Haladyna and Rodriguez’s (2013) viewpoints that a reading comprehension test should include at least three passages, with a total of 3–12 items each. A reading test that consists of three to six passages with a total of 30 items is reasonable, as it allows teachers to incorporate a variety of topics in a single test session. If a reading test only includes one passage, it is more likely to be influenced by content bias because students may perform well or poorly based on their familiarity with the topic (Bachman, 1990). It is also unrealistic to offer 30 items for a single passage. Similarly, a reading test that includes only two passages may be unduly influenced by content bias. Each passage must also be long enough to warrant 15 items, as the test criteria itself necessitates a total of 30 items. Thus, to effectively measure the students’ reading proficiency, a “3-6-30 test” (a reading test that consists of three to six passages and features 30 items in total) is more appropriate than one that offers only one or two passages. Regardless, this design may only be applicable to adult ESL students whose reading proficiency levels range from low- to high-intermediate.
Although it is suggested that this “3-6-30” design is most appropriate, it is important to note that the reading test in this study included vocabulary, factual, and inferential questions. Indeed, the question types varied in terms of difficulty and cognitive load. These varying types of questions may, for instance, ask students to locate key information in the text, combine key information across paragraphs, or integrate key information with their background knowledge. Most research studies confirm that more difficult question items are more likely to elicit higher level cognitive reading strategies (e.g., Alderson, 2000; Anderson et al., 1991; Davey, 1988; Liao, 2021; Liu, 2021; Rupp et al., 2006), especially those tapping implied information, such as inferential questions (Davey, 1988). Answering different types of questions requires a diversity of reading skills, which may affect the proportion of item distribution for each type of question. To more fully understand how question difficulty may affect test scores in reflecting students’ target language abilities, researchers should more thoroughly consider how question types influence L2 reading score reliability.
Implications and limitations
The findings of this study provide practical implications for L2 classroom assessment. Teachers, for example, in L2 reading classes often administer quizzes or tests as formative assessments to evaluate their students’ learning progress and determine whether their teaching and learning activities need some modification. The results of the study revealed that teachers who wish to use a formative assessment to reliably assess their students’ L2 reading comprehension must ensure that these reading tests include an appropriate number of passages and items. Teachers often include a variety of topics from different disciplines (e.g., education, psychology, and biology) in their reading tests in an effort to monitor students’ reading proficiency. The more topics that are covered in a reading test, the higher the number of passages and, consequently, the lower the number of items required per passage. Beyond considering those factors that are discussed earlier (e.g., the number of prescribed options per item and the plausibility of distractors), teachers tasked with developing test items should also attempt to strike a balance between the number of passages and items based on available testing time and overall intent.
Having established the implications for L2 assessment, it is worth discussing how the findings of this study may inform future research. Although the reading items that are used in this study include vocabulary, factual, and inferential questions, the impact of question types was not taken into account. Moreover, though the average length of the passages that appeared in this study was approximately 300 words, longer texts may pose more difficulty to students’ L2 reading comprehension (e.g., Alderson, 2000; Engineer, 1977; Newsom & Gaite, 1971) because processing longer texts requires more extensive vocabulary knowledge and advanced reading skills, such as summarizing (Grabe, 2009). Plakans (2009) indicated that students with lower level reading proficiency often have limited vocabulary knowledge, which affects their understanding of the reading texts. Thus, text length may impact the reading score reliability of the test, depending on the students’ L2 reading proficiency. Future research should therefore examine the possible ramifications of question types and text length while considering students’ target language reading abilities.
This study has a number of limitations that are worth noting. First, the variance components for the interaction effect in this study (i.e., person-by-item interaction, person-by-passage interaction, and person-by-item-nested-in-passage) include undifferentiated error, also called random error. As such, although the variance components for the interaction effect observed in this study were evident, it is unclear how much random error contributed to the total score variance. This confounding nature, which is also discussed in Lin and Zhang’s (2018) study, is a limitation of current G theory designs. Second, the majority of the participants in the study were French and Arabic speakers whose reading proficiency levels ranged from low- to high-intermediate. Future research should therefore include a wider variety of L1 speaking students at varying English reading proficiency levels to more fully generalize findings of the study to the larger ESL student population. Third, because the reading assessment in this study was conducted in college ESL classrooms, more studies are needed to investigate the score reliability of L2 reading tests in other contexts; indeed, other teaching contexts may introduce factors (for instance, reading proficiency level and age) that demand a different number of items and passages. Fourth, this study did not differentiate the test items based on question type. As such, future research should take into account how different question types affect the score reliability of L2 reading tests. Finally, the reading items in this study are constructed in the MC format. Although the selected response format is commonly adopted for reading tests, researchers have argued that it only assesses procedural or conceptual knowledge, both of which require such skills as recalling key information and recognizing the correct answer (e.g., Bennett et al., 1990; Frederiksen, 1984; Impara & Foster, 2006). In contrast, the constructed response format has grown in popularity in receptive language assessments because of its ability to assess higher order thinking skills (e.g., synthesizing, evaluating, and inferencing) and reduce the possibility of guessing (e.g., Haladyna, 1997; Sireci & Zenisky, 2006). Future studies should therefore consider investigating the impact of the constructed response format on L2 reading score reliability in relation to the selected response format.
Conclusion
With this study, I investigated the score generalizability of classroom-based L2 reading assessment. The results demonstrated that the relationship between the number of items and the number of passages within the assessment alters L2 reading score reliability. The results also revealed that maintaining a careful balance between the number of items and passages in a reading test is essential. Moving forward, I am hopeful that these results provide helpful guidance for practitioners tasked with developing L2 reading tests for classroom assessment.
Supplemental Material
sj-crd-1-ltj-10.1177_02655322211070840 – Supplemental material for The use of generalizability theory in investigating the score dependability of classroom-based L2 reading assessment
Supplemental material, sj-crd-1-ltj-10.1177_02655322211070840 for The use of generalizability theory in investigating the score dependability of classroom-based L2 reading assessment by Ray J. T. Liao in Language Testing
Supplemental Material
sj-crd-2-ltj-10.1177_02655322211070840 – Supplemental material for The use of generalizability theory in investigating the score dependability of classroom-based L2 reading assessment
Supplemental material, sj-crd-2-ltj-10.1177_02655322211070840 for The use of generalizability theory in investigating the score dependability of classroom-based L2 reading assessment by Ray J. T. Liao in Language Testing
Supplemental Material
sj-crd-3-ltj-10.1177_02655322211070840 – Supplemental material for The use of generalizability theory in investigating the score dependability of classroom-based L2 reading assessment
Supplemental material, sj-crd-3-ltj-10.1177_02655322211070840 for The use of generalizability theory in investigating the score dependability of classroom-based L2 reading assessment by Ray J. T. Liao in Language Testing
Footnotes
Acknowledgements
The author thanks Won-Chan Lee, Renka Ohta, and Robert L. Brennan for their suggestions to improve this study. He would also like to thank the anonymous reviewers and the editor, Paula Winke, for their comments, queries, and assistance in revisions of this article.
Declaration of conflicting interests
The author declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author received no financial support for the research, authorship, and/or publication of this article.
Supplemental material
Supplemental material for this article is available online.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
