Abstract
This study examined the use of artificial intelligence tools, which have garnered significant attention in recent years, in the assessment and evaluation processes of language education. For this purpose, student essays were scored by Turkish middle school teachers and artificial intelligence tools both with and without the use of a rubric, and the findings were evaluated based on generalizability theory. Additionally, the research findings were shared with participants to gather qualitative data, which were analysed using the inductive thematic analysis method to support the research results. The findings revealed that in evaluations conducted without a rubric, teachers were limited in their ability to distinguish individual differences and demonstrated low scoring consistency. In contrast, artificial intelligence tools were more effective in distinguishing individual differences and exhibited high consistency. In evaluations conducted using a rubric, scoring consistency increased in both groups, although, as in the first evaluation, artificial intelligence tools demonstrated a higher level of consistency. Regarding the research findings, teachers expressed that individual biases, mood, and professional experiences influenced their scoring processes and emphasized the potential of rubrics and artificial intelligence-supported feedback systems for achieving more consistent results. Artificial intelligence tools, on the other hand, highlighted their independence from subjective factors but stressed the need for more diverse and generalizable datasets to further enhance their evaluation capacities.
Keywords
1. Introduction
Writing has long held a special place in language education, not only as a means of communication but as a window into a learner’s developing proficiency. Unlike discrete-point tests that target isolated linguistic features, writing tasks demand the simultaneous management of content, structure, vocabulary, and grammatical accuracy, all shaped by purpose and audience (Hyland, 2003; Weigle, 2002). This cognitive complexity is precisely what makes writing so valuable as an assessment tool, and equally what makes it so difficult to evaluate with consistency. Decisions about what to measure, how to weight each quality, and how to ensure that scores carry the same meaning across different raters and contexts are far from straightforward (Bachman & Palmer, 1996; McNamara, 2000). The challenge is compounded by the fact that writing assessment must serve multiple purposes at once: it must generate defensible scores, inform instruction, and provide meaningful feedback to learners. These persistent and overlapping demands have kept writing assessment at the centre of debate in applied linguistics and educational measurement for several decades (Crusan, 2010; Hamp-Lyons, 1991).
Scoring rubrics emerged as one of the most practical responses to these difficulties. A rubric spells out what is being evaluated and describes what performance looks like at each level, turning what would otherwise be a largely subjective exercise into something more structured and transparent (Brookhart, 2013; Jonsson & Svingby, 2007). Analytic rubrics go a step further by breaking writing into separate dimensions, things like content, organization, vocabulary, grammar, and mechanics, and scoring each one independently. This approach is especially popular in research and high-stakes testing because it produces a diagnostic profile rather than a single number (Knoch, 2009; Weigle, 2002). Wang et al. (2021), for instance, found that analytic scoring rubrics yielded slightly better accuracy than holistic ones in automated scoring contexts, reinforcing the diagnostic value of separating writing into its component dimensions. In theory, a well-constructed rubric reduces the subjectivity of scoring by giving raters a shared frame of reference. In practice, however, the picture is more complicated. Research drawing on many-facet Rasch measurement and generalizability theory (G theory) has repeatedly shown that even trained raters working from the same rubric can differ substantially in their severity, in how consistently they apply criteria, and in the weight they assign to different aspects of writing quality (Eckes, 2012; Gebril & Plakans, 2014; Lumley, 2002). Wind et al. (2018) demonstrated that rater effects such as severity, centrality, and inaccuracy can even propagate into automated scoring systems when those systems are trained on human-scored data, underscoring how deeply rater variability is embedded in the assessment chain. These findings point to a fundamental tension at the heart of writing assessment: rubrics improve reliability in principle, yet human scoring remains a source of meaningful variability in practice. It is partly this tension that has driven interest in automated approaches to writing evaluation.
The use of computers to score written text is not a recent development, although the technology has changed considerably since its early iterations. The first major automated essay scoring (AES) system, Project Essay Grade (PEG), was introduced by Page (1966) as a way to mimic the holistic judgements of trained human raters through the statistical modelling of surface-level textual features (Dikli, 2006). Over subsequent decades, more sophisticated systems emerged. The e-rater engine at the Educational Testing Service brought natural language processing into the picture, modelling grammatical complexity, vocabulary sophistication, and how discourse was organized (Attali & Burstein, 2006; Ramineni & Williamson, 2013). The Intelligent Essay Assessor drew on latent semantic analysis to capture semantic relationships between ideas (Landauer et al., 2003). Taken together, these systems demonstrated that machine-generated scores could achieve strong correlations with human rater scores, particularly on well-defined writing tasks administered at scale (Dikli, 2006; Shermis & Burstein, 2013). What had seemed an implausible ambition in the 1960s became, by the early 2000s, a commercially viable assessment technology. More recently, deep learning and transformer-based architectures have pushed the boundaries further, with systems employing Bidirectional Encoder Representations from Transformers (BERT), convolutional neural network (CNN), and long short-term memory (LSTM) models to capture deeper semantic features and long-distance dependencies in text (Jin & Hu, 2026; Yuan et al., 2020).
Once AES systems moved from research labs into real testing programmes, questions about their validity and reliability naturally followed. On the reliability side, the argument for machines is fairly straightforward. Machine scores are perfectly consistent within a given system: the same text will always receive the same score, regardless of time of day, rater fatigue, or the quality of the preceding script. Human raters, by contrast, are susceptible to a range of well-documented biases including halo effects, sequential context effects, and topic-related preferences that can distort scores independently of writing quality (Daxenberger et al., 2023; Lumley, 2002). Research has shown that AES scores can match or exceed the agreement levels observed between pairs of trained human raters (Attali & Burstein, 2006; Voss et al., 2026; Williamson et al., 2012). Cui (2024), for example, reported a quadratic weighted kappa of 0.92 between automated and human scores in a university-level longitudinal study, while Gaggioli et al. (2025) found considerably lower agreement in higher education contexts involving subjective and context-sensitive criteria. Yet reliability, taken alone, is not sufficient evidence that a scoring system is measuring what it claims to measure. The more challenging question concerns construct validity: whether the features a machine detects and rewards correspond to the qualities of good writing that rubric criteria are designed to capture (Chapelle, 2010; Messick, 1996). Critics have noted that some AES systems respond strongly to essay length and surface-level lexical density while remaining relatively insensitive to logical coherence, argument quality, and rhetorical appropriateness (Deane, 2013; Perelman, 2014). Zhong et al. (2026) found that the e-rater® engine systematically assigned higher scores to artificial intelligence (AI)-generated essays than human raters did, revealing that automated systems may reward textual features that do not necessarily reflect genuine writing quality. More recent approaches employing transformer-based architectures and deep learning have made progress in capturing discourse-level and semantic properties of text, although questions about the alignment between machine scoring and rubric constructs have not been fully resolved (Rodriguez et al., 2019; Yang et al., 2020). Frameworks for evaluating the validity of automated scoring systems continue to develop, reflecting the recognition that standard psychometric approaches require adaptation when applied to machine-generated scores (Ercikan & McCaffrey, 2022; Ferrara & Qunbar, 2022).
Beyond these broader validity concerns, researchers have paid increasing attention to how well AI and human scores agree when examined at the level of individual analytic rubric dimensions. This dimension-level perspective is important because overall score agreement can mask meaningful divergences on specific criteria. Human raters, shaped by their professional training, linguistic backgrounds, and personal theories of writing quality, do not weight rubric criteria uniformly, and systematic differences in emphasis are well documented in the rater cognition literature (Barkaoui, 2010; Eckes, 2012). Automated systems, by contrast, process all texts through the same fixed algorithm, which ensures uniformity but may come at the cost of sensitivity to the kinds of contextual and interpretive judgements that experienced teachers routinely make. Vo et al. (2023) provided striking evidence of this divergence: in their statewide writing assessment, human raters appeared more likely to differentiate content-related traits than the automated scoring engine, especially in upper grades, and the shared variance among traits was notably lower for human-graded essays than for machine-graded ones. Karaçeper and Kıray (2026) reported a similar pattern in an English-as-a-foreign-language (EFL) context, finding that whereas human raters achieved almost perfect agreement on format components, ChatGPT-4o reached only fair agreement on the same dimension, and the gap widened further for style components, where human agreement was substantial, but AI agreement was slight. Studies comparing AI and human scores across specific analytic categories more broadly report divergent patterns: agreement tends to be highest for grammatical accuracy, where the features being judged are relatively discrete and rule-governed, whereas it is considerably lower for dimensions such as content relevance, argumentative depth, and organizational coherence, which call for interpretive rather than rule-based evaluation (Alharbi, 2023; Alshehri et al., 2025; Woo & Choi, 2021; Ataseven et al. (2025) found moderate agreement between GPT-based systems and human raters, noting that AI tended to assign higher scores and sometimes overlooked off-topic content. Hong and Dan (2025) further revealed that large language models (LLMs) produce scoring rationales that differ substantially from those of human raters, contributing to scoring inconsistencies on subjective dimensions. Aydın et al. (2025) observed that ChatGPT-4o tended to over-penalize linguistic errors compared to human raters, who were more tolerant of morphological and lexical mistakes, further illustrating how AI and human evaluators can diverge in their treatment of specific rubric criteria. These findings suggest that the question of whether AI can replace human scoring is not one that yields a simple yes or no answer; the more productive question is where, and under what conditions, automated scores offer reliable and valid information.
Moving from large-scale automated scoring research to the realities of classroom writing assessment in language learning settings is not a straightforward step. Most of the foundational research on automated scoring was conducted using essays written by native or near-native speakers of English in high-stakes standardized testing contexts, and the scoring models were trained accordingly (Knoch, 2009; Weigle, 2002, 2013). EFL learner writing presents a different profile of features. Interlanguage characteristics, transfer-related errors, and discourse conventions influenced by learners’ first languages and educational backgrounds are common, and it is not always clear that systems trained on native-speaker corpora handle these features in the same way that a trained EFL teacher would (Cumming, 2001; Hamp-Lyons, 1991). Weigle (2013) highlighted this concern directly, noting that although e-rater scores correlated as highly with human scores as two human ratings correlated with each other, the minor differences that did emerge provided negative evidence regarding the system’s ability to capture all features of non-native writing that human raters are sensitive to. Wang (2019) similarly observed that automated scoring systems designed primarily for native English speakers are not always suitable for EFL learners, whose error patterns differ in systematic ways. The available evidence from studies examining AI scoring in EFL writing contexts is promising but not conclusive: some report reasonable agreement with human raters, whereas others point to systematic discrepancies on particular criteria (Alharbi, 2023; Karaçeper & Kıray, 2026). Suhan and Wolf (2026), comparing GPT-4 with an existing automated writing evaluation (AWE) model and human raters on young EFL learners’ writing, found that GPT-4 achieved acceptable reliability overall but performed inconsistently across task types, with notably weaker results on source-based writing tasks. This underscores the need for context-specific evidence rather than generalization from standardized testing research. Furthermore, classroom writing assessment serves purposes that go beyond score assignment. Teachers use their evaluation of student writing to plan future instruction, to provide targeted developmental feedback, and to make nuanced judgements about learner progress that aggregate statistics do not easily capture. Wilson et al. (2021) found that elementary teachers perceived AWE as both facilitative and potentially counterproductive, noting that AWE functionality sometimes created new instructional challenges, particularly when feedback was misaligned with what teachers considered pedagogically appropriate. Whether AI tools can support these formative functions, rather than simply approximating summative scores, remains an open and pressing question (Akhter & Zaman, 2024; De la Vall & Araya, 2023; Selim, 2024; Steiss et al., 2024).
The literature reviewed above points to several unresolved questions that motivate the present study. AES has attracted a great deal of research attention over the past three decades, but most of that work has focused on holistic scoring in standardized testing environments. Far less is known about how well AI-generated scores line up with teacher judgements on individual analytic rubric dimensions in everyday language classrooms. Shermis (2025) evaluated ChatGPT-4o’s scoring on benchmark datasets and found that although the system came close to human-rater agreement on some essay sets, its consistency fell short on more complex or variable prompts, reinforcing the case for fine-grained, dimension-level investigation. Studies that do compare AI and teacher scores often treat the teacher score as a self-evident benchmark without examining the degree to which teacher ratings themselves are consistent or potentially biased (Eckes, 2012; Lumley, 2002). Wilson et al. (2026) demonstrated in the largest randomized controlled trial of AWE to date that the impact on student outcomes varied dramatically from one district to another, with positive effects where implementation was carried out faithfully and negative effects where it was not, a reminder that the conditions surrounding AI use matter just as much as the technology itself. Perhaps most importantly, few studies have placed AI-generated analytic scores and teacher scores side by side on the same learner texts, scored against the same rubric, in a way that allows for direct, criterion-by-criterion comparison. The scarcity of such evidence is especially pronounced in first language writing contexts, where the dynamics of scoring may differ from those observed in standardized testing of second or foreign language learners. This kind of evidence is precisely what practitioners need if they are to make informed decisions about incorporating AI tools into their assessment practice. The present study is designed to address these gaps. By comparing AI-generated and teacher-assigned scores on a shared analytic rubric applied to Turkish-learner writing in a first language context, it aims to shed light on the degree and nature of agreement across individual scoring dimensions and to consider what the findings suggest about the appropriate role of automated scoring in instructional writing assessment.
Based on the existing problem and studies in the literature, this study aims to explore the potential of AI tools in evaluating writing skills by comparing them with middle school teachers, seeking answers to the following research questions:
What is the distribution of scores assigned by teachers and AI tools in evaluations conducted without a rubric?
Are there differences between teachers and AI tools in evaluations conducted without a rubric?
What is the distribution of scores assigned by teachers and AI tools in evaluations conducted using a rubric?
Are there differences between teachers and AI tools in evaluations conducted using a rubric?
What are teachers’ and AI tools’ opinions on research findings?
2. Methodology
2.1. Research Design
In this study, the sequential transformative design, a type of mixed-methods design, was employed. This design involves the sequential collection and analysis of quantitative and qualitative data, guided by a specific theoretical framework or paradigm. Creswell and Plano Clark (2007) describe the sequential transformative design as an effective approach when the researcher aims to test, develop, or transform a primary paradigm. In line with this design, quantitative and qualitative data were collected sequentially and integrated within a complementary framework to ensure coherence and depth.
In the first phase, quantitative data were collected to examine the relationship between the scoring behaviours of instructors and those of AI tools. The findings derived from the quantitative analysis were further explored and interpreted in greater depth through qualitative methods in the second phase. During the qualitative data collection process, a structured opinion form, accompanied by a summary of the research findings, was distributed to instructors and AI tools. The qualitative data were analysed using the inductive thematic analysis technique to provide a more nuanced interpretation of the quantitative findings.
2.2. Study Group/Research Data
The study group consisted of middle school teachers and generative AI tools. Participants were recruited through a formal university–school collaboration initiative, where an open call was issued to middle schools within the partnership network. Participation was entirely voluntary. In Turkey, where the study was conducted, teachers are classified based on their years of experience and their status in the expert teacher training programme, which includes an examination for certification. Those who attain expert teacher status are referred to as head teachers if they serve in this role for 10 years after earning the title. All participants in the study had at least 10 years of professional experience and held the status of expert teacher. The demographic characteristics of the teachers are presented in detail in Table 1.
Demographic Characteristics of Participants.
The other study group consisted of generative AI tools. To support the scoring processes, tools aligned with the research objectives were carefully selected, focusing on those with multidimensional reasoning capabilities, the ability to interpret existing data, and draw inferences. Accordingly, the following tools were chosen for scoring student essays: Claude 3.5 Sonnet, developed by Anthropic and released in June 2024; ChatGPT 4o, developed by OpenAI and introduced in May 2024; Gemini 1.5 Pro, developed by Google and launched in 2024; and DeepSeek-V2, another generative AI tool released in 2024.
2.3. Data Collection and Evaluation Tools
The research utilized student writing tasks and a structured interview form to gather primary data. An analytic rubric was employed separately as a scoring framework to evaluate the collected writing tasks.
2.3.1. Student Writing Tasks
The primary data collection tool consisted of essays written by 6th, 7th, and 8th-grade middle school students. These essays were obtained from the portfolios of three Turkish-language teachers working in schools outside the study group. To ensure data privacy and ethical compliance, all selected essays were subjected to an anonymization (blinding) process before being uploaded to the generative AI tools. All identifying information, including student names, school IDs, and any specific locational markers, was removed to protect student identities.
From the initial pool, 18 essays (6 for per grade level) were submitted to 3 expert academics for evaluation. The experts were tasked with selecting two essays per grade level that aligned with the study’s objectives. Based on their feedback, the following topics were chosen for evaluation by teachers and AI tools:
6th Grade: “Who knows more: those who read a lot or those who travel a lot?” and “The importance of keeping our environment clean.”
7th Grade: “Global warming and my responsibilities” and “Values education and our prominent values.”
8th Grade: “Cyberbullying and its effects” and “If I were an inventor, what would I do?”
2.3.2. Evaluation Framework: Writing Skill Analytical Scoring Rubric
The Writing Skill Analytical Scoring Rubric developed by Asma (2024) functioned as the evaluation framework to investigate the impact of rubric use on scoring behaviours. The rubric consists of four sections: Spelling and Punctuation (0–5), Topic (0–5), Language Use (0–5), and Structure and Organization (0–5). The maximum score achievable with the rubric is 20 points, with each score from 0 to 5 indicating the performance level of the essay in the respective section (see Appendix A).
The rubric, originally developed by Asma (2024), was evaluated for content validity by Yıldız and Asma (2026) using the widely recognized method established by Lawshe (1975). Their evaluation process incorporated professional feedback from a panel of five academic experts. Consistent with this methodology, they calculated the content validity ratio (CVR) for individual scale items, followed by the computation of the Content Validity Index (CVI) for the entire instrument.
The critical threshold for determining acceptable content validity was based on the criteria established by Ayre and Scally (2014). For a panel consisting of five experts, the minimum acceptable threshold is 0.99. During their item analysis, the CVR for the sixth item in the language use category and the fourth item in the spelling and punctuation category dropped to 0.78, falling short of the required threshold because of revision suggestions provided by one expert.
Despite these specific item variances, the overall CVI for the instrument remained robust in their study. Yıldız and Asma (2026) reported that the spelling and punctuation and the language use subdimensions both yielded a CVI of 0.95, while the topic and the structure and organization subdimensions reached 0.99. Because the calculated CVI scores successfully exceed the critical threshold, the scale items exhibit statistically significant content validity, aligning with the benchmarks cited by Batdı (2013) and Lawshe (1975). Ultimately, their psychometric results confirm that the writing rubric possesses adequate content validity and is entirely suitable for practical implementation.
2.3.3. AI Evaluation Procedure and Prompt Formulation
The research utilized carefully constructed instructions to ensure standardized evaluations across the different generative AI models. The evaluation prompt was meticulously designed according to established prompt engineering principles to maximize scoring reliability. The instructions explicitly assigned an expert teacher persona to the models to establish a professional evaluation context. The prompt clearly defined the task boundaries and provided the exact rubric dimensions to prevent the generation of irrelevant criteria. The instructions also required a structured output format combining numerical scores with specific textual explanations for each rubric category.
The following exact prompt structure (See Figure 1) was utilized alongside the evaluation rubric to assess the student essays.

Prompt Structure Flow of AI Tools.
2.3.4. Structured Opinion Form
To ensure the internal validity of the study, the research findings were shared with both teachers and AI tools, and their opinions were solicited. Two structured questions were posed to examine the attitudes of both groups toward the findings, serving as a measure of internal consistency and as a complementary process to support the quantitative data:
It has been observed that AI tools are more consistent and reliable in the evaluation process. What are your thoughts on these findings, and what support do you need to improve your scoring practices?
What are your views on the impact of the rubric on the evaluation process? To what extent do you think such guides can enhance teacher performance?
2.4. Data Analysis
The analysis of the research data was conducted based on G theory, which is a statistical approach that evaluates the reliability and generalizability of measurement results. Going beyond traditional reliability analyses, this theory aims to separate different sources of error and examine the extent to which measurements can be generalized across different contexts. Detailed by Brennan (2001), G theory is widely applied in the fields of education and psychological measurement. By analysing different variance sources (e.g., individuals, raters, items), it provides strategies to enhance the reliability of measurements. Additionally, it guides the design of measurement processes, increasing the efficiency of data collection. At the core of the theory are “G” studies, which identify sources of error in the measurement process, and “D” studies, which aim to minimize the impact of these error sources on the measurement process.
The scoring data from middle school teachers and generative AI tools were analysed using both the SPSS 24.0 statistical software and the EduG 6.1 software, which is specifically designed for analyses based on G theory.
The process was conducted in four stages. In the first stage, the scoring of teachers (18) and AI tools without the use of a rubric was compared. For this design, individuals were denoted as “b,” raters as “p,” and items/tools as “m,” symbolized as a b × p × m design. In the G theory analysis for this design, there are three main effects: individuals, raters, and items/tools, as well as interaction effects including individual–rater (b × p), individual–item (b × m), rater–item (p × m), and individual–rater–item (b × p × m). Reliability coefficients were calculated using relative and absolute G coefficients.
In the second stage, the scoring of teachers (18) and AI tools with the use of a rubric was compared, and the same design was applied. In the third stage, data obtained from the scoring of middle school teachers and generative AI tools were interpreted using descriptive statistics and visualized with graphs. For this purpose, the processes were carried out in SPSS, and the data were visualized using graphs. In the fourth stage, the opinions of teachers and AI tools regarding the research findings were analysed using the inductive thematic analysis method. The manageable volume of qualitative data allowed for a manual analysis procedure without the need for specialized software. The coders engaged in a multi-step process: first, immersing themselves in the data for familiarization, followed by generating initial codes using a colour-coding system with highlighters to systematically categorize textual segments. These initial codes were then iteratively clustered into broader sub-themes and, ultimately, final themes.
Two independent coders analysed the data set to ensure analytical reliability. Cohen’s kappa analysis was conducted to determine the intercoder agreement. The analysis yielded a reliability score of 0.88 between the two coders. This specific value is considered highly reliable within the relevant literature and confirms the consistency of the analytical process. The final themes/codes derived from these opinions were presented in graphical form.
3. Findings
In this section, the findings obtained from quantitative and qualitative data are presented in the form of graphs and tables based on the research questions and are reported in an explanatory manner.
3.1. Findings Related to First Research Question
Regarding the first research question, “What is the distribution of scores given by teachers and AI tools in evaluations conducted without a rubric?”, the aim was to determine the scoring behaviours exhibited by teachers and AI tools in evaluations conducted without the use of a rubric, as well as the distribution of these scores. The data obtained within this framework are detailed in Figures 2 and 3.

Scatter plot of middle school teachers’ scoring without a rubric.

Scatter plot of artificial intelligence (AI) tools’ scoring without a rubric.
In the initial evaluations conducted without a rubric (see Figure 2), significant variability was observed in the scoring results of teachers. For 6th-grade essays, the minimum and maximum scores ranged from 13 to 19, with average scores varying between 15.67 and 16.72. The standard deviation values ranged from 1.49 to 2.93, indicating high variability among teachers’ scores. For 7th-grade essays, the minimum and maximum scores were between 14 and 20, with average scores concentrated between 15.89 and 16.67, and a standard deviation of 1.71. In the case of 8th-grade essays, the minimum and maximum scores ranged from 13 to 20, with average scores between 15.33 and 18.61, and the highest standard deviation (2.62) recorded in this group. These findings highlight significant differences in score ranges and consistency in evaluations conducted without the use of a rubric.
In the initial evaluations conducted without a rubric (see Figure 3), the scoring results of AI tools demonstrated consistent patterns. For 6th-grade essays, the minimum and maximum scores ranged from 15 to 18, with average scores concentrated between 16.00 and 18.75. The standard deviation values showed little variation among the tools, remaining between 0.50 and 0.82. For 7th-grade essays, the minimum and maximum scores fell within a narrow range of 17 to 18, with average scores ranging from 17.00 to 17.50 and standard deviations generally between 0.00 and 0.58. For 8th-grade essays, the minimum and maximum scores ranged from 18 to 19, with average scores concentrated between 18.50 and 18.75. The standard deviation values in this group also remained consistent, between 0.50 and 0.58. These findings indicate that, in evaluations conducted without a rubric, there was minimal variability among AI tools across grade levels.
3.2. Findings Related to Second Research Question
The second research question, “Are there differences between teachers and AI tools in evaluations conducted without a rubric?”, examined the effects of individual, rater, and item (essay) on the scoring of essays from students at different grade levels. The consistency between the two groups was analysed, and the findings are presented in Table 2.
Analysis of Variance (ANOVA) and Generalizability Theory Findings (Without a Scoring Rubric).
S = student; R =r; I = item. G coefficient: relative = 0.17, absolute = 0.16. Grand mean level = 16.64. **S = student; R = rater; I = item. G coefficient: relative = 0.97, absolute = 0.88. Grand mean level = 17.42. df = degrees of freedom. SS = Sum of squares. MS=Mean square.
The findings related to middle school teachers’ scoring revealed that the largest source of variance was the student–rater–item (SRI) interaction, accounting for 80.9% of the total variance. This was followed by rater (R) at 9.3% and student–item (SI) interaction at 8.4%. The student (S) effect was the smallest source of variance, contributing only 1.3%. This indicates that teachers had limited ability to differentiate among individuals during the evaluation process, with the rater effect playing a significant role.
In contrast, for generative AI tools, the largest source of variance was the student (S) effect, contributing 66.7%. This was followed by the student–rater–item (SRI) interaction at 16.3% and the item (I) effect at 12.1%. The rater (R) effect was limited to 5.0%, demonstrating that AI tools significantly reduced the impact of rater variability. In terms of reliability coefficients, significant differences were observed between teachers and AI tools. For teachers, the relative G coefficient was calculated as 0.17, and the absolute G coefficient as 0.16. In contrast, for generative AI tools, these values were 0.97 and 0.88, respectively. This highlights that the reliability of AI tools in evaluations was considerably higher than that of teachers.
When examining the overall average scores, the mean score for teachers was 16.64, and for generative AI tools it was 17.42. These findings indicate that generative AI tools were more effective at distinguishing between individual differences, eliminating the rater effect, and providing higher reliability in evaluations compared to teachers.
3.3. Findings Related to Third Research Question
The third research question, “What is the distribution of scores given by teachers and AI tools in evaluations conducted using a rubric?”, aimed to determine the scoring behaviours exhibited by teachers and AI tools, as well as the distribution of these scores, when a rubric was used. The data obtained within this framework are detailed in Figures 4 and 5.

Scatter plot of middle school teachers’ scoring with a rubric.

Scatter plot of artificial intelligence (AI) tools’ scoring with a rubric.
In the second evaluations conducted using a rubric (see Figure 4), teachers’ scores were concentrated within a narrower range. For 6th-grade essays, the minimum and maximum scores ranged from 9 to 17, with average scores recorded between 13.44 and 16.44. The standard deviation values were also limited to a narrower range, between 1.21 and 2.38. For 7th-grade essays, the minimum and maximum scores ranged from 12 to 18, with average scores concentrated between 15.33 and 16.44, and standard deviations ranging from 1.64 to 2.38. For 8th-grade essays, the minimum and maximum scores ranged from 9 to 17, with average scores between 13.44 and 15.33, and the standard deviation recorded at 1.21. Overall, the use of the rubric reduced the range between minimum and maximum scores, placed the average scores within a more balanced range, and decreased the standard deviation, thereby increasing the consistency of teachers’ scoring.
In the second evaluations conducted using a rubric (see Figure 5), AI tools demonstrated even higher consistency in their scoring results. For 6th-grade essays, the minimum and maximum scores ranged from 15 to 17, with average scores between 15.50 and 16.75. Standard deviation values remained low, between 0.00 and 0.58. For 7th-grade essays, the minimum and maximum scores ranged from 17 to 18, with average scores between 17.00 and 17.50, and standard deviations between 0.00 and 0.58. In particular, the standard deviation of 0.00 observed in the evaluation of S71 indicates that the tools assigned identical scores. For 8th-grade essays, the minimum and maximum scores ranged from 18 to 19, with average scores concentrated between 18.50 and 19.00. Standard deviation values also ranged from 0.00 to 0.58. These findings reveal that the use of a rubric increased the consistency of evaluations, providing more balanced results in terms of both minimum-maximum values and standard deviations.
3.4. Findings Related to Fourth Research Question
The fourth research question, “Are there differences between teachers and AI tools in evaluations conducted using a rubric?”, examined the effects of individual, rater, and item (essay) on the scoring of essays from students at different grade levels. The consistency between the two groups was analysed, and the findings are presented in Table 3.
Analysis of Variance (ANOVA) and Generalizability Theory Findings (With a Scoring Rubric).
S = student; R = rater; I = item. G coefficient: relative = 0.70, absolute = 0.60. Grand mean level = 15.07. **S = student; R = rater; I = item. G coefficient: relative = 0.99, absolute = 0.91. Grand mean level = 17.38. df = degrees of freedom. SS = Sum of squares. MS=Mean square.
The findings related to middle school teachers’ scoring revealed that the largest source of variance was the individual–rater (RI) interaction, accounting for 43.2% of the total variance. This was followed by the student–rater–item (SRI) interaction at 27.2% and the student–rater (SR) interaction at 24.7%. The student (S) effect contributed only 4.9% of the variance, and the rater (R) and item (I) effects were negligible, each contributing 0.0%. These results indicate that teachers had limited ability to distinguish individual differences, with individual–rater–item interactions playing a significant role in the scoring process.
In contrast, for generative AI tools, the largest source of variance was the student (S) effect, accounting for 74.9% of the total variance. This finding demonstrates that AI tools were much more effective in distinguishing individual differences. Following this, the student–rater–item (SRI) interaction contributed 9.0%, and the rater (R) effect accounted for 8.4%. Despite the use of a rubric, AI tools minimized the impact of rater variability in the evaluation process.
In terms of reliability coefficients, differences were observed between teachers and AI tools. For teachers, the relative G coefficient was calculated as 0.70, and the absolute G coefficient as 0.60. In contrast, for generative AI tools, these values were 0.99 and 0.91, respectively. These findings highlight that AI tools demonstrated significantly higher reliability in evaluations compared to teachers
The overall average scores were calculated as 15.07 for teachers and 17.38 for AI tools. These results suggest that AI tools outperformed teachers in distinguishing individual differences, enhancing consistency in scoring processes, and providing highly reliable evaluations.
3.5. Findings Related to Fifth Research Question
The fifth research question, “What are teachers’ views on the research findings?”, was addressed by posing two questions to middle school teachers and AI tools. The findings obtained through inductive thematic analysis of the responses are presented in Figure 6.

Teachers’ and artificial intelligence (AI) tools’ perspectives on research findings.
Teachers expressed that various external factors and subjective influences played a significant role in the scoring process. They noted that individual biases, mood, experience, and professional background impacted their evaluations. Additionally, teachers highlighted that focusing on different priorities during scoring also affected the results. From an improvement perspective, teachers emphasized the potential of rubrics to produce more consistent outcomes and stated the need for in-service training and practice-based workshops. They also suggested that AI-supported feedback and comparative methods could enhance the scoring process.
AI tools, on the other hand, attributed their consistency in evaluations to algorithmic and data-driven approaches, emphasizing that their scoring process was free from biases and external influences. From an improvement perspective, they highlighted the need for broader and more diverse datasets and emphasized that advancements in deep learning techniques and language models could enhance conceptual understanding. They also stressed the importance of future integration processes, indicating that they could play a more effective role in evaluation processes.
The findings suggest that whereas teachers focus on subjective and experience-based factors during the evaluation process, AI tools offer a more systematic and consistent approach. Both parties acknowledged the positive impact of rubrics on consistency and reliability; however, teachers placed greater emphasis on the applicability of these guides and their contribution to professional development. These results demonstrate that human and AI evaluations possess complementary characteristics, highlighting the potential for collaboration to achieve more consistent evaluation processes.
4. Discussion
The primary objective of this study was to compare the scoring behaviours, consistency, and reliability of middle school teachers and generative AI tools in the evaluation of student essays, both with and without the use of a scoring rubric. The findings obtained from this research offer noteworthy insights into the variability of human assessment versus the consistency of algorithmic scoring, and they point to the potential of AI to address persistent challenges in educational evaluation, including rater reliability and subjectivity.
One of the most significant findings concerns the evaluations conducted without a rubric, where teachers exhibited high variability in their scoring. The scatter plots and descriptive statistics revealed that teacher scores fluctuated widely, with standard deviations reaching as high as 2.93 for 8th-grade essays and average scores differed significantly. This pattern is well documented in the assessment literature; Barkaoui (2010) demonstrated that rater experience and scale interpretation contribute substantially to scoring variability, and Knoch (2009) similarly reported that even trained raters diverge when rubric guidance is absent. The G theory analysis further clarified this instability; for teachers, the largest source of variance (80.9%) was the interaction between student, rater, and item (SRI) rather than the differences between students themselves. This indicates that without a rubric, teacher scores were heavily influenced by undefined variables rather than the actual proficiency of the students, a finding consistent with Jonsson and Svingby’s (2007) observations on the subjective nature of manual scoring. In contrast, the generative AI tools demonstrated remarkable consistency even in the absence of a rubric, with standard deviations remaining between 0.50 and 0.82. The variance analysis showed that the primary determinant of AI scores was the “student” effect (66.7%), suggesting that the AI tools were distinguishing between varying levels of writing proficiency rather than introducing random error. This pattern is broadly consistent with research showing that automated scoring systems can reduce subjectivity and inter-rater variability (Wind et al., 2018). Wang et al. (2021) similarly found near-perfect agreement between human and computer-automated scoring in Chinese student writing, and Voss et al. (2026) reported that machine–human agreement rates equalled or exceeded human–human agreement in Spanish writing assessment.
The introduction of a scoring rubric notably improved the performance of the human raters, supporting the well-established pedagogical view that rubrics are essential for fair assessment (Brookhart, 2013; Jonsson & Svingby, 2007). When using a rubric, the teachers’ score ranges narrowed, and their standard deviations decreased, leading to a more balanced range of scores. The reliability analysis confirmed this improvement, with the teachers’ relative G coefficient rising from a negligible 0.17 without a rubric to a respectable 0.70 with one. However, even with the aid of a rubric, human consistency did not reach the levels achieved by the AI tools. The AI tools maintained near-perfect reliability in this condition, with a relative G coefficient of 0.99 and standard deviations as low as 0.00 in some specific cases. One observation that stands out in this regard is that when guided by the same rubric, teacher scoring appeared to show near-zero variance across certain rater sources, whereas AI scoring still demonstrated some degree of rater variance. This discrepancy suggests that rubrics may constrain human raters to the point of artificial uniformity on certain dimensions, possibly because teachers tend to anchor on the same salient rubric descriptors. AI tools, on the other hand, process each essay independently through their algorithms and produce slight but genuine variation that reflects actual differences in essay quality. Put differently, the near-zero teacher variance may not indicate true agreement but rather a ceiling effect of rubric-guided convergence. This finding warrants further investigation, as it raises questions about whether rubric-based human scoring conflates agreement with accuracy. Seßler et al. (2025) reported a similar pattern in their comparison of LLM and teacher ratings in multidimensional essay scoring, noting that AI tools can capture dimension-level differences that human raters tend to flatten. Ferrara and Qunbar (2022) also cautioned that high inter-rater agreement among humans does not necessarily constitute evidence of valid scoring, particularly when raters converge on surface-level features.
A critical aspect of assessment is the ability to differentiate between student abilities. The study found that teachers, particularly without rubrics, had a limited ability to differentiate among individuals, with the “student” effect contributing only 1.3% to the total variance. Instead, the “rater” effect and interactions dominated. This implies that a student’s score was less about their writing and more about who was grading it and under what specific conditions. Conversely, the AI tools proved highly effective at differentiation. In the rubric-supported evaluation, the “student” effect accounted for 74.9% of the total variance for AI tools. This demonstrates that AI is not just assigning random numbers but is sensitive to the qualitative differences in student essays. This parallels findings by Aydın et al. (2025) in the context of Turkish essays, where high correlations between human and AI scores suggested that AI can mirror valid assessment standards while maintaining higher consistency. The AI’s ability to minimize the “rater” effect, keeping it to a negligible contribution, suggests it can act as a powerful tool for ensuring equity in large-scale assessments. This is consistent with Jung et al. (2024), who found that automated scoring combined with machine translation can offer a consistent and resource-efficient approach for evaluating multilingual student responses.
The inductive thematic analysis of teachers’ and AI tools’ perspectives on the research findings revealed complementary yet distinct patterns. When asked about scoring discrepancies, middle school teachers pointed to a range of factors: the inclusion of external considerations in their evaluations, shifting focus across scoring criteria, the weight of subjective judgement, personal biases, mood at the time of grading, and their own professional background. As for improvement, teachers highlighted the value of rubrics in producing more consistent outcomes, the need for in-service training and hands-on workshops, and the benefits of comparing their scores with AI-generated feedback. The AI tools, for their part, attributed their outputs to algorithm and data-driven review processes and to their freedom from personal prejudice. They also suggested that future improvements could come from broader and more diverse datasets, deeper integration into classroom practice, and advances in deep learning and language modelling to strengthen conceptual understanding. These qualitative findings resonate with Wilson et al. (2021), who reported that teachers view AWE as both helpful and challenging, and that its uptake is shaped by the wider instructional context. The fact that teachers openly acknowledged the influence of their own biases and mood is worth noting. It confirms what the quantitative results already indicated: human scoring remains vulnerable to construct-irrelevant variance. At the same time, the AI tools’ recognition of their own limitations in grasping meaning at a deeper level echoes concerns raised by Jin and Hu (2026), who observed that current AI systems can still fall short when it comes to evaluating logic, coherence, and creativity.
The stronger reliability of AI observed in this study does not necessarily suggest the removal of teachers from the evaluation process. Rather, it highlights the potential of AI as a complementary tool to enhance efficiency and fairness. As noted by Jin and Hu (2025), although AI systems have advanced significantly, they may still struggle with nuanced aspects of writing such as logic, coherence, and creativity. Therefore, the high reliability of AI should be viewed as an asset for handling the mechanical or standardized aspects of scoring, freeing teachers to focus on providing qualitative feedback. The reduction of teacher workload is another significant implication. Zhu (2025) emphasizes that the efficiency of AI models can provide valuable auxiliary tools for teachers, thereby reducing their workload. By acting as a second reader or a calibration tool, AI can help mitigate the inconsistencies observed in the teacher evaluations in this study. This collaborative approach aligns with the perspective of Ruiz Herrero (2016), who advocates for a balanced view where the benefits of automated systems are weighed against the need for human contextual understanding. Additionally, Zhang and Yuan (2022) suggest that such scoring systems can be effectively applied in school settings to address the shortcomings of traditional scoring methods. Stоšić and Malyuga (2024) also highlight that AI-driven tools promise effective language learning and testing experiences by enhancing precision and personalization.
In conclusion, this study demonstrates that generative AI tools significantly outperformed middle school teachers in scoring consistency, reliability, and the capacity to differentiate between student proficiency levels, especially when no rubric was provided. The near-zero variance observed in rubric-guided teacher scoring, set against the residual but meaningful variance in AI scoring, is a finding that calls for further empirical attention. These findings support the integration of AI into writing assessment not as a substitute for human educators but as a means of promoting fairer, more accurate, and more efficient evaluation in language education.
5. Conclusion
This study examined the differences in scoring behaviours and reliability between middle school teachers and generative AI tools in the evaluation of student essays. The findings clearly indicate that generative AI tools offer a level of stability and precision that is difficult to achieve through manual evaluation alone. Although the use of a rubric significantly enhanced the consistency of teacher evaluations, it did not enable them to surpass the reliability coefficients observed in AI tools, which maintained high standards regardless of whether a scoring guide was provided.
A critical insight from this study lies in how scores are derived. The analysis revealed that AI scoring was primarily driven by differences in student performance, whereas teacher scoring was heavily shaped by rater interactions and external variables. This suggests that integrating AI into the assessment process can effectively mitigate the subjective limitations of human scoring. The ability of AI to consistently differentiate between proficiency levels further reinforces its potential as a reliable scoring mechanism. It is also worth noting that the near-zero variance observed in rubric-guided teacher scoring, when set against the residual but meaningful variance produced by AI, raises important questions about the nature of human agreement. What appears to be consistency among teachers may, in some cases, reflect a ceiling effect of rubric-guided convergence rather than genuine evaluative precision.
However, these results should not be interpreted as a call to replace educators. Instead, AI should be viewed as a valuable supporting mechanism that helps teachers by handling routine scoring tasks and ensuring fairness. Such integration aims to reduce teacher workload while improving the overall quality and efficiency of evaluation, without disregarding the pedagogical context in which assessment takes places.
In this study, grade level served as a practical and ecologically valid indicator of expected writing development, as it mirrors the natural progression outlined in the national curriculum. Although this approach perfectly aligns with how real classrooms structure writing instruction, it is worth noting that proficiency differences between middle school grades may not always be uniform across all students. Including an independent measure, like a standardized writing test, would have added another layer of validation and strengthened our overall conclusions. Even so, the fact that the scores generated by AI successfully captured systematic variations across grades suggests these tools are sensitive to real developmental differences. Future studies can build on this finding by pairing grade level comparisons with external benchmarks. This will clarify exactly what differences the AI is detecting and boost confidence in the diagnostic abilities of automated scoring.
Footnotes
Appendix
Writing Skill Analytical Scoring Rubric.*
| Score | Spelling and punctuation | Topic | Language use | Structure and organization |
|---|---|---|---|---|
| 0 | Unreadable; contains a significant number of spelling and punctuation errors. | The topic is not fully present or has not been developed with details. | Grammar and vocabulary usage are inadequate. | The text lacks any structure or organization. |
| 1 | Contains numerous spelling and punctuation errors. | The topic is unclear, and the details are not entirely relevant to the topic. | Grammar and word choices negatively affect the meaning. | The text exhibits weak organization and structural features. |
| 2 | Spelling and punctuation errors are common, but the text is still understandable. | The topic is clear but has not been developed with relevant details. | Grammar and vocabulary usage are generally correct, but some errors are present that slightly affect the meaning. | The text does not fully include the basic structural elements (introduction, body, and conclusion). |
| 3 | Contains some spelling and punctuation errors. | The topic is clear but has been partially developed with the given details. | There are recurring errors in grammar and vocabulary usage, but they do not affect the meaning. | The text includes the basic structural elements (introduction, body, and conclusion) but has some issues. |
| 4 | Contains few spelling and punctuation errors. | The topic is clear and developed with relevant details. | Grammar and vocabulary usage are nearly flawless, with only a few minor errors. | The text is well-organized, fluent, and uses a structure/organization with few errors. |
| 5 | Exhibits excellent spelling and punctuation. | The topic is clear and developed with strong and relevant details. | Grammar structures and vocabulary appropriate to the level are used flawlessly. | The text is excellently organized, impactful, and uses a flawless structure. |
The English version of the rubric is provided as a recommendation. The researchers intending to use the English version are advised to reassess the rater (inter/intra) reliability of the rubric.
Ethical Considerations
In the preparation of this manuscript, the authors adhered to the ethical guidelines and principles outlined by Budapest Ethical Guidelines, which include the principles of honesty, objectivity, openness, accountability, respect and 1964 Helsinki Declaration. Informed consent was obtained from all participants involved in the study. These forms clearly outlined the nature of the research, the procedures involved, the potential risks and benefits, and the rights of the participants, including the right to withdraw from the study at any point without any penalty.
Funding
The author disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: Open access funding provided by the ANKOS.
Declaration of Conflicting Interests
The author declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data Availability Statement
Data used within the scope of this research is available on request.
Use of Generative AI
Because of the subject of the manuscript, generative AI tools were used to score student essays.
