Abstract
It remains unclear if perceptions of academic dishonesty concerning artificial intelligence writing technologies (AIWTs) present new challenges or if they reflect prior, non-AI concerns. To structure this problem, we used a randomized control survey experiment. We compared student (n = 603) and instructor (n = 312) attitudes toward dishonesty in collaborations involving humans versus AIWT in 10 writing-related scenarios. Results suggest similar perception patterns among students and instructors, with both populations expressing significant differences in perceived dishonesty between AI and human collaborators in some scenarios. This experiment structures the problem of AI writing and academic dishonesty for future research in this emerging field.
The emergence of artificial intelligence writing technology (AIWT), such as OpenAI's GPT series, has led to a flurry of publication and pedagogical activity about writing. A question that arises is whether, and to what degree, AIWT presents a new challenge in academic writing integrity or if it is a new tool that reflects older concerns. One way to answer this question is to structure the problem. Problem-structuring is the “lack of information in start states, goals states, and transformation functions [that] will require extensive problem structuring before problem solving can commence” (Goel & Pirolli, 1992, p. 405). Problem structuring, then, generates and solicits “information that further structure[s] the problem” (p. 410). In other words, problem structuring in the context of AIWT involves identifying perceptions of these technologies and distinguishing these perceptions from non-AI related perceptions. By systematically examining these factors with a baseline of comparison, researchers can develop strategies and guidelines about AIWT and writing more broadly. Structuring the intersection of academic dishonesty and AIWT is of particular importance due to the attention this intersection has received in the media and from educators.
Regarding AIWT, the topic of academic dishonesty is both (a) prevalent and (b) ill-structured. First, this topic is prevalent, in our view, because it is a concrete issue that writing instructors have encountered (Byrd et al., 2023; Hartwell & Aull, 2023) as well as a perceived problem hyped by an inordinate number of policies about AI (Prem, 2023). Students can use these technologies to bypass writing course outcomes (Dale & Viethen, 2021) although faculty experiences (Fyfe, 2023) and quasi-experimental research (Ernst, 2020) suggest that students find effective use of these technologies difficult. Second, the intersection of academic dishonesty and AIWT is ill-defined because it is not clear how these concerns differ in substance from previous concerns about dishonesty involving other forms of assistance, such as peer collaboration. To structure this problem, we set about to determine perceptions of classroom writing collaboration in scenarios that used either AIWTs or human agents. We used a fully experimental design that randomly exposed writing instructors and students to writing scenarios and measured perceived levels of academic dishonesty in order to determine differences between the groups’ perceptions.
The rest of this article has six parts. First, we review the intersection of AIWT and academic dishonesty. Second, we explain our methods, including the design of the study, data collection, and analyses performed. Third, we report our findings. Fourth, we offer discussion of the findings. Fifth, we describe the limits of the study and directions for future study. We conclude with comments about the potential for using the interest and hype around AI to motivate studies about writing more generally.
Literature Review
Academic dishonesty in writing, outside of AI use, is itself an ambiguous and difficult concept to study. Complexity arises in part because academic dishonesty is not a binary, for example, “cheating” or not. Scholars have argued that students never compose in isolation; writing is a deeply social, collaborative activity (e.g., Blakeslee, 2001; Howard, 1999; 2007; Prior, 1998). Therefore, the extent to which a collaborative interaction is indicative of academic dishonesty (“cheating”) is unclear. In cases where many would agree that dishonesty has occurred, frameworks of academic dishonesty must contend with a multitude of variables (e.g., Crown & Spiller, 1998; Ercegovac & Richardson, 2004) motivating the transgression. Reasons include motivation, social pressure, perceived cost (Murdock & Anderman, 2006, p. 139), and students seeing their writing and their education as a commodity (Ritter, 2005). This ambiguity in definition and complexity in motivation makes academic dishonesty particularly challenging to study empirically. In a meta-analysis of academic dishonesty and personality traits of students, researchers found “significant variability in the correlations across studies” (Giluk & Postlethwaite, 2015, p. 63). Further, the authors noted a lack of uniformity in studies and uncertainty in the validity of measures used (p. 64) that was reflected in previous (Davis et al., 1992) and subsequent meta-analyses (Marques et al., 2019).
Despite these challenges, a frequently addressed starting point to the study of academic honesty is the determination of student versus faculty perceptions of relevant scenarios. In some studies, these perceptions appear to align (McNair & Haynie, 2017) while disagreement between students and faculty tends to be the more common finding (Kidwell et al., 2003; Paullet, 2020; Schmelkin et al., 2008). Thus, an assumption about whether there are differences between students and instructors should not be made without evidence, as we aim to provide here.
AIWTs, such as ChatGPT, are not search engines. Instead, the fundamental architecture of AIWT—namely the transformer (Vaswani et al., 2017)—creates texts from prompts using predictive analytics based on massive datasets and machine learning methods. The content produced from these models does not retrieve information directly from websites, but rather is procedurally generated from the user's prompt based on training data. Procedural generation is a characteristic feature of AIWT, and it adds another layer of ambiguity to the above concerns because the status of its output is not clear in terms of academic dishonesty, such as originality or correctness (the latter of which is mislabeled in the media as hallucinations). Given these vagaries and the high potential for variability in response for a given user input, AIWT complicates notions of agentive responsibility (i.e., “who is responsible for this work?”).
Research about academic dishonesty has attempted to disambiguate this problem by identifying student perceptions of AIWTs. These studies typically begin with a descriptive survey that identifies how students from different disciplines perceive AI. Such studies include topics about students’ experiences and attitudes with using AI (Busch et al., 2023; Chan & Hu, 2023), students' beliefs about AI's impact on society (Jeffrey, 2020) and fears about AI in terms of students’ professional careers (Busch et al., 2023; Jeffrey, 2020). Thematic analysis of engineering student perceptions of ChatGPT identified consistent findings while noting that students did not see ChatGPT as a major threat to learning or academic integrity (Shoufan, 2023, p. 38814).
Conversely, studies have sought to examine instructor perceptions of AI, typically finding that instructors have a negative view of AI. In a questionnaire and interview study of 67 instructors who taught English as a foreign language, Mohammadkarimi (2023) found that these instructors were concerned about the negative possibilities of AIWTs: Unquestionably, teachers unanimously concurred regarding the adverse influence of AI on the academic integrity of their students. They held the belief that AI had amplified the accessibility and allure of academic dishonesty for students, impeding the cultivation of fundamental general and transferable skills. (p. 6)
Two steps are necessary to advance the study of AIWT and academic dishonesty. First, studies must go beyond preliminary perceptions of AIWT. The above studies generally are descriptive in nature without a control with which to compare their findings. Second, studies must concretize perceptions around AIWT by providing specific situations. To achieve these two advances, researchers need to design experiments that include both AI and non-AI scenarios, thereby allowing for clarification as to what degree AIWT's impact on academic dishonesty is new. Our research takes these necessary steps; it compares student and instructor perceptions of AI collaborators across a set of scenarios anchored on a control of a human collaborator, that is, an experimental approach with random assignment.
1
The scenarios are the same except for varying the agent, human or AI. Our research questions (RQs) are thus as follows: RQ1: For writing instructors, is there a differential perception of academic dishonesty in collaborative writing scenarios when the identity of the collaborator is varied between a human and an AI agent? If so, what are the differences? RQ2: For students, is there a differential perception of academic dishonesty in collaborative writing scenarios when the identity of the collaborator is varied between a human and an AI agent? If so, what are the differences? RQ3: How do the answers in RQ1 and RQ2 compare?
Method
This study uses a fully experimental approach with random assignment. We used a Qualtrics survey because a survey approach gauges a population (Krosnick, 1999) and thus enables us to answer our RQs. After entering demographic and related information, participants were exposed to 10 scenarios.
2
Each scenario had two versions: the human and AIWT (see Appendix). The design of questions was drawn, in part, from a modified version of Hidalgo et al. (2021). Participants were exposed to all ten scenarios, with random exposure to the human or AIWT version. The scenarios were varied only by the agent of collaboration. AI use was defined with the examples of OpenAI's GPT series and Google's Bard. The following example highlights the scenario-based approach (underlining indicates differences in the question but did not appear in the survey): Scenario G (version 1): A student in a course has been assigned a paper. They consult Scenario G (version 2): A student in a course has been assigned a paper. They consult
Demographic Information of Participants.
Recruitment
This study was approved (#23989) by the Institutional Review Board. Participants were not paid to take the survey, and participation was voluntary. To recruit students, the first author (Gallagher) visited classrooms in-person from across his home institution, securing instructor permission before visiting. The link to the survey was posted to this institution's reddit home page (reddit.com/r/UIUC). The first author emailed writing program administrators (WPA) to forward the recruitment email to undergraduate listservs at Texas Tech University, Michigan State University, and Purdue University. Student participants’ responses were recorded from April 25 through May 21, 2023. To recruit writing instructors, the first author emailed a variety of professional writing studies listservs and WPA directors, as well as the Association of Teachers of Technical Writing listserv. Links to the survey were posted to the Writing Studies reddit page (reddit.com/r/rhetcomp) and to the first author's Twitter (now X) page. Writing instructor responses were recorded from May 9 through May 25, 2023. The surveys were closed after no responses were recorded for a week.
Measurements and Analysis
We averaged ratings for each scenario, for each population. Figure 1 depicts writing instructor ratings. Figure 2 depicts student ratings. We compared respondent ratings of scenarios with AI collaborators to scenarios with human collaborators. T-tests were used to compare responses for writing instructors and students. We used a Bonferroni correction (Bonferroni, 1935) to account for multiple comparisons within each group. Thus, α (threshold for statistical significance) was set at .005 rather than .05.

Writing instructor ratings of scenarios with AI versus human collaborators.

Student ratings of scenarios with AI versus human collaborators.
Findings
We found that both populations—writing instructors and students—saw differences in academic dishonesty between human and AIWT agents of collaboration. These populations had consistent perceptions across most scenarios. Finding 1: For writing instructors, there is a statistically significant differential perception of academic dishonesty in half of the collaborative writing scenarios when the identity of the collaborator is varied between a human and an AI agent.
To answer the first RQ, in 50% of our scenarios (5/10) there was a statistically significant perceptual difference between a human and an AI agent for writing instructors (see Table 2, Figure 1, and Scenarios B, D, F, G, and H in Appendix). Five scenarios were perceived as not statistically different between human and AI agents of collaboration (A, E, I, J, and K). Finding 2: For students, there is a statistically significant differential perception of academic dishonesty in three of the collaborative writing scenarios when the identity of the collaborator is varied between a human and an AI agent.
Comparison of the Differences Between Writing Instructors’ Perceptions of Human Collaborator and AI Collaborator Scenarios.
*Statistical significance.
Note. df = degrees of freedom; d = Cohen's d, a measure of effect size; t = t statistic; CI = confidence interval. Scenario C was eliminated because it is a duplicate of Scenario D.
For the student population, three of the 10 scenarios met the threshold for statistical significance (see Table 3, Figure 2, and Scenarios B, G, and H in Appendix). Seven scenarios were perceived as not statistically different between human and AI agents of collaboration (A, D, E, F, I, J, and K). Finding 3: Three scenarios were perceived as statistically significantly different for AI compared to human collaborators for both instructors and students.
Comparison of the Differences Between Students’ Perceptions of Academic Dishonesty in Human Collaborator and AI Collaborator Scenarios.
*Statistical significance.
Note. df = degrees of freedom; d = Cohen's d, a measure of effect size; t = t statistic; CI = confidence interval. Scenario C was eliminated because it is a duplicate of Scenario D.
Scenarios with statistically significant differences between perceptions of AIWT and human agents of collaboration are generally shared between writing instructors and students. In terms of differences between human and AI agents of collaboration, there were more scenarios with statistically significant findings for writing instructors than for students (see Table 4 and Figure 3). Most of the statistically significant scenarios rated AIWT agents of collaboration as more academically dishonest. Writing instructors and students appear to share some perceptions of AIWTs in the same way, with variation about the degree to which a scenario is (or is not) academic dishonesty. There were no scenarios in which writing instructors did not perceive AIWTs as academically dishonest when students did. In Figure 3, the shaded regions indicate the scenarios in which we found significant differences for instructors and students. In most of the scenarios with statistically significant differences, AI collaborations were considered to be more academically dishonest (the black shaded region in Figure 3).

Comparison between instructors’ and students’ scenarios in which there were significant differences in their rankings of academic dishonesty between AI collaborations and human ones.
Comparison of Statistical Significance and Effect Size Across the Scenario Rankings of Students and Instructors.
*Statistical significance.
Note. d = Cohen's d, a measure of effect size. Negative effect sizes indicate a perception that an AI collaboration is more dishonest than a human one whereas positive effect sizes indicate a perception that a human collaboration is more dishonest than an AI one. Scenario C was eliminated because it is a duplicate of Scenario D.
One nuance in this finding is the relative difference in magnitude between instructor and student responses for Scenario B (see Table 4). We used Cohen's d to characterize the magnitude of differences. Cohen's d reframes the difference between two means in standard deviation units. Thus, d = .5 means that there is a difference between two means of one half of a standard deviation. Typically, d = .2 is a small difference, d = .5 is a medium one, and d = .8 is a large one. Scenario B is as follows (differences are underlined): Version 1: A student has Version 2: A student has
Discussion
Preliminary results of this study suggest that while there are perceptual differences between AI and human agents of collaboration, writing instructors and students share some of these perceptions. As a result, we suggest that separate training programs may not need to be developed for writing instructors and students about AIWT or non-AIWT with respect to academic dishonesty.
When compared to human collaborators, AI agents of collaboration are perceived as more academically dishonest for both writing instructors and students when the AI produces text. For policy and teaching practices, then, we suggest being explicit about AIWT use when it comes to the production of text. Policy makers and writing instructors might address situations that involve the production of text in concrete ways. Readers of this article could use our scenarios in the Appendix as a starting point.
But our findings suggest there are largely no differences in perceptions between AI and human agents of collaboration when it comes to brainstorming or invention strategies. For these situations of academic dishonesty, we believe that policies about AIWT may not need to be substantially different than current policies involving human collaboration.
Two scenarios that we believe are worth discussing together are G and K. Scenario G, which discusses academic phrasing, is an outlier in which the AI collaborator is considered less academically dishonest than a human collaborator. Both populations generally rated an AI and a human collaborator for academic phrasing as not academic dishonesty (Figures 1 and 2). The difference in ratings for scenario G is statistically significant for both the writing instructors and students. This finding could indicate a general acceptance of external input (AI or human) in the context of academic phrasing, meaning that assistance in smaller scale edits is perceived as less academically dishonest. Given that scenario K (about brainstorming using AI or a human) is not significantly different between AI or human and is generally not perceived as academically dishonest in terms of AI and human (see Figures 1 and 2), collaboration may be more acceptable for both AI and human collaborators during preliminary writing stages and at smaller scales of textual production.
Finally, we would call readers’ attention to the unsettled nature of the scenarios for both human and AI agents of collaboration (Figures 1 and 2). In these figures, any score between 3 and 4 are ratings in which the participant reported “not sure” to some degree. These figures depict the complexity and nuanced nature of classroom writing and related literate activities. There is uncertainty and disagreement in participant responses from instructors and students for both AI and human agents. Our scenario D provides an emblematic example of this uncertainty. In this scenario, the student has an agent of collaboration (human or AI) produce an outline. The student then writes a paper themselves. The ratings from writing instructors and students of this scenario were marked close to the “unsure” category for both human and AI agents of collaboration. In other words, respondents simply are uncertain about issues of academic dishonesty when a student has someone (a human) or something (an AIWT) create an outline for them. Such ambiguity underscores, we suggest, the need for addressing the challenge of academic dishonesty in tangible, real-world contexts.
Limitations and Suggestions for Future Study
This investigation is an early study attempting to structure the problem of AIWT and academic dishonesty. Even given the hype and the rapid pace of publication around AI, there is a great deal about the impact of AIWT that we do not know. We hope that our study here begins to formulate and demarcate the intersection of AIWT and academic dishonesty.
There are clear limits to this study. While our writing instructor survey had open-ended fields for response, we opted against these fields for the student survey due to issues of time and satisficing (Krosnick & Presser, 2010). Analysis of open-ended questions and follow-up interviews is left for future study. As a result, this study is limited to perceptions of scenarios at a general level. A second limitation is the age gap between the instructor and student populations (see Table 1). As a result, perceptual differences between writing instructors and students might be attributable to age rather than role of student or instructor. Future research, we suggest, could account for age differences. For other possible future research trajectories, we suggest the following:
Run a comparable survey to the one we deployed here while accounting for demographics in how respondents perceive humans versus AIWT. Run a comparable survey to the one we deployed but rather than focus on academic dishonesty, researchers could examine professional writing situations involving AIWT, such as in law, engineering, or computer programming. Reframe this study from dishonesty to more positive language and connotations, such as academic honesty or types of positive collaboration. Rerun the survey for instructors who do not teach writing explicitly, such as science or engineering instructors. Rerun the survey in the future, such as in 5 years, to determine how perceptions have changed since the original survey was implemented.
An important concern to note is the potential misuse of our research findings. To be clear, this study is not focused on catching students cheating or policing instructors or students. Our study focuses on population-level trends, which is typical in social science studies in which variation is expected. Our report highlights general trends and tendencies observed in the study.
Conclusion
In the 1980s, research on writing extensively examined the role of word processing due to the explosion in availability of such technology (e.g., Daiute, 1983; Haas, 1989; Hawisher, 1986). Haas (1989) found, for example, that students were less likely to plan their writing when drafting with word processing than they were when composing drafts by hand. Students’ sentences, too, grew longer when composing with word processing when compared to writing by hand. Daiute (1983) argued that the text editor…eliminates the spatial and aesthetic barriers that are special inhibitors of revising activity. Writers are often reluctant to mess up a carefully written page by crossing out words or cramping inserts between lines and in the margins. Each change made with the text editor is neatly incorporated into the text, and writers feel pride at seeing their texts in this professional-looking format. When writers make corrections by hand, they have trouble considering them as an integral part of the piece. (pp. 136–137)
Here, Daiute observed the aesthetic and practical aspects of writing with a text editor as compared to writing by hand. The text editor, Daiute noted, removes the physical limitations of paper, such as the difficulty of inserting additional text or making revisions on a handwritten page.
Thus, the emergence of word processing made it easier for students to see their work as a nonlinear, recursive project rather than a static piece of text. The ease of editing encouraged more frequent and extensive revisions. The technological innovation of text editors and word processors allowed scholars of this period to make comparisons between how students previously composed and how they composed with this new technology. The emergence of word processing provided opportunities for investigating computer-mediated writing and writing more broadly.
Current researchers of writing are now at a similar juncture with AI wherein comparisons between pre-AI writing and current AI writing can be investigated. The interest—including the hype—around AI is an opportunity for opening lines of empirical inquiry into not only AI-related writing but also writing more generally. Our study here has focused on writing and academic dishonesty but, as we suggested in our future research section, there is much to be gained from comparing AI-related writing with non-AI related writing. We need these comparisons to form better foundations for the study of writing writ large. Without these kinds of studies, any pedagogical strategies or policies risk being based on anecdotes and assumptions about the products created by these AI technologies and their impact on writing processes.
Footnotes
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
Notes
Author Biographies
Appendix
Labels, Descriptions, and Scale for the Human and Artificial Intelligence Scenarios. (Scenario C was eliminated because it was a duplicate of Scenario D.)
