Abstract
The demand for a creative workforce has never been higher, yet schools struggle to teach and assess creativity among students efficiently. Compositions are an effective way to incorporate creativity across the curriculum; however, essays are time consuming to evaluate for quality or creativity. This study explored (a) if high creativity scores are related to high quality and sophistication in academic writing, and (b) if extant text-mining tools effectively identify quality, sophistication, and creativity in academic essays. Four teacher raters analyzed quality, sophistication, and creativity of 230 essays written by students aged 15–17 for Advanced Placement Language and Composition. We also used text-mining tools (e.g., semantic distance, Shannon's entropy, idea density) to score these essays. Teacher-rated creativity scores correlated with quality and sophistication scores, as well as with some of the text-mining tools, suggesting that these tools can capture quality and sophistication in addition to creativity. Implications for educational practice are discussed.
The story of human survival and progress has developed through a capacity for creativity (Puccio, 2017), which is a trainable skill (Fryer & Collings, 1991; Scott et al., 2004). Due to this trainability, the onus for creativity instruction logically settles on the educational sphere. Creative writing is a widely accepted form of creative expression, yet creativity within academic writing is not generally lauded. Teachers instruct using prescribed topics, mechanics over style, and prescribed vocabulary from a unit of study (Cheung et al., 2003). As a result, teachers and students often fail to consider creativity as part of the academic writing process. Whether teachers intentionally identify and measure creativity, students demonstrate it in myriad ways during a typical school day, and teachers may already unwittingly assess creativity in extant academic structures. Therefore, this study explores if existing methods of essay quality capture creativity in writing through Advanced Placement (AP) student essays. Further, given the importance of creativity in education, we explore if creativity of students’ academic writing can be measured by using computational methods.
Defining Creativity in the Domain of Academic Writing
Although variations on the definition of creativity exist, the field has primarily converged on two key characteristics: novelty and usefulness (e.g., Amabile et al., 2018; Runco & Jaeger, 2012; Stein, 1953). Plucker et al. (2004) expanded this definition as follows: “Creativity is the interaction among aptitude, process, and environment by which an individual or group produces a perceptible product that is both novel and useful as defined within a social context” (p. 90).
Here we unpack and contextualize this definition within the domain of students’ academic writing. Creativity has long been associated with the concept of novelty, which most definitions include (e.g., Acar et al., 2017; Diedrich et al., 2015; Guilford, 1962; Runco & Jaeger, 2012). Although it is necessary yet insufficient to capture the entire essence of the definition, newness adds authenticity to creative works. In prompt-based academic writing, novelty is observed through responses providing a unique approach to the prompt by making remote connections to various ideas. In this study, novelty is assessed by semantic distance, where the semantic distance between the response and prompt may signify a mental leap accomplished by making remote associations or conceptual combinations (Acar et al., 2023; Johnson et al., 2022; Mednick, 1962; Scott et al., 2005).
One way to get at novelty is non-repetition, which can be facilitated by the ability to make mental leaps and switch from one mindset to another. In the creativity literature, this is referred to as flexibility (Torrance, 2008) and allows bridging distinct domains by abstract relations and distant analogies (Vendetti et al., 2014). Flexibility is a subset of divergent thinking, which is defined as thinking in multiple directions (Acar & Runco, 2015). Flexibility reflects the divergence and diversification of ideas or possibilities within the creative process (Runco, 1986). This study explores flexibility through idea density, the amount of new information progressively presented, which indicates a lack of repetition (Covington, 2009).
Useful academic writing achieves task-appropriateness by attending to a prompt. Students must balance creativity with quality, to make their writing interesting to the reader while addressing the task at hand. Rather than establish a standard criterion of what is novel or useful, when students write essays, they address a prompt that constrains them to specific parameters (i.e., usefulness) while writing a novel response. Thus, they write under the creative constraint of task appropriateness (Costello & Keane, 2000; Novitz, 1997; Stokes, 2014). Usefulness for this study is ascertained by the human judges.
Creativity's context is important to consider. Students must learn the balance between self-expression and writing for an audience. Creativity is not apparent when it remains inside someone's thoughts; it must interact and intersect within the sociocultural context (Csikszentmihalyi, 1996). The social context of creativity in writing is critical because academic writers respond to prompts with the audience in mind. In this study, teachers are ideal judges and experts for assessing creativity in student writing, as they understand the context in which it is written. Importantly, academic essays in the present study were produced in an AP class, which influences the nature of teacher-student interactions as well as the essays produced.
AP Language and Composition and Advanced Academics
In the present study, we focused on the student essays that were produced for an AP Language and Composition class. Such academic services in the district from which the sample was taken use an advanced academics (Peters et al., 2014) approach wherein students self-select into advanced courses. Advanced academics is a model rather than an identification status and is used to provide differentiated academic opportunities for students who are not appropriately challenged in the regular curriculum (Peters & Matthews, 2016). Within this frame, this study explores whether a relationship exists between sophisticated writing in advanced students and what educational psychologists call “constrained creativity” (Boden, 2004; Novitz, 1997). One measure of such sophisticated writing is available through the College Board's AP free-response questions. Several content area rubrics include a point for writing that is superior in quality (i.e., sophistication, complexity, or depth of analysis), depending on the particular assessment rubric. Advanced students comprise the sample due to their potential for advanced writing; in one study by Kettler and Bower (2017), teachers rated gifted students’ writing samples higher than general education students.
Creativity in Academic Writing
Composition instruction is often singularly focused on mechanics, grammar, and clarity of argument. The attributes of academic writing are of utmost importance for clear communication, but to explore writing talent, educators must also examine creativity and productivity (Olthouse, 2014). As writers move up a continuum from novice to expert, aspects of writing, such as word choice, spelling, and advanced syntax, become more automatic. Advanced writers who have achieved automaticity have mental space free to focus on creative aspects of writing, such as fluency and originality (Chenoweth & Hayes, 2001), as well as humor, visual imagery, and sophisticated syntax (Piirto, 1992).
Although writing is an important way for students to exercise their creativity (Olthouse, 2014), high-level English courses often focus on academic writing at the expense of creative writing (Olthouse, 2012). However, if a link exists between advanced academic writing and creativity, teachers could potentially foster one with the other (Bruning et al., 2013). Students could learn to incorporate creativity into their academic writing to exercise creativity through academic work they already accomplish in the classroom.
Creativity in academic writing has become more recognized in recent years (Cheung et al., 2003), helping students to increase the quality of their essay writing (Hasnudin et al., 2015). For example, divergent thinking models increase students’ efficacy in writing fact-based essays (Ayob, 2020), and aspects of creativity (i.e., productivity, sentence complexity, and lexical diversity) have been correlated significantly with a higher overall score (Koutsoftas & Gray, 2012). Thus, creativity is not a luxury skill that teachers can treat as an addition to the lesson cycle; it must be integrated into the fiber of all thinking and products in the classroom.
Creativity Assessment in the Domain of Writing
Academic writing is an authentic way to conduct assessments because writing empowers students to demonstrate the acquisition of knowledge and skills. Whereas multiple-choice questions may be biased in their limited ability to assess knowledge, open-ended essays allow students to explain what they know rather than be penalized for what they do not. Written assignments allow teachers more flexibility to consider style or nuanced knowledge. Although writing is a qualitative form of self-expression, it can also be quantified to determine differences in quality and creativity. Assessment methods such as rubrics or the Consensual Assessment Technique (CAT; Amabile, 1982) aid teachers in objectively measuring creativity. However, essays must be read and scored to determine these differences. Human raters have historically engaged in this time-consuming endeavor.
Teacher Ratings
When teachers are expected to rate student essays, they are assumed to know what creativity is. Yet some educators misunderstand creativity or call it something else (e.g., Kettler et al., 2018; Mullet et al., 2016) which means they miss opportunities to award knowledge from different angles and to meaningfully train students to engage with the essential skill of creativity. Keeping this in mind, some researchers specified the criteria to look for in evaluating creativity such as productivity, novelty, figures of speech, flexibility, and elaboration (Malgady & Barcher, 1979) or originality and elaboration (Kettler & Bower, 2017), whereas others adopted a more holistic approach relying on their domain expertise. These two approaches are discussed below.
Teacher Rating Options
Rubrics quantify a qualitative skill such as writing. If creativity is subconsciously part of the quality rubric, teachers can be trained to acknowledge and reward creativity in academic writing. In AP Language and Composition, creativity is not an explicit part of the evaluation rubric. However, that does not necessarily mean creativity is out of consideration. For example, one aspect of quality essays is a high level of sophistication. Students who produced the essays used in this study were in an advanced course that prepares them to take the AP Language and Composition exam. Although points are awarded for a strong, defensible thesis, as well as evidence and commentary, the sophistication point is the point of interest for this study. Students can earn the sophistication point in one of four ways: (a) crafting a nuanced argument by consistently identifying and exploring complexities or tensions, (b) articulating the implications or limitations of an argument (either the students’ argument or an argument related to the prompt) by situating it within a broader context, (c) making effective rhetorical choices that consistently strengthen the force and impact of the student's argument, or (d) employing a style that is consistently vivid and persuasive.
Although it is part of a quality rubric, a potential relationship between complexity and creativity was hypothesized to exist for this high-level point. In Bloom's revised taxonomy, creativity is the highest form of thinking (Krathwohl, 2002). Additionally, creative people have complex personalities that manifest in high-level work (Csikszentmihalyi, 1996), which would explain how complex thinking or sophistication might indicate creativity. Thus, one objective of the present study is to examine if sophistication scores are predictive of creativity of student essays.
Consensual Assessment Technique
Another way to assess creativity of products (e.g., artwork, essays, poems, collages) is the Consensual Assessment Technique (Amabile, 1982). Expert raters independently assess works without relying on a rubric, a definition of creativity, or input from other judges; their expertise in the content area allows them to detect variations from the norm. Expert raters are better at creativity evaluations (e.g., poems) than nonexpert raters, and experts have high interrater reliability (Kaufman et al., 2008). In this method, judges compare the set of works within the pool they are given. At least two judges are necessary to establish interrater reliability, and more are ideal, as studies with two or more raters show interrater reliability rates between .70 and .90 (e.g., Amabile, 1996; Kaufman et al., 2004; Runco, 1989). Since this technique relies on experts’ judgment to assess creativity, the judges in the present study were experts in language and composition, not creativity. This method was previously used with poems (Kaufman et al., 2008), short stories (Kaufman et al., 2004), and creative writing (Dollinger, 2003). Typical study samples were college students although it was also applied to younger groups including high school students (Pretz & Kaufman, 2017).
Shift From Human Raters to Text-Mining
Both methods described above rely on human judges, specifically and typically teachers. The ideal assessor of a human brain is, according to Aristotle, the only other rational animal: another human. However, the human brain possesses the ability to engage in intentional and unintentional bias. The “halo effect” (Thorndike, 1920) is a bias toward students wherein teachers score student work (in this case, creativity) higher based on previously held notions of ability, even if it is in other content areas or skills. These biases can hinder a teacher's ability to judge their students’ creative work fairly. Combined with the sheer volume of time required for essay grading, teachers need tools to expedite the process.
Teachers may benefit significantly from automated essay scoring for quality and creativity to combat the dual issue of the halo effect and time constraints. Programs like the one used by the Educational Testing Service (2022) exist to measure essay quality. This software, called Criterion, scores essays using syntax, vocabulary, and organization based on similarity to the extant corpus of essays within its database. This program is not designed for creativity: it cannot determine whether the essay is off-topic or exceptionally creative. The present study aims to extend automated essay scoring to creativity based on computational methods such as text-mining. Fortunately, such methods proved successful in classic creativity assessment tasks, namely divergent thinking.
Automated Scoring of Creativity via Text-Mining
Researchers have explored the capabilities of text-mining for creativity in recent years (e.g., Acar & Runco, 2014; Devlin et al., 2018; Radford et al., 2019; Wang et al., 2022). While most creativity assessments, such as the Alternate Uses and the Just Suppose tests, are based on comparisons of single-word or short prompts and the responses given (Acar et al., 2023; Beaty et al., 2022; Dumas et al., 2021;), text-mining software to assess the creativity (in addition to the quality) of an entire essay may encourage teachers to recognize, honor, and reward creativity, as the speed of automated essay scoring makes creativity through essays more attainable.
The tension between variability and appropriateness is paramount in assessing creativity in essays. Ke and Ng (2019) identified a list of common dimensions of essay quality from a 50-year survey of automated essay scoring and computer-aided text analysis, including grammaticality, usage, mechanics, style, relevance, organization, development, cohesion, coherence, thesis clarity, and persuasiveness. In this list, relevance and thesis clarity linked to constraints (i.e., usefulness), while the dimensions of style, development, and persuasiveness paralleled originality. Notably, computer programs can already measure these constructs, which overlap with writing quality and creativity. Increasing the speed of scoring compositions with computer-aided text analysis and automated essay scoring could ease barriers such as time constraints for teachers to use composition in the classroom.
Semantic Distance
Semantic distance is a measure of the distance from one word to another. It is a valid way to measure divergent thinking because it objectively measures ideas’ conceptual remoteness (Chrysikou et al., 2021; Miura &Takagi, 2015). Forster and Dunbar (2009) were the first researchers to apply semantic distance to creativity in an Alternate Uses task. Other studies have also shown the power of semantic distance to score creativity automatically and objectively (e.g., Acar et al., 2023; Dumas et al., 2020; Green et al., 2012; Heinen & Johnson, 2018; Prabhakaran et al., 2014). Useful methods to measure semantic distance include latent semantic analysis and Global Vectors for Word Representation (GloVe; Pennington et al., 2014). Latent semantic analysis mathematically represents semantic space by comparing texts to a set of training texts to assess relevant content (Chen et al., 2011). Researchers have also used it to measure novelty and stylistic essay qualities related to higher-order thought (Dumas & Dunbar, 2014; Landauer & Dumais, 1997; Latifi & Gierl, 2021). GloVe is similar to latent semantic analysis but uses co-occurrences in a word matrix and compares texts against an aggregated corpus. GloVe is vital in assessing complex language skills (e.g., sarcasm; Eke et al., 2020) and approximating highly reliable human-rated scores of fluency, elaboration, openness, intellect, and self-reported creative activities (Dumas et al., 2021).
Idea Density
Idea density is another measure of creativity in writing that measures the occurrences of new ideas generated within a work. Writers use certain words or phrases to introduce original ideas, so these words or phrases translate to predictors of original ideas (Turkman & Runco, 2019). Idea density programs measure propositions that correspond to verbs, adjectives, adverbs, prepositions, and conjunctions and then reject repetitions to find the number of unique occurrences of an idea. In Turkman and Runco's study, idea density and keywords were significantly correlated. In a series of four studies, Runco et al. (2017) used the Computerized Propositional Idea Density Rater (CPIDR5.1; pronounced “spider”) to measure idea density related to citation impact, eminence, and divergent thinking, as well as in a set of TED Talks about creativity. The researchers found that idea density was a strong predictor of creative potential.
Entropy
The foundational principle of entropy is that uncertainty and disorder in a system are more likely than order. This cognitive uncertainty can be experienced negatively (Hirsh et al., 2012) or positively (Gabora, 2017). In the negative, entropy causes anxiety, while in the positive, it evokes creativity. Claude Shannon's conception of entropy measures the information contained in a message instead of the portion of the message that is determined, or predictable (Shannon, 1951). In writing, entropy is measured by the repetitiveness or complexity of words (Vähäkangas & Pyykkö, 2012). Therefore, a balance of a high entropy score with the ability to write academically could be viewed as the height of creative academic writing.
Entropy allows for an examination of redundancy in language structures or statistical properties measuring frequencies of letter or word repetitions. Due to this ability to measure for redundancy, entropy has been employed in varied subjects, including assessing morals in primary school textbooks (Aqili, 2021) and predicting sentences with high fluency and coherence for story infilling (Ippolito et al., 2019). Entropy's ability to measure redundancy may also make it effective for assessing creativity in academic writing. Lack of repetition is related to flexibility and surprise, both conceptually related to creativity. Thus, essays in which deviations are evident from sentence to sentence might also be highly creative.
The Present Study
Students frequently produce essays that can be assessed for creativity; however, this is not a routine or established practice in schools. We argued that existing methods of essay quality assessment might have been capturing creativity. Specifically, the sophistication point in the AP Language and Composition evaluation rubric might help distinguish highly creative essays from those that are not. Even if the sophistication point can successfully detect creativity, practical value of this information may be limited. Considering the amount of work teachers are expected to perform, removing some of the duties that are amenable to automation can allow for more space for better use of time and resources. Text-mining tools are promising developments to supplement or replace human labor to capture creativity. Further, such feedback from text-mining methods can inform teachers on specific qualities and aspects of writing that are not explicit in a manual evaluation. Thus, it is worth measuring whether text-mining tools can capture the nuances of the sophistication point, essay quality, or essay creativity at comparable levels to human raters. Therefore, this study explored potential relationships between teacher ratings of creativity and quality with automated ratings of creativity in academic essays, guided by the following research questions:
Is the teacher-rated creativity of student essays related to teacher-rated quality ratings and the teacher-rated sophistication level of the essays? Can text-mining tools such as semantic distance, idea density, and Shannon's entropy capture the teacher-rated creativity of student essays? Can text-mining tools such as semantic distance, idea density, and Shannon's entropy capture the teacher-rated quality and sophistication of student essays? If both teacher-rated and text-mining-based creativity scores correlate with the teacher-rated quality and teacher-rated sophistication level of the essays, which method better predicts the quality and sophistication ratings?
Method
We collected essay samples from academically advanced students and recruited four teacher raters to assess the papers for creativity, quality, and the sophistication point. We used a combination of correlations and exploratory regression to determine whether creativity was related to quality in academic essays and if text-mining tools that were intended for creativity assessment could capture essay quality. Since text-mining programs measure creativity, we explored whether teacher-raters or text-mining methods better predicted quality in academic essays. We also assessed which text-mining tools best predicted teacher ratings of creativity, overall quality, and sophistication point.
Sample
We gathered secondary data for students, age 15–17, in an advanced writing course in a large suburban district in a southwestern U.S. state. Each sample and set of demographic data were matched to a numerical identification number, and all identifying information was then removed. This double-blind method, combined with typewritten responses, ensured that the essays remained anonymous and removed bias based on handwriting.
The students who wrote the essays were coded as 51.7% female and 48.3% male and were 47.4% GT-identified and 52.6% general education students. These were advanced students as evidenced by the fact that 84.8% had a composite score of ≥ 120, while 27.4% had a composite score of ≥ 130 on the Cognitive Abilities Test (CogAT). The race or ethnicity of the students who wrote the essays was coded as 42.2% Asian, 36.5% White, 9.1% Hispanic, 6.1% Black, and 6.1% two or more races.
Students responded to one of three prompts (see Table 1) for a total of 230 essays, one per student. Because the text-mining programs used extant corpora and pre-trained models, different prompts were not an issue. All prompts required students to assert a position and then defend their argument with evidence.
Essay Prompt Choices.
Measures
We collected ten measures for each essay: (a) an average of the four teacher-rated creativity (i.e., creativity), (b) an average of the four teacher-rated holistic quality (i.e., quality), (c) an average of the four teacher-rated sophistication (i.e., sophistication), (d) semantic distance, (e) originality, (f) elaboration, (g) idea density, (h) the number of unique words, (i) Shannon's entropy index, and (j) the total number of words. These measures formed the data for related correlations and exploratory regression analyses.
Procedure
We collected the sample from a pool of student essays, which means students were in their natural school rather than a manipulated experimental setting. Students wrote essays based on a structure from the AP Language and Composition exam—the argument essay. The College Board's AP tests are widely administered to advanced students nationwide. Thus, prompts assigned by the College Board in the past added validity because raters had released samples and scoring to use as baselines with the essay quality rubric. The students wrote essays on the topics presented to them as a routine graded assignment, which was treated as a first draft for structures and mechanics. They had 40 minutes, and this time limit allowed for the comparison of creativity, as the written content supplied a raw reflection of idea generation.
Teacher Ratings
The standard number of raters for evaluations of this sort is two to three, but more raters with a high intraclass correlation (ICC) provide more validity (Cohen et al., 2018; Latifi & Gierl, 2021; Powers et al., 2015). Four teacher raters scored the essays for a general quality score, on a scale of 1–6, based on the College Board rubric (College Board, 2019) that included (a) an arguable thesis statement (one possible point), (b) evidence and commentary that follows a line of reasoning (four possible points), and (c) sophistication (one possible point).
Raters then waited one month before rating the papers for creativity. Previous studies had graders use a two-week (Kettler & Bower, 2017) to a two-month (Johnson, 1975) delay between teacher ratings for quality and creativity. This delay, coupled with the blind nature of the essays, allowed graders to score the essays for creativity while minimizing the halo effect.
The teacher raters scored each essay for creativity based on their experience with argument compositions (Amabile, 1982). Since these teachers had extensive experience (i.e., 20-, 18-, 10-, and 7-years’ teaching experience and 15-, 13-, 5-, and 5-years’ experience, respectively, in this particular course), they had insight into what constituted a more creative essay than the norm. Whereas the teacher-graders have routinely taught this course, they were not the teachers of the students who wrote these essays. Following the Consensual Assessment Technique guidelines, raters were instructed to use the pool of essays as the reference point rather than external criteria and they reviewed a subset of the essays before starting to score them. Teacher raters scored creativity on a scale of 1–5: 1 was not at all creative, 3 represented the expected level of creativity for this type of essay, and 5 meant highly creative.
To create single metrics for each essay, we averaged the four raters’ creativity, overall quality, and sophistication point scores. For sophistication, this changed the metric from four binary scores of 0 or 1 (i.e., earned or did not earn the point) into a scaled score from 0–1. These teacher-rated scores were the basis for all correlations and exploratory regressions.
MeanSim
We used the Covington Vector Semantics Tools (CoVec; Covington, 2016) to measure semantic distance. First, CoVec checked relationships from word to word within a text, using the GloVe (Pennington et al., 2014) dataset to calculate a mean similarity. Then, since this metric, called Mean Similarity (MeanSim), was a measure of similarity, its opposite (i.e., 1–x if x is a value between 0 and 1) was a measure of semantic distance. CoVec ignored stop words such as articles, pronouns, modal verbs, and the letters following the apostrophe in contractions. Since these stop words did not indicate the subject matter, they were useless in analyzing similarity. In the CoVec program, we ran the code line “CoVec -vec GloVe.840B.300d.txt -in t*.txt -wordseq.” This -wordseq option calculated the mean similarity between consecutive words in each file or, in this case, each essay. We used the resulting semantic distance metric in the correlation analyses and as an independent variable (IV) in the regression analyses.
Open Creativity Scoring (OCS)
The OCS program is a GloVe-based freeware created by Organisciak and Dumas (2020). OCS compared a prompt to a response to provide an output of originality (as a proxy for semantic distance) and elaboration. Fewer words are typically better as a prompt with this program because a higher number of words dilute the program's ability to compare the response against a baseline (Acar et al., 2023; Dumas et al., 2020; Forthmann et al., 2018). Therefore, we conducted a pilot run of the essays to determine which combination of words produced a simplified prompt with the most normality, with 30 essays from each prompt to determine the combination of words that provided data with a normal curve. Once the prompts were appropriately reduced, we ran all 230 essays to obtain scores for metrics average semantic distance (avg_orig), total semantic distance (total_orig), minimum semantic distance (min_orig), maximum semantic distance (max_orig), and average elaboration (avg_elab). We used elaboration and originality as predictors in the exploratory regressions and originality in the correlations.
CPIDR51
We used the CPIDR5.1 AI program (Covington, 2012) to find the idea density of each essay. Idea density resulted from dividing part-of-speech tags (i.e., propositions) by the total number of words. The CPIDR software measured idea density by dividing the propositions verbs, adjectives, adverbs, prepositions, and conjunctions by the word count (Snowdon et al., 1996). We uploaded the files (one essay per file; n = 230) and clicked Analyze Files to obtain the output. We used the idea density metric in correlations and as a predictor in exploratory regressions.
Planet Calc
We obtained the number of unique words by entering each essay into a unique-words calculator from an online calculator catalog (Planet Calc, 2015). Then, we calculated each essay (n = 230) with the online calculator one at a time, copying and pasting the text into the dialogue box. Once we obtained the output data from the unique words data processor, we used the data as part of the regression analyses.
Shannon's Entropy
Entropy was a measure of surprise, or the “average amount of information needed to represent an event drawn from a probability distribution for a random variable” (Brownlee, 2019). More certain events held less information. As an essay writer produced new content beyond what had been produced in the preceding sections, the surprise score (unpredictability) and Shannon's entropy score rose. For this study, more creative essays held more information. In Python, the log2() function used the probability of an event to provide the amount of Shannon information contained therein. This calculation was represented as:
Linguistic Inquiry and Word Count (LIWC)
We used LIWC (Pennebaker et al., 2015) to obtain a word count for each essay, which it determined against an established dictionary that highlighted psychological phenomena. The original use of LIWC was to determine particular psychological characteristics from written texts but also had applications for academic and creative assessment (Sadler-Smith et al., 2021).
LIWC detected aspects of linguistic expression and related psychological processes using parts of speech commonly given a psychological state, including those of creativity-relevant behaviors (Pennebaker et al., 2015; Tausczik & Pennebaker, 2010). Each essay (n = 230) was in a separate Word document, which we uploaded to LIWC. This measure was included in determining whether the length of the essay correlated with any other measures while including it as an IV in the exploratory regressions.
Results
Interrater reliability was assessed via ICC (1,4; Shrout & Fleiss, 1979). Interrater reliability revealed strong ICC among the four teacher rater scores: creativity ratings (.811), sophistication (.828), and holistic quality (.927). Due to this ICC, each essay was given a composite score of the four raters for creativity, sophistication, and quality.
We ran descriptive statistics for all major variables (see Table 2). Values for skewness and kurtosis were in the normal range between −2 and +2 (George & Mallery, 2010) or between −1.5 and +1.5 (Tabachnick & Fidell, 2013), which meant the variables could be effectively used in further statistical analyses. Moreover, correlations of all major variables were analyzed (see Table 3). These correlations captured possible relationships among the teacher-rated variables (i.e., average creativity, average quality, and average sophistication) and the text-mining variables (i.e., CoVec MeanSim, OCS Originality, OCS Elaboration, CPIDR Idea Density, # of Unique Words, Entropy, and LIWC Word Count), as well as across the teacher-rated and text-mining variables.
Descriptive Statistics (n = 230).
Note. Creativity was measured on a scale of 1–5, and quality was measured on a scale of 1–6. CoVec = the Covington vector semantics tools; MeanSim = mean similarity; CPIDR = computerized propositional idea density rater; OCS = open creativity scoring; LIWC = linguistic inquiry and word count.
All Major Correlations (n = 230).
Note. CoVec = the Covington vector semantics tools; MeanSim = mean similarity; CPIDR = computerized propositional idea density rater; OCS = open creativity scoring; LIWC = linguistic inquiry and word count.
*p < .05; **p < .001.
Research Question 1–Correlation Between Human Ratings of Creativity and Essay Quality
To answer research question one, we ran two correlations (see Table 3). The first examined a potential relationship between teacher-rated creativity and quality scores. The correlation of average teacher-rated creativity with average teacher-rated quality was statistically significant (r = .418, p < .001). The second correlation was teacher-rated creativity with sophistication, a subset of the quality score. The correlation of average teacher-rated creativity with average teacher-rated sophistication was statistically significant (r = .321, p < .001).
Sophistication points were part of overall quality assessment rubric. To assess to what extent including sophistication in the measure of teacher-rated quality confounded the measure of quality, we ran exploratory correlations with and without sophistication from the quality measure (see Table 4). Average teacher-rated quality was still moderately correlated (r = .407, p < .001) with average teacher-rated creativity, regardless of whether sophistication was included in the quality score.
All Possible Correlations of Quality With Creativity.
p < .001.
Research Question 2–Predicting Human Ratings of Creativity With Text-Mining Tools
We examined four correlations for research question two (Table 3). The first correlation of semantic distance (CoVec MeanSim) and average creativity ratings was statistically significant (r = –.131, p = .024). MeanSim, as a measure of similarity, provided the opposite (1–x) of semantic distance, so the correlation of semantic distance was r = .131 and p = .024. The second correlation of semantic distance (OCS originality) with teacher-rated creativity was statistically significant (r = .359, p < . 001). The third correlation of idea density (i.e., CPIDR) with average teacher-rated creativity was statistically significant (r = .368, p < .001). The fourth correlation was between Shannon's entropy and teacher-rated creativity. The correlation of Shannon's entropy with teacher-rated creativity was statistically significant (r = .388, p < .001).
To summarize, semantic distance (OCS), idea density, and entropy were moderately positively correlated with teacher-rated creativity.
Research Question 3–Assessing the Usefulness of the Text-Mining Tools for Essay Quality
For research question three, we examined eight correlations, four for text-mining with teacher-rated quality and four for text-mining and teacher-rated sophistication (see Table 3). The correlation between MeanSim and teacher-rated quality was statistically significant (r = –.277, p < .001). Since MeanSim was a measure of similarity, the opposite (1–x) was a proxy for a measure of semantic distance. After the signs were switched, the correlation of semantic distance and overall quality was r = .277, p < .001. The correlation between semantic distance (OCS originality) and teacher-rated quality was statistically significant (r = .511, p < .001). The correlation between idea density and teacher-rated quality was statistically significant (r = .548, p < .001). The correlation between Shannon's entropy and teacher-rated quality was statistically significant (r = .521, p < .001). Semantic distance (OCS), idea density, and entropy were strongly positively correlated with teacher-rated quality.
The correlation between MeanSim and teacher-rated sophistication was statistically significant (r = –.165, p = .006). Since semantic distance was the opposite of MeanSim, the correlation of semantic distance and sophistication was r = .165, p = .006 after the signs were switched. The correlation between semantic distance (OCS originality) and teacher-rated sophistication was statistically significant (r = .452, p < .001). The correlation between idea density and teacher-rated sophistication was statistically significant (r = .492, p < .001). The correlation between Shannon's entropy and teacher-rated sophistication was statistically significant (r = .414, p < .001). Essays awarded the sophistication point by teacher raters were moderately positively correlated with originality, idea density, and entropy as determined by text-mining tools.
Research Question 4
For research question four, we ran a series of exploratory regression analyses to determine whether teacher-rated creativity or text-mining creativity scores accounted for the most amount of variance in teacher-rated quality and teacher-rated sophistication scores (see Table 5). We also investigated which of the text-mining creativity tools best predicted quality, sophistication, and creativity. To do these, we first compared the R2 of the regressions to determine which of the methods (i.e., teacher versus text-mining) accounted for the most variance in quality and sophistication. We then ran a series of exploratory regressions to find the best model fits for text-mining tools’ ability to predict teacher ratings (see Table 6).
Regressions: Teacher Raters Versus Text-Mining.
Note. IV = independent variable; DV = dependent variable; CoVec = the covington vector semantics tools; MeanSim = mean similarity; CPIDR = computerized propositional idea density rater; OCS = open creativity scoring; LIWC = linguistic inquiry and word count.
Exploratory Regressions: Text-Mining Tools as Predictors of Teacher-Rated Essay Quality, Sophistication, and Creativity.
Note. MeanSim = mean similarity; DV = dependent variable; CoVec = the Covington vector semantics tools; MeanSim = mean similarity; CPIDR = computerized propositional idea density rater; OCS = open creativity scoring; LIWC = linguistic inquiry and word count.
p < .001.
Predicting Teacher-Rated Quality With Teacher-Rated Creativity
In the first regression, the dependent variable (DV) was the teacher-rated quality of essays, and the IV was the average of the four teachers’ creativity ratings. The regression model was statistically significant (R2 = .175, F(1, 228) = 48.382, p < .001). The average of the four raters’ creativity scores significantly predicted quality (β = .418, p < .001).
Predicting Teacher-Rated Sophistication With Teacher-Rated Creativity
The DV was teacher-rated sophistication, and the IV was the average of the four teachers’ creativity ratings. The model was statistically significant (R2 = .103, F(1, 228) = 26.201, p < .001). The average of the four raters’ creativity scores significantly predicted sophistication (β = .321, p < .001).
Predicting Teacher-Rated Quality With Text-Mining Tools for Creativity
Four additional regressions focused on the text-mining tools’ ability to predict quality and sophistication in student essays. For the first regression, the DV was teacher-rated overall quality, and the predictors were MeanSim (i.e., one measure of semantic distance), originality (i.e., another measure of semantic distance), elaboration, idea density, number of unique words, entropy, and total word count. The first regression was statistically significant (R2 = .445, F(7, 222) = 25.469, p < .001), so 45% of the variance of teacher-rated overall quality was explained by the text-mining tools.
Predicting Teacher-Rated Sophistication With Text-Mining Tools for Creativity
We conducted a second regression to test whether text-mining tools for creativity can predict the teacher-rated sophistication point. The DV was teacher-rated sophistication, and the IVs were MeanSim (i.e., one measure of semantic distance), originality (i.e., another measure of semantic distance), elaboration, idea density, number of unique words, entropy, and total word count. (R2 = .373, F(7, 222) = 18.836, p < .001). Overall, the text-mining creativity models predicted teacher-rated quality and sophistication better than teacher-rated creativity. Teacher-rated creativity predicted 18% of the variance in quality and 10% of the variance in sophistication, while text-mining tools predicted 45% of the variance in quality and 37% in sophistication.
Which Text-Mining Tools Best Predict Essay Quality?
The model regressing text-mining tools on teacher-rated quality was statistically significant (R2 = .445, F(7, 222) = 25.469, p < .001); however, most individual predictors were not, which could be a result of the high multicollinearity among predictors. Multicollinearity is less of a concern when the focus is on the effectiveness of an entire model (Alin, 2010; Haitovsky, 1969), yet it is worth consideration for future research (e.g., to create a tool incorporating the strongest text-mining predictors). To determine which predictors account for the most variance in the models, we conducted exploratory analyses. We examined multicollinearity in the model to avoid redundant predictors, misleading inferences, and artificially inflated R2 (see Table 6).
Although generally accepted cutoff values for variable inflation factor (VIF) are often VIF ≥ 5 or VIF ≥ 10, combining VIF cutoff values with other factors, such as correlation of variables, makes the VIF cutoff model-specific (Craney & Surles, 2002). After a series of removals, accounting for VIF and multicollinearity, the fourth exploratory model was statistically significant (R2 = .409, F(4, 225) = 38.967, p < .001) and explained 41% of the variance in teacher-rated quality. MeanSim (β = -.150, p = .005), OCS originality (β = .426, p < .001), OCS elaboration (β = .244, p < .001, and entropy (β = .192, p = .008) significantly predicted quality.
Which Text-Mining Tools Best Predict Essay Sophistication?
The multicollinearity problem was also evident in the model for text-mining tools to predict sophistication. Using the same justification as for the quality regressions, we ran exploratory regressions to reduce multicollinearity and to determine which text-mining tool or tools account for the most variability in the sophistication point. The first model was statistically significant (R2 = .373, F(7, 222) = 18.836, p < .001) but had the same issues with multicollinearity. After reducing the model with the same justification as used for the quality models, the new model was also statistically significant (R2 = .319, F(4, 225) = 27.777, p < .001) and explained 32% of the variance in teacher-rated quality. The difference was that two of the final variables were not significantly significant predictors in the model. MeanSim (β = -.032, p = .565) and entropy (β = .037, p = .507) did not significantly predict sophistication, but OCS originality (β = .502, p < .001) and OCS elaboration (β = .344, p < .001) did significantly predict sophistication.
Which Text-Mining Tools Best Predict Essay Creativity?
The focus of this study was on using creativity to predict quality in essays. For future research, it would also be beneficial to note which tool best predicts creativity in student essays. We ran one more regression that assessed the amount of variance the text-mining tools explained in teacher-rated creativity scores. This model was statistically significant (R2 = .215, F(7, 222) = 8.674, p < .001) and explained 22% of the variance in teacher-rated creativity. However, the same multicollinearity issue was present. After removing the highly correlated variables, the model was statistically significant (R2 = .183, F(4, 225) = 12.589, p < .001) and explained 18% of the variance. MeanSim did not significantly predict creativity (β = –.053, p = .396). OCS originality did significantly predict creativity (β = .233, p = .008). OCS elaboration did not significantly predict creativity (β = .085, p = .222). Entropy did significantly predict creativity (β = .221, p = .010).
Discussion
The purpose of this study was to explore existing manual and alternative computational essay scoring methods for teachers to use for objective quality or subjective creativity. To this end, we sought to discover whether creativity scores were related to quality and sophistication scores in academic writing and if text-mining tools effectively identify creativity and quality in academic essays. Both teacher-rated and text-mining based creativity correlated with the sophistication level and overall quality ratings of the essays. Exploratory analyses revealed that MeanSim, OCS originality, OCS elaboration, and entropy best contributed to the quality prediction model, while OCS originality and OCS elaboration best contributed to the sophistication prediction model. Additionally, OCS originality and entropy also predicted rated creativity.
Research Question 1
Teacher-rated creativity was related to both teacher-rated quality and teacher-rated sophistication in essays. The positive medium effect sizes of creativity with quality (r = .418, p < .001) and creativity with sophistication (r = .321, p < .001) suggest that teachers account for creative production in student essays when they use the College Board's rubric to evaluate the quality of academic essays.
The thesis subpoint of the quality rubric had a weaker correlation with both creativity and quality than did the sophistication and evidence and commentary subpoints. Teachers often emphasize the thesis point for clarity, assertion of argument, and effectively addressing the prompt, so higher correlation strengths elsewhere in the essay make sense. Students establish constraint (Boden, 2004) with an arguable thesis, which frees them to be creative in the body of the essay. Here, sophistication had a moderate correlation with creativity. This correlation supports the original drive of this study: to determine whether creativity is part of the underlying metric of sophistication. The positive relationship between creativity and essay quality suggests teachers unwittingly assess creative production to a certain degree when they use quality rubrics, although there is room for improvement. Although creativity is not explicitly listed on the College Board rubric (College Board, 2019), the results of this study support previous research that indicates teachers award higher academic scores to creative essays (e.g., Allen et al., 2016; Ayob, 2020; Chamorro-Premuzic 2006). The link found between creativity and quality in writing supports other studies, as well. Hasnudin et al.'s (2015) assertion that creativity can increase quality in academic writing differs from the metrics here because they studied creativity as the tool used to increase quality; however, the relationship found here adds support to their findings.
Combined with previous research, the findings of this study support that creative writing and academic writing have a common thread. Although students view creative and academic writing as disparate entities (Olthouse, 2012; Williams, 2019), teachers can use compositions as a tool to help them become creative producers (Kim & Chae, 2019; Rose & Lin, 1984; Stolaki & Economides, 2018). Teachers can bring style into step with mechanics by embracing creativity and avoiding prescribed topics and vocabulary, as Cheung et al. (2003) found.
Research Question 2
Semantic distance was related to creativity with both opposite of MeanSim (r = .165, p = < .05) and OCS originality (r = .359, p = < .001), but OCS originality had a stronger relationship. Nonetheless, both assess semantic distance based on the GloVe corpus. Semantic distance measures divergent thinking through remoteness of ideas (Chrysikou et al., 2021). The results suggest a difference between tools that compare word to word within a text, such as MeanSim and tools that compare a response to the prompt, such as OCS originality.
Idea density (r = .368, p < .001) and entropy (r = .388, p < .001) had similar strengths of relationship to creativity. Although similar strengths do not necessarily equate with similar constructs, idea density measures the unique occurrences of ideas and entropy measures the amount of information in a text. These combined factors suggest a conceptual similarity of the number of ideas with the amount of information contained in an essay.
Idea density and entropy are newer tools to measure creativity in writing, but the correlations support studies that exist. The correlation of idea density and creativity supports Turkman and Runco's (2019) study where keywords and idea density were correlated, as well as Runco et al.'s (2017) study in which idea density was a predictor of creative potential. Entropy findings support the study in which sentences high in fluency and coherence were used for story infilling (Ippolito et al., 2019). The findings of this study add to the limited set of extant studies that assess writing quality with tools originally designed to assess creativity.
Recent studies have successfully shown technology-based tools measure performance on divergent-thinking tasks (Acar & Runco, 2014; Beketayev & Runco, 2016; Dumas & Dunbar, 2014). The present study supports the assertion that automated forms of assessment provide the required speed and reliability (Sadler-Smith et al., 2021). Extending text-mining tools to reliable automated essay scoring is essential to provide teachers with efficient and effective means to assess compositions. This finding provides evidence of alternative methods of assessment for creativity essays, making authentic assessment methods more feasible.
Research Question 3
Building on Research Question 2 that examined the relationship between text-mining tools and creativity, we hypothesized that sophistication would also be significantly correlated with text-mining tools because both creativity and sophistication involve complexity of thinking. A significant relationship between text-mining tools and quality of student essays would provide evidence of concurrent validity of text-mining tools as a potential measure of student creativity and quality in essays.
The text-mining tools correlated with quality and sophistication, but the tools differed in the strength of relationship. Semantic distance assesses quality (Foltz et al., 1999); yet, in the present study, MeanSim had a weak positive correlation with both quality (r = .277, p = < .001) and sophistication (r = .165, p = .006). MeanSim reliably captured quality and sophistication, but a weak relationship leaves room for improvement if teachers were to rely exclusively on this tool. OCS originality, another measure of semantic distance, had a strong correlation with quality (r = .511, p < .001) and sophistication (r = .452, p < .001). OCS originality compares each essay to the prompt; combined with these strong correlations, OCS originality is a better tool for teachers to use to assess quality in academic writing.
Idea density is a valid indicator of creativity (Runco et al., 2017), which is also supported by research question 2. Idea density also had strong correlations with quality (r = .548, p < .001) and sophistication (r = .492, p < .001). The strong correlation of idea density with quality further supports the relationship between creativity and quality. Since idea density measures the amount of information (Chand et al., 2012), it is sensible that this information-driven metric would correlate with quality of academic essays. The disparate amount of information used to convey academic content translates to divergent thinking, or creativity, so idea density, although a creativity tool, is a strong option for teachers to use to assess quality in academic writing.
Entropy was strongly correlated with quality (r = .521, p < .001) and sophistication (r = .414, p < .001). Entropy measures the information in a message compared to what is predicted to be in the message, considering redundancies. As a newer metric for assessing essays, this study supports that entropy is a strong measure of quality in academic essays.
Text-mining tools designed to capture creativity also captured quality and sophistication in academic essay writing. Significant relationships between text-mining tools with quality and text-mining tools with sophistication provide evidence of concurrent validity of text-mining tools for creativity as useful tools for busy teachers to increase the use of essay writing in their classes. The sophistication correlations, which were of foundational interest in this study, were moderate for OCS Originality, Idea Density, and Entropy. This lends credence to the idea that sophistication in essays relates to creativity in essays.
Research Question 4
The results of the previously discussed correlation studies supported that the text-mining tools for semantic distance, idea density, and entropy correlated with quality and sophistication. This supported Heinen and Johnson's (2018) study showing that semantic distance measures both novelty and appropriateness while avoiding the subjectivity of teacher-rated creativity. The correlation studies further supported the assertion that the style, development, and persuasiveness dimensions of quality as assessed by computer-aided text analysis and automated essay scoring methods (Ke & Ng, 2019) relate to creativity. Clearly, creativity and quality overlap in academic essays, but the correlation studies were insufficient alone to explain which methods work best to assess quality in academic essays. The regression analyses then followed.
Predicting Teacher-Rated Quality and Sophistication With Teacher-Rated Creativity
For predicting teacher-rated quality with teacher-rated creativity, we conducted regressions using average ratings. Teacher-rated creativity explained 18% and 10% of the variation in quality and sophistication scores, respectively.
These explained variances are not sufficiently strong to argue for creativity ratings of teachers to be the exclusive explanation of variance in quality. Despite the relatively low amount of variance that can be explained by creativity, 18% of the variance in quality is decent since the essays are meant to be academic and much of the quality rating should stem from the clarity and strength of argument. It would be interesting to see what change would occur in creativity scores if students were told to be creative, as in some studies (e.g., Green et al., 2012).
Sophistication had a lower R2 than quality with only 10% of the variance of creativity explained, but this is understandable because the sophistication point only counts for one of the six points on the quality rubric. Sophistication can be awarded for a variety of reasons: nuance of argument, implications or limitations in the broader context, effective use of rhetoric, or vivid and persuasive style. The vivid and persuasive style option shares the most conceptual overlap with creativity, but the point can be awarded for other reasons, as well.
Text-Mining Tools for Quality and Sophistication
The text-mining tools model accounted for more variance in quality (45%) and sophistication (37%) than the teacher raters accounted for in quality (18–20%) and sophistication (10–13%; see Table 5). Although objective tools and subjective teacher ratings correlate (e.g., Orwig et al., 2021; Zhu et al., 2009), teachers’ creativity ratings do not score for quality as much as text-mining tools do. The present finding contradicts Dumas et al. (2020), in which human raters’ scores on the Alternate Uses Test were more reliable and valid than the text-mining tools. The researchers also used GloVe, albeit with a different series of processes than in this study. The results of this study, however, support Forster and Dunbar (2009), who measured semantic distance in a Uses of Objects task and found an objective measure could outperform human raters. Although entropy contributed to creativity measures in previous research questions, it did not contribute to sophistication; this supports entropy as an effective tool for creativity measurement, but it does not support the hypothesis that sophistication would be highly correlated with creativity. However, the text-mining model was a better predictor of both quality and sophistication than the teacher-rater model. The efficiency and reliability of the text-mining models could provide teachers with options to use such tools to score academic essays.
Exploratory Regressions
Although the text-mining tools capture quality and sophistication better than teachers, a future research interest is to create a text-mining toolbox that can reliably assess creativity in student essays. While a higher R2 gives a best model fit, we also wanted to explore which text-mining tools worked best. We removed variables based on VIF and high correlations to find the best model fit while reducing multicollinearity. The final exploratory models were statistically significant for quality (i.e., 41% of variance explained), sophistication (i.e., 32% of variance explained), and creativity (i.e., 18% of the variance explained). All VIFs were below 3 and the remaining four variables were not as strongly correlated to one another as some of the variables from the previous models (See Table 6). While the number of words equates to lexical diversity (Malvern et al., 2004) and seemed a clear choice for a text-mining tool, the number of words and unique words were also correlated. Idea density, although highly correlated with sophistication and quality, was also highly correlated with entropy. Since idea density had a high VIF, it was also removed from the model. The final model fit reduced multicollinearity and maintained the model's integrity. The tools most useful for assessing quality were MeanSim, OCS originality, OCS elaboration, and entropy; for sophistication, they were OCS originality and OCS elaboration; for creativity, they were OCS originality and entropy. These final model fits are excellent starting points for future research.
Limitations, Future Directions, and Conclusion
A larger sample size with representational demographics across multiple regions of the country could have more power and thus more generalizability. A replication study with non-AP student essays would also add to evidence of generalizability. The available demographic data limited race and ethnicity to five categories and gender to the two biological sexes, while the U.S. Census Bureau reports data using at least eight categories of race and ethnicity and four for gender identity (U.S. Census Bureau, 2021). Although other studies effectively used the same raters for quality and creativity with time elapsed in between, using separate raters for quality and creativity or counter-balancing the order of the rating tasks would eliminate any potential for the halo effect.
The moderate correlation between creativity and quality suggests some overlap, but more research can be done to capture creativity beyond the lens of quality. Creativity could be included as an explicit part of the rubric so more of it could be captured. Future research could also examine the role creativity plays in other forms of writing, such as other English essay types or across disciplines. In the present work, we measured essay quality based on teacher ratings. Future research could examine how they overlap with the automated tools of writing quality indicators would be very helpful. Another relevant avenue of future research would be to determine whether computerized assessment programs that assess for quality based on similarity to a corpus penalize creative papers.
The exploratory regressions to determine which text-mining tool best captures quality, sophistication, and creativity provided introductory views into parsing out the creativity tools. However, to create a useful tool for teachers, this aspect of the study needs further research. Many other tools and metrics exist and should be explored to move forward with creating a usable computer-aided quality and creativity assessment tool. Subsequently, such a tool may be used to examine how creativity of the essays can change in response to creativity training in experimental designs.
To conclude, teacher-rated creativity of student essays moderately correlated with teacher-rated quality and sophistication. Text-mining tools for creativity moderately correlated with teacher-rated creativity. Teacher-rated quality weakly correlated with the opposite of MeanSim (i.e., semantic distance) and moderately correlated with OCS Originality, Idea Density, and Entropy. Teacher-rated sophistication weakly correlated with the opposite of MeanSim and strongly correlated with OCS originality, idea density, and entropy. MeanSim, OCS originality, OCS elaboration, and entropy accounted for the most variance in teacher-rated quality when accounting for multicollinearity while maintaining the integrity of the model fit. OCS originality and OCS elaboration accounted for the most variance in teacher-rated sophistication when accounting for multicollinearity while maintaining model fit. Importantly, the total variance explained by the text-mining tools is still smaller in creativity compared to quality and sophistication. There is ample room for improvement in measuring creativity in essays, and additional tools beyond those tested in the present work could prove valuable. Recent research has provided some insights. For instance, Zedelius et al. (2019) investigated two tools (Coh-Metrix and LIWC) to predict three indicators of writing creativity: imagery, voice, and originality. While Coh-Metrix components (such as narrativity, syntactic simplicity, word concreteness, referential cohesion, and deep cohesion) predicted imagery and voice but not originality. Weinstein et al. (2022) identified general word frequency, infrequency of word combinations, context-specific word uniqueness, syntax uniqueness, rhymes, and phonetic similarity as predictors of higher ratings for sentence creativity. Johnson et al. (2022) compared different models of semantic diversity in explaining the human-rated creativity of short narratives. They found that Bidirectional Encoder Representations Transformer, known as BERT, was the most successful in explaining up to 72% of the rated creativity. Future research could investigate whether these tools can also effectively predict the creativity of academic essays by AP students.
Overall, our findings supported the view that creativity and quality are related in academic essays. That means, quality metrics can capture some of the creativity expressed in essays, but not entirely. Based on these findings, educators and students should incorporate creativity, as well as its assessment, into the K-12 academic writing experience. While this study provides evidence of text-mining tools’ efficiency and accuracy over teacher raters, future research could extend this evidence to create a technology-based tool for teachers to efficiency and effectively assess both creativity and quality in academic essays.
Footnotes
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
