Abstract
The accuracy of machine learning-based automated text classification systems, such as spam filters and search engine results, heavily depends on the quality of manual text classification. However, the cognitive demands of manual text classification tasks, particularly when dealing with challenging or difficult-to-comprehend texts, have not been extensively explored in previous studies. This research aims to address this gap by investigating the cognitive load associated with manual text classification tasks through analyzing eye tracking data. In this study, 30 participants performed manual text classification tasks while their ocular parameters were recorded using an eye tracker. The findings of this study revealed that ocular parameters recorded through eye tracking provided valuable insights into the cognitive load experienced during manual text classification tasks. Furthermore, it was observed that complex narratives led to higher cognitive load estimation. Moreover, native English-speaking participants exhibited lower cognitive load, compared to non-native English speakers.
Keywords
Introduction
Automated text classification has a variety of applications in daily life including content organization, spam email filtering, improving the quality of results of search engines, movie recommendations, sentiment analysis, and so on. Machine learning (ML) models are used to perform automated text classification for the above-mentioned practical applications as they can classify a large number of records very efficiently. These ML models are typically trained on human-labelled training datasets, and the accuracy of these ML models is heavily dependent on the accuracy of the training dataset. Human-labelled training data is generated through the manual text classification process, in which people assign the most applicable category to a textual record selecting from a list of possible categories. The textual record can be in different forms such as a document, narrative, or short text.
From a cognitive demand perspective, the manual text classification task can be significantly challenging if the text is difficult to comprehend (e.g., complicated, and noisy), or the person must select the most-applicable category from a long list of possible categories [Nanda, et al., 2019]. Given the impact that accuracy and efficiency of manual text classification can have on a multitude of practical applications, it is important to understand the human factor challenges associated with the manual text classification task. However, to the best of our knowledge, no previous study has examined the human factor aspects of manual text classification task. We aim to fill this gap in literature by studying the cognitive load associated with manual text classification task by analyzing the eye tracking data recorded during the task and associating it with the background data of the person and their task performance. Eye tracking has previously been used as a reliable measure of cognitive load for reading tasks by previous studies [Mishra, et al., 2018].
In this context, we designed a manual text classification user study where 30 participants were asked to perform text classification task of assigning injury event cause categories to accident narratives recorded at hospitals, which are relatively short and noisy in nature. The participants’ eye tracking data were recorded while they performed the task, and they were asked to submit a survey about their background and task experience after task completion. The various data points collected during the study were used to understand (a) the correlation between ocular parameters recorded from eye tracking and different aspects of manual text classification task and, (b) whether English proficiency of participants had any impact on the performance or experience of manual text classification task. A schematic diagram summarizing the study design is shown in Figure 1.

Schematic Diagram of Study.
Background
Eye tracking is the process of measuring either the point of gaze (where one is looking) or the motion of an eye relative to the head. An eye tracker is a device for measuring eye positions and eye movement [Eye Tracking, 2023]. Eye tracking allows researchers to study eye movements of users while performing a wide range of activities. Analyzing eye movements gives insight into visual assessment [Kooiker, 2016], analyze visual behavior [King, 2019], identify emotional state and mental occupancy [Li, 2020], estimating cognitive workload [Babu, 2019], diagnose learning disabilities [Rello, 2015; Saluja, 2019] and other user behaviors [Tzafilkou, 2017] while performing a particular task. Analyzing ocular parameters in context of the activity performed can lead to more accurate design and evaluation of Human Computer Interaction aspects for task analysis in general and for visual systems in particular [Stephane, 2017]. Recording eye movements while reading is one of the finest ways to understand human language processing [Frazier, 1982] at different levels of analysis within a sentence. They provide the advantage of measuring reading behavior relative to measuring the reading times for an entire sentence or paragraph. During normal, skilled reading, the eyes move sequentially through the text, fixating one word at a time. Word-based eye movement metrics have proven to be strongly correlated with high-level text processing [Barrett, 2020].
Recently, the field of Natural Language Processing has started considering gaze data for improving text analytics using machine learning. Mathias et al. (2018) improved quality of text evaluation using gaze data to represent words. They predicted text quality attributes like quality, organization, coherence, and cohesion from ocular parameters like fixation and regression along with textual features using a single-layer feed-forward neural network. Their best model combined gaze features with textual features.
Majooni et al. [2018] investigated the effect of layout on the comprehension and cognitive load of the viewers in information graphics. They analyzed eye-tracking data to provide quantitative evidence concerning the change of layout and its effect on the comprehension of participants and variation of their cognitive load. From results, they claimed that the comprehension from the zigzag form of the layout was higher with a less imposed cognitive load. Tomanek et al. [2010] analyzed eye tracking data during their annotation of named entities in texts. They recorded number of fixations, search time and fixation duration values as a promising means to get a better understanding of nature of the linguistic annotation processes with the goal of identifying predictive factors for annotation cost models. Results noted that the authors defined the difficulty of named entity instances based on the cognitive load estimated using gaze data. Mishra et al [2018] proposed and evaluated different approaches for cognitive load modelling associated with text comprehension using scan path complexity of gaze while reading. They tried to model readers’ eye-movement behavior to quantify the cognitive effort associated with reading processes. Results from their study showed that the measurement of complexity of scan paths recorded using eye tracking can lead to better cognitive models to explain better reading of users. Another study by Joshi et al [2014] tried to model manual sentiment annotation complexity using eye tracking data. The authors noted that saccade duration was not significant for annotation of short text, rather, the sentiment annotation complexity was obtained using the fixation duration values with appropriate normalization.
To the best of our knowledge, previous studies have not analyzed the manual text classification task using eye tracking to examine if the ocular parameters recorded during the task are related to cognitive load, task performance, and whether a person’s proficiency in language plays a role. We explored these aspects in this study. The detailed study design is explained in the next section.
Methods
An eye tracking study was conducted to classify injury related narrative texts into specific injury event cause groups. The study was approved by the Institutional Review Board (IRB) at Purdue University. We collected data from 30 participants, consisting of 18 males and 12 females aged between 23 to 30 years, and a mix of native and non-native English speakers. In the rest of the paper, we refer the native English speakers as ‘native English-speaking’ participants and the non-native English speakers as ‘non-English speaking’ participants respectively. Taking part in the study was voluntary, and participants were allowed to choose to quit the study at any point in time. They were given a consent form which was required to be read and signed before starting the study if they wished to participate. Additionally, participants were compensated with cash for their time and effort upon completing the study.
The dataset used in the study was generously provided by the Queensland Injury Surveillance Unit (QISU), Australia. The study was carried out in a well-lit indoor room. Each participant was called individually to the room and asked to position themselves comfortably towards the screen in a way they could read content on screen and access the mouse. The Tobii Pro Fusion Eye Tracker was placed along the lower edge of the monitor that displayed the injury classification software (Figure 2). The software consisted of an instructions page including meanings to each injury event cause group (Figure 2(a)) and 12 prompts. Each prompt displayed one injury narrative text and 6 possible cause-of-injury code groups (Figure 2(b)) from which the participant had to select one. These six injury groups included ‘Fall’, ‘Struck’, ‘Cut’, ‘Burn’, ‘Motor Vehicle’ and ‘Other’. The 12 prompts consisted of a set of two unique narratives texts belonging to each of the injury event cause groups. The ‘NEXT’ and ‘UNDO’ buttons on each prompt were used to navigate to next prompt and clear selection of injury event cause group to reselect the appropriate group respectively. Participants could read the prompt and use the mouse to select the desired injury event cause group. We created 17 unique sets; each set consisting of 12 unique injury narratives which were given as prompts to the 30 participants in randomized fashion. Each of the 12 narratives were internally classified into varying levels of complexity as Low, Medium, and High. These levels were based on the difficulty of narrative text comprehension and selecting appropriate injury event cause group. Information about the complexity was not provided to the participant.

(a) Experiment setup & instructions page (b) Injury classification software interface.
Participants were initially briefed about the process of the study and a trial session was given for them to be familiarized with the injury classification software. Then, the eye tracker was calibrated using the Tobii nine-point calibration routine. After the eye tracker was successfully calibrated, the injury classification software was displayed on screen. Participants were asked to read the instructions carefully and select the ‘START’ button when ready. Participants read the injury narrative text displayed on screen and selected one of the six injury event cause groups that they felt was most appropriate. Once the group was selected, the participant was required to think-aloud the reason for selecting the group before proceeding to the next prompt using the ‘NEXT’ button to continue with the study. An EVIDA digital voice recorder was used to record participants responses during the think-aloud in course of the study. Analysis for the think-aloud is not within the scope of this paper and will be explored in later studies. Participant could use the ‘UNDO’ button to clear existing group selection and reselect appropriate group if needed. Each participant completed 12 such prompts. Once the study was completed, participants were asked to submit a survey consisting of biographical information like gender, language, major, ethnicity, and NASA Task Load Index (TLX) and System Usability Scale (SUS) questionnaire. The questionnaires were used to study the mental workload and system usability for each participant.
We used eye tracking to record ocular parameters like fixation count, average fixation duration, average pupil diameter values for left and right eyes respectively. The injury classification software recorded task completion times to complete the task involving 12 prompts, average response times to read injury narrative text and select appropriate injury event cause group, and selected injury event cause group.
Results & Discussion
All participants completed all 12 prompts. We recorded 360 selections of injury event cause groups in total, including 60 selections for each of the 6 groups. Two participants used the ‘UNDO’ button to reselect the desired injury event cause group. We recorded English proficiencies of all 30 participants in and found that, 21 participants were non-English speakers, i.e., these participants speak a language other than English at home [Christen, 2008] and 9 participants were native English speakers.
Summary of Ocular Parameters, Time, and Accuracy
Analysis of recorded eye tracking data about ocular parameters indicated that the average number of fixations for all 30 participants was 1175 with an average fixation duration of 291.90 seconds. 12 out of 30 participants reported higher than average values of number of fixations and average fixation duration values, indicating that these participants spent more time to read the narrative text and select the appropriate injury event cause group. The average pupil diameter values for both left and right eyes were observed to be 3.03 and 3.05 respectively. 14 out of 30 participants were observed to have higher than the average pupil diameter values, out of which 6 participants showed higher (values greater than the average) fixation count and fixation duration values, estimating to have relatively higher cognitive load when compared to other participants.
Task performance was evaluated using two measures, accuracy, and task completion time. We noted that the average time taken to read injury narrative text and select the appropriate injury event cause group for all the 12 prompts (task completion time) across all 30 participants was 376.25 seconds. The average response time for reading the narrative text and selecting appropriate injury event cause group for a single prompt was observed to be 15.66 seconds. 12 out of the 30 participants showed less than average task completion and response times.
To calculate the accuracy, we compared the injury cause category selected by participants against the original E-code assigned in the QISU dataset to check if the participant’s selection agreed with the original code or not. We found that the average number of correct selections of injury event cause group across all participants was 9 out of total 12 prompts across all participants, hereafter referred as coding accuracy. This level of coding accuracy (75% =9/12) indicated that the task of noisy narrative text comprehension and selecting most appropriate injury event cause group was not very confusing for participants. Further, we noted the coding accuracy for ‘Burn’ injury event cause group was highest, indicating that all participants were able to categorize narrative texts that belonged to group ‘Burn’ without much indecision, maybe because of the unique nature of the injury category. In decreasing order of coding accuracies, “Burn” was followed by ‘Motor Vehicle’, ‘Fall’, ‘Cut’, ‘Struck’ and ‘Other’ respectively. It can be observed that many participants had difficulty in correctly categorizing narrative texts into ‘Other’ and ‘Struck’ injury event cause groups. “Other” was often incorrectly categorized into ‘Fall’, ‘Struck’, Cut’ injury groups and “Struck” was miscategorized into ‘Fall’, Cut’, and ‘Other’ injury event cause groups by participants. This may be because the definitions for “Burn” and “Motor Vehicle” injury groups are relatively unique in nature, while the injury groups “Fall”, “Struck”, and “Other” are not as distinctive and may be confusing for participants while selecting the most appropriate category.
It is to be noted that the narratives were internally classified into varying levels of complexity which would have influenced the coding accuracy of participants. We categorized each of the 12 narratives for each participant as Low, Medium, and High levels of complexity based on the level of difficulty of narrative text comprehension and selecting appropriate injury event cause group. It was noted that 15 out of 30 participants encountered most prompts with narratives of ‘Low’ complexity level, 12 participants encountered prompts with narratives of ‘Medium’ level of complexity and 3 participants encountered prompts with narratives of ‘High’ complexity levels. Overall, we can say that participants encountered prompts with narratives of ‘Low’ level of complexity with an average level of difficulty of 1.65.
Survey Results
We calculated performance and mental workload scores of participants by interpreting results from SUS and NASA TLX survey. The average performance (SUS score) and mental workload (TLX score) for all participants was 85.67 and 112.67 with standard deviation of 12.52 and 90.25 respectively. Based on general guidelines on interpreting SUS score [Will, 2021], it was observed that 27 participants showed Excellent/Good performance reporting the system to be easy to use, while 3 participants found the system to be difficult to use showing low performance scores (Poor/Awful). Then, we assigned the mental workload scores based on the interpretation illustrated in Chen et al [2022] and found that 18 participants showed high mental workload, 9 participants showed low mental workload and 3 participants had medium mental workload. The distribution of mental workload scores indicates that while most participants were able to successfully complete the task with reasonable accuracy, it required considerable time and effort to comprehend and decide on selecting appropriate injury event cause groups.
Correlation Analysis
Given the nature of the dataset, a non-parametric Spearman’s Rho correlation test was performed to examine any association between ocular parameters, coding accuracy, average level of complexity of narrative texts, and English proficiency of participants. The results shown in Figure 3(1) indicated the following:
positive correlation between average response times and (a) task completion times [r (29) = .56, p < 0.01], (b) fixation count [r (29) = .36, p < 0.05], and (c) fixation duration [r (29) = .39, p < 0.05]
positive correlation between average pupil diameter and fixation count [r (29) = .37, p < 0.05]
negative correlation between level of complexity of narrative texts and coding accuracy
negative correlation between native English speakers and average pupil diameter values.

Results from Spearman’s Rho Test for Prompt-Wise Correlation.
Further, we evaluated the association between ocular parameters, coding accuracy, and average level of complexity of narrative texts, for native English speakers and non-English speakers. For non-English speakers we found a positive correlation between task completion times and average response times [r (20) = .58, p < 0.01], fixation count [r (20) = .88, p < 0.01], fixation duration [r (20) = .91, p < 0.01]. For native English speakers, we found-
negative correlation between average response times and (a) coding accuracy [r (8) = -.68, p < 0.05], and (b) average pupil diameter [r (8) = -.72, p < 0.05].
positive correlation between fixation count and fixation duration [r (8) = .93, p < 0.01]
Next, we evaluated the association between ocular parameters, coding accuracy, and average level of complexity of narrative texts across 360 prompts for all 30 participants. The results shown in Figure 3(2) indicated the following:
positive correlation between fixation count and (a) average pupil diameter [r (359) = .25, p < 0.01], and (b) level of difficulty [r (359) = .34, p < 0.01]
negative correlation between fixation count and coding accuracy [r (359) = -.19, p < 0.01]
negative correlation between coding accuracy and level of difficulty [r (359) = -.31, p < 0.01]
Overall, the correlation analysis suggests that the ocular parameters recorded from eye tracker can be used to estimate cognitive load of participants while performing the text classification task. As the average task completion and response times increased for each prompt, the participant showed greater values of fixation count and average fixation duration. Participants with greater number of fixations had higher values of average pupil diameter indicating to have higher cognitive load estimation. The negative correlation between level of complexity of narrative texts and coding accuracy, may be indicative that people who encountered narratives with higher level of complexity found the text classification task to be challenging and were not sure about the applicable injury code group, thus resulting in lower coding accuracy.
The average pupil diameters values were observed to be relatively lower for participants who were native English speakers (average of 2.99 mm; standard deviation of 0.44) as compared to non-English speaking participants (average of 3.03mm; standard deviation of 0.45). This suggests that non-English speaking participants experienced higher cognitive workload. The possible explanation may be that native English speakers found the task of narrative text comprehension easier due to language proficiency, when compared to the non-English speakers. Additionally, non-English speakers showed longer task completion times and average response times in addition to higher number of fixations and fixation duration. This can also be indicative that non-English speaking participants experienced higher cognitive load while performing the injury classification task. On the other hand, native English-speaking participants showed a negative correlation between average response times and coding accuracy and average pupil diameter values indicating better performance along with reduced cognitive workload estimation. This may be because the participants carefully processed the text and made the selection. Finally, prompt-wise analysis indicated that for cases with higher levels of complexity of narrative text leading to lower coding accuracy, most participants showed relatively higher cognitive workload with increased fixation count and average pupil diameter values. It may be noted that, the think-aloud protocol that was used for recording the reason for selecting the group for each user can also lead to higher cognitive load estimation as it can be overwhelming to some participants. As mentioned earlier, analysis of the think-aloud is not within the scope of this paper and will be explored in later studies.
Conclusions
In this pilot study, our aim was to understand the cognitive load involved in manual text classification task using eye tracking and analyze the impact of language proficiency on the task performance. The results from this study involving categorizing noisy accident narratives into six cause of injury codes suggest that (a) ocular parameters recorded through eye tracking such as fixation count, fixation duration, and pupil diameter can be indicative of cognitive load experienced during manual text classification task, (b) participants who were native English speakers experienced relatively lower cognitive load during the task, and (c) participants showed higher cognitive load while categorizing narratives with higher complexity. We acknowledge that these findings were observed from a sample of 30 participants for a specialized case of text classification task, and therefore, may not be generalizable in all circumstances. Future work will aim towards collecting data from more participants for this study and conducting more studies for other types of text classification tasks and examining whether these findings are generalizable.
