Abstract
Eye metrics have been used to objectively evaluate proficiency levels in psychomotor tasks, including aviation, driving, sports, and surgery. Despite extensive research devoted to developing and utilizing eye metrics, the literature does not contain any explicit comparison between scene independent and dependent eye metrics for understanding skill acquisition or characterizing expertise. This study collected eye tracking data from medical students practicing the peg transfer task and computed both scene independent and dependent eye metrics to indicate proficiency. K-means clustering analysis on the eye metrics yielded three clusters corresponding to three proficiency levels which showed significantly different trial completion time. The box plots of the eye-gaze metrics illustrated different patterns of scene independent and dependent eye metrics with respect to proficiency levels, highlighting the need for further examination of these metrics for accurate and useful applications.
Introduction
Eye tracking technology has been widely used in studies of psychomotor tasks, as vision is a primary source of sensory information in controlling our limbs (Causer et al., 2013). By analyzing eye movements, researchers can gain valuable insights into how visual information is processed and guides subsequent actions. Eye tracking methods have proven particularly valuable in the assessment of the proficiency and design of training programs for psychomotor tasks, including those in aviation (Peißl et al., 2018), sports (Panchuk et al., 2015), driving (Palinko et al., 2010), and surgery (Tien et al., 2014). Specifically, researchers have developed and used various eye-tracking metrics to differentiate between experts and novices in psychomotor tasks.
Eye metrics can be generally classified into scene independent and scene dependent eye metrics (Deng et al., 2021). Scene independent eye metrics are independent of where eye gazes target, and thus are computed without specifying any areas of interest (AOIs). Scene independent metrics include fixation rates, blink rates, gaze entropy, and pupillometric data that can be directly computed from the quantitative data collected by an eye tracker. Statistical analyses can be used to compare these metrics between groups of different skill levels, thereby revealing differences in eye gaze behaviors. In this way, scene independent eye metrics can serve as an objective and quantitative measure of performance in psychomotor tasks (Peißl et al., 2018; Richstone et al., 2010).
Previous research has demonstrated that experts tend to have higher fixation rates because they are presumed to identify and locate useful information more rapidly than novices (Deng et al., 2022; Ottati et al., 1999; Richstone et al., 2010; Tien et al., 2015; Wickens et al., 2001; Xiong et al., 2016). In addition, due to their ability to extract useful information efficiently, experts tend to have shorter average fixation durations than novices (Wickens et al., 2001; Xiong et al., 2016). Further, given extensive practice can lead to automatic coordinated movements and reduced mental effort (Erridge et al., 2018; Tien et al., 2015; Zheng et al., 2010), more experienced professionals tend to complete psychomotor tasks exhibiting smaller change in pupil diameter (Beatty, 1982; Xiong et al., 2016), lower index of cognitive activity (Richstone et al., 2010), higher blink rates (Zheng et al., 2012), and lower gaze entropy (Deng et al., 2022; di Stasi et al., 2016; Wu et al., 2019).
While numerous studies focused on scene independent eye metrics, other researchers have explored scene dependent eye metrics for skill assessment and development of training programs for psychomotor tasks (Kulkarni et al., 2023; Law et al., 2004; Mourant & Rockwell, 1972; Wilson et al., 2010; Xiong et al., 2016). By investigating eye metrics within the context of a specific task and visual environment, scene dependent eye metrics illustrate the level of attention on AOIs with respect to how the task is being performed. Analyzing scene dependent eye metrics can thus highlight differences between experts and novices in terms of their gaze patterns or visual attention on AOIs. The use of scene dependent eye metrics also enables trainers or experimenters to provide trainees with meaningful feedback on where they are looking and should look to improved training and performance (Wilson et al., 2011).
Scene dependent gaze metrics have revealed that less experienced drivers tend to visually attend on the area immediately in front of the vehicle more frequently but check the mirrors less frequently than experienced drivers (Mourant & Rockwell, 1972). Compared to novices, expert pilots demonstrated a higher frequency of fixations on aimpoint and airspeed, and a lower frequency of fixations on altimeter, leading to better airspeed maintenance and landing performance under visual flight rules (Wickens et al., 2001). A separate study revealed that expert pilots with better landing performance frequently shifted their gaze between the cockpit instruments and the runway, while novice pilots almost never pay any attention to the cockpit instruments (Xiong et al., 2016). Research in surgery has demonstrated that experienced surgeons exhibit a prolonged dwell time on critical AOIs during certain operative segments (Erridge et al., 2018). Furthermore, attending physicians and residents displayed significant differences in dwell time within the experts’ AOIs when viewing surgical videos (Fichtel et al., 2019).
Despite the significant amount of research that has been devoted to developing and utilizing eye metrics, the literature does not contain any explicit comparison between scene independent and dependent eye metrics for assessing skill acquisition or characterizing expertise. It would be invaluable to understand how both types of eye tracking metrics complement each other or whether one outperforms the other.
To address the paucity of research on comparing scene independent and dependent eye metrics, we examined an eye-gaze dataset collected from medical students who self-trained on the peg transfer task in the Fundamentals of Laparoscopic Surgery (FLS) curriculum. Laparoscopic surgery has been employing eye metrics for skills assessment extensively and is thus a suitable application domain for this research (Law et al., 2004; Richstone et al., 2010). This study collected eye tracking data during self-training on the peg transfer task, and computed scene independent and dependent eye metrics. K-means clustering of the eye metrics generated three clusters that corresponded to three proficiency levels of trainees based on cross-referencing with completion time. By examining changes in both scene independent and dependent eye metrics across proficiency levels, we can obtain insights into the differences between these two types of metrics. This exploratory study shed light on the appropriate application and interpretation of eye metrics in skills assessment and training for psychomotor tasks.
Method
Participants
The study protocol was approved by the Carilion Clinic IRB (Carilion 19-423/VT 19-51). Thirteen medical students who had limited to no prior experience in laparoscopic surgery (7 males, 6 females; aged between 23-32) provided informed consent to participate in this study. All participants were right-handed and had normal or corrected-to-normal vision.
Apparatus
The experimental apparatus was built around an FLS trainer box and LCD monitor that was placed on an adjustable desk and mounted on a tripod to accommodate the varying heights of each participant (Figure 1). The trainer box was equipped with a monocular optical camera that captured and displayed the scene inside the trainer box on a 19” LCD monitor. The camera’s field of view was centered on the pegboard and adjusted to ensure that the entire pegboard was visible. The instruments required to complete the task were inserted through a trocar, limiting the degrees of freedom of the instruments, and simulating the fulcrum effect experienced by surgeons during laparoscopic surgery. Further, a Tobii X3-120 remote eye tracker was attached to the monitor and connected to the computer for collecting gaze data.

Apparatus setup with a FLS trainer box setup and a remote eye-tracker.
Procedure and task
Participants practiced the five FLS tasks until they met the passing criteria for each task. The experiment consisted of a total of sixteen sessions, each lasting approximately one hour. This study specifically focused on the peg transfer task that were practiced only in the first three sessions. During the first session, participants received an introduction to the study and provided informed consent. They then completed a demographic questionnaire and calibration of the remote eye tracker. Next, participants watched an instructional video of an FLS task and performed the task to collect baseline data. This process was repeated until baseline data were collected for all five FLS tasks (although only the first task is relevant for this study).
In the second and third sessions, participants practiced the peg transfer task. All participant met the passing criterial in the third session. They began by calibrating the remote eye tracker and reviewing the instructional video of the task. For the remaining time, participants practiced the task as many times as possible with a 2-minute break every 10 minutes. The peg transfer task required participants to use a grasper to lift six colored rubber ring objects, transfer them midair from their non-dominant hand to their dominant hand, and place them on pegs on the other side of the board. Once all six ring objects were transferred to the other side, the process was reversed until all objects were transferred back to the original side. Each practice trial was timed, and a penalty was assessed for any object dropped outside the field of view, based on the experimenter’s observation on a separate computer monitor. After the third session, the participants practiced the other four FLS tasks that are beyond the scope of this article.
Measures and eye metrics
The study collected gaze data and trainer box video for each practice trial. The eye tracking data files contained information about the position and size of both eyes at specific times, while the videos recorded tool and object movements inside the trainer box as participants practiced the peg transfer task. These data were used to compute trial completion time, scene independent and dependent eye metrics (Table 1).
Performance metric and eye metrics.
The time taken to perform a procedure is one of the most accepted objective measures of surgical skills that consistently differentiate between novice and expert surgeons (Law et al., 2004; Sánchez-Margallo et al., 2017; Yamaguchi et al., 2007). Thus, this study determined the level of proficiency achieved by participants based on trial completion time. Shorter completion times indicated better proficiency.
Scene independent eye metrics were computed solely on raw data from the eye tracker. For scene dependent eye metrics, we employed computer vision to process videos of the activities inside the trainer box to track AOIs, such as moving objects and grasping tools, and then calculated fixation rates on these AOIs/locations (Kulkarni et al., 2023).
Data Analysis & Results
Data analysis
A total of 13 participants completed 799 practice trials of the peg transfer task in our study. However, we excluded trials that had insufficient eye tracking data and task errors. After excluding these trials, four participants had fewer than ten trials per session that would be too few to demonstrate their skill acquisition, thus, their data were removed from this analysis. The final dataset consisted of 498 practice trials from 9 participants. Before proceeding with our analysis, we conducted exploratory data analysis to check for outliers and violations of model assumptions. We detected two outliers with erroneous values for standardized pupil diameter and stationary gaze entropy, respectively. As a result, the final analysis only included data from 496 practice trials.
K-means clustering analysis was used to identify groups/clusters of practice trials exhibiting similar behaviors with respect to eye metrics. Prior to clustering, all eye metrics were standardized into z-scores across practice trials except for pupil diameters, which were standardized into z-scores by participants. Pearson correlation coefficients between eye metrics were computed for feature selection. Eye metrics that had high correlations (r>0.69) with other eye metrics were excluded from subsequent data analysis to avoid redundancy and ensure the accuracy of the model. As a result, saccade rates and percentage of saccade duration were excluded due to high correlation with percentage of fixation duration with r equals to -0.768 and -0.693, respectively.
Results
K-means clustering of the eye metrics produced within-cluster-sum-of-squared error that suggested three clusters of practice trials corresponding to three proficiency levels. Each cluster contained multiple trials of the 9 participants who were considered performing the task at the same proficiency level. Specifically, cluster 1 contained 90 practice trials; cluster 2 contained 186 practice trials; and cluster 3 contained 220 practice trials. Figure 2 presents the boxplots for comparisons across the eye metrics with respect to trial completion time for the three clusters or trials or proficiency levels.

Boxplots of completion time (reference measure) and eye metrics for the three clusters/proficiency levels.
For scene independent eye metrics, the plots (Figure 2) illustrated that fixation rates and percentage of fixation duration for medium and high proficiency levels are relatively similar to each other and tend to be higher than those in trials of low proficiency level. Gaze entropy and pupil diameter did not exhibit a clear linear relationship with proficiency level. Gaze entropy was the highest in trials of medium proficiency level, and lowest in trials of low proficiency level. Pupil diameters were the largest in trials of high proficiency level, and smallest in trials of medium proficiency level.
For scene dependent eye metrics, as proficiency improved, participants tended to have increased fixation rates on objects, single-tool-holding-an-object, and future objects. Fixation rates on both-tools-holding-an-object were the highest in trials of medium proficiency level, and lowest in trials of low proficiency level. Fixation rates on areas outside the moving objects and tools (i.e., two AOIs) tended to be the highest in trials of high proficiency level, and lowest in trials of medium proficiency level.
Discussion
The study revealed clear linear relationships between proficiency level and several scene dependent eye metrics including fixation rates on objects, a-single-tool-holding-an-object, and future objects. At higher proficiency levels, fixation rates on these specific AOIs also increased in corroboration with extant research that individuals learn to prioritize task-relevant information with expertise (Haider & Frensch, 1999). This prioritization could lead to more efficient and focused visual attention that improves task performance.
The relationships appeared non-linear between proficiency level and fixation rates on both-tools-holding-an-object and areas outside tool-holding-an-object. Nevertheless, high proficiency trials had higher fixation rates on these AOIs than low proficiency trials that is also consistent with the literature. However, for holding an object with both tools, high proficiency trials had lower fixation rates than medium proficiency trials, likely because trainees can complete this step more quickly as they acquired proficiency in passing objects between their hands. Further investigation is necessary to understand fixation rates outside the AOIs that have not been investigated in the literature given the lack of correspondence to expectation.
The relationship between proficiency and scene independent eye metrics is less clear and more challenging to interpret. Fixation rates tended to increase as proficiency level increased, although the differences between medium and high proficiency trials were minimal. This finding generally aligns with previous studies indicating experts having higher fixation rates because they can identify and locate useful information more rapidly than novices (Deng et al., 2022; Ottati et al., 1999; Richstone et al., 2010; Tien et al., 2015; Wickens et al., 2001; Xiong et al., 2016). Fixation duration is not consistent with previous research, which indicates experienced professionals having shorter average fixation durations (Wickens et al., 2001; Xiong et al., 2016). Further research is necessary to understand the changes in fixation durations at the early training stage.
Gaze entropy and pupil diameter typically indicate mental workload that is expected to decrease as proficiency increase. However, the trends observed in the plots did not align with this expectation. One possible explanation is that the practice trials were timed to indicate performance thereby inducing sustained mental workload for continuous learning and improvement.
In summary, the qualitative findings based on the plots suggested that scene dependent eye metrics have a clearer linear relationship with proficiency level and align better with the literature than scene independent eye metrics. Therefore, scene dependent metrics may be more effective in differentiating proficiency levels during the early stage of training. Quantitative analysis is necessary to compare the two types of eye metrics to draw formal conclusion.
A limitation with this study is that clustering algorithms can be very sensitive to the dataset and score standardization procedures (e.g., by participants, or sessions etc.). Further, among all participants, there were two individuals with varying levels of laparoscopic experience, with one having 5 hours and the other having 20 hours. Although these participants were categorized as novices, the impact of their experience on the findings remains uncertain. A larger dataset would help mitigate the sensitivity to this issue and more sophisticated data preprocessing method may be necessary to account individual differences of the participants (i.e., subject effect) for ensuring the robustness of the findings.
Conclusion
This study collected eye tracking data to measure the proficiency level of medical students practicing the FLS peg transfer task. Both scene independent and dependent eye metrics were computed and K-means clustering identified three proficiency levels based on these eye metrics. The boxplots showed that greater portion of scene dependent eye metrics (3 out of 5) displayed a clear linear trend with proficiency, while scene independent eye metrics were less clear. The visual comparison of eye metrics in surgical training in differentiating proficiency levels highlights differences with these metrics and the need for further exploration into these metrics for more accurate and precise applications.
Footnotes
Acknowledgements
This research was supported in part by the National Center for Advancing Translational Sciences of the National Institutes of Health under Award Number UL1TR003015. We thanked the Carilion Clinic personnel who volunteered to support our data collection.
