Abstract
Objective:
The objective was to determine whether the scanpaths of air traffic controllers (ATCs) could be used to improve the performance of novices in a conflict detection task.
Background:
Studies in other domains show that novice performance can be improved by exposure to experts’ scanpaths. Whether this effect can be found for an aircraft conflict detection task is unknown.
Method:
Scanpaths of 25 professional ATCs (“experts”) were recorded using a medium-fidelity air traffic control simulation with realistic scripted traffic that included aircraft pairs that would lose separation. A total of 20 novices were exposed to experts’ scanpaths (“treatment”), and their performance (for both loss of separation detection rates and false alarm rates) was compared to that of 20 novices given no treatment or instructions (“control”) and 20 novices who were verbally instructed to attend to altitude (“instruction-only”). Interviews were held about the helpfulness of the exposure. The scanpaths were analyzed to find pattern differences among the three groups.
Results:
Chi-square tests showed significant differences for false alarm rates across the three groups (p = .001). Pairwise Mann–Whitney tests showed that the number of false alarms for the treatment group was significantly lower than that for the control group (p = .005), and trended lower than the instruction-only group (p = .08). Treatment group participants responded that experts’ scanpaths helped. Analysis of scanpaths showed an increased tendency of the scanpath treatment group to follow the experts’ scanpath.
Conclusion:
The scanpath training intervention improved novice performance by reducing false alarms.
Application:
Implementing experts’ scanpaths into novices’ active learning process shows promise in enhancing training effectiveness and reducing training time.
Introduction
Due to the high cost of training new air traffic controllers (ATCs) and the large number of projected hires, the Federal Aviation Administration (FAA) has been interested in less costly ways to train ATCs. If the scanpaths of expert ATCs could be leveraged to train novices, they could help reduce the time and cost of training new controllers. To investigate this possibility, human subject experiments were conducted to evaluate the effectiveness of exposing novices to the scanpaths of professional ATCs. The performance of novices exposed to the experts’ scanpaths was compared with that of novices in a control group who were not exposed to expert scanpaths and that of novices who were given only verbal instructions to attend to the altitudes of aircraft. The scanpath exposure was hypothesized to improve the performance of novices. Interviews were held about the helpfulness of the exposure, and novices’ scanpaths were analyzed to identify possible pattern differences.
This paper is organized as follows. First, some background on scanpaths and the aircraft conflict detection task is presented. Next, the method is described, followed by the results and discussion. Finally, limitations, future work, and a conclusion are provided.
Background
Scanpaths and Their Use in Training
A scanpath is an ordered sequence of eye movements (Noton & Stark, 1971a, 1971b). Specifically, the term scanpath refers to the eye movement path that is formed by fixations and their linkages (or saccades; Goldberg & Kotval, 1999; Goldberg & Schryver, 1993). Visual scanning strategies are important for efficiently carrying out certain tasks. For example, in an aircraft conflict detection task, when many targets are moving on screen, there are several possible visual scanning strategies. Some of these strategies are efficient at accomplishing the task of detecting potential future conflicts and some are not, particularly when considering that the task must be accomplished in a timely manner. If 12 aircraft are on the operator’s radar display, there are 66 different pairs of aircraft to consider for possible conflicts; however, it is unlikely that an observer will try to investigate all combinations one at a time, due to time and effort. Instead, the controller must devise an efficient means of scanning to find possible conflict pairs.
Consequently, experts’ scanpaths may be different from those of novices, which may help explain superior expert performance. Scanpath differences due to expertise have been investigated for landing aircraft under visual flight rules (Kasarskis, Stehwien, Hickox, Aretz, & Wickens, 2001), driving (Underwood, 2007; Underwood, Chapman, Brocklehurst, Underwood, & Crundall, 2003), playing computer games (Underwood, 2005), drawing (Tchalenko, 2009), and searching the web (Aula, Majaranta, & Raiha, 2005). Kang (2012) showed that experts’ and novices’ scanpaths are different for an aircraft conflict detection task for multiple moving targets.
If experts’ scanpaths are more efficient than those of novices, exposing novices to them may improve their performance. Such effects have been verified in other tasks. An experiment with an aircraft inspection task showed that novices who had observed an expert’s scanpath mimicked that scanpath and, as a result, performed the task better than the nontreatment group (Sadasivan, Greenstein, Gramopadhye, & Duchowski, 2005). Also, for a medical diagnostic decision making task, the novice group who observed experts’ eye movement features outperformed the nontreatment group in terms of decision accuracy (Dempere-Marco, Hu, MacDonald, Ellis, Hansell, & Yang, 2002). Based on these studies, we can hypothesize that novice ATCs’ performance may improve if they are exposed to the scanpaths of expert ATCs.
Although expert ATCs’ scanpaths are not completely homogeneous, there are several features that may differ from those of novices including the prevalence of a “circular” (clockwise or counterclockwise) visual scanning pattern and the tendency to fixate on altitudes during the scanning. However, such observations are largely anecdotal or subjectively reported; it is not entirely clear what features of an expert scanpath may be helpful. As a first step toward understanding the utility of being exposed to experts’ scanpaths, this work focuses on determining whether basic exposure to the scanpaths themselves is helpful. If so, the results may be useful to the FAA for training ATCs, and researchers could then try to understand why exposure to the scanpath may be helpful.
Aircraft Conflict Detection Task
One of the primary roles of an ATC is to view a radar display and to detect possible conflicts among multiple aircraft. For en route air traffic control, a loss of separation (LOS) occurs when two aircraft are separated by less than 5 nautical miles (nmi) horizontally and 1,000 feet vertically. LOS rules change when the aircraft are close to the airport and when radar stations are close together. LOS can be pictured as a thin disc with a radius of 5 nmi and a height of 2,000 feet, with an aircraft at the center. Only one aircraft can be in any one disc. Having two aircraft in one disc would represent an LOS.
Among the elements typically considered relevant for detecting an LOS are (a) aircraft altitude, (b) bearing from the other aircraft in the pair, and (c) speed. First, if an aircraft pair will remain separated by an altitude of 1,000 feet or more, then there is no LOS, regardless of bearing or speed, and there has been research indicating that altitude is one of the features on which ATCs fixate (Rantanen & Nunes, 2005). Second, if aircraft are not converging, then there is no possibility of an LOS. Last, even if aircraft are at similar altitudes and are converging, differences in distance to the point of convergence and/or in the respective aircraft speeds can ensure that proper lateral separation is retained. Studies have identified other conflict detection strategies (Neal, 2009; Niessen & Eyferth, 2001; Rantanen & Nunes, 2005), and scanpaths have been examined for conflict detection among a small number of aircraft (Hunter & Parush, 2009; Kang & Landry, 2010).
Method
This research investigates whether performance, in terms of LOS detection rates and false alarm rates, will be different for novices exposed to experts’ scanpaths as compared to a control group of novices without the treatment and to novices given instructions to attend primarily to altitude. Evaluation of both measurements would reflect improved performance for the conflict detection task from the scanpath treatment. In addition, the subjective ratings and opinions of the participants are provided with respect to the scanpath exposure. Finally, post hoc analyses of the scanpaths seek to identify scanpath differences among the groups.
Data Collection and Displaying Experts’ Scanpaths
To obtain expert scanpaths, 25 professional, FAA-certified ATCs at Indianapolis ARTCC were recruited. They had an average of 20.7 years (SD = 7.1) of experience, ranging from 3 to 30 years. These experts completed two practice scenarios to familiarize themselves with the equipment and display; their eye movements were recorded using a Tobii X60 eye tracker while performing conflict detection for three other scenarios. The task was to verbally identify the aircraft call signs of LOS pairs as soon as detected. Each scenario ended when the participant stated there were no remaining LOS pairs. (When presented with a new traffic situation, ATCs will typically scan the display and configure all traffic pairs to eliminate LOSs; the remaining task is monitoring. Moreover, after initial detection, controllers would normally apply control, which could induce additional conflicts and severely complicate the analysis.)
The three recorded scenarios were presented to each expert in a random order. As part of the agreement with respect to obtaining scanpaths from these controllers, no performance data were recorded. Detailed procedures for obtaining experts’ scanpaths were similar to how the novices’ scanpaths were obtained, as outlined in the procedure section.
After recording experts’ scanpaths, a short open-ended interview was conducted by instructing the experts to write their visual strategies. The strategies were categorized based on the terms that the experts expressed (shown in Table 1). A dominant search strategy was identified, which was a circular sweep, and the experts articulated that they were attending to the altitudes. Recorded scanpaths for the dominant search strategy were visually observed to identify whether circular scan actually occurred. Based on the dominant search strategy from the interview results and scanpath observations, three scanpaths, one from each scenario, were selected to be used as the treatment for novices. It was decided that showing a raw recorded scanpath would be better than generating an artificially “good” one since there was some complexity of scanpaths among experts.
Interview Results of Experts’ General Search Patterns
When displaying each scanpath as a treatment for novices, the recorded scanpath was shown in real time overlaid on the radar screen. Scanpaths were set at 50% transparency since the overlaid scanpaths could interfere with reading the data on the radar screen. An example of an expert scanpath used for treatment is shown in Figure 1. The circles indicate fixations, the lines indicate saccades, and the radius of each circle is proportional to the fixation duration. Figure 1 shows the scanpath capture over the first 5 s. In this case, the first fixation starts from the top-left corner, and then the participant scans in a clockwise movement back to the original starting point. As time progressed, experts would begin identifying LOS pairs; therefore, actual scanpaths would also typically contain repetitive movements between the possible LOS pairs.

Expert’s scanpath overlaid in real time on radar display.
Participants
Undergraduate and graduate engineering students were recruited at Purdue University. Participants were recruited through advertisements posted on boards and through distributed e-mails sent through the administrators at Purdue University. Initially, 40 participants were recruited and randomly assigned to the “scanpath treatment” group and the “control” group. For follow-up research, 20 additional participants were later recruited in the same manner and assigned to the instruction-only group. The average ages of the three groups were 23.1 years (SD = 2.4), 23.8 years (SD = 3.2), and 24.4 years (SD = 1.9), respectively. The participants had all seen a radar display before but indicated little knowledge of how to perform conflict detection.
Apparatus
The experiment was conducted at Purdue University using a Tobii X60 eye tracker and a 19-inch LCD monitor. The accuracy of the eye tracker was 0.5° of visual angle. A participant’s eye is approximately 1 m from the monitor; therefore, it is possible that the fixation error could be up to 1 cm on the screen. The size of the data block was 1.1 cm (height) × 1.5 cm (length); thus, the visual angle of the data block was approximately 0.6° considering the height. The data collection rate was set at 60 Hz.
A software suite, Simscope/Simtarget, was used to simulate the en route air traffic radar display. The radar mode was set to digital surveillance radar (DSR) mode, which is a high-fidelity representation of the actual radar display used in en route air traffic control facilities. The refresh rate of aircraft movement was set at 5 s.
The symbology for aircraft was standard for actual air traffic control displays. Figure 2 shows an example of an aircraft and its data tag is shown on the radar display. The small diamond shape is the actual aircraft, and the vector line that stems out from it shows the aircraft’s current direction. In the example, the aircraft name is British Airways 179 (BAW179), its altitude is currently constant at 37,000 feet (370C), and its speed is constant at 410 nmi/hr (or knots). If an aircraft were to change its altitude, the second data line (altitude data) in Figure 2 would be shown in the format as “300↓367,” meaning that the current altitude is 36,700 feet and it is descending to 30,000 feet.

An aircraft and its data tag.
Scenarios
Scenario information is provided in Table 2. Five scenarios, consisting of the two sample scenarios and the three expert scan scenarios (used to record experts’ scanpaths), were used as practice scenarios for the scanpath treatment, control, and instruction-only groups. Note that the recorded experts’ scanpaths from the expert scan scenarios were used as training for the scanpath treatment group, as described later.
Scenario Details
Note. LOS = loss of separation.
The four performance evaluation scenarios and five practice scenarios included LOSs to occur approximately 8 min (or more) after the start of the scenario. For the performance evaluation scenarios, the perfect score for each participant to identify all LOS pairs was 8 (2 LOSs each over 4 scenarios), and the perfect score for each group was 160 since there were 20 participants per group.
An example expert scan scenario is shown in Figure 3. (The image is inverted for legibility in printed form; on the actual display the background is black and the aircraft and data tags are green. In addition, the size of the data block is doubled for legibility.) As time progresses, each aircraft makes discrete “jump” movements along its flight path, which is consistent with the type of movement that occurs on the actual radar display. In Figure 3, there are three LOS pairs: (UAL120, AAL833), (AWE2585, AAL123), (FLG915, TRS988). Notice that each LOS pair has the same altitude, whereas other converging aircraft pairs, such as (UAL535, TRS988), do not.

Initial aircraft layout for an expert scan scenario.
Initial direction, speed, and altitude for all aircraft were carefully configured to create aircraft pairs that would have LOSs in the near future. All other scenarios were similar in appearance to the scenario shown in Figure 3. Configurations do not change as time progresses except for the aircraft that change altitude. If aircraft were climbing or descending, which was assigned randomly, such change of altitude was indicated immediately at the start of each scenario. Note that the initial starting position, aircraft name, and the direction/speed/altitude of each aircraft were varied to prevent the participants from making guesses based on previous scenarios without actually paying attention to the multiple aircraft that they were currently observing.
In addition, the performance evaluation scenarios differed from each other in two primary ways: the type of conflict in the scenario (two angled conflicts or overtaking/head-on) and whether altitude changes occurred (yes or no). For the two angled conflicts scenarios, the conflicts occurred due to converging paths at similar altitude. For the overtaking/head-on scenarios, one of the conflicts was a head-on conflict whereas the other was an overtaking conflict. These differences were crossed, resulting in the four scenarios.
Procedure
Participants were trained on how to observe and identify LOS pairs from the simulated DSR en route radar display. The participants were informed of how the scenarios were set up (that there were no abrupt changes of direction, speed, or altitude after a scenario was initiated) as well as practiced through the sample scenarios. Upon spotting a possible LOS pair that would occur in the future, they were instructed to quickly verbally state the call signs of the pair. After the participant believed there were no remaining LOS pairs, the participant would answer “done” to end each scenario. Participants were told that they could identify any pair again without penalty.
All three groups (scanpath treatment, control, and instruction-only) performed the five practice scenarios. Then, all three groups were shown the initial layout of the three recorded expert scan scenarios for self-review.
Next, the scanpath treatment group was told to observe only the experts’ scanpaths overlaid on the radar screen in real time (Figure 1). Three scanpaths, one from each expert scan scenario, were shown sequentially. When showing a scanpath, a fixation and its associated saccade were set to disappear after 5 s to avoid any clutter of fixations and saccades on the screen if the whole scanpath were to be shown. The time it took to show all three experts’ scanpaths was approximately 10 to 12 min. Although this treatment was for training, we did not explicitly inform them that they are obligated to follow the experts’ scanpath pattern to consider a possibility that the participants could find the experts’ scanpaths to be merely confusing to follow after viewing them.
When showing the scanpaths to the treatment group, the participants were told that the expert ATCs from whom the scanpaths were recorded explained that the fixations were primarily on the altitudes rather than the targets themselves. The reason for this is that the accuracy of the eye tracker was 0.5°, which, at the viewing distance of both the experts and novices, was insufficient to distinguish between different elements of the data tag.
The 20 instruction-only group participants were not shown the expert scanpaths but instead were told, “When observing all aircraft on screen, attend to the altitudes first.” Control group participants were given no instructions, nor were they shown the expert scanpaths. Afterward, all three groups performed four performance evaluation scenarios (S1–S4) in random order.
A Likert-type scale questionnaire was administered that asked how helpful the scanpath training was based on the following scale: 1 (no help), 2 (slightly helpful), 3 (somewhat helpful), 4 (much helpful), and 5 (extremely helpful). A follow-up question requested that the participant provide a reason for the rating. The same type of questionnaire was also provided to the instruction-only group in terms of emphasis on altitude instruction.
Measurements
The dependent variables were LOS detection rates and false alarm rates. A false alarm occurs when an operator identifies an LOS pair but no LOS is projected to occur for that pair at the time of detection. False alarms are undesirable since controllers typically apply some control to prevent the LOS from occurring, such as path-stretching maneuvers, altitude changes, or speed changes. Each of these maneuvers incurs a fuel burn penalty and a possible schedule impact. In addition, the controller expends unneeded effort in determining and executing the resolution.
For the participants’ subjective evaluation of the scanpath training, ratings of how much the training was helpful were collected, as indicated earlier.
The eye-tracking data taken on each participant were visually observed by overlaying the scanpaths onto the scenarios to identify general scanpath patterns.
Data Analysis
A chi-square two-sided test was used for statistical tests to determine whether the measurements differed by group (scanpath treatment, control, or instruction-only). For statistical analysis purposes, however, the four scenarios were treated as replications and not subdivided. This decision was made to preserve statistical power and because these differences were considered normal random variation. Post hoc tests were used to identify possible effects of these differences for focused future study.
Results
Performance Results
The LOS detection rate was 0%, 50%, or 100% (0, 1, or 2 correctly detected LOS), so the data were not normally distributed. A chi-square test of whether the distributions of LOS detections (0, 1, or 2; for all four scenarios) differed by group was performed, but the differences seen in Figure 4 were not statistically significant, χ2(4) = 4.181, p = .382. (A significance level of .05 was used for all analyses.) The modes and medians of the scanpath treatment, instruction-only, and control groups were 2, 2, and 1 (modes) and 1, 1, and 0.5 (medians), respectively.

Loss of separation (LOS) detections by group.
To provide some additional insight into the differences in correct LOS detections for future research, a nonparametric bootstrapping method (2,000 repetitions) was used to compute an estimate of mean proportion of correct LOS detection and a 95% confidence interval on that estimate for each group. For the scanpath treatment group, the mean LOS detection rate was 0.743, with a confidence interval of 0.669 to 0.819. For the instruction-only group, the mean LOS detection rate was 0.694, with a confidence interval of 0.613 to 0.769. For the control group, the mean was 0.681, with a confidence interval of 0.613 to 0.750.
Since the number of false alarms was, in all cases, 0, 1, or 2, the data were again not normally distributed, and it was decided to use the same procedure as used for the LOS detections. A chi-square test of whether the distributions of false alarms (for all four scenarios) differed by group was performed. The differences seen in Figure 5 were statistically significant, χ2(4) = 18.132, p = .001. The modes and medians of the scanpath treatment, instruction-only, and control groups were all zero. The total numbers of false alarms for the scanpath treatment, instruction-only, and control groups were 10, 21, and 35, respectively.

False alarms by group.
The trend in Figure 5 seems to be that there were fewer false alarms for the treatment group as compared to either the instruction-only or the control group. To test this, pairwise two-way Mann–Whitney tests were done on the raw false alarm counts against group. Between the treatment and control groups, the difference was significant (W = 5860, p = .005, adjusted for ties). Between the treatment and instruction-only groups, the difference was marginally significant (W = 6105, p = .08, adjusted for ties). Between the instruction-only and control groups, the difference was not statistically significant (W = 6166.5, p = .22).
Again, to provide some additional insight into the differences in false alarms for future research, a nonparametric bootstrapping method (2,000 repetitions) was used to compute an estimate of mean number of false alarms and a 95% confidence interval on that estimate for each group. For the treatment group, the mean number of false alarms per trial was 0.12, with a confidence interval of 0.05 to 0.20. For the instruction-only group, the mean was 0.26 per trial, with a confidence interval of 0.15 to 0.39. For the control group, the mean was 0.44 per trial, with a confidence interval of 0.29 to 0.61.
To see if any of the underlying factors had a strong influence on false alarms, the false alarms were investigated in more detail and are categorized as shown in Tables 3 and 4. In Table 3, the number of false alarms are divided into two categories: (a) the pair is converging, but would not have had an LOS due to having proper altitude separation and (b) the pair was at the same altitude, but would not have had an LOS since the pair was not converging. “Not converging” means that the aircraft were on diverging courses or parallel courses or one aircraft was following the other aircraft but not overtaking it.
False Alarm Details (Converging/Co-altitude Variations)
False Alarm Details (Altitude Change Variations)
Results of Scanpath Observation to Identify Scanpath Differences
To help clarify what effect the treatment may have had, scanpath behaviors for each group was obtained by observing the scanpath data from the performance scenarios. The characterizations and the number of participants who appeared to follow that behavior for each scenario are provided in Table 5. Individual runs were not analyzed separately to prevent misinterpreting individual data points for trends.
Scanpath Observations for All Performance Scenarios
The categorization of the scanpaths was performed by one rater, and a broad categorization, using roughly orthogonal classifications, was used to enable simpler and more accurate classification with fewer possible errors. The categories were developed by observing the data. For example, when a scanpath was observed, the scanpath clearly created a circular motion similar to Figure 1; therefore, it was classified as a “circular scan.” The scanpaths that were difficult to characterize due to their complexity were classified as “others.”
On average, a large majority (71%) of the participants in the treatment group appeared to follow a circular scanpath, which was consistent with the experts’ scanpath. However, relatively fewer in the instruction and control group (8% and 5%, respectively) followed this type of scanpath. Instead, approximately half of the participants in the instruction and the control groups (50% and 55%, respectively) showed “trajectory-based” scanpath, meaning that they were largely making transitions among aircraft that shared a converging point.
Questionnaire Summary—Subjective Scanpath Rating and Its Reason
The treatment group’s mean rating for the scanpath treatment was 4.2 on a scale of 1 (not helpful) to 5 (very helpful), with an SD of 0.93. The majority of the participants stated that (a) experts’ scanpaths have authority and they were inclined to follow them, (b) they began to observe all the altitudes first, and (c) they tried to perform a circular scan. The others who said that scanpaths were not helpful answered that they tried to follow the movement but found it difficult to adapt.
The instruction-only group’s mean rating for the emphasis on altitude instruction was 3.8 on a scale of 1 (not helpful) to 5 (very helpful), with an SD of 0.96. The majority of the participants stated that it was easier to estimate LOS pairs by reading the altitudes first; however, others stated that it was still difficult to attend to and keep track of the altitudes due to so many moving aircraft.
General Discussion
Performance
In summary, the effect of the treatment appears to be, primarily, a strong (approximately 73%) reduction in the number of false alarms. This effect is substantially larger than when the participants were given only verbal instructions, suggesting that the graphical depiction of scanpath provides a benefit over and above such verbal instructions. The effect was consistent, regardless of whether the false alarm was due to mistaking altitude differences or lack of convergence, or whether the conflict occurred in a pair where at least one aircraft’s altitude was changing. The nature of the effect, on false alarms rather than direct performance, was not expected prior to the commencement of the research, and may be instructive for future work.
A relatively small, not statistically significant, effect on LOS detection was found as well; if nonzero, this difference appears to be in the 5% range. If such an LOS detection performance difference is of sufficient practical meaning for the FAA, additional tests with more power should be done to check and refine the estimate. The improvement provided by the scanpath training treatment largely agrees with the performance increase on Sadasivan et al.’s (2005) aircraft inspection task and Dempere-Marco et al.’s (2002) medical diagnosis task.
The visual analysis of scanpaths indicates that the participants in the scanpath treatment group were predominantly using a circular sweep strategy, which was more consistent with the strategy of professional ATCs. This was not found in the other two groups in such a predominant fashion. In contrast, the majority of the control and instruction-only group seemed to try to project trajectories, trying to identify whether one aircraft was converging with another.
The increased efficiency, as indicated by having fewer false alarms with at least as good performance, might therefore be related to the effect of the scanpath training in causing the novices to adopt a circular sweep pattern, which, when combined with attending primarily to altitude, resulted in reduced false alarms. Specifically whether this is the underlying causal mechanism, along with a theory for why such a mechanism can result in improve efficiency, is not known and should be studied further; no clear insight is provided from the results of this experiment.
It is noted that the control and instruction groups’ trajectory-oriented scanpath finding differs from those of Rantanen and Nunes (2005) and Kang and Landry (2010), perhaps due to the different numbers of aircraft on screen. The results of that previous research showed that novices were well aware of the importance of altitudes. Specifically, Rantanen and Nunes stated that novices first viewed altitudes. However, both experiments were based on using one or two pairs of aircraft that were on a converging course to some extent, whereas the experiment in this paper was based on a more realistic traffic scenario that included 12 aircraft, not all of which converged, and some of which were changing altitude. It is therefore recommended that future studies use a more realistic traffic configuration to ensure that results are consistent with likely behavior.
It is also possible that the control group participants were saturated by having to observe many converging pairs, which led them to either misread or misunderstand the altitudes. As shown in an expert scan scenario in Figure 3, there were in some cases as many as 20 converging aircraft pairs among the 12 aircraft on the radar screen. This, however, is not unusual in real air traffic control scenarios.
The interviews with the scanpath trained group indicated another possible reason. Some of the participants answered that they tried to follow expert ATCs’ scanpath but found it difficult to adapt. It may be possible that the participants who said the scanpath training was helpful also had difficulties to some degree. That is, the treatment group may have been still undergoing a learning phase due to the training time being short. Since no practice time was given after the scanpath training, it is possible that the novices were still practicing when they participated in the four performance evaluation scenarios.
The treatment group mimicked experts’ scanpaths even though they were not given any instructions to follow it. On one hand, this may not be surprising, as there was no other logical conclusion for why the group was presented with the scanpath other than that they should follow it. When asked, many of the participants answered that the reason they followed the depicted scanpath was due to authority of the experts’ scanpaths. Specifically, the treatment group novices stated that if the experts whose scanpaths they observed were much better at detecting LOSs, then they should probably follow the same scanning behavior. Therefore, it seems that the experts’ scanpaths not only might provide motivation, but also might enforce the participants to follow a specific search pattern.
The analysis of experts’ and novices’ scanpaths suggests two important visual search strategies that may contribute to efficient conflict detections. First, the circular scanning seems to ensure that the observer does not miss any aircraft on screen. The organized, repetitive pattern would seem to promote more consistent and thorough coverage than a pairwise search; in the latter case it would seem to be easy to lose track of which pairs had and had not been interrogated.
Second, focusing on reading aircraft altitudes during the circular sweep seems to help reduce the observer’s workload during LOS detections. For example, if the observer reads only the altitudes first, then that observer needs to investigate only the co-altitude pairs; however, if the observer first searches for converging aircraft pairs based on their trajectories, then there may be many aircraft pairs to investigate, including many that will not lose separation.
However, it is important to state that although the “circular full scan of altitudes” strategy is favorable, it is not a single absolute strategy. As shown earlier, experts showed different types of search strategies, such as density-based, quadrant-based, and zigzag searches. Therefore, in addition to showing experts’ dominant scanpaths to novices, we may also need to identify other expert visual strategies from which novices might also benefit.
Limitations and Future Research
One of the major issues for identifying the effects of scanpath training is that it is difficult to isolate the effect of the scanpath itself from the effect of the underlying characteristics of the scanpath, such as the general nature of the scanpath (e.g., a circular sweep) and the information to which the controller primarily attends (e.g., altitude information). If such information could be clearly identified, then perhaps it could be transferred verbally to controllers, as was done here with the altitude instructions, instead of showing a scanpath to trainees. However, there are likely many characteristics of expert scanpaths, some of which may be hard to identify or explain, and simply showing the image of a scanpath may have benefits beyond those of simple verbal instruction (see, e.g., Chervinskaya & Wasserman, 2000).
Nonetheless, a better understanding of why aspects of the scan, such as the circular sweep, are utilized by experts would be of great utility, but was not investigated in this study, which was simply focused on whether the portrayal of expert scanpaths to novices would result in improved performance. We would be very interested in any insight other researchers could provide on why certain aspects of the expert scanpaths are commonly observed and whether those particular aspects have a clear connection with improved performance.
Another limitation is the length of the experiment. ATC training takes several years, whereas this experiment lasted only one period. The effect of repeated or prolonged training with scanpaths should be investigated. Moreover, this experiment, although showing a benefit of scanpath training, did not examine the effect of scanpath training on the required duration of training, which may be of more direct interest to the FAA. Therefore, a full-scale implementation and longer testing of using scanpaths at ATC training centers are required. This implementation would enable additional experiments to refine scanpath-based training and to investigate the limitations further.
Last, several of the findings were not significant or were of marginal significance. As noted in the discussion section, the performance effect (5%–6%) is probably of practical significance, and additional observations are needed to see if this effect actually exists or not. Similarly, the false alarm effects, which are in the 50% to 70% reduction range, are also likely of practical significance, but would need to be confirmed with additional observations.
Conclusion
Presentation of expert scanpaths was used to find out whether novices could increase their performance on detecting aircraft conflicts. The results indicate that scanpath training appears to be effective at substantially reducing the number of false alarms produced by novices after a short training period with no substantial reduction in LOS detection performance. A circular pattern for the scanpath, in addition to primarily attending to aircraft altitudes during the scan, seems to result in the task being accomplished more efficiently. It is suggested that this intervention may assist in training novice ATCs.
Key Points
Experts’ scanpaths were presented to novices in an aircraft conflict detection task as part of their training.
The group treated with the scanpath presentation had approximately 70% fewer false alarms than the control group.
This false alarms effect was substantially larger than just providing a group of novices with verbal instructions to attend primarily to altitude.
The group treated with the scanpath presentation strongly tended to follow the same scanpath pattern as experts, whereas the control and instruction-only groups did not.
Footnotes
Acknowledgements
We thank the Air Traffic Control Association (ATCA) and the ATCs at the Indianapolis ARTCC for fully supporting the experiments. Many thanks to Tom Dury at the Indianapolis ARTCC and Barry Frazier at the Washington, D.C., ARTCC for playing key roles in making the experiment possible. We also thank Ellen Bass for providing insightful comments on this paper.
Ziho Kang is a postdoctoral researcher in the College of Computing and Informatics at Drexel University in Philadelphia, Pennsylvania, specializing in decision making and eye-tracking research. He earned his PhD in the School of Industrial Engineering at Purdue University in West Lafayette, Indiana, in 2012.
Steven J. Landry is an associate professor and associate head in the School of Industrial Engineering at Purdue University in West Lafayette, Indiana, specializing in the area of human factors and air transportation systems engineering. He earned his PhD in the H. Milton Stewart School of Industrial & Systems Engineering at Georgia Institute of Technology, Atlanta, Georgia, in 2004.
