Abstract
Objective
To examine evidence of sensitivity, predictiveness, and methodological concerns regarding direct, objective measures of situation awareness (SA).
Background
The ability to objectively measure SA is important to the evaluation of user interfaces and displays, training programs, and automation initiatives, as well as for studies that seek to better understand SA in both individuals and teams. A number of methodological criticisms have been raised creating significant confusion in the research field.
Method
A meta-analysis of 243 studies was conducted to examine evidence of sensitivity and predictiveness, and to address methodological questions regarding Situation Awareness Global Assessment Technique (SAGAT), Situation Present Assessment Technique (SPAM), and their variants.
Results
SAGAT and SPAM were found to be equally predictive of performance. SPAM (64%) and real-time probes (73%) were found to have significantly lower sensitivity in comparison to SAGAT (94%). While SAGAT was found not to be overly memory reliant nor intrusive into operator performance, SPAM resulted in problems with intrusiveness in 40% of the studies examined, as well as problems with speed-accuracy tradeoffs, sampling bias, and confounds with workload. Concerns about memory reliance, the utility of these measures for assessing Team SA, and other issues are also addressed.
Conclusion
SAGAT was found to be a highly sensitive, reliable, and predictive measure of SA that is useful across a wide variety of domains and experimental settings.
Application
Direct, objective SA measurement provides useful and diagnostic insights for research and design in a wide variety of domains and study objectives.
Introduction
The study of situation awareness (SA) has grown over the past three decades to encompass investigations of not only the construct, but also the evaluation of new system designs and training programs in a wide variety of domains. To accommodate these efforts, the need to measure SA has remained fundamental to progress in this area of research. A number of approaches have been proposed and used in research on SA, including process measures, performance measures, and measures that seek to assess a person’s level of knowledge and understanding about the situation via direct questioning of the individual. These techniques are summarized in Table 1, along with some advantages and disadvantages of each. See Endsley (1995b) or Endsley and Jones (2012) for a review. SA measurement approaches will be briefly described, followed by a focus on the direct measurement of SA as an ongoing state of knowledge about relevant information in the environment that is needed to support decision-making in complex and dynamic domains.
Comparison of SA Measurement Approaches
Note. SA = situation awareness; SART = Situational Awareness Rating Technique; SAGAT = Situation Awareness Global Assessment Technique; SPAM = Situation Present Assessment Technique.
Some researchers have examined the processes that people use to develop SA using measures such as eye tracking (de Winter, Eisma, Cabrall, Hancock, & Stanton, 2019; Ikuma, Harvey, Taylor, & Handal, 2014; Smolensky, 1993), communications (Bolstad et al., 2007; Gorman, Cooke, Pederson, Connor, & DeJoode, 2005; Orasanu, 2000; Prince, Salas, & Stout, 1995), verbal protocols (Hall & Phelps, 1983; Rose, Bearman, Naweed, & Dorrian, 2019; Sullivan & Blackman, 1991; Walker, Stanton, & Young, 2008), and physiological measurement (Wilson, 2000). While process measures can provide insights into how people develop SA, they can only be used to indirectly infer the quality and completeness of the resulting SA, as a state of knowledge about the situation, obtained by the individuals involved. An individual’s knowledge and capabilities, as well as the system interfaces available, mediate the degree to which the SA processes used are successful in creating accurate SA; so even two people using the same process may not arrive at the same situation understanding. Furthermore, these techniques generally provide only partial insights into what information is attended to and how it is processed to form situation comprehension and projection. Verbal protocols have not been found to be effective for measuring SA (Rose et al., 2019; Walker et al., 2008), and physiological measures have had little research to establish validity for this purpose. Therefore, process measures are inadequate for measuring SA on their own.
Other researchers have focused on trying to infer whether or not people have good SA based on their actions and performance; however, this approach can be severely hampered when people do not act as expected (Farley, Hansman, Amonlirdviman, & Endsley, 2000; Jones & Endsley, 2000b; Pritchett, Hansman, & Johnson, 1995). While collecting data on performance is always desirable, the main limitation of this class of measures is that it only infers SA indirectly, in what can be a circular argument, without providing sufficient detail or diagnosticity on what the operator really thought was going on. In addition, Wickens (2000) pointed out that SA for normal and emergency events may be quite different, thus inferences about SA are highly constrained by the scenarios that are tested. In that many other factors can effect the decisions and actions people make based on their SA (such as tactics, procedures, and execution skills), behavioral and performance measures also offer only limited insights into SA.
Largely due to the limitations of the above approaches, the majority of research employing SA measures has focused on directly determining a person’s SA, as a state of knowledge, through either subjective or objective assessment (Endsley & Garland, 2000). Subjective measures of SA are easy to obtain; however, they suffer from the fact that people may be unaware of what they do not know and may be unduly influenced by observations of performance (Endsley, 1995b; Endsley & Jones, 2012). Research indicates that subjective SA ratings more likely assess an individual’s confidence level (Endsley, Selcon, Hardiman, & Croft, 1998; Hamilton, Mancuso, Mohammed, Tesler, & McNeese, 2017; Sulistyawati, Wickens, & Chui, 2009). Observer ratings of SA overcome some problems with self-assessment, but still suffer from becoming a proxy for subjective performance ratings in that observers have limited information on a person’s mental representation of the situation. See also Endsley (in press).
Objective SA measures assess the operator’s knowledge of the current situation by asking the operator relevant questions about the situation that can be objectively scored as correct or incorrect. Two commonly used metrics in this category are the Situation Awareness Global Assessment Technique (SAGAT) (Endsley, 1988, 1995b) and the Situation Present Assessment Technique (SPAM) (Durso et al., 1998). However, these methods have been the subject of considerable debate, which has created confusion in the field (Chiappe, Rorie, Moran, & Vu, 2012; Durso et al., 1998; Endsley, 1995a; Jones & Endsley, 2004; Salmon et al., 2008; Salmon, Stanton, & Young, 2011; Sarter & Woods, 1991).
The objective of this article is to examine the SAGAT and SPAM methods for objective measurement of SA and to review and compare the success of these metrics. I first provide an overview of these two SA measurement approaches along with a summary of their advantages and disadvantages. I then focus on questions of sensitivity and predictiveness, performing a meta-analysis of the research literature that has been conducted on and with these SA metrics. In addition, a systematic review of the available literature is conducted to address a number of methodological issues that have been raised concerning the construct validity of the two techniques.
SAGAT
SAGAT is one of the earliest and most widely used measures of SA. When using SAGAT, simulations of representative tasks or scenarios are frozen at randomly selected times, and system displays are blanked while people quickly answer questions about their current perceptions of the situation. Queries can be provided verbally, via pencil and paper, or on a computer or tablet for ease of administration. SAGAT queries correspond to an individual’s SA requirements as determined from the results of a SA requirements’ analysis, such as a Goal Directed Task Analysis (Endsley, 1993b; Endsley & Jones, 2012), for a given domain and role. People’s perceptions are then compared to the real situation, based on simulation computer databases, to provide an objective measure of SA. Scoring some queries, such as those related to situation comprehension, may be provided by subject matter experts with perfect knowledge of the situation at the time of the freeze. Multiple “snapshots” of a person’s SA are acquired in this way, providing an assessment of the quality of SA provided by a particular system design.
SAGAT provides an objective, unbiased assessment of SA by providing the queries at random times across a scenario, collecting data during both high-workload and low-workload periods. By assessing SA at different points in time, SAGAT avoids the problem of relying on memory of events after the scenario is completed (Nisbett & Wilson, 1977). As a global measurement tool, SAGAT includes queries across a wide range of SA requirements for a given job, including Level 1 (perception of data), Level 2 (comprehension of meaning), and Level 3 (projection of the near future). This includes a consideration of system functioning and status, as well as relevant features of the external environment and team as appropriate. Since SAGAT includes queries across the full spectrum of an individual’s SA requirements, this approach minimizes the possible biasing of attention, as people cannot prepare for the queries in advance (Endsley, 1995b). Examples of queries for air traffic control (ATC) are shown in Table 2.
Sample SAGAT Queries for Air Traffic Control
Note. SAGAT = Situation Awareness Global Assessment Technique; SA = situation awareness.
SAGAT scores are normally expressed as percent correct for each query, based on operationally relevant tolerance bands. Many re-searchers have varied from this recommended approach, instead combining the scores on all SAGAT queries into a combined overall score, or into three combined scores that represent Level 1, Level 2, and Level 3 SA. Other scoring variants include (a) the SA Control Room Inventory (SACRI), which computes a d′ sensitivity and bias score based on signal detection theory (Hogg, Follesø, Strand-Volden, & Torralba, 1995); (b) the Qualitative Assessment of SA (QUASA), which poses probes in the form of true/false statements and adds an assessment of confidence in each answer (McGuinness, 2004); (c) SALSA, which attempts to weight each query type in computing an overall score (Hauss & Eyferth, 2003); and (d) the SA Verification and Analysis Tool (SAVANT), which provides a partial display to ask questions about missing information, and includes an assessment of time to answer each question along with accuracy (Willems & Heiney, 2001).
SPAM
SPAM and real-time probes also provide queries to assess SA; however, the queries are provided in real time, usually verbally, while the individual is carrying out his or her normal operational tasks. In addition to response accuracy, the time to respond to each SA probe is collected as an index of how readily available the information is. This form of measurement is referred to as a real-time probe (Endsley, 1995b; Jones & Endsley, 2000a). SPAM also provides a “ready” prompt prior to each SA probe that allows operators to delay receiving it until they are ready for the probe. SA probes may be staged as an embedded probe, with a confederate playing the role of the questioner, so that they appear to be natural to the scenario. SPAM provides queries that correspond to past, present, and future aspects of the situation.
Methodological Issues and Concerns
A number of studies have attempted to directly compare SAGAT and SPAM, focusing on their relative abilities to predict performance in the simulation, with differing conclusions (Durso, Bleckley, & Dattel, 2006; Durso et al., 1998; Jones & Endsley, 2004; Loft, Bowden, et al., 2015; Pierce, Strybel, & Vu, 2008; Strybel, Vu, Kraft, & Minakata, 2008). In addition, many researchers have sought to determine how sensitive these techniques are to the independent manipulation provided in the experiment (Alexander & Wickens, 2005; Jones & Endsley, 2004; Silva, Grigoleit, Ann Burress, & Fitzpatrick, 2017; Vidulich, 2000). A goal of the present study is to examine the existing research base to address these questions and to determine how well these two approaches to objective SA measurement fare in terms of sensitivity and predictive ability. Rather than making an assessment based on any one study, a meta-analysis across the available research literature is conducted to provide a clearer picture of the utility and predictive validity of these two metrics.
The research literature is also rife with a number of criticisms of SAGAT and SPAM that need to be addressed. First, a number of researchers have criticized SAGAT claiming that the freezes to collect data are intrusive and that it relies too much on working memory (Chiappe et al., 2012; Durso et al., 1998; Salmon et al., 2011; Sarter & Woods, 1991). They advocate for the SPAM approach, claiming that it produces a picture of SA that is more natural and “situated” (Chiappe, Strybel, & Vu, 2015).
On the other hand, concerns have also been raised about the intrusiveness of real–time probes and SPAM, in that they require people to multi-task to answer questions while performing operational tasks, which could negatively affect primary task performance (Endsley, 1995b; Jones & Endsley, 2000a; Pierce, 2012). Because of this, SPAM could actually be considered a secondary workload task, even though operators are allowed to delay answering until they feel able. The ability to wait to answer probes until the participant is ready also creates a problem, as it systematically biases SPAM results toward lower workload periods. Another concern is that the ability to look at displays while answering questions fails to capture SA as an ongoing understanding of the world, but rather measures people’s ability to look up information (Endsley, 2015a).
The literature base is full of these logical, but differing, viewpoints, and many of the claims have been offered without evidence. Therefore, empirical research addressing these methodological issues will be examined, including the potential intrusiveness of the measures, their reliance on working memory, speed accuracy trade-offs that may affect outcomes, sampling bias, potential confounds with workload, and query design. The utility of these measures for assessing Team SA is also considered.
Method
A literature search was conducted on Google Scholar, Scopus, and Science Direct with the key words SA, probes, SAGAT, SPAM, SACRI, QUASA, SAVANT, and SALSA. Papers that collected experimental data with one of these techniques (or variants of them), with sufficient data reported, were included in this review. Duplicates of the same study were excluded where identified. A total of 243 papers were tabulated that include data on metrics, experimental environment, domain area, type of participants, number of participants, and experimental results of the study, including statistical analyses.
Given that SACRI, QUASA, SAVANT, and SALSA are all variants of SAGAT, based on the same freeze and query technique, these measures are reported on together. Differences in results associated with the effect of these different analysis approaches are then broken out in the meta-analysis. Similarly, both real-time probes and SPAM are based on the same technique of providing probes to operators during task performance, although SPAM also provides a ready prompt prior to probe administration. These two techniques are also reported on together; however, differences in findings for these two variations are broken-out in the meta-analysis.
Sensitivity
To determine the sensitivity of a measure, typically an attempt is made to independently vary the construct of interest and to assess how well the measure detects this change. In that there is no way to independently vary SA directly, the best proxy in the SA research literature is the variation provided by the experimental manipulation in the study (e.g., changes in displays, automation, training condition, or operator experience). The sensitivity meta-analysis, therefore, examined the sensitivity of the measures to the experimental manipulation in each study, that is, the likelihood that the measures would find a difference in SA between conditions, assuming that true differences exist. While it is almost impossible to say that every study manipulation should result in an effect on SA, the hypothesis of each study was that such a difference should theoretically be found. Furthermore, when performance or workload measures in the same study found differences between conditions, this increases the expectation that the SA measures should be sensitive to study manipulations.
Other approaches have also been applied to examining the sensitivity of metrics in meta-analyses, including Cohen’s d (Cohen, 1995), Cohen’s h (Cohen, 1995), and Cramer’s V (Cramer, 1946), which determine comparative effects sizes. Unfortunately, calculating the comparative sensitivity of the SA measures using any of these approaches was not possible, due to the differing types of output measures provided by the SA metrics. Cohen’s d is often used to compare relative differences between the means of two different groups. So Cohen’s d could be used to show the effects size of a measure such as SPAM that produces response time as its outcome measure. However, this metric is inappropriate for measures where the outcome is expressed as a proportion (i.e., percent correct), such as SAGAT. In such a case, Cohen’s h is more appropriate. Furthermore, some of the studies examined provided only chi-square statistics for frequency counts, in which the appropriate metrics for an effects size calculation is Cramer’s V. Therefore, it was not possible to rely on the calculation of effect sizes in each study as a common method for comparison across these different types of SA measures, since these analyses rely on different statistics that are not comparable.
Instead, sensitivity of each measure was calculated as the percentage of studies where the metric was employed in which it detected a difference between study conditions (i.e., probability of detection). This approach has been used by Vidulich (2000) and Bushman (1994), and is based on considering the independent variable in the study to have been a means of manipulating SA. The sensitivity of the metric can then be assessed as its ability to detect that change. As Wickens (1998) has argued in favor of considering larger p-values than .05, the analysis also considered studies where the metrics showed a significant difference between conditions at a p-value of up to .10. If the measure was found sensitive (p < .05) it received a score of 1, if it was marginally significant (p < .10) it received a score of .5, and if it was not significant (p > .10) it received a score of 0. This provided for an assessment of the proportion of studies using each metric that showed sensitivity to the studies’ manipulation of SA. The relative sensitivity of the SA metrics across different domains, test environments, subject types, and methodological differences were also compared to determine if these factors were relevant to metric sensitivity. In addition, studies that employed both SAGAT and SPAM in the same study were analyzed separately to determine relative sensitivity between the measures.
Predictiveness
Another issue of concern has been whether the SA measure employed is predictive of performance and if so, whether SAGAT or SPAM is more predictive than the other. Assessing the predictiveness of SA metrics against performance measures is rather complicated. First, it assumes that SA metrics should be predictive of performance. While this is true at a high level, it neglects the fact that both SA and performance are multi-dimensional. For example, Wickens (1995) shows that the SA needed for routine performance may not be the same as that needed for emergency performance. And SA of some information may only be relevant to some performance outcomes. For example, knowing how fast the driver of a car is going may be relevant for getting speeding tickets, but not for lane tracking accuracy. Not all SA is relevant to all performance measures. Furthermore, most studies are limited in the number of performance measures assessed, increasing the likelihood that some SA metrics may not have the relevant performance metrics for comparison. Therefore, this meta-analysis assesses whether any SA measure was predictive of any performance measure in each study.
Following the same approach as the sensitivity meta-analysis, predictiveness was calculated as the proportion of studies using a metric in which the SA measure was predictive of at least one performance measure. If the measure was found to be predictive (p < .05) of performance in the study, it received a score of 1, if it was marginally predictive (p < .10) it received a score of 0.5, and if it was not predictive it was assigned a score of 0.
In addition, the degree of strength of the correlations between SA measures and the performance measures across studies was also calculated, by converting each Pearson’s r to Fisher’s Z. This provided for a direct comparison between the mean correlation strength (performance predictiveness) and a 95% confidence interval across metrics. The relative predictiveness of the SA metrics across different domains, test environments, subject types, and methodological differences were directly compared to see if these factors affected predictiveness. Papers that employed both SAGAT and SPAM in the same study were analyzed separately to determine relative predictiveness of the measures.
Intrusiveness
Studies that directly compared performance and or workload in trials in which the SA measures were employed to trials that did not include a SA measure were analyzed to determine evidence of the intrusiveness of the techniques. Intrusiveness was determined by whether the SA measure significantly (p < .05) affected a performance metric in the study.
Memory Dependence
Whether or not SAGAT is overly reliant on working memory was analyzed by qualitatively examining the studies that also independently assessed working memory to determine the relationship between these two measures. Differences between expert and novice subjects were considered. Because SPAM allows operators to see all information when answering queries, it was not considered in this analysis.
Speed Accuracy Trade-offs
As the predominant measure for SPAM is the time to respond to the SA query, this analysis focused on studies that provided evidence of speed-accuracy trade-offs that potentially confound this measure. Because SAGAT does not involve timed performance, but only accuracy, it was not the focus of this analysis.
Sampling Bias
Both SAGAT and SPAM should provide an unbiased estimate of SA by sampling operator knowledge across a wide range of conditions and events during a scenario. Studies were examined to determine whether the provision of a “ready” prompt with the SPAM technique, allowing people to defer answering, created a bias toward lower workload periods.
SA vs. Workload
Studies that examined the relationship between the SA measures and workload were also assessed to determine whether the measures assess SA as an independent construct, or are confounded by workload.
Query Design
Differences in the content and nature of the SA queries provided in SAGAT and SPAM studies were qualitatively assessed.
Team SA
The SA of teams (as an aggregate or in terms of the shared SA among team members) is often a focus of research interest. Studies that used either SAGAT or SPAM to measure team SA were analyzed separately to determine the degree to which these techniques support SA research at the team level in addition to at the individual level.
Results and Discussion
The presentation of results and discussion will be considered together for each aspect of the meta-analysis: sensitivity of the metrics to the study manipulation, and predictiveness of the metrics for the performance measures provided in the study. This is followed by a discussion of findings with regard to intrusiveness of the techniques, memory dependence, the presence of speed-accuracy tradeoffs, the potential for sampling bias, distinguishability from measures of workload, query design, and ability to measure team SA.
Sensitivity
In all, 171 papers were found that used a direct, objective SA measure to evaluate a system design (hardware or displays), automation, operational concept, attention allocation, training manipulation, or to assess differences in participants and the effects of other factors such as workload. Of those, 150 papers involved SAGAT or a variant, and 27 involved SPAM or a real-time probe (with 6 papers including both measures).
SAGAT and variants
Of the 150 papers that included SAGAT or a variant of it, 5 papers were duplicates, 1 paper included 3 studies, and 5 papers included 2 studies, making for 152 unique studies in all. These papers are listed in Appendix A, which is available with the manuscript on the Human Factors website. The proportion of studies in which SAGAT or one of its variants was sensitive to the experimental manipulation in the study is listed in Table 3. The sensitivity score across all 152 studies is 85.5%.
Studies Examining the Sensitivity of SAGAT (and Variants)
Note. SAGAT = Situation Awareness Global Assessment Technique; ATC = air traffic control; SACRI = SA Control Room Inventory; QUASA = Qualitative Assessment of SA; SALSA = Situation awareness of en-route air traffic controllers in the context of automation; SAVANT = SA Verification and Analysis Tool; SA = situation awareness.
p < .05.
Studies were conducted in a number of domains including aviation, ATC, driving, military, medical (including nursing, surgery, and other specialty areas), process control, robotics and unmanned air vehicle control, artificial tasks created specifically for experiments, and other areas including firefighting (4), train driving (2), maritime (2), teleoperations (2), emergency management, weather forecasting, cyber security, and maintenance. Overall, sensitivity of the measure was not significantly different across domains (χ2 = 2.47, p = .96).
The majority of studies were conducted in simulations or microworlds, but there were also studies that employed the viewing of videos and live exercises. There were no significant differences (χ2 = 0.38, p = .94) in the sensitivity of the technique between testing environments, nor in whether the test subjects were experienced domain practitioners, novices in the domain, or were completely naïve (either students or from the general population) (χ2 = 2.26, p = .69).
Significant differences in sensitivity were found based on how the measure was administered. Studies that only measured SA at the end of trials (73% sensitivity) experienced lower sensitivity than those using the freeze technique prescribed by SAGAT (89% sensitivity), (z = 2.02, p = .02). Studies that collected SA only at the end of trials generally suffered from a low sample size as well as only capturing SA at the end of the trial, which may not be indicative of SA throughout.
A second source of variance occurred due to the method of scoring used. The SAGAT methodology recommends scoring each query separately because many design interventions can effect changes in just one or two elements of SA, often in different directions, and if the queries are combined these differences can cancel each other out (Endsley, 2000). This is consistent with Landry and Yoo (2012) who recommend scoring by query rather than a combined score to avoid the problem of assumed equal likelihood of each element. Analysis of SAGAT queries has also shown a lack of internal consistency (Endsley, 1990b, 2000), arguing against combining across queries. Nonetheless, many researchers analyzed SAGAT data as a single score combined across all queries, or as a separate score for each level of SA. It was found that studies using a single combined SA score (81% sensitivity) were less likely to be sensitive than those that were analyzed by query (91% sensitivity) or by SA level (92% sensitivity; z = 1.77, p = .04).
When just the 68 studies that used SAGAT as designed (with scenario freezes and analysis by level or by query) are considered, overall sensitivity rises to 90%. An additional problem was found in that some studies collected only 1 or 2 SAGAT freezes across conditions, resulting in very small sample sizes, well below the 60 samples per query per condition recommended by Endsley (2000). When the four studies with very low sample sizes are removed, overall sensitivity of SAGAT rises to 94%.
It should be noted that the two studies using the SALSA scoring method performed well (100%), three out of four SACRI studies were sensitive (75%), and four of the six studies using the QUASA method were sensitive (67%). However, the numbers of studies using these techniques are so few that it is difficult to conclude anything about them.
SPAM and real-time probes
Of the 27 papers that included SPAM or real-time probes to measure SA, 2 were duplicates and 1 paper involved 2 studies, making for a total of 26 studies. These papers are listed in Appendix B, which is available with the manuscript on the Human Factors website. The proportion of studies in which SPAM or a real-time probe was sensitive to the experimental manipulation in the study is listed in Table 4. The mean sensitivity for these studies overall is 69.2%.
Studies Examining the Sensitivity of SPAM and Real-Time Probes
Note. SPAM = Situation Present Assessment Technique; ATC = air traffic control.
The majority of the studies were conducted in aviation, ATC, driving, military, process control, and submarine management. There were not sufficient numbers of studies to compare sensitivity across these domains. The majority of the studies were conducted in simulations, although a few were also conducted in microworlds, watching videos, or in live exercises. Participants were primarily either experienced operators or students. Neither testing environment nor participant type significantly affected the sensitivity of the metric (p > .1). SPAM probes (sensitivity = 64%) were actually slightly less sensitive than real-time probes (sensitivity = 73%), although this difference was not significant (z = 0.53; p = .30).
Limitations
The sensitivity meta-analyses were based on whether each measure was sensitive in a wide variety of studies, each with different numbers of participants. The use of p-values for conducting hypothesis testing has been criticized on the basis that it does not necessarily indicate the magnitude of effects (i.e., it is subject to the effects of sample size) (Kline, 2004).
However, a p-value threshold for hypothesis testing has been the standard in the field and is generally reported across the literature spanning some 30 years. To examine the possibility that variations in sample size across studies unduly affected the sensitivity analyses, an analysis of variance (ANOVA) was conducted that found no significant effect of number of subjects on the sensitivity score of the studies in this meta-analysis, F(2, 175) = 0.05, p = .95.
Several other caveats also need to be noted with regard to this meta-analysis of sensitivity. First, it is certainly possible that a metric could find a difference between study conditions that does not exist (i.e., false positives). It is also possible that the study itself was poorly constructed with insufficient power, leading to a lack of sensitivity due to study design rather than the metric itself. While there was no obvious data to support that either situation was more or less likely to exist with regard to the studies that used either SAGAT or SPAM, a further assessment was made of the studies that employed both measures within the same study as a means of more carefully assessing their comparative sensitivity to SA differences.
Comparison of technique sensitivity
Figure 1 shows the proportion of studies in which each metric was sensitive to the study manipulation. When conducted with freezes according to the prescribed methodology and scored by query, SAGAT is significantly more sensitive than both SPAM (z = 14.88, p < .001) and real-time probes (z = 3.05, p = .001).

Sensitivity of SA Metrics: Probability of detection of SA difference due to study manipulation.
Six studies included a direct comparison of the sensitivity of SAGAT and either SPAM or real-time queries in the same study. Two studies found SAGAT more sensitive than real-time probes (Endsley, Sollenberger, & Stein, 2000; Jones & Endsley, 2004), three studies found SPAM more sensitive than SAGAT (Alexander & Wickens, 2005; Cummings & Guerlain, 2007; Silva et al., 2017), and one study found mixed results with SAGAT and real-time probes equally sensitive to the study manipulations (Burns et al., 2008).
To further examine these findings, Jones and Endsley (2004) found that SAGAT was more sensitive than real-time probes because the simulation freeze allows more queries to be collected, yielding a more reliable measure of SA (which is important when periodic sampling is involved). As SPAM and real-time probes can only be asked one at a time and since excessive interruptions are not possible, it is at a disadvantage. Of the three studies where SPAM was more sensitive than SAGAT, one measured SAGAT only at the end of two trials, which created insufficient data (Cummings & Guerlain, 2007). In Silva et al. (2017), while SAGAT was not sensitive to operator experience level, neither were the performance measures captured. Experienced operators had better accuracy on SPAM however.
Predictiveness
A total of 53 comparisons were found across 46 papers that examined the predictiveness of an objective measure of SA, as shown in Appendix C, which is available with the manuscript on the Human Factors website. Overall, 89.6% of the analyses found SA measures predictive of the performance or decision-making measures collected in the study.
Table 5 shows that 35 analyses involved SAGAT (predictiveness = 89%), 4 examined end-of-trial SA assessment (predictiveness = 100%), 1 involved SALSA (predictiveness = 50%), 1 involved a real-time probe (predictiveness = 0%), and 12 involved SPAM (predictiveness = 100%). The sensitivity of SAGAT/SALSA/end-of-trial assessments compared to SPAM/real-time probes was not significantly different (z = .37, p = .35).
Predictiveness of SA Metrics
Note. SA = situation awareness; SAGAT = Situation Awareness Global Assessment Technique; SPAM = Situation Present Assessment Technique; SALSA = Situation awareness of en-route air traffic controllers in the context of automation.
A total of 43 studies provided correlation data between the SA measure and performance measures. In comparing the mean Pearson’s r calculated for studies using each method, the end-of-trial measures were most highly predictive of performance (r = .533), followed by SAGAT (r = .459), and then SPAM (r = .411). There was partial overlap in the 95% confidence intervals that were calculated from the converted Z-scores for SAGAT and SPAM, however.
Seven papers examined the relative predictiveness of these measurement techniques within the same study. Two studies found SAGAT to be more predictive than real-time probes or SPAM (Jones & Endsley, 2004; Loft, Bowden, et al., 2015). Two studies found SPAM to be more predictive than SAGAT (Durso et al., 2006; Kraemer & Süß, 2015). The Durso et al. (2006) study, however, artificially restricted SAGAT to only one query per freeze, significantly reducing its sample size unnecessarily. In the Loft, Bowden, et al. (2015) study, which allowed more queries during freezes, this handicap disappeared resulting in SAGAT accounting for twice the variance in performance compared to SPAM.
Three studies found mixed results. Durso et al. (1998) found that SPAM was more sensitive to one performance measure and SAGAT was more sensitive to another. Pierce, Strybel, and Vu (2008) similarly found SAGAT and SPAM to each be more sensitive to different outcome measures. Strybel et al. (2008) showed that SPAM latency was a slightly better predictor of pilot airspeed variability than SAGAT (r2 = .08 vs. r2 = .07), although they used a combined score for SAGAT which is generally less sensitive. Based on these findings, it would appear that both SPAM and SAGAT are predictive of performance.
In one counter-intuitive finding, Loft et al. (2018) found that while SAGAT was predictive of performance in a submarine management study, this was only at the between subjects level, but not at the within subjects level. This would indicate that the differences in SA between subjects correlated with performance, but not changes in an individual’s SA within a trial, which is hard to explain on any theoretical level. SAGAT scores have been shown to be highly stable across individuals (Endsley & Bolstad, 1994), so the finding of between subjects predictiveness is not unexpected.
In examining the lack of a within subjects effect, the most likely explanation is that in this study the queries were all referenced to ship contact numbers which creates a significant memory burden. Previous research has shown that air traffic controllers do not remember aircraft call signs, for example, but rather rely on spatial memory of where air traffic is located (Endsley & Rodgers, 1998). By pinning each query to the ship contact number in the Loft et al. (2018) study, rather than a spatial map as is typical with SAGAT administration, it is possible that the researchers created an artificial memory requirement that interfered with their results. This matter will require further research. The authors also indicated that it is possible that different performance windows need to be considered to link SA to performance in their domain.
Intrusiveness
A common concern about both measures is whether they are intrusive on operator performance, SAGAT by freezing the scenario, and SPAM by dual tasking the operator while conducting operations. A total of 18 studies were found that examined the issue of intrusiveness of one or both of the techniques, shown in Table 6.
Intrusiveness of SA Metrics
Note. SAGAT = Situation Awareness Global Assessment Technique; SPAM = Situation Present Assessment Technique; ATC = air traffic control; RT = response time; SACRI = SA Control Room Inventory; SA = situation awareness.
Eleven of the papers examined the effect of SAGAT on operator performance and or workload. No negative effect of SAGAT administration on operator performance was found in any of these studies. Endsley (1995b) showed that stopping the simulator to collect SAGAT data for 30, 60, or 120 s and either 1, 2, or 3 times in a 20 min scenario had no effect on participant performance as compared to scenarios in which there were no SAGAT freezes. In a second study, I examined whether just the possibility that a SAGAT freeze would be administered created a change in performance (Endsley, 2000). In that study, I told participants there would definitely not be a freeze in one-third of the trials and that there might be a freeze in the other two-thirds (of which half actually experienced a SAGAT freeze). Again, there was no effect on performance from either the SAGAT freeze or the threat of a SAGAT freeze, showing no “pop quiz” effect.
Since those studies, an additional nine studies have compared trials in which SAGAT freezes were administered to trials without SAGAT freezes and found no negative effect on performance indicating intrusiveness. In one study, a small increase in workload (2.4%) was reported in the SAGAT trials (Loft, Bowden, et al., 2015), but the authors reported that this may have been due to the fact that twice as many queries were provided in the SAGAT trials than in SPAM trials. In another study (Kraemer & Süß, 2015), an improvement in performance was shown in the trial that included SAGAT; however, this effect was confounded with trial order and likely showed a learning effect for the very inexperienced participants in that study.
In contrast, 6 of 15 studies (40%) that used SPAM or real-time probes reported a negative effect on operator performance or workload. Pierce (2012) found a negative affect of SPAM on both operator performance and workload. Loft et al. (2016) also found that experts in their study (submarine operators) took almost 20 s to accept SPAM queries, presumably due to workload issues, calling into question its validity. Reportedly subjects delayed answering the probes when SA was lower and uncertainty higher. Similarly Shelton, Kinston, Molyneux, and Ambrose (2013) reported that their participants (physicians) found the verbal probes intrusive and interfered with team dialog. Some queries were not answered at all, presumably due to workload issues. Four studies reported a negative effect of SPAM on concurrent performance, as compared to trials in which SPAM was not administered (Pierce, 2012; Pierce, Strybel, & Vu, 2008; Pierce, Vu, et al., 2008; Strybel et al., 2008).
Based on these studies, SPAM appears to create significant interference with performance and increases operator workload. Keeler et al. (2015) implemented SPAM with computerized queries rather than verbal queries, similarly to Bacon and Strybel (2013) and Silva et al. (2013). None of these 3 studies found SPAM to be intrusive. However, their instructions were to only answer the SPAM queries if the operator felt able. This results in a skewing of SA probes into low workload periods, as will be discussed in a later section.
These results indicate that concerns over the potential intrusiveness of SAGAT are unwarranted, with no negative effects on performance in any of the studies that examined it. It should be noted that the reason SAGAT is not intrusive on performance is that during the freeze period operators are actively working with their situation knowledge to answer the queries (Endsley, 1995b). This keeps the situational information active in memory.
This is a very different situation than interruptions in which a person conducts unrelated tasks or answers unrelated questions, which has been shown to be detrimental to both SA and performance (Gartenberg, Breslow, McCurry, & Trafton, 2014; Loft, Sadler, et al., 2015; McGowan & Banbury, 2004; Ratwani & Trafton, 2008). Negative effects can also be incurred with other types of interference that are not consistent with the SAGAT methodology. For example, Tremblay, Vachon, Lafond, & Kramer (2012) required people to recover SA of changed situations by failing to pause the simulation while displays were blanked, negatively effecting task performance. McGowan and Banbury (2004) provided 10-s interruptions in 20- to 25-s-long video clips that were passively viewed, showing both that irrelevant questions negatively affected performance and that relevant questions improved performance (even when the trial was not interrupted). This very short, highly artificial task is not representative of the normal types of domain relevant and engaging simulations that SA data are collected in. It likely became a task of memorization for the very narrow set of items that were questioned during the study. In contrast, SAGAT is generally only collected two to three times, no closer than 3 min apart in a 20-min interactive scenario where participants are engaged in normal task performance, and provides a wide range of queries that cannot be effectively prepared for in advance.
Memory Dependence
Another concern about SAGAT has been whether it is overly reliant on working memory. This discussion rests partially on whether people in complex domains need to gather and integrate information in memory to make relevant, timely decisions, or whether one believes that simply looking up information as needed is sufficient for expert performance. In time-critical domains like driving, medicine, aviation, and ATC, it is clear that the former is the case.
Endsley (2015b) provides a detailed discussion on the role of memory in SA, showing that inexperienced operators’ SA is often constrained by working memory limitations, but that experts have access to long-term memory stores that largely circumvent these limits. The question for SA measurement is whether operators can report SA data during freezes, or whether their knowledge will rapidly decay over the period of the freeze. Working memory generally lasts about 20 to 30 s unless it is kept activated (Baddeley, 1986).
To investigate this question, Endsley (1990a) provided pilots with SAGAT queries in random order during freezes. This study showed that pilots were equally accurate in answering SAGAT queries across a 6-min period after a freeze, indicating no memory decay. Endsley (1990a, 2000) concluded that this result supports a model of cognition in which working memory is an activated subset of long-term memory (Cowan, 1988). This model is consistent with many others, including Durso and Gronlund (1999), Adams, Tenney, and Pew (1995), and Sarter and Woods (1991).
A number of studies were found in this review that refute the charge that SAGAT is overly dependent on working memory, shown in Appendix D, which is available with the manuscript on the Human Factors website. Endsley and Bolstad (1994) showed that primary working memory ability did not predict SAGAT scores in experienced pilots. Sulistyawati, Wickens, and Chui (2011) also found that experienced pilots answered SAGAT questions equally well, irrespective of their scores on memory span tests. Jipp and Ackerman (2016), similarly, found no relationship between primary working memory scores and SAGAT in an ATC task with student participants.
Gonzalez and Wimisberg (2007) showed that Level 1 SA, as measured by SAGAT, improved over time with experience on a task, and its relationship to a visual working-memory measure decreased accordingly. Gutzwiller and Clegg (2012) found no relationship between working memory and Level 1 SA in their study, based on an end-of-trial SAGAT measure; however, it was correlated with Level 3 SA which would be required for complex reasoning and projection per the Endsley (1995c) model of SA.
Two studies show somewhat different conclusions. Durso et al. (2006) reported that SAGAT scores correlated with measures of complex working memory (Operational Span and Reading Span) as well as fluid intelligence, but not measures of primary working memory or visual memory. It should be noted that they used inexperienced subjects from the general population performing in an ATC simulation. They also provided questions about the situation that were not necessarily reflective of the SA needed to perform the task, such as “at what level will the next airplane to have a heading change be?” This type of question requires extensive manipulation of information in working memory to answer, as compared to typical SAGAT queries that ask for information that should already be known to perform the task (e.g., what is the next sector of the indicated aircraft?).
Cak, Say, and Misirlisoy (2019) found that both SAGAT and SPAM were correlated with an Operational Span complex working memory test in a study with experienced pilots. (In addition, both measures were correlated with the number of hours the pilots had spent in the simulator and SPAM was also correlated with a measure of divided attention.) The fact that SPAM as well as SAGAT correlated with this complex working memory measure is puzzling in that all information is visually present when answering SPAM probes (although since SPAM includes queries that focus on the past, it introduces its own memory component).
Cak, Say, and Misirlisoy explained that their finding runs counter to previous studies due to the different measure of working memory employed in their study—the Operation Span test. This test requires people to actively maintain information in memory while performing secondary processing tasks and has been shown to be significantly correlated with the ability to control attention in numerous studies (Kane, Bleckley, Conway, & Engle, 2001; Poole & Kane, 2009; Unsworth & Engle, 2005; Unsworth & Spillers, 2010).
Primary memory (the ability to hold information in memory), attentional control, and the ability to recall information from secondary memory have been shown to be separate aspects of cognition that all contribute to Operation Span—a measure of complex working memory (Shipstead, Lindsey, Marshall, & Engle, 2014). As concluded by Shipstead et al (2014) these mechanisms are not similarly represented by all working memory tasks. Running memory span performance reflects primary memory more strongly than either complex span or visual arrays tasks. The performance of these later tasks is more closely associated with a persons’ attentional control and retrieval abilities. (p. 138)
Therefore, it appears that both Durso et al. (2006) and Cak et al. (2019) employed a complex span working memory measure that is more closely representative of factors such as attentional control and secondary memory aspects, whereas the other studies that examined the role of working memory on SA have operationalized it via tests which more closely reflect primary memory.
From the standpoint of SA, therefore, it is not too surprising that both SAGAT and SPAM performance correlated with the Operation Span test, reflecting the importance of attentional control on SA. What is surprising is that Cak et al. (2019) only found SPAM to be correlated with a measure of attention sharing ability, whereas Endsley and Bolstad (1994) found attention sharing to be highly predictive of SAGAT in pilots. From the standpoint of the present question, it does not appear that SAGAT is overly reliant on the need for people to hold information in primary (working) memory with experienced participants; however both SAGAT and SPAM do have correlations with attentional measures, reflecting its importance to SA.
Therefore, with experienced participants collection of SAGAT data is not working memory constrained when the information is collected immediately following a freeze, for at least 2 to 3 min afterwards, as they have access to relevant information in longer term memory representations. Novices and inexperienced participants are likely more reliant on working memory for processing information and typically have lower SAGAT scores than experienced participants (Hogan, Pace, Hapgood, & Boone, 2006; Jackson, Chapman, & Crundall, 2009; Kass, Cole, & Stanny, 2007; Randel, Pugh, & Reed, 1996; Soliman & Mathna, 2009; Strater, Endsley, Pleban, & Matthews, 2001). Even so, it is interesting to find that there were not significant differences in the sensitivity of SAGAT based on the experience level of the participants in the 152 studies reviewed here (p >.05). SAGAT was sensitive with both experienced and inexperienced participants. Overall these findings are consistent with recent work showing that with increasing levels of domain expertise people rely more on running memory, actively updating and processing new information as the task proceeds, as opposed to passive working memory that simply tries to memorize information without further processing (Anderson-Montoya, Scerbo, Ramirez, & Hubbard, 2017).
It should also be noted that a few studies using SAGAT applied it to very short, artificial tasks that may not be suitable. Gugerty (1998), for example, provided very short scenarios (18–35 s) and then assessed driver’s ability to recall the location of other vehicles. deWinter et al. (2018) provided participants with a set of six gages to monitor with instructions to report when a gage exceeded a limit. Even though SAGAT was sensitive in these studies, they are not realistic examples of real-world tasks. Such artificial tasks likely leave participants with little to do but attempt to memorize information.
In contrast, SAGAT was designed for realistic, ecologically valid simulations, and it is recommended that no SAGAT freezes occur for at least the first 3 min (and with at least 3 min in between freezes) in order that participants become engaged in performing their tasks (Endsley, 1995b, 2000). Participants are instructed to perform their tasks as usual and to simply answer the SAGAT questions to the best of their ability. Because SAGAT queries are representative across a wide sample of normal SA requirements, this precludes trying to memorize answers. In a meta-analysis of 65 studies, Vidulich (2000) found that probe measures are most sensitive when a wide range of queries are provided, rather than just a few. This is likely because it precludes trying to memorize answers in advance of a freeze. The unpredictability of when the freezes will occur also avoids memorization and effects on performance (Endsley, 1995b, 2000).
Speed-Accuracy Tradeoffs
A number of studies found that answers to SPAM probes were often more accurate than similarly worded SAGAT queries (Durso et al., 2006; Pierce, Strybel, & Vu, 2008; Strybel et al., 2008), which is not surprising considering that the correct answer is clearly visible when responding to a SPAM probe.
While SPAM reaction time theoretically measures how cognitively available information is, a number of studies reviewed show that it may instead be measuring look up time. SPAM and SAGAT differ in two main respects: SAGAT pauses the simulation and hides displays during query administration, while SPAM leaves them visible and does not pause the simulation. C. A. Morgan, Chiappe, Kraut, Strybel, and Vu (2012) performed a study comparing the same questions provided when the display was visible versus not visible and when the simulation was paused or not paused. They found that SA queries were more accurate when the displays remained visible, as would be expected, but that response time was almost 2 s slower, revealing a significant speed-accuracy tradeoff. Several other studies have also found speed-accuracy tradeoffs for SPAM and real-time probe measures (Alexander & Wickens, 2004, 2005; Jones & Endsley, 2004; Taber, McCabe, Klein, & Pelot, 2013). In general, it appears that people will look to check before answering queries if the information is available, showing that SPAM more accurately assesses the time for people to look-up information on a display, regardless of whether they know the answer.
Sampling Bias
An additional concern is whether the methods contain any sampling bias. Because SAGAT recommends conducting the freezes at random times, the likelihood of sampling bias (e.g., by collecting during only high workload or low periods, or just during interesting events) is minimized. While Lau, Jamieson, and Skraaning (2016) advocate for collecting SA at only important times as indicated by a subject matter expert, this can cue subjects and provide artificial information affecting their performance (Endsley, 1995b, 2000).
SPAM provides probes at random intervals as well; however, participants are able to delay answering, or may even neglect to answer. Durso et al. (1998) noted that participants could take as long as 10 s to answer the call line to get a SPAM probe in their study, and as long as 4 s to answer a query. Loft et al. (2016) found their participants took as long as 20 s to accept queries, taking longer when SA was low or uncertainty was high. Trapsilawati, Wickens, Qu, and Chen (2016) reported means latencies in accepting a SPAM probe of between 9 and 23 s. Strybel, Vu, Battiste, and Johnson (2013) reported that 5% of probes went unanswered, and Cunningham et al. (2015) found that pilots did not respond the ready prompt 16% of the time in their study due to workload. Further two additional studies showed that SPAM latency measures are correlated with NASA TLX scores, indicating that it is associated with workload (Strybel et al., 2010; Strybel et al., 2008).
These results indicate that participants may delay attending to SPAM probes until lower workload periods. While the time to accept a probe is a secondary workload indicator, this ability to delay answering probes, or to ignore them completely, creates a bias in the measure in favor of lower workload periods. While Bacon and Strybel (2013), Silva et al. (2013), and Keeler et al., (2015) reported that instructing participants to only answer SPAM queries if they felt able to and providing computerized SPAM queries rather than verbal queries could reduce SPAM intrusiveness, this method also systematically skews SA probes into low-workload periods.
SA vs. Workload
Another theoretical concern is the extent to which SPAM is measuring SA or workload. While high workload can lead to low SA, it has also been shown that SA and workload can be independent across much of the demand scale, with low SA possible in low-workload periods (vigilance) and both high and low SA possible across moderate workload levels (Endsley, 1993a; Vidulich & Tsang, 2015). It is therefore important that SA measures be clearly distinguishable from workload measures.
The fact that operators must answer real-time probes while conducting operational tasks creates a situation in which it serves as a secondary workload measure. Jones and Endsley (2004) found a weak correlation between real-time probes and a concurrent measure of workload. This issue is somewhat reduced with the introduction of a ready prompt with the SPAM technique, which allows participants to answer the probes while they have more spare capacity, but it may not be completely eliminated.
Pierce, Strybel, and Vu (2008), for example, found that SPAM latency measures were not predictive when participants were under high workload because they neglected the SPAM probes while meeting task demands. Pierce, Strybel, and Vu (2008) also found correlations between workload subscale ratings on trials that used SPAM and those that involved other types of secondary tasks. Given that Pierce (2012) found a negative effect of SPAM on workload, and that 40% of the studies reported on here that used SPAM found negative effects on performance, it seems likely that SPAM is providing secondary task loading. Coupled with the finding that people frequently look up the answer to SPAM probes, which is a secondary task, there is considerable evidence that SPAM likely measures workload to some degree, rather than SA.
Query Design
Whichever measure is used, the actual queries provided are critical to obtaining valid research results. Variability between the studies reviewed in this study may be due to (a) whether they employed a systematic analysis of SA requirements for the domain under study, (b) construction of clear queries that can be consistently scored, (c) differences between SAGAT and SPAM in the types of information solicited, and (d) methods for soliciting participant responses.
Reflection of SA requirements
One observed issue with SPAM, and with some researchers use of SAGAT, has been the development of queries that may in fact have little to do with the operator’s SA requirements. For example, Durso et al. (1998) posed “Which has the lower altitude, TWA799 or AAL957?” While that probe undoubtedly requires knowledge of situational information to answer (either from memory or from displays), it is not something that air traffic controllers typically need to know to do their jobs (Endsley & Rodgers, 1994b). Similarly, Chiappe et al. (2015) report providing queries such as “are the majority of the aircraft in your sector westbound?” Someone looking at a display could answer that question, but it is quite extraneous to what is required for good SA in ATC (Endsley & Jones, 1995; Endsley & Rodgers, 1994a). Asking questions about the situation is not the same as asking questions that are pertinent to SA. Whichever technique is used, a detailed analysis of SA requirements for the operational role being examined is extremely important (Endsley, 1995b). While many of the studies reviewed reported conducting a goal-directed task analysis (GDTA) or other analysis of the task domain to determine their questions, others did not.
Query construction
Second, crafting clear, concise questions that can be objectively scored as correct or incorrect is important. The questions should not be subjective or ambiguous, and should produce the same answer from experts with full knowledge of the situation. For example, Lau, Jamieson, and Skraaning (2014) had trouble with ill-defined queries that could not be consistently scored by three experts who were actually looking at the same displays, even with all the data visible and all the time needed, resulting in low inter-rater reliability in evaluating their queries. While typically the query scoring key is determined via data in the simulation computer, this study points to the importance of creating clear and relevant queries.
Query focus
A notable difference between SPAM and SAGAT is that SPAM creates questions that test a participant’s knowledge of the past, present, and the future, while SAGAT questions are directed at the perception, comprehension, and projection components of SA. SPAM-future and SAGAT-projection queries seem clearly aligned. SPAM-present and SAGAT-perception also seem to be aligned, although SPAM-present may also potentially include SAGAT-comprehension questions.
SPAM includes questions on the past under the belief that remembering past events is important for projecting future events (Durso et al., 2006). This, however, adds a clear memory component to SPAM that may or may not be relevant to current operations. SA is generally considered to be a person’s knowledge and understanding of the present, and its reliance on what has occurred in the past is only relevant to the extent that it affects the present or future, for example, understanding the trend in some variable (Endsley, 1995c). SAGAT may directly ask whether some relevant variable (e.g., altitude or speed) is increasing, decreasing, or staying level, for example, but would not ask what it was at some prior point in time. In other words, while SPAM asks specifically for memory of past events or status, SAGAT asks for how information is changing or how it will change in the future, not specifically requiring reporting on memory of the past. Thus, SPAM may be more reflective of memory than SAGAT in this respect.
Query presentation format
A final issue has to do with the referents for queries that are provided. It is recommended that SAGAT queries start with a domain relevant map that allows participants to input the relevant objects (e.g., aircraft, automobiles, friendly, and enemy forces) that they know about in an intuitive spatial layout (Endsley, 1995b, 2000). Subsequent queries refer back to the objects entered on that map. This is because identifiers (e.g., aircraft call signs, track numbers, train numbers) are often not known, but rather are peripheral information that serve a functional task, but have little value otherwise in terms of situation understanding (Endsley, 2015a; Endsley & Rodgers, 1998). Implementations of SAGAT measures that rely heavily on such identifiers to answer queries create an artificial memory requirement that is likely to lead to difficulties such as experienced by Golightly, Wilson, Lowe, and Sharples (2010) in train operations, or Loft et al. (2018) in submarine management. Also Ratwani and Trafton (2008) found that spatial memory (such as that provided with maps in SAGAT) is important for resumption from an interruption.
Conversely, SPAM is heavily reliant on identifiers, but this does not cause a problem because they are readily available on displays. And Strybel et al. (2016) found that participants performed more poorly on SPAM with graphically presented probes as compared to verbal probes, most likely due to physical interference with the ongoing task, or potentially because they were presented in a way that made guessing less easy. In contrast, Shelton et al. (2013) created a text based display for SPAM probes on a small hand-held computer because verbally presented queries interfered with normal discussions in a study with health care teams.
Team SA
In addition to assessing individual SA, many researchers are also interested in the SA of teams. A total of 27 studies were found that used either SAGAT or SPAM to assess SA at the team level, shown in Appendix F, which is available with the manuscript on the Human Factors website. Because SPAM does not elicit data from multiple operators at the same time, and queries and responses are generally provided verbally, an assessment of shared or team SA is logistically challenging with the technique, and only two papers were found that used SPAM for assessing team SA.
A key advantage of SAGAT is that it allows for the assessment of team SA (Bolstad & Endsley, 2003) because it is administered to all participants in a simulation at the same time. Problems with tools and bottlenecks in team processes where information is not passed, or where different interpretations are made from data that could lead to poor team performance can be readily identified by comparing the SA of different team members or sub-teams at the same point in time. A total of 24 studies were found that used SAGAT to evaluate team SA, 1 used SACRI, and 1 used QUASA.
These measures were used to assess team SA in several ways including
Creating a combined or average SA score across the team as an overall Team SA score (11 studies);
Allowing for a collaborative team response to queries (two studies);
Assessing Shared SA (or SA similarity) as the degree of concurrence of teammates on information elements that are relevant to both roles (14 studies);
Examining the degree to which team members are aware of the SA of each other (Team Meta-SA) (three studies);
Assessing the correlation between the SA of different team members, or sub-teams (six studies).
The SA of teams was assessed in a variety of domains including military (10), health care (5), aviation (4), emergency management (4), process control (3), and business (1) settings. While most studies were conducted in simulators or microworlds, two studies were conducted in live exercises.
Although the exact methods for assessing SA in teams varied among these studies, a few general statements can be made. Shared SA as measured by SAGAT was shown to be predictive of overall team performance (Bonney, Davis-Sramek, & Cadotte, 2016; Coolen, Draaisma, & Loeffen, 2019; Rosenman et al., 2018), as well as when it was measured by SPAM (Cooke, Kiekel, & Helm, 2001). Combining SA scores into an overall Team SA score generated mixed results, with five studies showing it to be predictive of overall team performance (Cooke et al., 2001; Crozier et al., 2015; Gardner, Kosemund, & Martinez, 2017; Parush et al., 2017; Prince, Ellis, Brannick, & Salas, 2007), and three that did not (Brooks, Switzer, & Gugerty, 2003; P. Morgan et al., 2015; Sorensen, Stanton, & Banks, 2010). This body of research also demonstrated the effect of team membership on the SA of individuals (Bolstad & Endsley, 2003; Cuevas & Bolstad, 2010; Leggatt, 2004; Sætrevik, 2012; Seet, Teh, Soo, & Teo, 2004), as well as the significant effect of the team leader (Cuevas & Bolstad, 2010), although one study did not find a significant effect of team membership on Shared SA (Sætrevik, 2012).
In addition, measures of combined Team SA and Shared SA were shown to be sensitive to other aspects of teams, including team knowledge (Bolstad, Cuevas, Gonzalez, & Schneider, 2005; Cooke, Stout, Rivera, & Salas, 1998), team cohesion and coordination (Hallbert, 1997), team experience (Crozier et al., 2015), organizational hub distance (Bolstad et al., 2005; Saner, Bolstad, Gonzalez, & Cuevas, 2009), team displays (Javed, Norris, & Johnston, 2012), information flows (Artman, 1999), and stress (Price & LaFiandra, 2017).
Only a few studies have examined team meta-awareness, with Sulistyawati, Chui, and Wickens (2008) and Sulistyawati et al. (2009) showing that fighter pilots’ SA was moderately correlated with their teammates, but awareness of the teammates’ SA was not predictive of performance. Yuan, She, Li, Zhang, and Wu (2016) found that awareness of teammate SA was negatively correlated with own SA, however, likely due to competing task demands.
While the SA of teams has often been investigated by examining team communications or coordination metrics, or by an examination of the types of processes they engage in, this work shows that objective measurement of the overall Team SA or Shared SA that people have when working in teams can provide useful complementary information to this body of research.
Conclusions
In conclusion, this analysis examined 243 papers that measured SA using SAGAT, SPAM, or one of their variants, and provided a useful examination of concerns regarding these techniques. While both SAGAT and SPAM were found to be equally predictive of performance, this analysis shows that SPAM has significantly lower sensitivity than SAGAT (64% compared to 94%) and is far more intrusive, with 40% of studies showing an effect of SPAM on task workload or performance. In addition, issues with speed-accuracy trade-offs, correlations with workload, and a sampling bias in favor of lower workload periods create considerable concern over SPAM’s validity as a measure of SA. Evidence of any benefit of “situatedness” during the SA probe is lacking.
When administered during simulation freezes and with adequate sample size, SAGAT was shown to have a very high level of sensitivity to study manipulations at 94%. Sensitivity was shown to be much higher when the analysis was based on accuracy of individual queries or SA levels, rather than an overall combined SA score. Studies that collected only a few samples or that collected data only at the end of the trial had lower sensitivity.
This analysis shows that claims that SAGAT is intrusive (Chiappe et al., 2012; de Winter et al., 2019; Durso et al., 1998; Salmon et al., 2011; Sarter & Woods, 1991) are unwarranted, with no negative effects on performance found in the 11 studies that examined it. The conditions of SAGAT administration are significantly different than interruptions with unrelated tasks. This analysis also lays to rest many other potential criticisms of SAGAT. Concerns about the over-reliance of SAGAT on working memory (de Winter et al., 2019; Gugerty, 1998; Jeannot, Kelly, & Thompson, 2003; Langan-Fox, Sankey, & Canty, 2009; Stanton, Salmon, Walker, & Jenkins, 2010) were found to be unsupported.
Most SA measurement examined in the present analysis was conducted in medium to high fidelity simulations, or in computer microworlds, both to good effect. While the ability to administer SA queries in real-world situations is often cited as a potential benefit of SPAM as compared to SAGAT, it was interesting to find that SPAM was only used in two live exercises, while SAGAT was used in nine. All of these were controlled studies, rather than actual real world activities.
SAGAT also was found to provide valuable insights into the nature of Team SA and Shared SA in a number of studies. The ability to determine the outcome of team communications, coordination processes, or compositions provides a useful addition to studies in this area. This is in contrast to claims that SAGAT is not suited to study of team SA (Langan-Fox et al., 2009; Stanton et al., 2017).
A few other spurious criticisms of SAGAT have been made that also deserve comment. Claims that SAGAT does not capture situation dynamics (Langan-Fox et al., 2009; Salmon et al., 2008) are perplexing in that many queries do indeed cover issues of how key information is changing (e.g., whether aircraft are ascending or descending in Table 1), as well as how information is projected to change over time as represented in Level 3 SA questions.
Additional concerns have been raised over whether SA queries, which tap into operator’s explicit awareness, are in some way unreflective of implicit awareness (de Winter et al., 2019; Gugerty, 1997; Lo, Sehic, Brookhuis, & Meijer, 2016). Endsley (1995c) describes ways in which automatized performance may be accomplished without much conscious awareness (e.g., a driver following a well known route). While this allows for effective performance in some circumstances, the resultant lower conscious awareness can be a problem if something novel occurs (e.g., the driver would miss a new stop sign, or a police officer looking for speeders). I argue that this represents one of the ways in which SA (as an explicit measure of situation knowledge) can sometimes diverge from performance outcomes. Despite this possibility, SAGAT was shown to be highly predictive of performance.
Other domain specific criticisms of SAGAT are also largely unsupported. Langan-Fox et al. (2009) and Jeannot et al. (2003), for example, claim that SAGAT for ATC (a) considers all aircraft equal, which is not the case (for example, see Endsley, Mogford, & Stein, 1997; Endsley, Sollenberger, & Stein, 1999), (b) requires expensive simulators (see also Salmon, Stanton, Walker, & Green, 2006), even though many studies use inexpensive microworlds or computer games, and (c) is not suited to multi-sector studies, which is not accurate as many studies involve multiple sectors (e.g., Endsley & Rodgers, 1998). Overall SAGAT had 94% sensitivity in ATC studies.
SAGAT was successfully used in 20 military studies, including 5 in distributed team command and control, despite claims that it is unsuitable for this domain (Salmon et al., 2006; Stanton et al., 2010). Similarly, Golightly et al. (2010) concluded SAGAT was not suited for SA measurement in the train driving domain, based on a short study in which they only collected a small amount of data at the end of the trial, significantly reducing its sensitivity. Rose et al. (2019) found it sensitive in train driving, however.
As a caveat, the present meta-analysis is constrained by the available data and statistics reported on in the various studies reviewed. Some have criticized the use of hypothesis testing reliant on p-values on the basis that such statistics are often misinterpreted, do not indicate the magnitude of effects, and that statistical assumptions are rarely tested, instead favoring the reporting of confidence intervals and effects sizes (Cumming, 2014; Kline, 2004). However, others have pointed out that these statistical approaches are reliant on the same statistics as p-values and that many concerns are over-generalized (Kennedy, 2015; van der Linden & Chryst, 2015). In the present research, a consistent p-value of .05 was used in comparing studies, as well as a wider p-value of .10 to consider trend data, although very few studies were added with this more lenient criterion.
In addition, variance in research results may have been introduced by the manner in which different researchers have employed the SA methods, which was specifically examined in the analyses to the extent possible and reported on. Finally, variance in research results may be due to factors associated with study designs. The analysis considered the effects of domain, experimental condition, participant type, and number of subjects, which were largely insignificant; however, there could be other unknown sources of variances between the studies.
In conclusion, many of the methodological concerns regarding objective measurement of SA via SAGAT can be laid to rest based on this meta-analysis of research on SA over the past 30 years. When implemented as designed, SAGAT can provide valuable insights into the effects of display designs, training implementations, automation effects, individual differences, and expertise, and into the construct of SA in both individuals and teams.
Key Points
A meta-analysis of studies employing direct, objective measures of SA found both SAGAT and SPAM to be predictive of performance.
SPAM was found to be less sensitive than SAGAT, intrusive on primary task performance and to suffer from speed-accuracy trade-offs, correlations with workload, and a sampling bias in favor of lower workload periods, which make it problematic for the measurement of SA.
SAGAT was found to have a high level of sensitivity (94%) when employed using the freeze technique and analyzed by query or SA level, not to suffer from primary memory constraints, and not to be intrusive on task performance across a large number of studies. It provides useful insights into both individual and team SA in a wide variety of domains and experimental settings.
Supplemental Material
Appendices – Supplemental material for A Systematic Review and Meta-Analysis of Direct Objective Measures of Situation Awareness: A Comparison of SAGAT and SPAM
Supplemental material, Appendices_Meta-analysis_of_SA_measures_Final._docx_1 for A Systematic Review and Meta-Analysis of Direct Objective Measures of Situation Awareness: A Comparison of SAGAT and SPAM by Mica R. Endsley in Human Factors: The Journal of Human Factors and Ergonomics Society
Footnotes
Supplemental Material
The online supplemental material is available with the manuscript on the Human Factors website.
Mica Endsley is President of SA Technologies and is the former Chief Scientist of the U.S. Air Force. She received a PhD in Industrial and Systems Engineering from the University of Southern California. She has published extensively on situation awareness, automation, and system design.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
