Abstract
Objective
An experiment used workload capacity analysis to quantify automation usage strategy across different task difficulty and display format types in a speeded task.
Background
Workload capacity measures the efficiency of concurrent information processing and can serve as a gauge of automation usage strategy in speeded decision tasks. The present study used workload capacity analysis to investigate automation usage strategy while information display format and task difficulty were manipulated.
Method
Subjects performed a speeded judgment task assisted by an automated aid that issued decision cues at varying onset times. Response time distributions were converted to measures of workload capacity.
Results
Two variants of a workload capacity measure, CzOR and CzAND, gave evidence that operators moderated their own decision times both in anticipation of and following the arrival of the aid’s diagnosis under difficult task conditions regardless of display format.
Conclusion
Assistance from an automated decision aid may cause operators to delay their own responses in a speeded decision task, producing joint response time distributions that are slower than optimal.
Application
Even when it renders its own judgments quickly and with high accuracy, an automated decision aid may slow responses from a user. Automation designers should consider the relative costs and benefits of response accuracy and time when choosing whether and how to implement an automated decision aid.
Automated systems assist human operators with various forms of information processing, including data acquisition, decision making, and response selection (Parasuraman, Sheridan, & Wickens, 2000). Human-automation interaction (HAI) occurs in domains including aviation (e.g., Billings, 1997) and air traffic control (e.g., Durso & Sethumadhavan, 2008), process control (e.g., Moray, Inagaki, & Itoh, 2000), and surface transportation (Lees & Lee, 2007). Designed and implemented well, automation can improve the human operators’ decision making (Sarter & Schroeder, 2001), reduce their response time (RT) (Rovira, McGarry, & Parasuraman, 2007), and decrease their cognitive load (Weiner, 1988).
Unfortunately, operators’ strategies for working with automated aids often fail to optimize performance (Parasuraman & Riley, 1997). Robinson and Sorkin (1985), for example, compared human-automation performance in a statistical decision task to the predictions of a normative model. The human-automation teams outperformed unaided human subjects but fell well short of ideal performance. Numerous other experiments have produced similar results (e.g., Bartlett & McCarley, 2017; Corcoran, Dennett, & Carpenter, 1970; Meyer, Wiczorek, & Günzler, 2013; for a review, see Wickens & Dixon, 2007).
Unsurprisingly, an operator’s willingness to depend on an automated aid varies with the aid’s reliability (Wickens & Dixon, 2007). Less obviously, automation usage may also depend on the perceptual characteristics of the aid’s interface. Object-based theories of attention (Duncan, 1984) propose that attention selects objects or perceptual groups as wholes, allowing multiple properties of an attended item to be processed in parallel. These theories have found application to display design through Wickens and Carswell’s (1995) proximity compatibility principle (PCP). The PCP advises that the perceptual proximity of information channels within a display comport with their processing proximity. Perceptual proximity refers to the strength of perceived grouping (Wertheimer, 1923) between display elements, and processing proximity refers to the degree that different sources of information must be integrated for an operator to perform a task. According to the PCP, high perceptual proximity between display channels facilitates information integration, and low perceptual proximity encourages selective attention to individual channels, as is useful when task demands do not require mental integration.
The PCP implies that operators may depend more on an aid’s guidance when a display integrates the aid’s cues with raw task data than when it presents raw data and cues in isolation. Meyer (2001) tested this idea by examining the effects of cue format on operators’ responses in an automation-assisted signal detection task. Subjects judged the height of a rectangle each trial, deciding whether it was sampled from a distribution of short or long items. In some blocks, subjects were assisted by an automated aid. Cues from the aid were presented in either of two forms. In the integrated cue condition, the stimulus rectangle appeared in green to signal a judgment of “short” from the aid and in red to signal a judgment of “long” from the aid. In the separated cue condition, the aid’s judgments were presented as red or green bars above the stimulus rectangle. When the aid’s judgment was incorrect, integrated cues produced less accurate human judgments than separated cues. Meyer concluded that operators had difficulty ignoring a cue when it was embedded within the raw data themselves.
But in addition to changing response accuracy, the format in which an aid’s cues are displayed might also influence operators’ response speed. By encouraging parallel processing, a display that integrates raw data and automation cues might support faster judgments than a display that separates them. Measuring the effect of automation on an operator’s RT, however, is less straightforward than it may seem. In some circumstances, support from an aid slows operators’ responses to a time-sensitive signal without evident improvements in other aspects of performance (Abe & Richardson, 2006). This is obviously a performance loss. In more auspicious cases, the aid allows operators to respond more quickly to critical signals than they do unaided while maintaining or improving response accuracy (e.g., Rovira et al., 2007; Yeh, Wickens, & Seagull, 1999). This pattern is clearly a performance gain.
But by itself, a change in mean RT provides only a rough gauge of the human-automation team’s temporal efficiency. The team may produce faster responses on average than either the operator or aid alone but may not match the speed that is achievable based on the agents’ individual RT distributions (Miller, 1982; Raab, 1962). RT gains resulting from the assistance of an automated aid may therefore mask performance inefficiencies. To circumvent this limitation, Yamani and McCarley (2016) recommended the analysis of workload capacity to gauge the temporal efficiency of human-automation teams. A conceptual element of Townsend and Nozawa’s (1995) systems factorial technology (SFT), workload capacity is the efficiency with which multiple channels operate together, benchmarked to their operating rates in isolation. Workload capacity analysis is most commonly used to study information processing within the black-box cognitive system of an individual decision maker, measuring changes in perceptual or cognitive processing rates that occur as the number of redundant target signals to be processed in a task increases (e.g., Hawkins, Houpt, Eidels, & Townsend, 2016; Yamani, McCarley, & Kramer, 2015).
An operator and aid working concurrently on a task that either could in principle complete by itself, however, constitute a system of parallel redundant processing channels, and system efficiency can be measured in the same way as that for a black-box system. In the simplest case, each channel will process information at the same rate under parallel operating conditions as it does in isolation. In other words, the presence of one channel will have no effect on the other channel’s processing speed. The system in this case is said to operate with unlimited capacity. When parallel redundant channels are stochastically independent and capacity is unlimited, the processing time needed for the system to reach a judgment can be predicted from the individual channels’ RT distributions (Townsend & Ashby, 1983). Unlimited capacity independent parallel (UCIP) model performance therefore provides a useful benchmark of system performance. Performance is said to be limited capacity when the system responds more slowly than expected from the UCIP model and supercapacity when the system responds faster than expected from the UCIP model (Townsend & Nozawa, 1995). In the context of HAI, analysis of capacity measures the response speed of an aided human operator relative to the operator’s speed in isolation (Yamani & McCarley, 2016). Holding the processing rate of the automated aid constant, limited capacity would indicate a tendency for the operator to process information more slowly when aided by the automation than when unaided. Supercapacity would indicate a tendency for the operator to process information faster when aided.
Note that the concept of workload within SFT differs from the conventional idea of workload in human factors. In the analysis of black-box cognitive systems, SFT equates workload with processing load, the number of targets to be processed by an observer. In human factors, contrastingly, workload generally refers to the level of effort or resources that an operator expends to perform a task (Gopher & Donchin, 1986), which may not increase with processing load. The analogy of redundant human and automated agents to redundant channels in a black-box system is also imperfect. SFT’s notion of workload as processing load presumes that redundant targets may inhibit (Eidels, Houpt, Altieri, Pei, & Townsend, 2011) or compete for resources (Wickens & McCarley, 2007) with one another. This is unlikely to occur between an operator and automated aid, which are separate agents that most often do not directly exchange suppressive signals or draw on a common, limited resource pool. Nonetheless, capacity analysis reveals changes in channel processing rates within a human-automation system just as it does within a black-box cognitive system. A measure of capacity is thus a description of system efficiency, not a mechanistic explanation of system performance.
A measure of capacity can be calculated as a function of time (Townsend & Nozawa, 1995) where useful but for many purposes can be collapsed over time into a summary statistic, Cz (Houpt & Townsend, 2012). Cz scores follow a standard normal distribution, where 0 denotes unlimited capacity, negative values denote limited capacity, and positive values denote supercapacity. Two variants of Cz can be calculated, corresponding to two stopping rules (van Zandt & Townsend, 1993). Under an OR stopping rule, the RT of the human-automation team is determined by the first of the two channels, human operator or aid, to respond to an event. The corresponding capacity score, CzOR, therefore reflects the operator’s behavior before the arrival of the cue: Limited capacity indicates a slowdown of operators’ responses in anticipation of a cue from the automation, whereas supercapacity indicates a speedup. Under an AND stopping rule, team-level RT is determined by the second of the two agents, human or aid, to respond to an event. The capacity score CzAND thus represents operators’ behavior after the arrival of the cue: Limited capacity indicates a slowdown following a cue from the aid, as compared to unaided performance, while supercapacity processing indicates a speedup. By either measure, unlimited capacity indicates that the human operator’s response speed is the same under assisted and unassisted conditions.
In the current study, we employed workload capacity analysis to measure automation use strategy in a speeded judgment task, with raw data and cues presented either in separated or integrated format. Subjects performed a task similar to that used by Meyer (2001), judging whether a stimulus in each trial was sampled from a long or short distribution. Here, however, subjects were asked to render judgments in a speeded manner. They were assisted by an aid whose response appeared either above the stimulus (the separated condition) or as a property of the stimulus itself (the integrated condition). Cues from the aid arrived with stochastic onset times following the stimulus, allowing us to examine temporal properties of the human subjects’ automation interactions. Note that separate processing times for parallel redundant channels are typically unobservable in a black-box cognitive system (e.g., Townsend & Nozawa, 1995). Because cue onset times here were generated randomly, however, the current approach provided an empirical RT from the human subject and a randomly sampled cue-onset time for every trial. This allowed us to calculate both variants of Cz for human-automation teams, using the faster response each trial as the RT generated by the OR rule and using the slower response as the RT generated by the AND rule.
We hypothesized that integrated displays, which allow the operator to efficiently process raw data and the automated aid’s cues in parallel, would encourage greater deference to the aid, as evidenced by lower CzOR scores and higher CzAND scores, than separated displays. Additionally, we manipulated task difficulty, by sampling stimuli from distributions with either small or large variability, to examine generalizability of the observed effects.
Method
Subjects
Sixty subjects (42 female, 18 male; mean age = 20.5 years, SD = 5.1; range, 18–47 years) were recruited from the community of Old Dominion University, Virginia. All reported normal or corrected-to-normal visual acuity and normal color vision. Subjects were remunerated with class credit for participation.
Apparatus
Stimuli were presented on a Samsung T24C550 23.6” LED monitor with 1920 × 1080 resolution. The experiment was controlled by a Dell Optiplex 9020 running E-prime 2.0 (Psychology Software Tools, Inc., Pittsburgh) on Windows 7. Subjects viewed the monitor at the distance of approximately 57 cm without a chin rest. Experiments were conducted in a quiet room with dim light.
Procedure
Subjects performed a task modeled after that used by Meyer (2001). Instructions asked subjects to imagine themselves as factory production managers and make speeded judgments of whether or not to attempt production based on the measurement of raw materials, represented by the height of a stimulus rectangle each trial. The height of the rectangle was sampled from a pseudorandom Gaussian distribution of small or large rectangles (mean = 2.2° and 2.8°, respectively), and the difficulty of the task was manipulated through changes to the variance of the distributions (SD = 0.15° and 0.3° for easy and difficult conditions, respectively). Subjects were informed that shorter rectangles were to be approved for use in production and longer rectangles were to be rejected. Subjects were not informed of the exact target lengths of rectangles to discourage them from explicitly and visually marking the cutoff point.
Each trial began with a 500-millisecond blank screen, followed by the stimulus display. Subjects reported their decision each trial by clicking either the left (approve production) or the right (stop production) mouse button as quickly as possible. The stimulus display remained visible until the subject responded or a timeout period of 1 second elapsed. A failure to respond was treated as an error. A 750-millisecond feedback display followed each trial, reading + for a correct response and x for an error. The next trial began automatically.
Subjects performed the task with or without the assistance of the aid, alternating across blocks. Figure 1 illustrates the trial sequence for the two cue formats. On aided trials, the aid provided a cue either as the color of the stimulus rectangle itself (integrated format) or the color of a separate indicator (.5° × 2.5°) located 5° from the center of the display, above the stimulus rectangle (separated format). A green cue recommended a judgment to approve production, and a red cue recommended a judgment to stop production. In the unaided conditions, the cues were not presented, and the separate indicator remained gray. For half of the subjects, the aid’s cues and the raw stimulus data were presented in the integrated format. For the other half, the cues and data were presented in the separated format.

A time course of events in a typical trial for the integrated and separated display format conditions.
The stimulus onset asynchrony (SOA) of the cue relative to the stimulus rectangle was sampled from a pseudorandom exponential distribution with a mean of 676 milliseconds each trial. The mean SOA was chosen to match the mean RT of human subjects who performed the task unaided in a pilot experiment. The exponential distribution was chosen because it is memoryless (Townsend & Ashby, 1983), giving subjects no incentive to either speed or delay their responses in anticipation that the cue would arrive at a particular delay. The aid’s reliability was set at 95% (d’ = 3.28), a value chosen to encourage use of the aid while simulating imperfect automation.
Subjects were randomly assigned in equal numbers to one of four groups in a 2 × 2 design with Format (integrated vs. separated) and Difficulty (easy vs. difficult) as between-subject factors. Post hoc analysis confirmed that the groups were similar in demographic characteristics. Each subject performed two blocks of 20 practice trials each, followed by four blocks of 50 experimental trials each. Practice and experimental blocks alternated between aided and unaided conditions, with order counterbalanced across subjects. Subjects were allowed a break between blocks. The experiment lasted approximately 40 minutes.
This research complied with the American Psychological Association Code of Ethics and was approved by the Institutional Review Board at Old Dominion University. Informed consent was obtained from each subject.
Data Analysis
Mean RTs and sensitivity (d’) were calculated separately for the aided and unaided trials. RTs for incorrect trials were excluded from analysis.
Capacity analysis requires responses from the two channels (i.e., the human and aid) every trial to compute capacity coefficients (see Yamani & McCarley, 2016) and assumes that the RTs are stochastic. A cue onset time for the automated aid was therefore sampled randomly and recorded every trial. The onset time determined the SOA for the cue presentation in the aided trials but had no influence in the unaided trials. Subjects’ RTs and the randomly sampled cue onset times from the unaided trials served as the single-channel data. RTs and cue onset times from the aided trials served as redundant-channel data. CzOR was calculated by treating the faster of the two responses, human or aid, as the system response on aided trials, and CzAND was calculated by treating the slower of the two responses as the system response. Capacity analyses were conducted using the SFT package in R (Houpt, Blaha, Mclntire, Havig, & Townsend, 2014).
Default Bayesian tests (Rouder & Morey, 2012; Rouder, Morey, Speckman, & Province, 2012) were employed instead of null-hypothesis significance tests (NHSTs). In this analysis, Bayes factors are the measure of evidence for an effect of interest. Briefly, a Bayes factor is the likelihood ratio with which the data favor one hypothesis relative to the other. One advantage of Bayesian analysis is that it allows evidence in favor of the null hypothesis, while NHST does not. Another is that Bayes factors provide a more meaningful measure of evidence strength. For example, a Bayes factor of 100 in favor of an alternative model relative to a null model indicates that the data are 100 times more likely to have come from the alternative model than the null model. Bayes factors reported in the following are ratios of likelihood of the data favoring a model including an effect of interest to likelihood excluding the effect. Values are denoted as B10 (Rouder & Morey, 2012), and terminology for describing evidence (anecdotal, substantive, strong, very strong, decisive) is drawn from Jeffreys (1961; see also Wetzels et al., 2011).
Results
Mean RTs
RTs were submitted to a 2 × 2 × 2 mixed-effects Bayesian analysis with Block (aided vs. unaided) as a within-subject factor and Difficulty (difficult vs. easy) and Format (integrated vs. separated) as between-subject factors.
Figure 2 presents mean RTs as a function of Format and Automation, plotted separately for the two task difficulty levels. RTs were decisively longer under difficult task conditions, F(1, 56) = 24.31, η2 G = .28, B10 = 1.3 × 103, and contrary to expectations, were longer in the automation-aided blocks of trials than in unaided blocks, F(1, 56) = 9.91, η2 G = .01, B10 = 10.08. However, these main effects were qualified by a two-way interaction, F(1, 56) = 19.72, η2 G = .02, B10 = 330.77, indicating very strong evidence that effects of the aid in the difficult condition (M = 807 vs. 922 ms for the unaided and the aided conditions, respectively), paired-sample t(29) = 4.07, B10 = 90.46, than in the easy condition (M = 619 vs. 600 ms), paired-sample t(29) = 1.53, B10 = 1/1.79.

Mean response times (RTs) as a function of Format and Aid conditions separately for the easy condition (top panel) and the difficult condition (bottom panel). Error bars represent 95% within-subject confidence intervals (Cousineau, 2005; Morey, 2008).
Subjects in the separated displays condition produced substantially longer RTs than those using integrated displays (M = 661 vs. 813 ms), F(1, 56) = 8.64, η2 G = .12, B10 = 6.90. Surprisingly, though, the data gave substantial evidence that the effect of display format was comparable between the aided and unaided conditions, F(1, 56) = 1.19, η2 G = .001, B10 = 1/3.75. This suggests that the effect of format on RT did not reflect the relative difficulty of processing integrated and separated cues from the aid but was spurious. The data gave no substantial evidence for a Format by Task Difficulty interaction, F(1, 56) = 3.54, η2 G = .001, B10 = 1.16, or for a three-way interaction, F(1, 56) = 3.29, η2 G = .004, B10 = 1/1.23.
Sensitivity
Values of d’ were entered to an analysis identical to that used for mean RTs. Figure 3 presents mean d’ as a function of Format and Automation, plotted separately for the two Task Difficulty conditions. Sensitivity was decisively lower in the difficult condition than the easy condition (M = 1.78 vs. 1.27), F(1, 56) = 46.66, η2 G = .36, B10 = 1.8 × 106, as expected. Assistance from the aid decisively improved sensitivity, F(1, 56) = 41.79, η2 G = .19, B10 = 1.9 × 106, giving evidence for a speed-accuracy tradeoff between the aided and unaided task conditions. The benefit of automation was decisively greater for the difficult than the easy condition, F(1, 56) = 16.88, η2 G = .08, B10 = 279.29. Data substantially favored the null models for the remaining effects, all B10 < 1/4.03, suggesting that format did not modulate subjects’ use of the automated aid’s cues.

Mean d’ as a function of Format and Aid conditions separately for the easy condition (above) and the difficult condition (below). Error bars represent 95% within-subject confidence intervals (Cousineau, 2005; Morey, 2008).
Workload Capacity
Cz is the average of workload capacity scores over time, inversely weighted by the data variability at each timepoint (Houpt & Townsend, 2012). As noted previously, scores follow a standard normal distribution, where values greater than zero indicate supercapacity and values less than zero indicate limited capacity.
CzOR and CzAND were analyzed separately in 2 × 2 between-subject Bayesian analyses with Task Difficulty and Format as factors. Figures 4 and 5 present mean CzOR and CzAND scores, respectively. Data gave decisive evidence for a main effect of task difficulty on CzOR, F(1, 56) = 21.74, η2 G = .27, B10 = 1,122.24, indicating that when the stimuli were less discriminable, the subjects’ decision making slowed prior to cue arrival. Data gave no substantial evidence either for or against a main effect of format, F(1, 56) = 1.24, η2 G = .02, B10 = 1/2.35, or an interaction, F(1, 56) = 1.05, η2 G = .01, B10 = 1/2.03. Mean CzOR value did not differ substantially from zero in the easy condition, one-sample t(29) = 2.11, B10 = 1.13, but fell decisively below zero in the difficult condition, one-sample t(29) = −4.17, B10 = 115.19.

Mean CzOR scores as a function of Task Difficulty and Display Format. Error bars represent 95% between-subject confidence intervals.

Mean CzAND scores as a function of Task Difficulty and Display Format. Error bars represent 95% between-subject confidence intervals.
Analysis of CzAND gave decisive evidence for a main effect of task difficulty, F(1, 56) = 24.52, η2 G = .30, B10 = 2,975.83, indicating lower mean capacity under difficult conditions. Data were close to indifferent toward the main effect of Format, F(1, 56) = 3.30, η2 G = .05, B10 = 1.01, and the interaction, F(1, 56) = 0.93, η2 G = .01, B10 = 1/2.04. CzAND scores were decisively greater than zero in the easy condition, one-sample t(29) = 4.78, B10 = 523.01, indicating that responses following a cue from the aid were faster than expected based on unaided performance. On the other hand, in the difficult condition, CzAND scores fell substantially below zero in the difficult condition, one-sample t(29) = −2.66, B10 = 3.71, indicating that responses following a cue were slower than expected based on unaided performance.
Discussion
Contrary to expectations, displays that integrated an automated aid’s cues with raw data did not alter automation use strategies compared to displays that separated cues and raw data. Operators’ RTs were shorter when using integrated rather than separated displays, but the benefit did not arise due to better mental integration of the aid’s cue and raw data, as evinced by the comparable effect of the aid across the different display formats. Workload capacity scores, further, gave no evidence that integrated displays eased integration of raw data and the aid’s recommendation.
Operators did interact with the aid differently between the easy and difficult conditions, however. In the easy condition, operators showed no differences in either sensitivity or mean RT between aided and unaided conditions. In the difficult condition, on the other hand, the aid improved sensitivity at the cost of response speed, a speed-accuracy tradeoff. Workload capacity scores lent insight into exactly how automation dependence slowed the operators’ response in the difficult discrimination condition. As noted previously, the two variants of workload capacity, CzOR and CzAND, reflect different facets of automation usage strategy in a speeded task: CzOR represents operators’ behavior prior to the arrival of the aid’s cue, and CzAND represents behavior following the arrival of the cue. In the current data, both CzOR and CzAND indicated limited capacity in the difficult condition. As compared to their performance in the unaided control condition, participants in the aided blocks produced slower responses both in the interval preceding the cue’s arrival and the interval following it. The former effect suggests that participants may have delayed their own responses, waiting for assistance from the aid. The latter effect suggests that participants required time to check the aid’s diagnosis against the raw data, or weigh it against their own judgments, before acting on it. In contrast, CzAND scores for the easy condition revealed supercapacity performance, suggesting that when stimulus discriminability in the easy task was high, participants reacted to the aid’s cues immediately (cf. Yamani & McCarley, 2016).
In application, the current findings imply that designers cannot assume that an automated aid, even if it renders judgments quickly and with high accuracy, will improve human performance without costs. In some cases, an operator and aid may operate with unlimited capacity even while automated assistance improves human response accuracy. Under those circumstances, the decision to implement the aid may be easy. In cases like the current task, where the aid improves discrimination but with an increase to RT or to the operator’s own processing rate, an expected-value analysis considering the costs and benefits of response accuracy and speed will be necessary to inform design choices (cf. Sheridan & Parasuraman, 2000). Speed-accuracy tradeoffs between aided and unaided conditions might be mitigated by allowing the aid itself to execute a response (Parasuraman et al., 2000) without input or approval from the human operator. Such a system would effectively implement an OR stopping rule, but one under which the aid’s diagnosis triggered a response directly, bypassing the human operator. Where practical concerns—technical, logistical, ethical, or otherwise—make it necessary for the human operator to approve or execute a response, the gains that response accuracy purchases with assistance from an automated aid might or might not outweigh the RT costs with which they are purchased.
Contrasting with the results of Meyer (2001), the present study did not find differences in automation use strategies across display formats. It should be noted that this is not strictly a failure to replicate, as the experimental method and analysis used here differed from Meyer’s in multiple ways. In Meyer’s experiment, for example, the automated aid was of relatively low accuracy, and the effects of display format were manifest on trials in which the aid’s judgment was wrong. Here, the aid operated with near-ceiling accuracy, precluding an analysis that isolated trials on which the aid erred. Moreover, the cues used in this experiment arrived at variable onset times relative to the raw stimulus. Visual transients produced by cue onsets (Folk, Remington, & Wright, 1994; Steelman, McCarley, & Wickens, 2011; Yantis & Jonides, 1990) might therefore have attracted attention, overriding the effects of integration. The effect of display format would have been stronger if the spatial distance between the target and the separate cue were larger. A further possibility is that a change of display semantics (Bennett & Flach, 1992), mapping the automation cues more meaningfully to the recommended responses or to underlying physical properties of the stimulus set, might better facilitate automation use strategies.
Conclusion: Studying Automation Usage With Workload Capacity
Analysis of workload capacity provides a method to characterize operators’ response behaviors when interacting with an automated aid in various speeded perceptual-cognitive tasks. As stated earlier, the workload capacity analysis can reveal operators’ strategic deference to the aid that is often masked in more conventional analyses of mean RTs and error rates. One caveat is that the workload capacity analysis requires a distribution of RTs from the decision aid and enough observations to construct empirical RT distributions for an operator and the human-automation team. The analysis is applicable to various speeded choice tasks, though, and alongside existing methods for detecting automation misuse or disuse (e.g., Bartlett & McCarley, 2017; Wang, Jamieson, & Hollands, 2008, 2009), offers a converging measure for characterizing human-automation performance. For a more comprehensive characterization of automation usage, future work might extend the current approach through the application of workload assessment functions (Townsend & Altieri, 2012). Assessment functions capture performance of a multichannel system as manifest in both accuracy and RT, allowing analysis of capacity conditional on accuracy. This methodology would expand the study of human-automation workload capacity to more difficult probabilistic tasks in which neither the human nor the aid approaches ceiling-level accuracy.
Key Points
An experiment used workload capacity analysis to assess operators’ automation usage strategy while perceptual proximity between an automated cue and raw information and task difficulty were manipulated in a speeded decision task.
The workload capacity summary measures, CzOR and CzAND, indicate evidence that operators moderated their own decision times both in anticipation of and following the arrival of the aid’s diagnosis under difficult task conditions regardless of display format.
Assistance from an automated decision aid may improve sensitivity in a time-stressed signal detection task at the cost of response speed.
Footnotes
Acknowledgements
We thank Arianna White for assistance with data collection and Mark Draper and three anonymous reviewers for helpful comments on an earlier draft of the manuscript.
Yusuke Yamani is an assistant professor in the Department of Psychology at Old Dominion University. He earned his PhD in psychology at the University of Illinois at Urbana-Champaign in 2013.
Jason S. McCarley is a professor in the School of Psychological Science at Oregon State University. He received his PhD in experimental psychology from the University of Louisville in 1998.
