Abstract
The focus of the current study was on how the dismounted soldiers’ decision cycle is affected by the use of a display device for utilizing intelligence from an unmanned ground vehicle during a patrol mission. Via a handheld monocular display, participants received a route map and sensor imagery from the vehicle that was ~20–50 m ahead. Twenty-two male participants were divided into two groups, with or without the sensor imagery. Each participant navigated for 2 km in a military urban terrain training facility, while encountering civilians, moving and stationary suspects, and improvised explosive devices. The OODA loop (observe–orient–decide–act) framework was used to examine soldiers’ decisions. The experimental group was slower to respond to threats and to orient. They also reported higher workload, more difficulties in allocating their attention to the environment, and more frustration. These can be partially attributed to the novelty of the technological capability, but also to its implementation in the study. The breakdown of performance metrics into the OODA loop components enabled analysis of the major difficulties in the decision-making process. This evaluation highlights the need for new roles in combat-team setups and for additional training when unmanned vehicle sensor imagery is introduced.
Keywords
Most of the current military operations happen in and around urban environments. Lack of situation awareness (SA) and uncertainty in identification of foes increase demand and complexity (Grau & Kipp, 1999; Phillips et al., 2001; Strater, Endsley, Pleban, & Matthews, 2001). Furthermore, the intensity level of the operations varies, and levels of conflict call for drastically different responses and tactics. Hence, military operations in urban terrain pose many cognitive challenges to dismounted soldiers (Helmus & Glenn, 2005). Nevertheless, there are limited experimental studies looking at soldier performance in such environments (Durlach, 2007).
One common notion in combat performance is known as Boyd’s OODA (observe, orient, decide, act) loop and is drawn from military strategies to present real-time decision-making processes. The OODA loop describes four basic phases: (1) Observation of an environment, gathering raw information about the surroundings through an array of human senses; (2) orientation by analyzing the gathered information and converting it into conclusions; (3) decision on what to do based on weighing potential courses of action (COAs) and the accumulated knowledge; and (4) acting or execution of the decision. A more elaborative model “of winning and losing” was offered by Boyd to comply with the various forms of combat and feedback loops available in the decision-making process (see Brehmer, 2005). In this elaboration, the observe stage consists of multiple information sources, both internal, derived from the unfolding of the circumstances and the immediate environment, and external, outside information provided by others (e.g., video feed provided by unmanned aerial or ground vehicles that are surveilling the area). Observation is often influenced by implicit guidance of higher echelons or sources whomay have broader perspective of the mission (e.g., look for a suspect who was seen exiting a compound …). In the orient stage, synthesis of information is accomplished—from representing the physical location of various elements in the immediate environment to a mental model of the situation. Although not using the same terms, this process is somewhat similar to what is being described as generating the mental simulationin Kleins’ recognition primed decision making (RPD) model (Klein & Crandall, 1995).
Several modified versions have stemmed from Boyd’s original sketch. The Dynamic OODA loop (DOODA loop; Brehmer, 2005) is a generic model of Command and Control (C2) that was formulated as functions that must be accomplished for effective C2, and these functions can be analyzed in terms of (partially overlapping) processes. The DOODA loop represents potential sources of delays in the decision process and thus broadens the emphasis on speed of decision making for studies aimed at finding ways of improving C2 effectiveness (Brehmer, 2005). To illustrate, sensemaking is influenced by several factors such as the mission, the information collected (and thus by the sensors available for the decision maker), and the command concept. The term command concept refers to the quality and ideas of the dismounted commanders, which enable exploitation of the technical C2 systems and utilize the information to the benefit of the mission (Builder, Banks, & Nordin, 1999). Hence, there is a separation between the technical capabilities of the C2 systems per se, the training andbackground of the soldiers and how they take advantage of the information at hand.
Generally, integrated visualization concepts have been shown to aid sensemakers (Ntuen, Park, & Gwang-Myung, 2010) in complex information environments. Here the focus is on ground forces performing Close Target Reconnaissance (CTR) missions. In CTR missions, the focus is on survey of a targeted area or path (patrol). Patrols exemplify the need to develop processes of sensemaking quickly and within a limited range. Baber, Fulthorpe, and Houghton (2010) proposed a framework representing two parallel cycles: a short CTR OODA loop cycle and a broader cycle of recording, communicating and interpreting information that feeds into the short cycle and improves sensemaking, SA, and situation understanding in context (e.g., imagery provided by cameras to supplement technologies such as night vision goggles or binoculars). Hence, the input fed into the short CTR loop should be relevant, accurate, and timely, and the notion that information is “a means to an end, not an end in itself” must be continuously stated.
As their level of autonomy increases, small unmanned ground and aerial vehicles are becoming an inseperable part of the platoon. Their utilization for airborne and ground reconnaissance is constantly increasing. Small dedicated unmanned systems aim precisely at providing immediate, accurate, relevant, and timely information for the infantry commanders and soldiers at the battalion, platoon, and squad level. Generally, the use of video feeds obtained from unmanned vehicles facilitates observation and orientation (Barnes & Jentsch, 2010). However, the introduction of unmanned systems has also raised concerns as to the tolls imposed on soldiers from the availability of additional visual information during the observe and orient components of their decision-making process. The visual attention requisites of additional information sources may diminish performance related to the immediate environment in real-time operations. McGuirl, Sarter, and Woods (2009) added video feed from an unmanned aerial system (UAS) to incident commanders whose main task was to correctly identify hazards and apply appropriate responses. They found that the presence of the video feed caused the commanders to narrow their assesments of the situation, ignore, or less frequently search for other sources of data, generating a nonbalanced monitoring strategy of the information at hand. Nevertheless, there is a general agreement that higher SA can be obtained by this type of information (e.g., Barnes & Jentsch, 2010).
Interests and information needs vary among echelons: Dismounted soliders at the lowest levels of the hierarchy are mostly concerned with threats that are in their immediate environment (~50 m; Redden, 2002), whereas their commanders are most interested in the location of enemies and friendly forces (Evans & Baus 2006). It has been shown that dismounted soldiers prefer ground views over aerial ones when those can providethe necessary information (e.g., Oron-Gilad,Redden, & Minkov, 2011). They can also benefit from simultanous presentation of aerial and ground feeds (e.g., Ophir-Arbelle, Oron-Gilad, Borowsky, & Parmet, 2012). Based on postexperiment interviews with platoon leaders, Durlach (2007) found that the addition of tele-operated unmanned ground vehicles (UGVs) lead to loss of attention of the immediate environment for the UGV operators who then needed to be supervised by others. Similarly, Lif, Jander, and Borgvall (2006) reported that utilization of a UGV provided more possibilities to gather reconnaissance but slowed the advance pace of the team due to the need of the UGV to scout the terrain and of the team to interpret these additional data. They too mentioned that UGV operators could not simultaneously perform ordinary soldier duties, such as scouting, while operating the UGV. In both of these studies, the UGV was tele-operated by a soldier operator in the team, and it is not clear to what extent these patterns will change when the UGV has more autonomous capabilities or is led by a remote operator. Additionally, there is the issue of the display device itself. There are limitations in size, type, and availability of display devices and those may interact with the type of information provided and mission characteristics (see Oron-Gilad et al., 2011).
Finally, it is important to note that, in reality, CTR operations are conducted not at the level of individuals but rather in teams. It is not clear yet how the technology would be set up, whether one person will attend to the display or all team members will have access to the same information. Furthermore, it is not clear how such new capabilities will influence dynamically occurring team processes related to communication, coordination, and information sharing among team members. Evidently, roles can be allocated within a team, to reduce workload on each individual operator, and to compensate for an individual’s lack of SA due to the use of the technology. However, in order to do this, one needs to first understand how an individual operator is impacted by the introduction of the technology. As noted by Salas, Prince, Baker, and Shrestha (1995), measurement of team SA may be premature until there is a clear understanding ofwhat the concept represents. Hence, as a first step, it is important to measure and understand the individual’s SA, particularly here, to understand how the introduction of new technological capabilities affects a single operator’s ability to perceive and decide how to act to various events.
Hypotheses Development
This field study was designed to look at responses to operational events by experienced dismounted soldiers conducting a patrol mission. Responses were analyzed as sequential processes utilizing the OODA framework. For the experimental group, video feed from a fronting UGV was presented via a handheld monocular display (HHMD). While patrolling, the video feed could provide critical information that was not available to the soldier, as the UGV preceded the soldier at a range of 20–50 m at all times. Yet there were situations where the UGV could not provide the necessary critical information due to limitations of its operation and payload range. The control group could use the display device to look at the map and respond to events but without the UGV video feed. Table 1 presents the hypotheses on the relevant stages of the OODA loop. Additional considerations are given below.
OODA Loop Stages (Observe, Orient, Decide, Act) and Expected Impact
Note. UGV = unmanned ground vehicle; HHMD = handheld monocular display.
Observe—Problem Detection: Events That Can Be Seen Only by Own Eyes/Events That Can Be Seen Both Ways
Dismounted soldiers are limited in the type of display device that they can carry. The two most common display device types for them are near eye displays (NEDs) or handheld displays (HHDs). Many in the military advocate the use of NEDs during dismounted movement, stating that they allow soldiers to observe information presented on the display while they are conducting dynamic operations (Redden, Pettitt, Carstens, & Elliott, 2010, but see also Oron-Gilad, Minkov, & Goshen, 2014, on by-foot navigation). NEDs have a disadvantage with regard to perceptual issues, attention allocation, movement speed, and nausea (e.g., Elliott, Duistermaat, Redden, & van Erp, 2007; Oron-Gilad et al., 2011). HHDs are not hands free and canbe cumbersome to handle, particularly when larger screens are necessary. The screens of many HHDs are difficult to see in daylight becauseof glare. Most individuals who attempt to walk and look at the screen of their device have compromised SA of their surroundings. During stationary operations, HHDs were equally effective or more effective than helmet-mounted ones (e.g., Scribner, Wiley, Harper, & Kelley, 2007). The HHMD is a monocular high-quality resolution nonglare display that incorporates the advantages of a HHD with regard to orientation and obstruction without the drawbacks of traditional NEDs. The use of the HHMD itself is not novel. The novelty here is in the type of information that is provided via the display. The preview of the environment as obtained from the UGV sensor feed (i.e., where the soldier will be in 20–50 m) may generate overreliance (Mosier & Skitka, 1996) in the display, as participants will be less attuned to their immediate environment and more attracted to the happenings seen in the display. As such, it was hypothesized that events that can be observed only by own eyes will be handled faster and more accurately without the additional information input, whereas events that can be seen via the display (and via own eyes a bit later) will be facilitated by the new capability enabling faster responses and higher accuracy.
Orient
The addition of the new technological capability may distract the soldiers from their immediate environment. It was hypothesized that soldiers using the HHMD in the experimental condition with sensor imagery will have reduced awareness of the environment, resulting in fewer reports of waypoints and/or slower responses toward designated waypoints in their surroundings. This component of SA is the basis for planning and replanning dimensions that were not examined in the current field evaluation.
Development of Mental Simulation—Anticipation of Events (Level of Certainty)
All participants in the field study were trained soldiers experienced in urban terrain combat techniques. All have the acquired skills to encounter suspects and to recognize improvised explosive devices (IEDs) and trip wires, which are typical threatening events in operations of this sort. All threatening events in this field study fall under the categories of familiar routine events and unfamiliar but anticipated events (Ntuen, 2006). For this level of expectancy, it is hypothesized that once the problem is identified, common ground and decision on COA will not be affected by the addition of the technological capability. Thus, once an event is identified, soldiers will know what decision to make or what COA to take.
Method
The study was conducted in a military urban terrain training facility, familiar to all participants from their previous military training. This site holds buildings, roads, paths, and allies similar to a typical Arab village, including a central Kasbah with a more densely built area and higher buildings of up to eight floors. The study was not part of any training routine. In a between-participants design, participants were equipped with a HHMD that enabled them to register responses to elements in the environment and (in the experimental condition) to view video feed from an UGV that was leading their path. Paid actors were playing the roles of human civilians and suspects, and potential threats were planted in advance by the experimenters along the patrol route.
Participants
The 22 male participants were aged 21 to 29 years. All were former Israel Defense Forces (IDF) infantry soldiers who had undergone military operations in urban terrain (MOUT) training in their mandatory service (3 years) and in their active duty reserve service (also mandatory for up to 1 month per year once released from mandatory service). All had experience in patrol missions. All were students at Ben-Gurion University who had been in active duty at least for one service-cycle/training in the past 12 months prior to the experiment. None of the participants had prior experience with using video feed from UGVs. Participants were recruited via forum postings and ads in the university. They received 100 NIS ($25) for their participation. To increase motivation, a bonus of 100 NIS was given to the participant with the highest scores.
Experimental Tasks
Task 1—Navigate a prescribed route. The task was to navigate toward a target destination via a specific designated route (patrol).
Task 2—Respond to threats. Report or respond to events along the route. An appropriate response to an event was defined through a specific set of rules, explained to the participants in advance (i.e., rules of engagement/COA).
Task 3—Report progress along the route. Orientation is a challenge in urban areas; closed areas and irregularity of building settings cause difficulties in orientation and failure of forces to follow routes. Participants were asked to report when the UGV (where applicable) and they themselves have reached two specific waypoints marked on the route map.
Operational Scenario
The operational scenario was a military activity in the village that required the soldiers to reach a specific destination while continuously searching for potential suspects and threats that may endanger them or their team along the way. The length of the route was 2 km, and in moderate pace, it took about 15 min to complete it. Similar to real-life combat in such areas, soldiers are exposed to unanticipated threats when they are moving and are trained to avoid and respond to such threats. Potential threats were IEDs mostly hidden within a pile of rubble or trash, with or without noticeable tripwires. They were placed in advance along the route at various locations. Actors, wearing green T shirts, were positioned as human suspects. Some of them were stationary and conspicuous, holding various types of weapons such as rocket launchers; some were static but partially hidden (e.g., hiding behind a building opening); and some were moving in the area. Actors with any other shirt colors were considered civilians. All actors were carrying communication devices so that the experimenters located in the command area could coordinate their appearance with the walking pace of the participant moving in the area. Figure 1 presents examples of events that required action.

Examples of events along the route. Top left is a stationary suspect holding a rocket launcher, bottom left is a moving suspect, top right is an improvised explosive device (IED), and bottom right is a tripwire connected to an IED.
Apparatus
HHMD
Participants were equipped with a HHMD, a standard military equipment device by Na-Or Navigation & Orientation Systems Ltd. The HHMD has a resolution of 1280 x 1024 pixels and two input buttons, as shown in Figure 2. Response buttons were used to report upon events and orientation. The HHMD is common gear in some infantry units, used mainly to view C2 maps. The novelty here was on the use of the HHMD to view video feed from an unmanned vehicle.

Left is the handheld monocular display (HHMD); right is a participant using the HHMD.
UGV video feed
A video feed derived from a UGV was generated and transmitted (when applicable) to the HHMD. The UGV drove approximately 20–50 m in front of the participant, filming elements in the environment that the participant was about to approach before the participant has actually arrived to that particular location. As noted by Redden (2002), dismounted soldiers are most interested in threats in close proximity of a 50-m radius. Participants had no means to control or direct the UGV payload, the speed of the ground vehicle, or the viewing angle of its camera. The UGV’s camera rotated from left to center to right at a fixed pace and at a fixed height. It was not able to cover elements that were above its camera view, as shown in Figure 3.

An example of a sequence of snapshots taken from the unmanned ground vehicle (UGV) video feed. From left to right: The UGV camera faces upfront and then turns left into an entrance and back upfront (there are about 10 seconds between each two snapshots).
Experimental user interface
The experimental user interface in the HHMD display consisted of three types of elements, as shown in Figure 4: (a) a map of the route to be used for orientation and navigation, including the markings of waypoints that the participant had to register (as noted in the secondary task); (b) the video feed of the UGV (when applicable; otherwise this area was kept empty); and (c) response buttons used to register responses to events along the route (Event button) and when reaching designated waypoints (WP) along it, as shown in Figure 4. Once pressing on the Event button, the operator could choose two modes of action—report or act—as shown in Figure 5. The interface did not provide information as to how far the UGV was from the participant. The HHMD was connected to a laptop computer that was carried in a backpack, which was part of theparticipant’s gear. All button presses were registered and logged into a log file.

The screen of the handheld monocular display (HHMD) by experimental condition. Left = with unmanned ground vehicle (UGV) video feed (experimental condition); right = without (control). The remaining screen components were identical. The bottom left area in each screen consisted of the route map. Red markings indicate the waypoints. On the right side of the screen, there were response buttons that could be selected by pressing the physical buttons on the HHMD. Event button is on the top right corner of the screen; Waypoint report button is on the bottom right.

The screen of the handheld monocular display (HHMD). Once the Event button was pressed, the participant had to select the courses of action (COA) and two new buttons appeared: Act and Report. They were operated in a similar way to the original Event and Waypoint buttons, via the HHMD physical buttons.
Experimental Design, Event Categorization, and Objective Measures
The study was a between-participant design. Participants were randomly allocated to one of the two experimental groups: with UGV video feed (experimental condition) and without (control). There were 14 events in the scenario. Those varied in several dimensions: the type of event (moving human/stationary human/IED), the way they could be seen by the participants (via the HHMD, own eyes, or both), and their level of certainty (easy to identify vs. ambiguous/less obvious). Seeing an event only via the HHMD meant that participants could not see the event at the time it appeared in their own eyes. The expected COA as determined by subject matter experts (SMEs) is specified in Table 2. Participants could decide to act on their own or report to a higher echelon commander for assistance.
Scenario Events and Dimensions
Note. COA = courses of action; IED = improvised explosive device; HHMD = handheld monocular display.
Objective Assessment of Performance
Table 3 provides details on the performance measures and how they were obtained from the collected data.
Objective Assessment of Performance
Note. HHMD = handheld monocular display; UGV = unmanned ground vehicle; WP = waypoints; COA = courses of action.
Posttask Evaluation
Regaring common-ground and sense-making interpretation, using a retrospective process, the experimenter was verifying with each participant why he chose to respond in a particular way to each one of the events that he registered. This was primarily conducted to verify whether there were gaps between the master solution (prepared by SMEs experts; Table 2) and the individual responses and, if there were, what the cause was for the different intrepretation.
Closing Questionnaire
A closing questionnaire was administered at the end of the experimental session. It had two parts: a brief usability evaluation of the HHMD interface and a subjective evaluation of performance and of the impact of the UGV feed on performance (see Appendix).
Subjective Workload
Mental workload was assessed using theraw NASA Task Load scale (RTLX), an unweighted average of the subscale values (Hart & Staveland, 1988; Nygren, 1991). RTLX was administered twice: once before and once after the operational mission.
Experimental Procedure
Participants arrived at the MOUT facility, one at a time, early in the morning. Upon arrival, they were briefed about the operational mission and the rules of engagement. Then, they tried the HHMD, depending on their assignment (experimental condition/control), and once they had basic understanding of the operational tasks and the use of the device, they underwent a practice trial. Because there are no common guidelines for use of HHMDs to view video feed from unmanned vehicles, participants were not given explicit instruction as to how to use the display; that i, they were not told when or how to look for information in the display. In both the practice and experimental route, participants were followed by an experimenter, who was responsible for synchronizing the actors to their walking pace. While walking, the participant and the experimenter did not interact. In the practice trial, which was shorter in length (~500 m), they practiced the use of the HHMD and were familiarized with the events and potential threats. Once verification of the expected COA was obtained and participants felt comfortable with the mission and the gear, the practice session ended. The NASA TLX questionnaire was administered before the experimental route began. Following the completion of the experimental route, they filled the NASA TLX questionnaire again. Then they were questioned about common-ground and sense-making interpretation, and finally they filled the closing questionnaire. Lastly, they were debriefed and compensated for their time.
Results
Objective Performance
Statistical analyses and descriptive results are summarized in Tables 4 and 5. All analyses were conducted at a .05 significance level utilizing a general linear mixed model (GLMM) and a sequential backward elimination procedure to reach the final model. For all analyses, the variable “participant” was added to the model as a random effect to account for multiple observations of each participant. The independent variables were as follow: the experimental condition, the type of event, and the way the event could be seen (Table 2). Second-order interactions were not included in the model (as not all were possible due to the experimental design). To calculate the degrees of freedom for each test, we used the Satterthwaite approximation, which is useful if the sample size is smaller or data are unbalanced. With this approximation, the degrees of freedom can be lower than the degrees dictated by the numbers of observations. Details for each measure are provided in the following sections.
Summary of Performance Results: Main Effects
Note. GLMM = general linear mixed model; WP = waypoint.
Summary of Performance Results: Descriptive Statistics
Note. Exp. = experimental; UGV = unmanned ground vehicle; IED = improvised explosive device; MS = moving suspect; SS = stationary suspect; HHMD = handheld monocular display.
Exp./control.
IED/MS/SS.
HHMD/own eyes/both.
Respond to Threats – Problem Detection – Observe
Accuracy
A GLMM analysis within the logit framework was conducted on the accuracy of response (0 = missed; 1 = detected). The independent variables were the experimental condition, the type of event, and the way the event could be seen (see Table 2). All in all, there were 363 observations (14 × 11 for the control group and [14 + 5] × 11 for the experimental group due to the fact that Events 2, 3, 6, 8, and 13 could be registered in any one of two ways). Using a backward elimination procedure, the final model included two main effects (see Table 4). The estimated probability to detect an event was lower for the experimental group than for the control group. There were also differences in the ability to detect events. As expected, the probability to detect an IED was lower than for moving or stationary suspects that were not statistically significant from one another. Detection rates for events that could be detected only from the video feed (relevant to the experimental group) were lower than when events could be seen with own eyes or in both ways.
Response time
A GLMM analysis was conducted on the normally distributed log transformation of the response time. The independent variables were the experimental condition, the type of event, and the way the event could be seen. The final model included all three main effects (see Table 4). Estimated means for response times were 980 ms higher for the experimental group than for the control group. Moving suspects were detected faster than IED or stationary suspects. Finally, events that could be detected only from the sensor feed yielded the slowest responses, 3 s or more than when events could be seen with own eyes or seen both ways, as seen in Table 5.
A following analysis conducted only on the events that could be seen with own eyes or both ways with the experimental group and the way it could be seen revealed a statistically significant interaction, F(1, 257) = 10.62, p < .002, as shown in Figure 6. Notable, participants in the control group could not distinguish between events, as they had no sensor feed. Participants in the experimental group responded faster to events that were seen in the sensor feed and with their own eyes but responded slower to events that were seen only in their own eyes.

The interaction between response time to an event and the experimental condition (with/without video feed). In the experimental condition, events that could be seen both ways, from the video feed and own eyes, were identified faster than those that could be seen only with own eyes.
Respond to Threats – Response – Orient, Decide, and Act
Once an event was detected, participants had to determine the appropriate COA, as specified in Table 2. Responses were collected via the HHMD buttons. The level of certainty of the event was included in the analysis model.
Accuracy
A GLMM analysis within the logit framework was conducted on the accuracy of response (0 = if wrong; 1 = if correct). There were 287 observations. The independent variables were the experimental condition, the type of event, the way the event could be seen, and certainty level. Using a backward elimination procedure, the final model included two main effects for type of event and certainty level (Table 4). The estimated probabilities to respond correctly to an event were relatively high (Table 5), indicating that once an event was identified, it was not important how it was detected. Yet the type of event and the ambiguity were relevant for the decision.
Response time
A GLMM analysis wasconducted on the normally distributed log transformation of the response time (i.e., from the time the event was identified until the decision on COA was made). The independent variables were the experimental condition, the type of event, the way the event could be seen, and certainty level. Using a backward elimination procedure, the final model included two main effects (Table 4). The decision toward humans was determined faster than for IEDs (Table 5), which is reasonable as humans are more distinct than IEDs. The COA was determined faster when it could be seen by own eyes or in both ways than when it could only be seen in the sensor video (Table 5). Participants were slower to make a decision when events were identified only in the sensor video, but note that there were only two events of this type (Table 2).
Report Progress Along the Route – Observe and Orient
Participants were asked to register two waypoints: once the UGV (where applicable) had passed them and once they had passed them. The dependent variable was the correct identification of the waypoint location (1 = correct; 0 = no report or incorrect). The independent variable was the experimental condition. Because the UGV was ahead of the participant, participants in the experimental condition could press the waypoint button twice: once when the UGV reached the waypoint and then again when they reached the waypoint. Comparisons among the three possibilities—report of self-location in the control group, report of self-location in the experimental group, and report of the location of the UGV in the experimental group—were made using a GLMM analysis with a logit link function. The main effect was statistically significant (Table 4). Participants were able to identify that they have reached the waypoint equally well for the experimental and control group, respectively (Table 5). However, the experimental group failed to acknowledge that the UGV reached the waypoint location, possibly because they had no indication of its location on the HHMD map. Furthermore, participants in the control group identified faster that they have reached the designated waypoint (Table 5).
Posttask Evaluation
Participants’ posttask evaluations of the various situations did not differ between the experimental group and the control. The vast majority of answers were consistent with the master solution (Table 2).
Subjective Performance Evaluation
Table 6 provides qualitative subjective evaluations of performance.
Summary of Subjective Performance Evaluation
Note. HHMD = handheld monocular display; UGV = unmanned ground vehicle.
Mental Workload
The raw NASA TLX questionnaire was administered twice: before and after the operational mission. Participants in the experimental group reported higher perceived mental workload (71% and 90% for the control and experimental group, respectively) and frustration (45% and 71% for the control and experimental group, respectively). Other workload items did not vary between the groups.
Discussion
The use of the OODA naturalistic decision-making paradigm was shown to be an effective way to break down the complex conditions of the current decision-making problem and highlight particular mission components where performance was degraded, as seen in Tables 4 and 5 and Figure 5. What remains is to attempt to identify why performance degradation occurred and how it can be eliminated or mitigated viabetter designs, changes in roles and team configurations, or improved training. Nevertheless, it is important to note that the mission in this field evaluation used a constrained choice of simple COAs (Table 2) and that the OODA framework may not be suitable for more complex missions, where there are more degrees of freedom. Possibly there, it may be more useful to employ NDM models such as RPD with more elaborative macrocognitive functions.
The main task in this field study was to patrol while correctly identifying threats along the route and making appropriate decisions (act or report to a higher echelon commander; Table 2). As orientation in urban environments is difficult, participants were also required to report upon reaching two preassigned waypoints. The events encountered along the route were “familiar routine” events that dismounted soldiers experience in urban terrain, such as the presence of civilians and rubble piles that could conceal IEDs, or “less familiar but anticipated events,’ such as an encounter with an armed suspect. These types of events fall within the categories of events with which trained soldiers should know how to deal (Ntuen et al., 2010). Indeed, as seen from the orient–decide–act component of the response to threat task, once identified, soldiers in the experimental group and control group were able to handle the events equally well, depending on the uncertainty of the event. Thus, as confirmed also by the posttask evaluations, participants’ understanding of the appropriate COA was not harmed by the new technological capability.
However, it is evident from the results that the availability of the UGV video feed through the HHMD slowed participants when it came to detecting events that were not seen in the video feed (Figure 5). In urban terrain conflicts, slow responses may cost lives. It was also evident that participants in the experimental group had difficulties in realizing that the UGV reached a waypoint and were more than three times slower to register when they reached a waypoint themselves (Table 5), reinforcing that they were more attuned toward the sensor imagery and less aware of the UGV’s location in the environment or of their own. Because here the route was planned in advance and no changes in it were made during the experimental trial, poorer awareness toward their own location and toward the location of the UGV was not a critical component of their mission. However, in real-life operations, replanning may occur and can be severely affected by deficiencies in orientation. Possibly, some of these deficiencies could be reduced if the location of the UGV was dynamically marked on the map.
It is possible that the novelty of the technological capability of viewing UGV sensor feed, its display in the HHMD, or both made participants more likely to focus on the sensor video than on their old scouting skills (in line with previous studies that reported on a technological halo effect; e.g., McGuirl et al., 2009). It is also possible that the perceived value of the information received from the UGV was high, as it presented timely data preceding the participants. Because soldiers are most concerned with information that is within the radius of 50 m from them (Redden, 2002), it may have seemed to the participants that this newly available predicting information is more important. Coupled with a misjudgment that their well-trained MOUT scouting skills will not be affected by the new technology, it is possible that participants felt that they should focus more on the information from the HHMD than on their immediate environment. The subjective performance evaluations (Table 6) illustrate that participants in the experimental condition were aware of the difficulties that the sensor imagery had on their ability to allocate their attention between the display and their immediate environment. Hence, there is an inconsistency between perceived effectiveness and actual use, and the actual use was higher than its perceived utility. Possibly, the sensor imagery was dominating because it was perceived as a preview, a way to anticipate the comings, although in reality it only predicted a portion of the information. Empirical results provide confirmation to what has been reported previously in a more retrospective manner by platoon leaders and SMEs (e.g., Durlach, 2007; Lif et al., 2006), that the use of video feed tolls the soldier whois utilizing it and does affect his understandingof the immediate environment, making himmore vulnerable. It is possible though that more training with sensor imagery and better understanding of its limitations will reduce the amount of attention it attracts. However, the toll of using a visually demanding display cannot be completely eliminated.
There are two major limitations to this study: (a) With regard to the implementation of the UGV capabilities, it had a limited field of view (180 degrees) and was not capable of capturing elements above its sensor height. It did not have capabilities to directly communicate with the dismounted soldiers, hence participants could not control its pace or stop it at certain points along the route to gain a better look, which forced them to be attuned continuously tothe video feed. Finally, its movement was not dynamically marked on the map, making itdifficult for the participants to orient it. Possibly, alternative technological ways can be developed to indicate when the value of the information provided by the UGV is most critical (e.g., by sending cues from the UGV sensors or by a remote operator who is also monitoring the feed) and may be a better way to reduce tunneling of attention toward the video feed at all times (see also Katzman, Oron-Gilad, & Salzer, 2015). (b) With regard to the experimental design, this was not a full factorial design, and sample size was relatively small. A small sample size yields low statistical power, which can cause a miss of a real effect that could possibly become statically significant with a greater sample size. Future studies should focus on having more moving suspect events and more events that can be seen only from the display device and attempt to recruit more participants.
Lastly, participants in the experimental group stated in the final briefing that the device cannot be used by a sole soldier and should not be used by warfighters on duty but rather by a supporting team or by a designated team member in the platoon. Now, once the implications of sensor imagery on an individual have been looked at, it is time to examine such technological capabilities within a team setup, while taking into consideration team SA metrics in determining the roles and distribution of roles among its members.
Conclusion
Conclusions from this field evaluation can be separated as follows: (a) regarding the framework of investigation of the decision-making process in field evaluations, (b) practical implications regarding the suitability of the particular device (HHMD) and interface to the mission investigated (CTR-patrol), and (c) moving from an investigation of a single operator into team formation.
With regard to the framework of investigation, using the OODA framework for analysis of operator performance was beneficial. It improved the sensitivity of the evaluation, provided insights as to where (stages of decision) and how (response time and accuracy) performance degradations occurred, as well as provided insights as to what needs to be improved in the interface design.
Enhancements in the technical capabilities of the display device and the interface should improve orientation and detection of threats, enhance orientation, and must offer dismounted soldiers affordable ways to balance their attention allocation policy between the immediate environment and the display device. Specific modification areas that were identified include the following: (a) enhancing the capabilities of the unmanned vehicle’s sensor imagery (e.g., field of view), (b) displaying and synchronizing the location of the unmanned vehicle in the C2 map, (c) providing the dismounted operators with some level of control over the pace of the sensor imagery (e.g., stop or pause), (d) adding alerts and annotations retrieved from automatic target recognitions or human sources to supplement the presence of the sensor imagery, and lastly (e) allowing for bidirectional communication between the dismounted soldier and the unmanned vehicle (or its remote operator) to improve the payload’s search pattern and adapt it with changing mission needs.
To take advantage of the information at hand, as noted also by the soldier participants, the display device technology may be best facilitated in a new team-configuration setup where only one designated person attends to the information in the display while the other team members focus on the immediate environment. Notably, a change in team setup must be accompanied with a new training regime to improve team sense making.
Footnotes
Appendix
Acknowledgements
This work was supported by the U.S. Army Research Laboratory through the DCS Subcontract APX03-S010 (B.G. Negev Technologies and Applications Ltd.) under Prime Contract W911NF-10-D-0002 (Michael Barnes, technical monitor). The views expressed in this work are those of the authors and do not necessarily reflect official Army policy.
Tal Oron-Gilad is an associate professor and the chair of the Department of Industrial Engineering and Management at Ben-Gurion University. She holds a BS and MS degree in industrial engineering from the Technion and a PhD from Ben-Gurion University. She spent 3 years as a research associate at the Institute for Simulation and Training, the University of Central Florida, where she was part of a U.S. army multiuniversity research initiative. Her research interests are in human factors engineering and decision support systems. Some of her current experimental work concerns the evaluation bidirectional graphic communication between operators and remote unmanned vehicles. She has been continuously funded by extramural sources for every one of the years of her professional career, including support from U.S. Army Research Lab, IMOD, General Motors, Israeli Aerospace Industry, the Israeli road safety authority, Israeli ministry of science, and EU projects (SOCRATES).
Yisrael Parmet is a faculty member in the Department of Industrial Engineering and Management. He holds a BA in economics and statistics and MS and PhD degrees in statistics from Tel-Aviv University. He specializes in areas of design of experiments and statistical modeling. During his studies for his master and doctorate degrees, he served as a research assistant at the statistical laboratory at the Department of Statistics and OR, Tel Aviv University, which granted him knowledge in practical data analysis. In 2007–2008, he was a visiting professor in the Department of Dermatology and Cutaneous Surgery in the UM Miller School of Medicine.
