Abstract
Background
Red Teaming is widely used to discover vulnerabilities, test defensive measures, and anticipate emerging but novel threats. It has rarely been conducted both systematically and at scale, substantially limiting confidence in its results and the generalizability of its findings.
Aim
We introduce distributed, empirical, systematic, and scalable red teaming (DESSRT), a framework for translating tactical-level Red Teaming into a replicable research methodology. We apply DESSRT to address whether the information about and availability of computed tomography (CT) scanners influences adversary decision-making in aviation security.
Method
Using a convenience sample of 143 university students, participants role-played as adversaries in an eight-hour attack planning exercise. Via a custom instrument, participants were randomly assigned across three adversary profiles built on historical cases and then designed a simulated attack. Afterwards, one of three injects about CT scanners were randomly assigned, and participants were asked about potential changes in attack plans (including target changes). Differences among assigned profiles and CT scanner injects were evaluated using standard statistical tests of association.
Results
Although differences in explosive and weapon package selections were not statistically significant across profiles, security evasion methods were. Following injects, participants were equally as likely to change tactics across profiles, with the majority (53%) changing at least one tactical area. When asked, the majority (18) of those who changed targets (27/143) reported that the additional information on CT scanners did have some effect on their target change decision.
Conclusion
Overall, the DESSRT framework provides a novel mechanism for translating traditional Red Teaming exercises into a replicable and empirical research method. Although not a replacement for historical data, where available, DESSRT allows analysts and researchers to test theories about human decision-making, generate novel what-if insights to support planning efforts, and validate parameters within complex models.
Background
Red Teaming, or “the simulation of adversary (or adversarial) decisions or behaviors, where outputs are measured and utilized for the purpose of informing or improving defensive capabilities” (Ackerman & Clifford, 2021) 1 , is widely used within the national security community to aid security and threat assessments. It counters intrinsic cognitive biases and heuristics that distort decision-making and undermine effective planning and policymaking (Kahneman, 2011; Kardos & Dexter, 2017; Manoogian & Benson, 2017) by providing novel perspectives, intentionally challenging existing plans, and improving understanding of the operational environment (Gladman, 2007; Longbine, 2008; NWDC, 2011; U.K. Ministry of Defence, 2013). Red Teaming is especially helpful for discovering previously unidentified vulnerabilities (Defense Science Board Task Force, 2003; Landry, 2017; Longbine, 2008; Nettles, 2010); robustness testing of existing defensive measures (Kardos & Dexter, 2017; Zenko, 2015; Zhang & Gronvall, 2018); exploring emerging but novel threats (Zhang & Gronvall, 2018); and raising awareness of incipient security challenges (Wood & Duggan, 2000).
Methodologically, Red Teaming is better conceived as a simulation toolkit (NWDC, 2011), encompassing a variety of structured activities from field exercises and computational simulations to cyber penetration testing and tabletop exercises. Although widely implemented, Red Teaming (outside of the cyber domain) has rarely been conducted both systematically and at scale, limiting the generalizability of its results. This undermines its credibility among some decisionmakers, who are understandably reluctant to shape policy around the results of a single (or even a small number of) simulations. At the same time, choice experiments, whether in a laboratory setting or conducted as survey experiments, have long been used by social scientists to investigate human decision making (Hainmueller et al., 2014; Johnson, 2007; Street & Burgess, 2007; Wittink & Cattin, 1989). In recent years, such experiments are increasingly applied to security contexts (Gadarian, 2010; Huff & Kertzer, 2018; Kearns et al., 2020; Rousseau, 2021). However, even in the rare case when these studies ask participants to act as adversaries (e.g., Stotz et al., 2021), they are not conducted as true Red Teaming simulations, often lacking measures to properly acclimate participants to their roles and rarely including sufficient narrative context or inputs about the adversaries themselves. Such elements are especially critical when analyzing tactical choices by adversaries, where context, culture, and contingency shape decision-making (Brown, 2020; Hoffman, 1993; McCormick, 2003).
Detailed adversarial behavior, whether empirically coded or simulated, is of particular interest to operational agencies. Comprehensive information on how adversaries may act aids planning efforts, especially in relation to the emergence of new offensive capabilities or technologies. Moreover, impact assessments of security policies and practices often require behavioral data with and without the specific implementation. For new policies, data of this type can be scarce, and the lack of suitable validation data is often identified as a problem with extant risk models (Collier & Lambert, 2019; Lathrop & Ezell, 2016; Morral et al., 2012). Thus, methods are needed to investigate potential adversary behavior where empirical data is limited or challenging to obtain, while producing generalizable results that permit a robust impact assessment of interventions.
Intervention
To this end, this article introduces distributed, empirical, systematic, and scalable red teaming (DESSRT), a framework for translating tactical-level Red Teaming into a scalable, replicable research methodology. DESSRT falls into the family of scenario-based Red Teaming conducted in a tabletop or ideational fashion rather than actual penetration testing. Like many types of scenario-based Red Teaming, DESSRT participants adopt adversary roles and then plan a simulated attack. Unlike shorter survey instruments, participants are provided extensive background on the adversaries and primed (through a series of debiasing exercises) to role play adversary (Red) decisions under the given input conditions. Granular information regarding courses of action (such as target selection) capture which options are chosen, as well as the reasons why others were initially considered and then rejected. DESSRT can also be combined with other methodologies, including choice experiments and haptic measurements, to better understand the dynamics of adversary decision-making.
At its core, the DESSRT framework consists of four advances on existing Red Teaming methodologies, which when combined improve both the fidelity of the findings generated and their generalizability. First, the framework leverages distributed technologies to expand Red Team participation beyond the typical conference room, or tabletop, venue. Although not new to Red Teaming, previous distributed exercises have focused on collaboration via virtual sessions enabled by web conferencing technologies, such as Microsoft Teams or Zoom (e.g., Kodalle et al., 2021). The DESSRT framework moves beyond simultaneous collaboration to use asynchronous capabilities, such as the Qualtrics platform, to allow remote individuals to participate at their own pace. This expands the opportunity for Red Teaming exercises to diversify perspectives and avoid potential biases by including global participants whose availability falls outside regular business hours in a single time zone.
Second, the framework promotes systematic approaches to Red Teaming. One of the benefits of Red Teaming is its ability to incorporate flexible inputs from participants to encourage creative, exploratory thinking about threats (Zenko, 2015). Unfortunately this can lead to open-ended inputs and unstructured data that is difficult to compare across multiple iterations. The DESSRT framework, which places emphasis on the replicability of findings, standardizes data collection instruments while minimizing constraints on participant creativity. Thus, while allowing for open-ended inputs necessary for qualitative assessment, the DESSRT framework requires more structured participant deliberations (whether at the individual or group level). These can include standardizing instructions and data input requests, requiring follow-up questions to initial inputs, and constructing similar experiential requirements across iterations (e.g., the amount of time spent on a particular portion of the scenario). Systematizing also requires exercise designers to “script” the sequence of experiences as well as the format of scenario prompts and participant inputs, moving Red Teaming from an open-ended format to a more structured assessment instrument.
Third, the systematic nature of the framework allows for the generation of useful empirical data from the participants in the Red Teaming simulation. Standardizing the simulation instrument, incorporating structured sequences of interaction, and allowing for integration with other research methods (such as survey experiments) expands the empirical research opportunities regarding adversarial decision-making processes and adaptations. Although most Red Teaming exercises by their nature generate simulated data (as opposed to real-world observations), DESSRT simulations generate systematic observations of human behavior in an experiential or experimental setting. These qualities allow for DESSRT outputs to constitute usable empirical data, allowing advanced data analytic approaches to be applied.
Finally, the DESSRT framework allows—and indeed is intended—for scalable Red Teaming, where human simulation is consistently repeated across multiple iterations and participants. Red Teaming is typically employed in one-off fashion, often when prompted by a gathering crisis or emerging threat. The marginal cost of conducting additional iterations of a simulation has traditionally been quite high, as most if not all the logistics involved with conducting the initial simulation must be repeated. Focused on systematizing and distributing the Red Teaming format, the DESSRT framework reduces the marginal cost of additional iterations, thus making it highly scalable. In fact, once instruments are created, the number of simulations can be scaled up as needed, limiting additional costs to recruitment, management and (sometimes) compensation of additional human participants. The simulation itself can be easily iterated dozens and even hundreds of times, allowing for replicability of findings and the experimental manipulation of one or more factors (input variables) in the simulation. A Red Teaming simulation can also be scaled across teams or units, allowing for multiple institutional as well as individual perspectives regarding vulnerabilities or other concerns for the problem at hand.
Ultimately, the DESSRT framework is embedded in a set of principles (distributed, empirical, structured, and scalable) that can propel Red Teaming simulations from mere exercises to a structured research method, increasing their ecological validity (Lin-Greenberg et al., 2021) and expanding the problem sets to which they can be applied. This is not the first instance where Red Teaming-like activities have been leveraged as a research tool, since others (Romyn & Kebbell, 2013, 2017) have utilized simulated role-playing to understand potential terrorist target selection. Yet prior implementations have been localized into an in-person laboratory environment, which limited their ability to draw on diverse viewpoints. Moreover, scholars and practitioners (Hoffman, 2017; Scott, 2020; Zenko, 2015) have advocated for Red Teaming to become a regular, institutionalized process within organizations, yet without access to the structured, empirical, and distributed Red Teaming described here, those processes are difficult to deploy across an organization or repeat consistently over time.
In the remainder of this article, we describe one application of the DESSRT framework, focusing on tactical-level adversary decision-making within an aviation security context. Specifically, we investigate whether tactical choices in weapon, security evasion, and target decisions shift in response to new information about specific defensive postures; namely the introduction of computed tomography (CT) scanners (Greenberg et al., 2020; Zhang, 2019). In so doing, we demonstrate our contention that DESSRT has expanded and enhanced the utility of Red Teaming for both conducting research and informing policy decisions. We summarize the exercise design below, followed by a discussion of the initial empirical results and conclude by connecting back to the original DESSRT principles.
Methods
We demonstrate the DESSRT framework within the context of aviation security, evaluating how simulated adversaries would respond to governmental information disclosures about new planned security measures. It is hypothesized that modifications in security installations can influence adversary behavior in one of two ways: 1) information about where the new security measures are deployed influences the adversary’s perceived likelihood of apprehension; and 2) the depth of information about the security measures influences the adversary’s level of uncertainty about the effectiveness of the security measure (and affects their perceived likelihood of success). In our case, the specific focus was the installation of new CT scanners in airport passenger screening, and whether the simulated adversary would have access to either limited or detailed information about the scanning technology.
Instrument
The DESSRT approach used in this application was structured as a two-phase simulation of adversary decision-making (given the title Operation Chameleon Fire), with an experimental inject provided between the two phases. The activities were divided into five sequential segments: 1) Preliminary Activities, 2) Phase 1 Attack Planning, 3) Inject, 4) Phase 2 Attack Planning, and 5) Debrief. Overall, each participant was required to input at least 470 distinct pieces of information (data points) during the simulation (which ran approximately 8 hours). Given the complexity of the exercise design, and the length of time needed to complete all phases, participants were allowed up to two weeks from their starting date to complete the exercise. As an asynchronous, individualized process, participants were able to sign in and sign out via a unique code, allowing respondents from multiple geographic areas to complete the exercise while adapting to changing COVID-19 protocols in their area. A complete copy of the simulation instrument is provided in the Supplementary Materials.
Step 1: Preliminary Activities
Original Historical Cases.
Historical and Simulated Case Alignment.
Pre-Exercise Activities: Prior to attack planning, participants completed a short demographic survey, a series of awareness exercises regarding potential cognitive biases prevalent in Red Teaming, and a Cognitive Reflection Test (CRT; Frederick, 2005) to capture potential differences in decision-making processes between respondents. 3
Scenario Assignments: Participants were randomly assigned to one of the three constructed cases, followed by a cognitive association exercise to encourage perspective-taking 4 with their simulated adversary (Batson et al., 1997). In Red Teaming exercises that rely on small-group decision-making, individuals deliberate with other participants, discussing strategies and tactics for attack planning during structured (usually) tabletop sessions. Since the current study was conducted asynchronously, we regarded the participant as a unitary decision maker. This was the preferable approach for the unaffiliated adversary, but for the group-based adversaries we placed the participant in the role of a cell leader that was assumed to have autonomy in tactical decision-making so long as he/she abided by the constraints given in the scenario, which represented the larger goals of the organization’s leaders.
To support attack planning, participants were presented with additional information related to their specific scenario and then undertook a structured process for planning an attack on their assigned aviation target (i.e., airport). Specific scoping requirements were provided related to pathway (through a passenger screening checkpoint) and weapon choice (required to use at least one explosive) to examine the impact that the CT scanners would have on adversary decision making (see Supplementary Materials).
Step 2—Phase 1: Initial Attack Planning:
To aid attack planning, participants completed three 30-minute planning sessions, each coupled with an untimed journal entry (written in the first person from the perspective of their assigned perpetrator) to document their planning processes, information sources consulted and the status of their plot. Each session included note-taking space, asked respondents to list the top 10 websites that were most important for planning during that session, and presented questions on planning progress in three areas: explosive, weapon package, and circumvention efforts. Individuals were allowed to access the outside internet to aid in their planning.
Attack Plan: Participants were then asked to provide an overview of their attack plan (300–500 words), detailed information on key dimensions of operational significance and their justification both for these choices and for their rejection of other options considered, including: 1. Weapon Consideration and Selection
5
: First, participants identified all the explosive types they considered for use, the specific explosive type(s) they selected as well as the amount of explosive(s) selected. Second, participants recorded all the different weapon package combinations (explosive, trigger, detonator, power source, and housing) they considered for their attack, as well as those selected. Finally, participants were asked for their reasoning in selecting both the specific explosive and weapon packages. 2. Weapon Acquisition and Assembly: Regarding acquisition, participants were asked how they planned to acquire the weapons and materials, as well as what sources (books, websites, articles, or other communication channels) were used to identify potential acquisition pathways for the different components. Concerning assembly, participants described how they assembled the weapon, including any deviations or new expectations from additional knowledge gained during planning efforts. 3. Circumvention / Security Subversion Consideration and Selection: Since the exercise required conveyance of the weapon through passenger screening, participants had to consider their process for circumventing, subverting, or deceiving airport security measures. Scenarios provided to participants contained information on airport security measures equivalent to that which could be extracted via repeated reconnaissance visits to public areas of the airport. Using that information, participants were required to identify all considered options to circumvent, evade, or otherwise subvert airport security installations, as well as the selection and reasoning for their final selection of subversion method(s). 4. Additional Details: To capture other attack-related factors, participants were asked whether, and for what reason, any other weapons (outside of the required explosive weapon) were selected, including knives, firearms, and incendiary devices. Unlike explosives and weapon packages, other weapons that were not selected were not required to be recorded. Finally, participants were required to note what measures they would take to protect vulnerable critical information prior to the attack, such as electronic communications, surveillance activities, or bank account activity.
Step 3—Experimental Treatment:
Once initial attack planning was completed, participants were randomly assigned one of three injects providing distinct levels of information on a new screening method (CT scanning equipment) that could be disclosed by the government regarding the deployment and efficacy of CT scanners (see Supplementary Materials): a. Low-Specific (LS): Simulated information releases about airport business and/or aviation security that contained low levels of detail (two sentences) on CT scanning equipment and only mentioned its installation at the original targeted airport (with no information about any other airport installations). This was included among several unrelated pieces of aviation-related information (which functioned as “noise”). b. High-Specific (HS): Simulated press release with detailed information (about two pages) about the installation of new CT scanning equipment at the original targeted airport. The press release was modeled on an actual Transportation Security Administration (TSA) press release about CT scanning capabilities, and no mention was made regarding installations at other airports. c. High-All (HA): Same press release as the High-Specific treatment, but with additional information that the technology would be installed at all airports.
Random assignments were balanced across assigned adversary/scenario and experimental treatments to ensure equivalent sample sizes in each scenario-treatment combination. In addition, participants were given the option to change their targeted airport in Phase II, with additional briefing materials provided for all three of the simulated airports (Danbury, Meroxia, and Studebaker).
Step 4—Phase 2:
Revised Attack Planning: Following the inject, participants engaged in two subsequent planning sessions of 30 minutes each, coupled with journal entries, similar to those previously conducted. Participants were provided with their original attack plan, and then after completion of the planning sessions were asked whether they chose to modify their target, their choice of explosive, weapon package or security subversion technique. To address potential anchoring bias (Tversky & Kahneman, 1974), participants were required to supply their reasoning for each decision, including for those aspects where they decided not to make any changes.
In addition to many of the items collected for the initial attack plan, participants also needed to address the following areas: a. Target Change: Participants were able to change their initial airport target (assigned to align with the historical case studies) to one of the other airport targets if desired. Participants were also asked to justify their final airport target choice, as well as specify whether the additional information provided in the treatment encouraged them to change their target. b. Explosive and Weapon Package Changes: Participants were asked for their final explosive and weapon package selections, their justification for those choices, and whether they gathered additional information for the new selection(s). c. Security Subversion Changes: Questions related to their final selection of security evasion methods were also asked, as well as their justification of those choices and whether they gathered additional information for the new selection(s). d. Additional Changes and Justifications: Participants were also asked about their final weapon procurement and assembly decisions, operational security procedures, and other weapon selections, and had to provide their justifications for those choices and any changes made.
Step 5—Debrief:
After the conclusion of the simulation, participants were provided a short follow-up survey regarding their experience with the exercise, recommendations for improvement, and whether they believed that these types of exercises could be helpful to operational government agencies.
Sample
The simulation was administered asynchronously via the Qualtrics survey platform from March 3rd, 2020, to June 10th, 2020. Participants were drawn from a university in the Northeastern United States, recruited in three ways: 1) class credit was offered for majors and minors within domains relevant to security studies; 2) degree training hours were provided for those outside these courses who completed the red teaming exercise; and 3) general interest recruitment from other students on campus. As a pilot effort, sampling procedures were not intended to produce an otherwise generalizable sample. Unlike the historical cases that the scenarios build upon, our sample was not restricted to males, allowing attack planning processes to vary across gender identity.
Once identified, our initial pool included 189 potential participants. After initial receipt of the simulation instructions, roughly 10% of those recruited never opened the simulation (16 participants), one participant opened the simulation but did not progress further, and a further 9% started the initial simulation activities but did not progress to attack planning. 6 Overall, 79% of those initially recruited completed the exercise (150/189), with an additional seven removed for either insufficient responses (missing data across multiple variables, lack of explanations for decisions made) or for non-compliance with the exercise requirements (participants were required to proceed through passenger screening and several attacked prior to screening or chose checked baggage).
Sample Descriptives.
Participant Safety and Study Security
The study protocol was approved by the Institutional Review Board at the University at Albany (Protocol/Study #20E092). Two key concerns regarding participant safety emerge when conducting an asynchronous Red Teaming exercise. First, attack planning requires the use of the internet to research weapons, tactics, and technologies that participants may use in their plans. These searches may use words, phrases, or terms of concern to law enforcement and/or intelligence officials monitoring specific search strings used online, especially related to specifics about weapon construction and development. Due to COVID-19 protocols, in-person laboratory environments were not available to limit the internet research to a specific machine or IP address in a controlled environment. To minimize any risks to participants during the exercise, their participation in the exercise (names and email addresses), which were separated from any participant responses, were shared with a network of federal, state, and local law enforcement partners stating that these individuals were participating in an approved Red Teaming exercise. At no time were specific responses submitted during the exercise shared with the notification network, the participant identification information was never stored in the same file location as the de-identified responses, and participants were made aware prior to consent that their names and email addresses would be shared to ensure participant security during the exercise. In addition, project staff were also available via email to allay any concerns that participants had during the conduct of the exercise.
Second, there are concerns that simulating the role of an adversary could create psychological distress among participants, especially when planning a simulated attack. Although online surveys are easier to abandon than face-to-face studies due to reduced social pressures (Sproull & Kiesler, 1991), monitoring participant responses online during the exercise becomes more difficult (Kraut et al., 2004). To ameliorate these concerns, participants in the exercise were not penalized for early withdrawal from the exercise, and information was provided in the study notification for individuals feeling emotional distress to stop the exercise immediately and seek psychological counseling if needed.
Finally, for reasons related to national security, the research team did not assess the attack plans developed by study participants for viability or probability of success. Doing so could place the researchers in possession of potentially security-sensitive information, precluding the public dissemination of results. Given that one of the goals of the simulation was to demonstrate the utility of the DESSRT framework to a wider scholarly and policy audience, it was necessary to avoid these complications altogether by placing the viability of specific attack plans outside the scope of the study. Therefore, the plots resulting from the simulation are not necessarily viable or likely to be successful if executed; indeed, the opposite may be the case. It should be noted, however, that the results could be evaluated for their technical accuracy by appropriate authorities, or the experiment conducted within a classified environment, without changing the central dynamics of the DESSRT framework.
Statistical and Qualitative Analysis
Although a pilot effort, we do provide some analysis to highlight how Red Teaming according to the DESSRT framework could be used to generate empirical findings related to adversary decision-making. All data were downloaded from the Qualtrics platform at the close of the exercise, cleaned, and processed in Microsoft Excel to generate descriptive insights. Many variables were directly quantitative or categorical in nature, while for some free-text fields (specifically the reasoning fields), a team of project coders reviewed each of the reasons provided by participants and inductively determined a coding schema for specific reasons. 7 Participant responses could be associated with more than one reason for their selection, and coded variables were double-coded to ensure consistency across coders. These and other free-text fields were also analyzed qualitatively to identify specific themes associated with adversary decision-making that emerged from participant inputs. Finally, standard methods for testing the presence of statistically significant differences between observed and expected frequencies (chi-square tests) and two proportions (Z-test) were used.
Results
To demonstrate the potential value of the DESSRT framework, we present an illustrative sample of four types of substantive results that can help address the primary operational questions related to tactical choices and CT scanning posed earlier.
Descriptive Operational Results
The DESSRT simulation provided a sufficiently large sample to examine the relative frequencies of potential adversary choices across various tactical decisions. For example, concerning explosive type, participants overwhelmingly preferred secondary high explosives (64) over primary high explosives (35), low explosive materials (16), and other explosive types (9). Moreover, as Figure 1 reveals, the simulation was able to provide fine-grained tactical details about the choice of weapon components, in this example trigger mechanisms, where combustible fuses (37) were preferred more often than electronic switches (25), direct flames (20), or remote signals (20). When asked about security evasion approaches, less than 10 percent of participants chose to use TSA Pre-Check as a method to secure decreased screening requirements, and only 14 percent opted-out of scanning during passenger screening. As operational findings, results from the DESSRT approach can thus prove useful to practitioners who need to prioritize training, screening procedures and security measures against specific threat vectors. Frequency of Trigger Mechanism Selected by Participants. Note: Total number of trigger mechanisms can exceed number of participants, as multiple weapon packages could be used by a single participant.
Variation by Assigned Role
The DESSRT approach revealed significant variation in tactical decision making across the assigned adversary. For example, when it came to how participants sought to subvert existing security measures, significant differences were found across the three adversaries for tampering with screening devices (χ2 = 10.27, df = 2, p < 0.01), opting out of the scanning process (χ2 = 28.50, df = 2, p < 0.01), and bypassing the screening / scanner altogether (χ2 = 9.47, df = 2, p < 0.01). Specifically, those assigned the Unaffiliated adversary (an idiosyncratic lone actor) were more likely to select methods that opt-out or bypass specific screening requirements when compared to those assigned the other two (well-resourced) adversaries.
Treatment Effects
The simulation demonstrates that information about CT scanners had differential effects on adversary adaptation in Phase II of the exercise. The majority of participants (53%) changed one or more of the four key tactical decision areas (target; explosive; weapon package; subversion technique), which did not vary appreciably across treatment. For example, although there was some variation in the proportion of participants that changed their targeted airport, ranging between 17% for the HS treatment to 21% for the HA treatment, this difference was not statistically significant (χ2 = 0.229, df = 2, p = 0.89). Information differences did inform weapon selection, with 23 percent (33/143) of participants overall changing either weapon package or explosive. Those receiving the HA treatment were most likely (33%) to change, with 21% changing some component of their explosive and 29% changing some component of their weapon package. This is expected, as participants receiving this treatment would know the novel technology in depth and that it will be at all airports, reducing uncertainty about likelihood of detection with these new technologies. Those receiving the LS treatment were least likely to change explosive or weapon packages (17% overall), although the difference between all three treatments was not statistically significant (χ2 = 4.36, df = 2, p = 0.11). However, when HA was compared with a pooled sample receiving the HS and LS treatments, those differences (33% for HA; 18% for HS+LS) were statistically significant (z = 2.07, p = 0.04), suggesting that some aspect of the HA treatment was encouraging explosive or weapon package changes.
Sample Experimental Effects.
*Of those who changed targets
Note: % are cell percent
Qualitative Decision Insights
Qualitative analysis of why participants changed (or did not change) targets can provide additional insight into how individuals process new information when making tactical choices. Below, we present two responses by participants assigned the same adversary, receiving the same treatment, and who both changed their targets. First, Participant 1 stated: “After looking at the different security measures at each airport I have decided to change location because [the new airport] only uses explosive trace detectors at screening stations to test baggage that has already been flagged as suspicious on X-ray. They also do not use millimeter wave full body scanners for passengers on any international flight on a U.S.-based carrier. No explosive detection canines were seen at the airport, but it is not 100% that the airport does not use them. Although CCTV cameras are everywhere, this airport has procedures that I will be able to work around better than the initial airport of choice.”
In comparison, Participant 2 changed targets because: “The reason I changed the airport is because [the new airport] seemed to have similar measures as [the original airport] but it gave us more advantage because of their lack of use of millimeter wave full body scanners. So, trying to circumvent the supplies past the security would be easier / more likely to happen.”
For both participants, key details about the limitation of current millimeter wave full body scanners (not allowed for any international flight on a U.S.-based carrier) were cited as reasons for selecting the new airport. Yet one participant pulls in additional information (use of explosive trace detectors; explosive detection canines; presence of CCTV cameras) that another does not reference, even though both were presented with the same parameters and background information. Although these participants recorded the same response in the exercise, the similarities and differences in their reasoning and processing of scenario information remains an important avenue for additional research using the DESSRT framework.
Discussion
Overall, the DESSRT framework demonstrated here provides a novel mechanism for translating traditional Red Teaming exercises into a replicable and empirical research method. The illustrative study addressed an operational question of interest for aviation security, where historical data is non-existent and parameterized assumptions within advanced simulation models are unverifiable. By not only utilizing human participants to role-play as potential adversaries but also to study their decision-making processes as they engage, the application of Red Teaming at scale demonstrated here could contribute relevant and generalizable insights to many vexing questions. Although not a replacement for historical data on past behaviors, where this is absent or inaccessible Red Teaming at scale allows analysts and researchers to test theories about human decision-making, generate novel what-if insights to support planning efforts, and validate parameters within complex simulations and models.
These potential advances still require that the DESSRT framework is viable, replicable, and relevant. Regarding its viability, the above experiment involved almost 150 participants contributing up to eight hours per participant in a distributed online simulation occurring during a global pandemic. Participants engaged from multiple states and across different times of day, enabled by the asynchronous platform utilized in this study. Moreover, this was accomplished without a single reported negative human subjects outcome, which implies that the risks to participants, at least with the precautions described above, is no more than minimal, in line with most survey-type research. With respect to participants’ subjective experiences of the simulation, 94 percent of participants reported that this was their first interaction with this type of ideational, asynchronous Red Teaming and that the instructions were clear in the exercise. Most participants (79 percent) reported that the experience was positive, and an even higher proportion (84 percent) agreed that federal agencies (like DHS and TSA) would benefit from these exercises. Interestingly, 83 percent (119/143) reported that they had enough time to complete the exercise within the total 8-hour window.
Concerning its replicability, the approach provided here utilized a commercial-off-the-shelf technology (Qualtrics survey platform), which is FedRAMP compliant (so can be used within secure environments), accessible to governmental and industry users, and widely available in many academic institutions. The convenience sample of college students used in this pilot study is commonly employed in psychological and criminological research on decision-making, providing a familiar starting point for comparison and extension within future research. Although adhering to the best practice principles of Red Teaming, input item construction followed conventional standards for both qualitative (free text) and quantitative (numeric) data collected, aligning with general accepted practices within the survey research community.
Finally, regarding relevancy, although we discuss only a fraction of the results obtained from the study, those presented offer strong indications that the DESSRT framework provides operationally useful results. This includes both the large-N quantitative results and the individual qualitative innovations generated by participants. Indeed, informal feedback from project sponsors suggests that several of the results provided much-needed guidance to their risk assessment and policy deliberations. The results also indicate how the DESSRT approach can provide insights that are not available from traditional Red Teaming, such as exploring variation between different initial conditions (adversary characteristics), experimental treatments, and even diverse types of participants. All these analyses were possible without sacrificing the core advantages of Red Teaming, which include in-depth role-playing from the adversary’s point of view and careful consideration of the decision process.
Limitations and suggestions for further future research
There are some limitations to this study, however, which can be addressed by future research. First, although the simulations used here were grounded in historical cases of attacks on aviation transportation, questions remain whether findings from role-players are generalizable to actual adversarial behavior. Concerning wargames, Lin-Greenberg and colleagues (2021) note that their ecological validity is derived from “simulation conditions that reflect the types of pressures, incentives, and information environment” which decision-makers face in each situation (p. 6). Our simulation focused on emulating these characteristics for participants, including role immersion, deep background on adversaries, detailed schematics of the target environments, realistic requirements for planning operations, and interactive decision-making regarding the addition of CT scanning. Even so, questions remain whether the attack planning and tactical decisions would reflect actual adversarial behavior, whether generated from historical data or advanced decision-making models. Given the limited occurrence of explosive attacks against aviation targets in the United States, future research should compare the results of this simulation with both historical data and common utility models used by operations researchers to model adversary decision-making.
Second, beyond the general ecological validity of the simulations, there are still many questions regarding the optimal make-up of the DESSRT participant sample. As a novel experimental study, we used a convenience sample of college students, and we did not select for other attributes, in particular the demographic characteristics of the adversaries themselves, such as gender, age, ethnicity, or religion. Future research could investigate developing broader demographic samples that align with distributions of individuals known to act as violent adversaries. Datasets such as the Profiles of Individual Radicalization in the United States (PIRUS; LaFree et al., 2018) could be used to develop demographic distributions used for sample recruitment in future efforts.
Third, although the DESSRT exercise presented here was able to collect data from 143 participants, the 3 (ideology) x 3 (CT info/availability) nature of the experimental design meant that the number of participants in each experimental category was too small to conduct certain analyses. Future research could focus on either limiting the scope of the experimental design or obtaining a larger participant pool. Future research could also assess the optimal number of participants assigned to each scenario and treatment and thus might be able to cover more treatments with a given number of participants. Of course, this must be balanced against the need to achieve sufficient statistical power, depending on the analysis to be conducted.
Conclusion
Application of DESSRT Framework.
The successful completion of the simulation, including the generation of multiple operational insights and the investigation of three different experimental treatments, provides strong support for the framework’s viability and replicability. Although much work remains to be done in refining the technique and establishing optimal parameters for participant samples, this approach holds much promise for leveraging the benefits of large-sample studies, while retaining the essential advantages of Red Teaming for mitigating biases and providing alternative perspectives in simulating adaptive adversary decisions.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the U.S. Department of Homeland Security under Grant Award Number, 17STQAC00001-03-03.
