Abstract
Team skill metrics were operationalized by translating team constructs to metrics based on observable behaviors. As human coordination with autonomous agents turns to collaboration, humans may increasingly view agents as teammates. This transition will require agents to possess team skills and necessitate appropriate metrics for measuring team skills across human-agent and human-human teams. Thirty-eight teaming metrics were developed across five stages of teaming: preparation, execution, evaluation, adjustment, and team chemistry. Behaviors from 78 multiplayer gameplay videos were coded to establish which metrics could be measured via observable behaviors. An exploratory assessment demonstrated that the metrics captured teaming differences in team composition (human-human teams vs. human-agent teams) and three levels of team expertise. Results suggest that these team skill metrics could aid agent designers in anticipating the team dynamics of humans working with their agents.
Introduction
Increasingly capable autonomous systems take on roles with more authority, intention, and room for independent action. This has given rise to research on human-agent teaming, in which increasingly autonomous systems work in close coordination with humans. As this coordination turns to collaboration, humans may increasingly view the autonomous system or agent as a teammate. For this perception to arise, agents must be designed with team skills in addition to the more traditional task skills (McNeese et al., 2018). Accordingly, current research in HAT has focused on the impact on team performance of team skills such as team coordination (Klein et al., 2004), interaction paradigms (Chen et al., 2018; Tokadli & Dorneich, 2019), or shared awareness (Shively et al., 2018).
Recent efforts have focused on understanding the teamwork functions that an autonomous teammate needs (McNeese et al., 2018). Important factors include understanding the task, communicating with human teammates, and anticipating information needs. In addition, general status updates, repeated requests, inquiries about the status of other players, suggestions, and planning are potential behavioral features of an agent (McNeese et al., 2018). While assessing a team’s performance at their specific tasks is relatively straightforward, it is more difficult to measure the effectiveness of their team skills (Wiese et al., 2015).
The current paper provides an exploratory assessment of a set of team skill metrics. Specifically, team metrics were operationalized by translating team constructs and behaviors to metrics based on observable behavior. Work focused on developing useful team metrics with the greatest potential to inform and assess the development of intelligence-aided systems. These potential metrics were then applied to video gameplay recordings to establish which team skill characteristics could be measured via observable behaviors. Thanks to Twitch, YouTube, and eSports generally, the internet now offers an enormous freely available dataset of recordings. While generic lists of behavioral markers for team skills and emergent states (e.g., coordination, team cohesion, trust) can be drawn from a meta-analysis of academic research papers (Sottilare et al., 2017), this project tested those hypothesized behavioral markers on real data in the eSports domain within videogames involving teams with a mix of human and agents. Human-human teaming (HHT) refers to teams comprised of only humans, and human-agent teaming (HAT) refers to teams comprised of both humans and increasingly autonomous agents.
Related Work
Several issues arise when assessing the team behavior. At the highest level of abstraction, constructs describe teamwork components (Ostrander et al., 2019). Salas et al. (2005) suggested five dimensions of efficient teamwork: team leadership, performance monitoring, backup behavior, adaptability, and team orientation. Constructs are challenging to measure directly and often rely on subjective assessments (e.g., surveys, interviews), which can be subjective and biased (Sottilare et al., 2017; Wiese et al., 2015). Team skills can manifest as a mix of behaviors and emergent states. It may be possible to derive assessments of team skills through observing behaviors, which may offer a more objective, quantitative measurement approach (Sottilare et al., 2017). The unpredictable relationships between individual team members’ task-related skills, teamwork skills, and the team’s experience working together make behavioral markers complex to disambiguate (Salas et al., 2007). If behavioral markers of team constructs could be normalized and quantified in a reproducible manner, it could be possible to establish reliable team metrics for the study of team skills. Standardization of measures or methods would better support the comparison of research results across the community (Münstermann & Weitzel, 2008).
Therefore, behavior must be operationalized into specific metrics that capture team-related events to support team skill assessment. The challenge of understanding teaming becomes further complicated with HATs. Humans with good teamwork skills with other people may not have an appropriate understanding of how to interact with increasingly autonomous systems. Additionally, those struggling with teaming skills may find it challenging to work with simulated agents (Chen & Barnes, 2014). Furthermore, what team skills must autonomous systems have to meet human teammate expectations? The eventual goal of this project is to develop a team assessment framework that can be applied to both human-human teams (HHTs) and human-agent teams (HATs). Such a framework would enable researchers to assess teamwork in a reproducible manner, correlate teamwork and task work, and assess the efficacy of different design approaches to support team performance.
Methods
Figure 1 provides an overview of the approach to developing and assessing team metrics. The effort included developing criteria to select which gameplay videos would be most appropriate to analyze, developing team metrics, and developing gameplay analysis procedures.

The research approach to develop team skill metrics.
Game Selection
Develop game selection criteria
Figure 2 presents four dimensions used to determine relevance to HAT: teaming action stage, teamwork composition, difficulty, and feasibility. Criteria were developed to evaluate each game to rate its potential usefulness for teaming research. For each criteria, levels were assigned points to develop an overall score to rate the viability of a video for analysis.

Eleven game selection criteria with levels for each. The numbers are point values toward a total score to rate the viability of a video for analysis.
Select game videos
An initial set of games were identified based on literature searches, prior experience, game forums, top eSports prize earnings, and previous work (Sepich et al., 2021). Higher earnings usually imply more players play the game, which increases the likelihood that videos can be found to analyze (increased feasibility). This resulted in a list of 105 games, then down-selected to a final list of 31 unique game titles using the game selection criteria.
Collect game videos
Video gameplay footage was collected by searching YouTube. The data set consisted of 78 10-minute videos from the 31 game titles.
Team Metrics Development
Develop teaming metrics
Based on (Rousseau et al., 2016), a teaming action loop was developed to situate the teaming constructs, behaviors, and metrics. The loop consisted of four stages: preparation, execution, evaluation, and adjustment, as well as a parallel process to develop team chemistry. Figure 3 organizes the teaming behaviors, where the number of teaming metrics is denoted in parenthesis.

Teaming action loop to organize team constructs. The number of metrics per construct is in parentheses.
Develop video analysis procedure
The codebook was a collection of concrete, observable events that can be seen or heard within a video. Each event identifies a unique instance or period of a teaming behavior that can be identified and logged. For example, the code “Ask for Input” can be identified as an event when a player or agent asks a question to another player or agent, thus expecting some input or response. The codebook contained 38 codes across the five stages of the teaming action loop (see Figure 3). After a video was identified, it was assigned to a coder. The coder would then create a new file in BORIS’s video annotation tool (Friard & Gamba, 2016). Each video was coded for 10 minutes of gameplay, which led to an average of 201.3 (SD = 86.3) events per video.
Refine codebook
After the initial development of the codebook per Creswell & Creswell (2022), the coding process was refined through a process of double coding and discussion. Two or more coders would code the same video. The resulting logs were compared for interrater reliability (IRR). The coder team discussed any discrepancies and updated the codebook accordingly. Once the process had been established, a final round of multiple coders reviewing the same video was performed to establish IRR. The team used Krippendorf’s alpha because it accounted for both a code’s presence (or absence) and timing (Krippendorff, 2011). Three researchers coded the same four videos. Two researchers coded an additional six videos, resulting in two coders for each. A Krippendorf’s alpha of 0.7 is considered an acceptable agreement rate (Krippendorff, 2004). The IRR for all 10 videos was above 0.7, and together the IRR averaged 81%.
Team Gameplay Analysis
Game team play coding
After establishing the reliability of the coding process using 10 of the videos, the team transitioned to rating the final 68 videos with one coder each, for a total of 78 videos.
Exploratory Data Analysis
The game analysis resulted in a coded set of teaming metric data across a variety of games. A preliminary analysis of two factors is presented in this paper: 1) team composition (human-human teams vs. human-agent teams) and 2) team expertise (novice vs. intermediate vs. expert)
A game was classified as a HAT if one or more active team members were computer-controlled agents that operated independently of human control on similar tasks.
Team expertise was determined based on information from the video. For instance, the video title might say “first-time player.” Some games have player ratings, which inform the assessment of expertise. Teams with a mix of expertise among players were typically classified as intermediate; however most teams were the same skill level.
It is useful to analyze both counts of behaviors as well as the percentage distribution of behaviors because they answer different questions. The overall count indicates the number of teaming events, which indicates how much effort is required. The percentage indicates how different teaming behaviors changed as the team composition moved from HHT to HAT.
Results
Team Composition
Total Counts
Figure 4 shows the average count of all codes between human-human teams (HHT, n = 65) and human-agent teams (HAT, n = 13). The average count of teaming codes in HHT (M = 211.0, SD = 83.8) was significantly higher than in HAT (M = 152.8, SD = 84.0), F(1,76) = 5.20, p = .025, d = 0.69.

Average count by team composition. Error bars are standard error.
Counts by Stage
Table 1 provides the descriptive and inferential results of the average counts per stage. Execution and team chemistry were significantly higher in HHTs than in HATs.
Team composition for each stage by behavior counts. * denotes statistically different results. Effect size is Cohen’s d.
Percentage Distribution by Stage
Table 2 provides the descriptive and inferential results of the percentage distribution between human-human teams (HHT) and human-agent teams (HAT) in each stage. The average percentage of teaming codes in HHT was significantly higher than in HAT in the preparation stage.
Team composition for each stage percentage. * denotes statistically different results. Effect size is Cohen’s d.
The above differences make sense from several perspectives. The overall number of teaming events decreased from HHTs to HATs. Some behaviors involve capabilities agents do not yet possess and thus occur less often in HATs. However, the number of preparation events was maintained between HHT and HAT, indicating that while the overall number of events decreased, preparation required the same level of teaming to be successful. This is why in Table 2, the overall percentage of events spent in preparation increased in HAT; all other stages decreased. Second, agents are simpler in their capabilities than humans, and thus fewer human behaviors were required to communicate with the agent to check what it was “thinking.” Among humans, these regular check-in communications can improve team chemistry.
Team Expertise
Total Counts
Figure 5 shows the average count for levels of expertise. The main effect of expertise on the average count of teaming codes was not significant (F(1,76) = 1.45, p = .242.

Average count by team expertise level. Error bars are standard error.
Counts by Stage
Table 3 provides the descriptive and inferential results of the counts per stage. The main effect of team expertise on the overall average count of teaming codes was significant in the execution stage. Post hoc analyses showed that experts had more teaming events than novices (p = .049).
Team expertise level for each stage count. * denotes statistically different results. A Tukey’s letter report represents the difference between levels.
Percentage Distribution by Stage
Table 4 provides the results of the percentage distribution between the three levels of team expertise. The main effect of team expertise on the overall average count of teaming codes was significant in the execution and team chemistry. In the execution stage, post hoc analyses showed that intermediate teams had more teaming events than novice teams (p = .038). In team chemistry, post hoc analyses showed that novices had more teaming events than intermediates (p = .026).
Team expertise for each stage percentage. * denotes statistical difference. A Tukey’s letter report represents the difference between levels.
Overall, team skill behaviors were not significantly different by team expertise level (although they trended higher for higher team expertise). When looking specifically at the execution stage (which involves coordination among team members), novice teams did significantly fewer behaviors than expert teams. It is possible that novice teams did not understand the game and its complexity and, therefore, engaged in fewer teaming behaviors. Intermediate teams used a significantly greater percentage of their behaviors on execution than novice teams. This suggests that intermediate teams were learning the game complexities and coordinating more with their teammates. However, intermediate teams may not yet optimize such communication, so they may have required more communication than an expert team.
However, the percentage of team behaviors in the team chemistry was significantly higher for novices than intermediate teams. It is possible that novice teams, in their figuring out the game, still need to develop team chemistry to establish common ground, making sure everyone agrees with the answers to questions like “Do we all understand the goal and how we have to meet it?” (trust, cohesion) and “How well did we do just now?” (collective efficacy).
Team Composition by Team Expertise
Total Counts
Figure 8 illustrates the interaction plot of team composition by team expertise level for total behavior counts. The interaction between team composition and team expertise was not significant, F(7,52) = 0.04, p = .962.
Counts by Stage
For total counts per stage, the interaction between team composition and team expertise was not significant for any stage: preparation (p = .36), execution (p = .56), evaluation (p = .50), adjustment (p = .14), and team chemistry (p = .55). Figure 6 illustrates the interaction plots of team composition by team expertise level for each of the five stages.

Average count team composition and team expertise level per stage: (a) preparation, (b) execution, (c) evaluation, (d) adjustment, and (e) team chemistry.
Percentage Distribution by Stage
The interaction between team composition and expertise was significant at the stages of preparation (p = .03) and adjustment (p =.03). The interaction was not significant for stages execution (p = .17), evaluation (p = .18), and team chemistry (p = .23). Figure 7 illustrates team composition by team expertise level for each of the five stages.

Percent distribution for team composition and expertise per stage: (a) preparation, (b) execution, (c) evaluation, (d) adjustment, and (e) team chemistry.

The average count of teaming codes by team composition and team expertise level. Error bars represent standard error.
The distribution of teaming behaviors varied significantly for the preparation and adjustment phases, particularly for HAT teams. In the preparation stage, the novice human-agent team engaged in significantly more preparation behaviors than all other combinations of team composition and team expertise. For a novice team, it is possible that the unfamiliarity of the game and the agents’ abilities required more preparation. In the adjustment phase, intermediate human-agent teams engaged in more behaviors. As teams increase in expertise, they begin to recognize when adjustments are needed. In both cases, it is during stages that require planning or re-planning where the human-agent teams require more teaming behaviors at specific levels of team expertise. This result suggests that HATs may benefit from more initial preparation and training to familiarize humans with agent teammates’ abilities.
Discussion
This research aimed to establish a set of metrics to evaluate team skills and measure the characteristics of team dynamics. Well-established metrics based on observable behavior can evaluate HHTs and HATs similarly. The current work presents initial results to answer questions such as, “Do these metrics make sense? Are they interpretable?” The behavioral-based metrics show both HHTs and HATs carrying out all four stages of team actions and team chemistry development. This effort also offers a robust dataset of 38 team behaviors coded across 78 games. Results varied between variables such as team composition and expertise level.
While a final validation of these metrics will likely include task performance variables, comparisons of objective and subjective measures, and further validation efforts, the current work represents an initial step in developing a set of behavior-based teaming metrics. Such measures can be used to answer questions such as: What team skills does an agent need to enable a HAT to be as effective as an HHT?
The current paper provides an exploratory assessment of a set of team skill metrics. However, more work is needed to validate these metrics. Future work will assess the behavior-based observable metrics with more established subjective metrics to establish a level of consistency. Currently, these behaviors are coded manually, which limits their ability to be used rapidly. Future work will look at ways to automate the process using detection and classification approaches. Currently this work has focused on multi-player, competitive games. Future work will look to expand to other context or domains where human-agent teaming may occur. Teaming research would be accelerated with a set of validated, rapidly assessable teaming metrics based on observable behavior.
Footnotes
Acknowledgements
This project was funded by US Air Force Research Laboratory (AFRL). Statements and opinions expressed in this text do not necessarily reflect the position or the policy of the United States Government, and no official endorsement should be inferred.
