Abstract
Objective
This study’s purpose was to better understand the dynamics of trust attitude and behavior in human-agent interaction.
Background
Whereas past research provided evidence for a perfect automation schema, more recent research has provided contradictory evidence.
Method
To disentangle these conflicting findings, we conducted an online experiment using a simulated medical X-ray task. We manipulated the framing of support agents (i.e., artificial intelligence (AI) versus expert versus novice) between-subjects and failure experience (i.e., perfect support, imperfect support, back-to-perfect support) within subjects. Trust attitude and behavior as well as perceived reliability served as dependent variables.
Results
Trust attitude and perceived reliability were higher for the human expert than for the AI than for the human novice. Moreover, the results showed the typical pattern of trust formation, dissolution, and restoration for trust attitude and behavior as well as perceived reliability. Forgiveness after failure experience did not differ between agents.
Conclusion
The results strongly imply the existence of an imperfect automation schema. This illustrates the need to consider agent expertise for human-agent interaction.
Application
When replacing human experts with AI as support agents, the challenge of lower trust attitude towards the novel agent might arise.
Keywords
INTRODUCTION
Novel decision support systems based on artificial intelligence (AI) are making inroads to a plethora of application domains. When introducing such systems into any given work environment, the main goal is to improve overall safety and performance (e.g., Mosier & Manzey, 2020). As such, these systems become more and more sophisticated, and can sometimes even outperform human experts (e.g., Bejnordi et al., 2017; McKinney et al., 2020). Given the wide range of application domains of AI support (Islam et al., 2022), it is clear that AI support will in some cases replace support that was previously given by a human. Thus, it is important to understand human–AI trust in order to foster successful interaction dyads with respect to joint performance. Specifically, one important issue is the potential difference in trust towards AI compared to support by other humans.
Prior research has investigated to which degree trust attitude (Lee & See, 2004) differs towards automated support compared to human support (Madhavan & Wiegmann, 2007). One key finding from this earlier research is the so-called perfect automation schema, as brought forward by Dzindolet et al. (2002). The perfect automation schema proposes that without failure experience, humans have higher trust towards automated support compared to support by other humans. Moreover, Dzindolet et al. (2002) also suggest that after experiencing a failure by the respective support agent, trust decreases more for automation support than for human support (illustrated in Figure 1(a)). They argue that this difference is due to expectations of near-perfect performance when working together with automated systems, but that there is no such expectation when working together with another human due to the awareness of human fallibility. Key predictions from the perfect automation schema (a) and from recent results in line with an imperfect automation schema (b).
However, recent findings challenge this often referred-to effect. Specifically, in numerous experiments with relatively large sample sizes, Rieger et al. (2022), Roesler et al. (2022), and Appelganc et al. (2022) have consistently shown higher trust attitude towards humans than towards automated support systems, irrespective of failure experience. For instance, Rieger et al. (2022) conducted three experiments in different contexts (i.e., loan assignment, X-Ray evaluation, chemical plant) with the consistent finding of higher trust towards human support agents than towards decision support systems or AI. Moreover, in these studies, trust was reduced after failure experience, in line with theoretical frameworks of human-automation interaction (Hoff & Bashir, 2015). Crucially, though, this reduction in trust did not differ between automated and human support agents (Rieger et al., 2022). It seemed like participants were not more forgiving towards humans than towards technical support agents after failure experience in these studies, in contrast to the expectations from the perfect automation schema. Given these findings, Rieger et al. (2022) coined this pattern of results as an imperfect automation schema (illustrated in Figure 1(b)), given the fact that trust was consistently higher for human support agents.
To get a better idea about which subdimensions of trust are relevant for the above-mentioned differences, Roesler et al. (2022) further looked into the trust difference between humans and automation, using a multidimensional trust questionnaire. They found that the overall trust difference was mainly related to the performance dimension, that is, that participants perceived human support to be more reliable than automated support. Moreover, Appelganc et al. (2022) ruled out the possible explanation of differences in expectations towards what one perceives as “reliable” for the respective support agents. These combined findings by Rieger et al. (2022), Roesler et al. (2022), and Appelganc et al. (2022) therefore raise the question of why there was such a difference to the perfect automation schema as proposed by Dzindolet et al. (2002).
When taking a closer look at the experiments providing evidence for the perfect versus imperfect automation schema, it becomes clear that the framing of the human support notably differed between these two lines of research. Specifically, whereas Dzindolet et al. (2002) framed the human support to be the former participant (i.e., a novice), Rieger et al. (2022), Roesler et al. (2022), and Appelganc et al. (2022) framed the human support as an expert in the task (i.e., to be more comparable to the technical support conditions). Of course, in any professional setting, it seems more likely that a human support agent is an expert in the task. Although this rationale might have made sense to design the earlier studies by Rieger et al. (2022), it is of course also an inconsistency with the research by Dzindolet et al. (2002). Thus, the main goal of the present study was to investigate whether the discrepancy between the perfect and the imperfect automation schema comes from the framed expertise of the human support agent. Moreover, the inconsistent findings are not only theoretically puzzling but of high practical relevance, as the dynamics of trust attitude and especially behavior are crucial for practitioners seeking to implement AI agents effectively. If human expertise framing is the decisive point for the (im)perfect automation schema, practitioners can utilize this insight to make informed decisions on how to deploy AI support systems, ensuring that AI agents are appropriately introduced to the interaction dyad. So in the possible absence of the postulated advantage of trust in technologies (Dzindolet et al., 2002), practitioners might even need to implement trust-building measures when introducing AI technologies to jobs previously performed by another human.
In earlier research, the focus was mainly on questionnaire-based outcomes such as trust attitude. Therefore, another goal was to also look at trust behavior. However, given the proposed close link between trust attitude and behavior (Hoff & Bashir, 2015; Lee & See, 2004), one would expect the results to be extendable to a more behavioral outcome.
The Present Experiment
We conducted an experiment using the medical X-ray screening context that was also used in earlier research (Appelganc et al., 2022; Rieger et al., 2022; Roesler et al., 2022). We chose this task because novel AI systems can already compete with the performance of highly specialized human experts in this domain (Bejnordi et al., 2017; McKinney et al., 2020). Three groups of participants carried out the simulated X-ray task, with the support agent experimentally manipulated between the three groups to be framed as either an AI, a human expert, or as a human novice. The AI was framed to be a capable system based on deep learning, the human expert was framed to be a medical expert from a clinic, and the novice was framed to be a former participant (in line with Dzindolet et al., 2002). Moreover, we varied failure experience within-subjects to investigate whether the forgiveness towards different agents differed. Specifically, to understand the dynamics of trust, in the first block, participants always received correct recommendations. In the second block, the agents’ reliability dropped and the agent gave several imperfect recommendations. Then, in the third and final block, no more failures occurred and only correct recommendations were given by the agents. We used this design to be able to study trust prior to failure experience (i.e., trust formation), after failure experience (i.e., trust dissolution), and after having experienced a perfect agent again (i.e., trust restoration). Overall, this design allowed us to investigate overall effects of support agent on trust as well as potential differences in failure-induced trust changes between the agents. In order to consider both trust attitude and behavior, we assessed trust not only via subjective ratings but also by a behavioral measure (i.e., how strongly participants followed the support agents’ recommendations).
Our hypotheses were based on earlier research (Appelganc et al., 2022; Rieger et al., 2022; Roesler et al., 2022) as well as on research on the perfect automation schema (Dzindolet et al., 2002). Hypotheses were preregistered via the Open Science Framework (OSF, https://osf.io/z73j2). We had the same hypotheses for trust attitude and trust behavior, as trust attitude is commonly assumed to go hand in hand with behavior (Hoff & Bashir, 2015; Lee & See, 2004). Thus, we will introduce them jointly here. We hypothesized that trust would be higher for the human expert than for the AI than for the human novice support. Moreover, we hypothesized that trust would be higher for the initial, perfect block than for the subsequent block with failure experience, and that trust in the final, back-to-perfect block would be higher than the imperfect, but lower than in the initial perfect block. Further, we expected an interaction effect of agent and block, based on the idea that humans are more forgiving towards other humans (Dzindolet et al., 2002; Madhavan & Wiegmann, 2007). Specifically, we expected a stronger decrease after failure experience in the AI compared to the human expert condition, with only a slight decrease in the novice condition (as expectations are likely low in this condition anyway). Moreover, we expected a stronger increase in trust after the final, back-to-perfect block in the human conditions (i.e., expert and novice).
We also included perceived reliability as an additional dependent variable, with similar hypotheses (please refer to the OSF). Perceived reliability was mainly included to check whether trust attitude and behavior is also related to perceived reliability, as suggested by Roesler et al. (2022). In sum, the goal of the present study was to investigate whether the (im)perfect automation schema depends on the level of expertise of the human support in the task.
METHOD
This research complied with the tenets of the Declaration of Helsinki, and the experiment was approved by the local ethics committee at the Department of Psychology, Technische Universität Berlin, Germany. Informed consent was obtained from each participant. The experiment was preregistered (https://osf.io/z73j2), and the data and scripts are available under https://osf.io/75s2z/.
Participants
Overall, 110 participants (79 female, 28 male, 3 nonbinary) took part in the experiment and were considered for further analyses. Participants were recruited via the participant portal of the Department of Psychology at TU Berlin and online postings. They ranged in age from 18 to 40 (M = 26.16, SD = 4.36) and reported normal or corrected-to-normal vision as well as fluency in English. Target sample size of 108 was preregistered based on a power analysis to detect a small to medium effect size of .15 at the standard .05 alpha error probability for the within-between interaction effect of the two experimental factors.
In addition, 31 participants were also tested but excluded from further analyses. Out of those, 14 participants were excluded because they failed one or more attention checks. Moreover, 15 participants were excluded because they reported very good knowledge about AI, as this might have possibly undermined the believability of the experimental task. Finally, two participants were excluded because they provided the same initial rating in the task 20 times or more throughout the experiment.
Apparatus and Stimuli
The experiment ran on a JATOS (Lange et al., 2015) server and was built using jspsych (de Leeuw, 2015). Participants could individually run the experiment in their browser. The images used in the X-ray task contained simulated 1/f 3 noise, which resembles the power spectrum of mammograms (Burgess et al., 2001) and was used already in earlier research (e.g., Rieger et al., 2022; Rieger & Manzey, 2022; Sha et al., 2018).
Procedure
Participants gave their informed consent, and they were also briefed about the aim of the study. Subsequently, the experiment started with a general introduction to radiology and the task participants were about to perform. The task was the same task as in Rieger et al. (2022, Experiment 2). Specifically, the task was to evaluate which percentage of simulated X-rays were brighter than a given cutoff (brighter than grayscale value of 150). To allow participants to perform this task, they were first shown a grayscale continuum, with the critical cutoff threshold marked, along with an example image. Participants were instructed that parts brighter than the cutoff could potentially be malignant and that their task would be to estimate the percentage of potentially malignant tissue in X-ray samples. They were also instructed that when the percentage is lower than 15%, there is no reason for concern at all. Further, participants were instructed that in radiology, X-ray images are analyzed in dyadic teams, and that collaborative decision making has been found to improve the quality of health care assessments.
After the introduction to the task, participants were told that during the experiment, they would be supported by either an AI, an expert colleague working at a hypothetical clinic, or by the decisions of a former participant. In all groups, participants were again told that they would work in a dyadic team to improve their decision making. In the AI group, participants were instructed that they would be supported by an AI, which was described as follows: The system is an innovative artificial intelligence based on deep neural networks. It continuously analyzes previous decisions and outcomes. The intelligent system decides by using learned parameters to increase the probability of making the correct assessment. The combination of technologies like machine learning, data analysis, and natural language processing constantly improves the artificial intelligences’ assessments.
In the expert group, participants were instructed that they would be supported by an expert from a hypothetical clinic and the expert was described as follows: Your colleague has been working as a radiologist for many years, and has gained considerable experience. The colleague has worked on many previous cases and knows which areas to look at and rate the relevance of anomalies in a differentiated manner to increase the probability of making the correct assessment. Your colleagues’ decisions are based on knowledge, effort and experience.
In the novice group, participants were instructed that they would be supported by the decisions of a former participant and this was further described as follows: The participants had to make their decisions alone without the help of another participant. The participant from which the decisions are presented to you, has no specific experience in radiology assessment except for participating in the study. The decisions by the participant are influenced by the given instructions about health care evaluations of X-Rays and the given information of the presented case. The same information has been and will be presented to you.
Subsequently, participants were shown four examples (two correct assessments and two incorrect assessments). This was done to illustrate the importance of both correctly estimating low percentages (to avoid unnecessary treatment) as well as high percentages (to avoid missing cancer). Then, participants were again reminded about the type of support they would later on receive. After this reminder, the first attention check was included, asking participants about the type of support agent they would be supported by. Afterward, they were again reminded about the task and the support agent and were told that they would be working on a total of 60 cases, with some questions asked in between.
Following the extensive introduction and framing, the experiment consisting of three blocks with 20 trials each started. Blocks 1 and 3 were with perfect support from the support agent, while in the second block, participants received six erroneous recommendations from their support agent. Three of the six erroneous recommendations were underestimations and three were overestimations of the true value. The wrong recommendations always deviated 15%–25% from the true value (M = 20.42, SD = 3.72). The respective 20 trials per block were counterbalanced across participants.
On each trial, participants were first shown a persona along with an X-ray image. For the personas, we used AI-generated images from Karras et al. (2020). Participants were then asked to fill in their manual assessment (i.e., without the aid from the support agent). The minimum time to fill this out was set to 1.5 s. After entering their own assessment, participants were shown the recommendation of their respective support agent. Recommendations were color-coded (i.e., green: ≤15% malignant, orange: >15% ≤ 50% malignant, red: >50% malignant) and shown below the own initial assessment (see Figure 2). After seeing the advice, participants could then enter their final assessment after 1.5 s, which could be the same as before or adjusted in response to the recommendation of the agent. Exemplary trial from the AI condition.
After each block of 20 trials, participants were asked about their trust in their respective support agent and their perceived reliability of their support agent. These questions were embedded within two distractor items and one attention check item. Note that after the first block, the initial attention check (i.e., which support agent), was also repeated within this block of questions.
After completing the questions subsequent to the third block, participants were debriefed (i.e., that no real AI/expert/novice supported them, that no medical information was shown, and that all cases were made up) and thanked for their participation.
Design and Dependent Variables
Support agent (i.e., AI vs. expert vs. novice) was varied between-subjects, and failure experience (i.e., perfect support, imperfect support, back-to-perfect support) was varied within subjects, resulting in a 3 × 3 mixed design. Subjective dependent variables were trust attitude (single item, from 0 “Not at all” to 100 “Completely”) and perceived reliability (single item, from 0% to 100%). The single item was preferred against a more complex multi-item assessment as earlier research (Rieger et al., 2022) showed a high correlation between the single-item trust and one of the most widely used multi-item questionnaires in the field (Jian et al., 2000). For trust behavior, we calculated the absolute difference between the own initial assessment and the final assessment after seeing the recommendation. If, after seeing the recommendation, participants changed their assessment away from the recommendation (e.g., initial assessment 20, recommendation 29, final decision 15), we used negative scores (in this example, final trust behavior value: −5). If participants over-corrected their initial assessment (e.g., initial assessment 20, recommendation 29, final decision 35), we subtracted the over-correction from the adaptation (in this example, final trust behavior value: 9–6 = 3). Thus, higher values represent higher behavioral trust. Trials in which participants’ initial assessments equaled the upcoming recommendation were excluded from the behavioral trust analysis.
RESULTS
We ran separate 3 (support agent) x 3 (block) mixed-ANOVAs for our dependent variables. The corresponding results are visualized in Figure 3(a)–(c). Whenever the assumption of sphericity was violated for the within-subjects factor, we report Greenhouse–Geisser corrected results. For post hoc pairwise comparisons, we used the Bonferroni–Holm correction as preregistered, and the p-values were adjusted accordingly. Means and standard errors for (a) trust attitude, (b) perceived reliability, (c) trust behavior as an adjustment to the recommendation, and (d) trust behavior as full agreement with the agent, as a function of support agent (human expert versus articial intelligence (AI) versus human novice) and block (i.e., block 1: perfect support, block 2: imperfect support, block 3: back-to-perfect support).
Trust Attitude
For trust attitude, there was a main effect of support agent,
Perceived Reliability
The pattern of effects for perceived reliability directly corresponded to the effects of subjective trust. The main effect of support agent was again significant,
Trust Behavior
We ran a parallel analysis for trust behavior. In this analysis, there was again a significant main effect of support agent,
We included an additional behavioral trust measure (Figure 3(d)), which might be less prone to be affected by learning effects to further investigate and confirm the prior results. Specifically, we checked the percentage of trials where participants’ final decision equaled exactly the recommendation of the support agent, as this indicates full dependence on the recommendation. For this analysis, trials in which the initial assessment of the participant already perfectly matched the agent’s recommendation were excluded (3.30%). A perfect match between the participants’ final rating and the support agents’ recommendation was found in 12.46% of all remaining trials. However, these matches were not equally distributed across the different conditions as indicated by a main effect of support agent,
DISCUSSION
The introduction of AI support systems into various domains has the potential to improve safety and performance. Therefore, AI support is in some cases replacing human support. As trust towards a support agent is crucial for successful human-agent interaction, the main goal of the present study was to disentangle the inconsistencies in the literature concerning the findings supporting a perfect automation schema (Dzindolet et al., 2002) and findings supporting an imperfect automation schema (Rieger et al., 2022). To this end, in a simulated X-ray task, we experimentally varied the support agent to be either an AI, a human expert, or a human novice. At the same time, we also varied failure experience within-subjects (i.e., perfect recommendations in the first block, imperfect recommendations in the second block, and back-to-perfect recommendations in the third block).
The first key finding is that the degree of expertise of human support agents determines attitudes towards and perceptions of human support. In particular, trust was higher for human expert support than for human novice support. More crucially, trust in the AI condition was in between these two human conditions. For human experts, this is in line with findings supporting an imperfect automation schema in different task contexts (Rieger et al., 2022). For human novices, this is in line with findings supporting a perfect automation schema (Dzindolet et al., 2002). Consequently, human expertise seems to be a defining boundary condition for the existence of an (im)perfect automation schema. Given the fact that in most real-world applications, human experts are usually supported by other experts (e.g. two radiologists examining the same X-ray), the AI versus expert comparison seems to be the more relevant one. However, the strength of the (im)perfect automation schema may vary based on different influencing factors, such as task complexity, required user expertise, system reliability, and contextual variables.
The results for perceived reliability also match the trust attitude findings. This is in line with the findings by Roesler et al. (2022), who also reported higher trust towards human experts than towards AI on the performance subscale in their multidimensional trust questionnaire. Despite the consistency across subjective measures, there were no differences in trust behavior between the AI and expert conditions. Even though trust attitude has often been postulated to guide behavior (Hoff & Bashir, 2015; Lee & See, 2004), the present findings only partially support this link. Specifically, lower trust attitude only went hand-in-hand with lower trust behavior for the comparisons involving the novice condition. Perhaps, the lack of differences in trust behavior between the AI and expert conditions might be due to the fact that trust attitude and behavior are not perfectly correlated (e.g., Chavaillaz et al., 2019) and both attitude and behavior have other factors impacting the respective component (e.g., Hoesterey & Onnasch, 2022; Rice & Keller, 2009; Wiczorek & Meyer, 2019). This disconnection between trust attitude and behavior might be rooted in the performance-based components rather than attribute-based components of the framing, as both the AI and the human expert agents were framed as having expertise in the task (Langer et al., 2021; Rieger et al., 2022). However, as one reviewer suggested we also want to acknowledge the possibility of both Type I and Type II errors in the current results. Both types of errors may have contributed to the lack of association between trust and behavior, which should be therefore investigated more closely in future research.
The second key finding was that there were no differences in trust restoration between the support agents. We observed the typical pattern of trust formation, dissolution, and restoration (Lewis et al., 2018) for trust attitude and behavior as well as perceived reliability. However, this pattern did not differ between the different support agents. This result contributes to a growing body of research (Rieger et al., 2022; Roesler et al., 2022; Roesler & Onnasch, 2020) indicating that there are no differences in forgiveness towards different agents. The fact that the impact of failure experience did not differ between the support agents is again in contrast to the perfect automation schema. The perfect automation schema (Dzindolet et al., 2002) and other earlier research (Madhavan & Wiegmann, 2007) have argued that humans would be more forgiving towards mistakes by human support agents due to the awareness of human fallibility. Even though this assumption might be somewhat reasonable, we did not obtain any evidence in line with this assumption—neither for human experts nor for human novices. This finding may be explained by the idea that agent performance overshadows other features of the agent in this formation-dissolution-restoration process.
When integrating these two key findings, it seems like the predictions of the perfect automation schema (Dzindolet et al., 2002) are limited to situations prior to failure experience where an automated system is compared to a novice human. Moreover, we argue that having a human expert as a support agent is much more realistic in most settings—particularly in domains where AI is recently being introduced (e.g., banking (Bahrammirzaee, 2010; O’Neil, 2020), medicine (Bejnordi et al., 2017; McKinney et al., 2020), or process industry (Mao et al., 2019)). Thus, the imperfect automation schema seems to be more appropriate and makes the following predictions. First, the imperfect automation schema predicts lower trust (attitude and behavior) for technical systems such as AI than for human experts. Second, it predicts a decrease in trust after failure experience, with no differences in forgiveness towards the respective support agent. Third, it predicts trust restoration after receiving perfect support following agent failures, again irrespective of the support agent framing. In sum, given the apparent under-trust towards just-as-capable AI, the predictions of the imperfect automation schema challenge the approach of focusing on automation’s weaknesses instead of its high reliability to achieve adequate trust. Moreover, the substantial gap between perceived reliability (58.81%) and truly experienced reliability (90%) gives further reason to focus not on the weaknesses but on the strengths of a support agent. As this is a common issue in the interaction with highly reliable support (Parasuraman & Riley, 1997; Rieger et al., 2022), there is a need for research that explores ways to counteract the under-utilization of support agents. Moreover, the degree of reliability underestimation seems to be context-sensitive. Whereas earlier research (Rieger et al., 2022) already illustrated the context sensitivity of perceived reliability, the underlying reasons for this are still largely unexplored. One influencing factor seems to be the task difficulty, as perceived reliability is especially lower for tasks that are perceived more difficult by humans (e.g., the radiology task used in this study) compared to easier ones (e.g., loan assignment) (Appelganc et al., 2022). In terms of future research, it would be useful to extend the current findings by examining the influencing factors for the strength of the actual-perceived reliability gap.
Another aspect that should be addressed in future research regards the possible impact of the sequence of assessments on the finding of a perfect or imperfect automation schema. In the recent study, the participants’ assessments were always made before the recommendations of the support agents were presented. This marks another difference to the study of Dzindolet et al., 2002, where the sequence was reversed. Making own assessments prior to or after receiving recommendations from a support agent might have different impacts on information processing. While we are uncertain about how this difference could explain the disparities between our expert condition and the previous findings of Dzindolet et al., 2002, we cannot completely rule out the possibility that variations in the assessment sequence might have played a role in contributing to these differences in some manner. Therefore, further research is necessary to better understand the potential impact of the assessment sequence.
Limitations
The present study of course does not come without limitations. First, the experiment was conducted online and therefore has the typical limitations associated with conducting research online. However, we used strict, preregistered exclusion criteria to ensure data quality, and the effect sizes were relatively large. Second, it is of course impossible to know to which degree participants actually believed the support agent manipulation, a limitation that is shared, however, with a lot of earlier research in this field (Dijkstra, 1999; Dijkstra et al., 1998; Dzindolet et al., 2002, 2003; Madhavan & Wiegmann, 2007). Third, this study only included a first-party perspective (i.e., the perspective of the advice-taker, Langer & Landers, 2021), and earlier research (Rieger et al., 2022) has shown that changing to the second party perspective (i.e., the perspective of being assessed) might make a difference. However, it seems unrealistic to have one’s own X-ray case to be evaluated by a novice, and we, therefore, limited the present research to the first-party perspective.
CONCLUSION
In sum, the present study corroborates earlier findings (e.g., Rieger et al., 2022) that showed higher trust attitude towards human expert support agents than towards AI support agents. Failure experience reduced trust with a trust repair after subsequent failure-free interaction but with no differences in failure effects across the different support agents. Thus, the imperfect automation schema seems more appropriate to describe differences in trust towards human or technical support agents. As AI will ubiquitously be a changing force in the way we work together, human factors research will need to keep up to also consider AI-specific trust-related aspects such as explainability in order to foster human-technology symbiosis.
Author’s Note
The experiment was preregistered via the open science framework (OSF) at https://osf.io/z73j2. The data is obtainable via the OSF at https://osf.io/75s2z/.
KEY POINTS
Trust attitude towards and perceived reliability of the support agents were higher for human experts than for AI than for human novices. Failure experience reduces trust, which can be restored after failure-free interaction, but not differently for the different support agents. The findings are in line with an imperfect automation schema; however, the expertise of the human support seems to be the boundary condition for the (im)perfect automation schema. By acknowledging the existence of an imperfect automation schema, practitioners can address potential challenges of lower trust towards AI agents and implement strategies to bridge the gap between human and AI support agents.
Supplemental Material
Supplemental Material - The (Im)perfect Automation Schema: Who is Trusted More, Automated or Human Decision Support?
Supplemental Material for The (Im)perfect Automation Schema: Who Is Trusted More, Automated or Human Decision Support? by Tobias Rieger, Luisa Kugler, Dietrich Manzey and Eileen Roesler in Human Factors
Footnotes
Supplemental Material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
