Abstract
Trust in automation depends on more than just the automation itself, but the larger context in which the automation and the human operator are collaborating. This study takes a naturalistic approach to explore providers' trust in a Clinical Decision Support System. Primary Care Providers were shown simulated medical records and a prototype Clinical Reminder indicating that the patient should be titrated with recommended Beta Blockers to address the patient's Heart Failure with reduced ejection fraction. Analysis of responses showed three main themes: Concerns about the medical documentation used to generate the recommendation; Complexity of the patient condition and care delivery context (and how such factors limit possible courses of action); and Concerns about the Clinical Reminder and clinical guideline it is instantiating. These results align with the macrocognitive model of trust and reliance based on sensemaking and flexecution.
Introduction
Trust involves a relationship based on roles and responsibilities, and an implicit agreement (Middleton 2018) for the truster to accept being vulnerable in a specific situation with the expectation that the trustee will behave in a way that benefits the truster (Hall 2001). Thus, trust in automation is about the operator assessing opportunities for collaboration. The operator must be able to evaluate the current capabilities of the artificial agent with respect to the demands of the primary task itself, and to the demands of collaboration. Designing for collaboration with automation means facilitating the ability of the human operator to establish and update justified trust and justified mistrust in the automated agent (Hoffman 2017).
Thus, it is useful to explore trust in automation via actual practitioners performing representative tasks that involve realistic (i.e., imperfect) automated agents. However, most of the literature on trust in automation is theoretical or from experimental studies; with few exceptions (e.g., Dorton & Harper 2022), naturalistic studies of trust in automation have been neglected (Hoffman 2017).
In the domain of medicine, trust in an automated system is manifest in the context of other trust relationships (Montague & Lee 2012; Alexander 2006) such as trust between patient and provider, trust in the organization, trust in practice guidelines, and issues of responsibilities and liabilities. Clinical Decision Support Systems (CDSS) touch all of these different aspects of trust in medicine.
Work focused specifically on CDSS suggests it interacts with trust-associated factors like other cognitive support automation (such as those presented in Hoff 2015; Schaefer 2016). The types of explanations provided by the CDSS affect trust (Bussone 2015). The number of alerts and the rate of false alarms are recognized as issues (Ancker 2017). Operators following incorrect recommendations of CDSS has been identified as a risk (Goddard 2011, Sutton 2020). The operator must evaluate the CDSS for appropriateness of its recommendation in any given patient.
Beta Blocker CDSS as Case for Trust in Automation
This CDSS uses Natural Language Processing and other methods to identify patients who have a particular type of Heart Failure (HF) but are not receiving the type and dosage of beta blocker medication (BB) that is recommended (Stout et al 2019). For such a patient, a component of the CDSS system screens the patient’s chart in the Electronic Health Record system (EHR) for detectable contraindications (e.g., documented allergies, asthma, systolic blood pressure < 90 mmHg) (Smith et al 2018). If no contraindications are detected in the record, it will present a Clinical Reminder (CR) to the practitioner when the patient’s chart in the EHR is accessed. The CR provides information about the guideline and its clinical merit. It also shows the patient’s recent vital signs, weight, current beta blocker prescription (if any), and ejection fraction (which indicates if the patient has the particular type of HF for which this guideline applies).
For the following reasons, this presents an informative case with which to explore issues of trust and automation.
The performance of the automation is not independent of the rest of the system. It depends on the functioning of the EHR (such as how frequently the patient’s chart is automatically updated with information from other parts of the system), and on the documentation entered by providers (e.g., Is the allergy to medication X coded in the correct field or simply mentioned in a free text note?).
The decision to titrate is very complex, involving various risks and trade-offs. Titration of BB is a process that takes several weeks and requires careful (and time-consuming) monitoring of the patient. Even patients without contraindications can experience low energy and other side-effects. Depending on the organizational context, the practitioner may be at risk of over stepping by initiating BB titration. Yet failure to titrate can result in more frequent hospital admissions, and shorter lifespan (Murphy, Ibrahim & Januzzi 2020). This decision may have to be made in a short time frame, with little time for preparation, and with other competing priorities for the 10 minutes available for the encounter with the patient.
This CDSS is designed to facilitate appropriate application of the evidence-based guideline on BB for HF patients with reduced ejection fraction (Yancy 2013). The practitioner’s interpretation of the automation is interwoven with their interpretation of the guideline.
The automation in this CDSS encompasses multiple functions. Following the four automation information processing stages of Parasuraman et al (2000), the CDSS acquires information (extracting it from the patient’s chart), analyzes information (assesses if the patient qualifies for the Clinical Reminder), and does decision selection (making the recommendation for the provider to complete the assessment for BB titration). It even facilitates - but does not initiate - action implementation (titration and referral orders are available from the Clinical Reminder itself).
These functions are ones that the providers have also done themselves and can therefore have some first-hand basis for performance expectations. In contrast, one key function of the automation is to scan through several thousand medical charts to detect patients who meet the criteria for the Clinical Reminder. This function operates at a scale possible only with automation.
Methods
Participants
The participants were all members of primary care teams in two different VA medical centers from the Great Lakes regional network (including some participants from affiliated community-based outpatient centers). There were nine participants: four PCPs (one Nurse Practitioner, three physicians); four clinical pharmacists, one Registered Nurse care manager. Institutional review board approval was obtained for this study.
Protocol
The scenario was that the provider was opening the chart of a patient in preparation for seeing that patient soon. The Clinical Reminder appeared, reflecting the data for that particular patient. The participant was asked to decide how to resolve the Clinical Reminder (i.e., make a decision about what to do regarding the alleged under-treatment of the patient’s HF).
We used a modified Wizard-of-Oz technique (Kelley 2018) to provide a mixed fidelity simulation of the Clinical Reminder and the patient’s chart in CPRS (the EHR of the VA). The participants were asked to think aloud, and were prompted to describe their thinking processes.
Cases
There were two fictional cases, each concerning a patient who was readmitted to the hospital after exacerbation of HF. The cases were generated using de-identified medical records of real patients to reflect the types of challenges that can occur with HF patients in the VA. The cases were not intended to be representative of a typical case. One case was a true positive – the patient correctly qualified for titration with recommended BB but was not on them. The other case was a false positive – the patient had HF with reduced ejection fraction, and was not on recommended BB, but had a contraindication not detectable by the CDSS. In this way our sample includes a situation in which the automation fails, potentially increasing the challenge facing the human provider (Bainbridge 1983; Brauner 2019). The sequence of presentation of the two cases was alternated across participants.
Data Collection and Analysis
The sessions were audio-recorded and transcribed. The transcriptions were reviewed individually by members of the qualitative analysis team (which included specialists in clinical and public health informatics, human factors, and industrial/organizational psychology). Emergent codes were compared, and a consensus code set was established. Two pairs of coders each coded all of the transcripts. The transcripts were reviewed as a group to establish consensus on the coding. Two coders classified the interviews and discussion phrases using a coding rubric. We calculated a kappa statistic based on a systematic sample of interview and discussion group documents. Following four rounds of classification and discussion, the percent agreement was 87.2% and the kappa was 0.8469, representing substantial interrater reliability.
Results
Eleven codes emerged that related directly to trust, which yielded 81 statements. The codes and example quotations are presented in Table 1 (next page), organized by three overarching themes: concerns about documentation, complexity of patient condition and care delivery context, and concerns about clinical reminder and/or guideline.
Results.
Discussion
The results reaffirm how trust in a Clinical Reminder (or other CDSS) is not just affected by the tool itself (its appearance and screening behavior), but also by: the relationship between the provider and the guideline (familiarity with the guideline, with BB treatment for HF); the accuracy and completeness of the data in the patient’s chart; and the degree of complexity in the patient’s medical condition and in the care delivery system itself.
Our results correspond to the macrocognitive model of trust and reliance proposed by Hoffman et al (2015). Trust in automation is described as emerging from the sensemaking process (gathering and interpreting data to elaborate or revise one’s understanding (Klein et al 2006; Fallon 2010)) focused on two things: the situation or process in the world, and the automation itself.
Reliance on the automation is described as emerging from the flexecution process (the elaboration or revision of plans and goals as information or situations evolve (Klein 2007) executed upon two things: the situation or process in the world, and the automation itself.
Sensemaking
By providing details on the guideline and key data from the patient chart used to generate the recommendation, the CR enables providers to assess the recommendation, to compare their own thinking with that of the automation.
However, the documentation reliability is also a factor. To make sense of the patient’s situation, providers need not just data about the current state (e.g., vitals, diagnoses, medications), but also trends in various aspects of the patient’s health (e.g., is the patient’s kidney function stable?), and the decision making of past providers (e.g., did a cardiologist try the patient on BB and decide to stop?).
Medical documentation may sometimes have deficiencies in accuracy, thoroughness, or structure (Stetson 2012). Our results indicate that providers can be sensitive to this. They expressed concerns about out-of-hospital data being infrequently updated or involving uncertain contextual factors which make it hard to interpret. Their sensemaking activities included looking for additional information to fill-in missing elements or to resolve information that does not seem to fit.
Problems that affect trust in the documentation also affect trust in the CDSS, due to the dependency of the CDSS on its input source (Alexander 2006; Kilsdonk 2017), and via a halo effect (Hoff 2015).
An additional target of sensemaking is the CDSS itself. This involves interpretation of the clinical data extracted and presented in the Clinical Reminder, including assessments of the thoroughness of the search for information, and the recency of the information. Even with complete and accurate data in the record, a CDSS could fail to access or utilize that data correctly.
The sensemaking process would be supported by access to metadata about the clinical data. This would help providers establish and revise their trust or mistrust in the clinical data as mediated by the CDSS.
Sensemaking of the CDSS itself also involves establishing and revising an understanding of how the algorithm uses records to identify which patients should receive the CR and which should not. Some providers expressed perplexity and uncertainty about the algorithm’s exclusion of patients with detectable contraindications. They sought verification that the rule ‘if contraindication then do not trigger CR’ was being followed. This inadvertently put the providers in a situation similar to the classic ‘confirmation bias’ study paradigm (Wason, 1960). Filtering out patients with detectable contraindications is a background process, not visible to the providers. Furthermore, this process is not 100% accurate, resulting in some False Positives.
To better support sensemaking and adjustment of trust and mistrust, providers should be able to access a list of which of their patients were screened and the results. Additionally, the providers should be able to run tests on known patients to observe how the algorithm interprets those patients’ records.
Yet another target of sensemaking is the guideline in the context of this patient-specific clinical reminder. Clinical guidelines can be difficult to parse (Bracha 2011) and their interpretation and application involves cognitive challenges (Klein 2016). Our CDSS specifically flags HF patients who are likely at risk of under-treatment with BB, so it helps with the challenge of characterizing the medical problem and identifying the relevant guideline. However, our results show that lack of familiarity with the guideline, or with BB and HF treatment more generally, can be an issue with some providers, and therefore affects the processes of sensemaking and collaboration between the CDSS and the provider.
Flexecution
The CDSS is intended to help the provider make changes in the treatment of applicable patients. Providers can initiate titration, issue a referral to another provider, or simply document (in a structured way) the reason why BB titration is not indicated. There is more than one way to satisfy the CR, and the provider might start with one in mind but switch to another. The provider might plan to manage the titration themselves and investigate the patient’s other medications. After seeing the number and type of medications, they decide it would be better to have the clinical pharmacist manage it, so they make a referral to the pharmacist. This is one way that flexecution (Klein 2007) is supported by the CR.
The ability to make changes in treatment is affected by dependencies with the healthcare system and with the patient. Some of the BBs may not be on the institution's formulary. There may be ambiguity about responsibility between the PCP, cardiologist, and HF clinic. The patient may have some medical condition that is insufficiently stable to initiate BB titration.
The presence of comorbidities or other interacting factors can impose constraints on what treatments are possible (Klein 2016). There may be other applicable guidelines competing for attention (Hamm & Nagykaldi 2018). Each guideline (and any associated CDSS) is focused on one particular issue by itself, but the provider must view all the different issues as an interacting, integrated set (Horsky 2017).
Thus, even if the CDSS is correct that the patient could benefit from initiation of BB titration, there may be other competing priorities for that patient at that moment. In such situations, it may be necessary to replan or reprioritize what to address now and what to address later. Individual Clinical Reminders may help inform providers of particular medical needs they might otherwise be ignorant of, but there is no good support for the cognitive work of identifying which needs are most important to focus on right now (Xiao & Gorman 2018). The ability of the human-automation team to negotiate a sacrifice of local, short-term goals (such as the focus of a particular CR) for global, longer-term goals (e.g., integrated management of the patient) is fundamental to coordination (Woods & Hollnagel 2006).
If our CR informs the provider of this hitherto unknown problem of undertreated HF, which is then taken into account during the planning of when to address various issues, then the CR has positively contributed to the care of the patient, even if it does not directly lead to BB titration of the qualifying patient at this time.
Limitations
Our cases were designed to explore the role of the Clinical Reminder in the cognitive work of the provider in interpreting the presumptive need for BB titration. As such, they included more complexity than is typical of HF patients in the VA.
Unlike the typical scenario of a provider treating a patient, in our study the providers had never seen the patient before, nor could they contact colleagues who may have known the patient. The only source of information was the simulated EHR chart.
We only looked at the decision to start titration, not issues occurring during titration. And this decision was framed only with the provider and the Clinical Reminder (with the patient’s chart); decision-making with the involvement of the patient or caregiver was not included.
Footnotes
Acknowledgements
This work was supported by VA Health Services Research and Development (VA HSRD CRE 12-037). The views expressed in this article are those of the authors and do not necessarily reflect the position or policy of the Department of Veterans Affairs or the United States Government.
