Abstract
Causal knowledge has been shown to affect diagnostic decisions. It is unclear, however, how causal knowledge affects diagnosis. We hypothesized that it influences intuitive reasoning processes. More precisely, we speculated that people automatically assess the coherence between observed symptoms and an assumed causal model of a disorder, which in turn affects diagnostic classification. Intuitive causal reasoning was investigated in an experimental study. Participants were asked to read clinical reports before deciding on a diagnosis. Intuitive processing was studied by analyzing reading times. It turned out that reading times were slower when causally expected consequences of present symptoms were missing or effects of absent causes were present. This causal incoherence effect was predictive of participants’ later explicit diagnostic judgments. These and related findings suggest that diagnostic judgments rely on automatic reasoning processes based on the computation of causal coherence. Potential implications of these results for the training of clinicians are discussed.
Keywords
The Diagnostic and Statistical Manual of Mental Disorders (DSM–IV—American Psychiatric Association [APA], 2000; DSM–5—APA, 2013) and the International Statistical Classification of Disease (World Health Organization, 1992), which mirrors the DSM, provide normative frameworks for the diagnosis of mental disorders. Classifications are supposed to be based on diagnostic criteria, most of which are neither necessary nor sufficient. These diagnostic criteria are usually directly observable symptoms. Neither of the two classification systems is based on causal, etiological theories of mental disorders. Potential causal relations between the symptoms listed as criteria are not included in the definition. This means that clinicians should ignore causal considerations when making a diagnosis, but should rely on counting symptoms instead.
Research on diagnostic decision making, however, has shown that clinicians (both experienced and in training) take causal considerations into account when making a diagnosis (Kim & Ahn, 2002; Kim & Keil, 2003). For example, Kim and Ahn (2002) asked mental health professionals about their causal assumptions with respect to common disorders from the DSM. They found that clinicians assumed causal relations among the symptoms used as diagnostic criteria. More important, the researchers also showed that these causal assumptions affected diagnostic classifications when participants were later presented with case vignettes. Specifically, participants were more likely to diagnose a hypothetical client with a mental disorder when she or he presented with symptoms that were assumed to affect other symptoms (causally central symptoms) than when the person presented with symptoms that presumably had no impact on other symptoms (causally peripheral symptoms). Later studies replicated these findings (e.g., Kim & Keil, 2003).
Assumptions about causal relations have also been shown to affect judgments of treatment efficacy (Ahn, Proctor, & Flanagan, 2009; de Kwaadsteniet, Hagmayer, Krol, & Witteman, 2010; Yopchick & Kim, 2009), information seeking (de Kwaadsteniet, Kim, & Yopchick, 2013), the perception of patients as being normal versus pathological (Ahn, Novick, & Kim, 2003), and judgments of the need for psychological treatment (Kim & LoSavio, 2009).
Thus, clinicians’ diagnostic and treatment decision making is affected by causal assumptions. This finding is in line with recommendations from proponents of clinical case conception for treatment decision making (e.g., Eells, 2007; Persons, 2006). Clinicians are advised to deliberately consider the factors and mechanisms that cause and maintain a client’s problem when choosing a treatment. The finding, however, contradicts requirements from the DSM for diagnostic classifications. Clinicians’ causal beliefs seem to bias their diagnostic decisions, because diagnostic criteria are not weighted as prescribed by the manual. Research on the cognitive processes underlying diagnostic decision making can help us to understand why and how causal beliefs affect diagnostic judgments.
Theoretical Accounts
Probably the most influential psychological framework theory in judgment and decision making is dual process theories (Evans, 2008; Evans & Stanovich, 2013). Roughly speaking, these theories distinguish between Type 1 processes, which are assumed to be nondeliberate, rapid, automatic, and high capacity, and Type 2 processes, which are conceived as deliberate, slow, controlled, and capacity-limited (i.e., its capacity to deal with complex situations is limited by resources available from working memory). Type 1 processes quickly yield impressions and provide intuitive answers to judgment and decision problems. Type 2 processes are thought to monitor the quality of these proposals and, if considered necessary, correct or override the intuitive judgments.
Some theorists assume that thinking about causes and causal relations is, at least partially, a Type 1 process (Kahneman, 2011; Kahneman & Frederick, 2002; Kahneman & Miller, 1986; Weiner, 1986). Presumably, people automatically start to think about causes and causation when they make an observation. They also automatically activate domain specific causal knowledge and schemata. This activation may occur even in cases in which causal knowledge should be ignored to make a correct judgment (Tversky & Kahneman, 1980). Based on the intuitively derived causal assumptions, a judgment or decision is made, again using Type 1 processing.
There are two theoretical models from cognitive psychology, which explain how causal assumptions may affect diagnostic decision making. The first model is the so-called causal status hypothesis (Ahn, 1998), which assumes causally central symptoms (e.g., symptoms that affect many other symptoms) are weighted more in classification than causally peripheral symptoms (e.g., symptoms that do not affect other symptoms). This model predicts that a client should be rated as more likely to have a certain disorder when she or he presents with causally central rather than peripheral symptoms. Evidence for the causal status hypothesis comes from the cited studies of Ahn and Kim (2000) and Kim and Ahn (2002).
The second model is the causal model theory of categorization (Rehder & Kim, 2006, 2010), which assumes that an observed pattern of symptoms is compared to an expected pattern of symptoms derived from preexisting causal knowledge. The model predicts that a client should be rated as more likely to have a certain disorder when the observed symptoms are coherent with the expectation. Evidence for this model comes from research on categorization, which found the so-called causal coherence effect (Rehder & Kim, 2006, 2010), which is not predicted by the causal status hypothesis. Roughly speaking, an observation is causally coherent with an assumed causal relation if the cause and the effect are either both present or both absent; it is incoherent when either only the cause or only the effect is present. It is important to note that the causal model theory of categorization also predicts that clients presenting with causally central symptoms should receive higher diagnostic ratings than clients presenting with peripheral symptoms, because a client having a certain disorder is more likely to have causally central than peripheral symptoms. It is important to note that neither model makes a specific proposal regarding the nature of the processes responsible for the calculation of causal centrality or causal coherence. Our view is that these processes are Type 1 processes. Thus, we assume that causal assumptions affect diagnostic decision making by Type 1 processes.
Previous Findings on Type 1 Processes in Diagnostic Decision Making
Although there is ample evidence that causal knowledge affects judgment and decision making in general (Hagmayer & Fernbach, 2016; Hastie, 2014) and diagnostic classification in particular, there still is little empirical evidence for the involvement of Type 1 processes. The first evidence comes from two recent studies by Flores and colleagues (Flores, Cobos, López, & Godoy, 2014; Flores, Cobos, López, Godoy, & González-Martín, 2014). In their experiment, Flores, Cobos, López, Godoy, and González-Martín (2014) presented clinicians with fictitious reports about individual clients and asked for a diagnostic judgment. Each report stated an initial diagnosis and three symptoms. The researchers manipulated the temporal order in which the three symptoms developed over time. According to participants’ causal beliefs about presented mental disorders (e.g., anorexia nervosa, specific phobia), the symptoms formed a causal chain Symptom 1 → Symptom 2 → Symptom 3. Therefore, participants should expect symptoms to have occurred in the S1-S2-S3 temporal order (consistent condition) rather than in the reverse order S3-S2-S1 (inconsistent condition). Two dependent variables were collected: reading times (RTs) for sentences referring to the symptoms and final diagnostic judgments. RTs are particularly interesting because they provide an online measure of cognitive processing. It turned out that the clinicians’ RTs were longer when symptoms occurred in the inconsistent, unexpected temporal order than when they occurred in the consistent temporal order. Corresponding to this, the clinicians tended to give higher diagnostic ratings in the consistent than in the inconsistent condition although clients presented with exactly the same symptoms.
The inconsistency effect on RTs found by Flores, Cobos, López, Godoy, and González-Martín (2014) and also by Flores, Cobos, López, and Godoy (2014) is very well known in the field of reading comprehension (Albrecht & O’Brien, 1993; Long & Chong, 2001; Peracchi & O’Brien, 2004). It is interesting that there is additional evidence showing that causal inferences during reading are made in a fast and spontaneous manner (Hassin, Bargh, & Uleman, 2002). Violations of expected causal sequences result in longer RTs, indicating that causal incoherence is detected. There is also ample neuroscientific evidence showing that local- and discourse-level inconsistencies of very different kinds are detected in about 250 ms (see Hagoort & Van Berkum, 2007, for a review). This extremely fast detection of inconsistencies indicates that the underlying processes are likely to be Type 1 processes according to dual process theories.
Inspired by theories of reading comprehension (Graesser, Singer, & Trabasso, 1994), Flores, Cobos, López, Godoy, and González-Martín (2014) explained their results by postulating that reasoners engage in Type 1 processes to generate a mental model, which tries to integrate the diagnostic information provided at the beginning of the clinical reports, respective preexisting causal knowledge, and the information about symptoms provided later on. Depending on the coherence of the resulting mental model with the given diagnosis, reasoners may come up with a feeling of coherence. This feeling of coherence might then be used to make a diagnostic judgment after reading the clinical report. In a similar way, coherence-driven processes have been frequently assumed to be a source for intuitive judgments and biases (e.g., Morewedge & Kahneman, 2010; see also Glöckner & Witteman, 2010).
Aim and Outline of Present Study
In the study presented here, we aim to provide new and more compelling evidence for Type 1 processes involving the computation of causal coherence in diagnostic decision making. The presence of Type 1 processes would explain why causal assumptions automatically affect clinicians’ judgments without any deliberate causal reasoning process. The computation of causal coherence would explain and predict under which conditions clinicians would be less likely to diagnose a patient with a specific disorder even when such conditions are balanced regarding the number of diagnostic symptoms present.
The current study builds on the research by Flores et al. (2014a, 2014b). Participants had to read clinical reports providing preliminary information about the diagnosis received by hypothetical clients followed by information about three symptoms that participants should assume to form part of a causal chain of the sort S1 → S2 → S3. As before, the symptoms were part of the DSM–IV diagnostic criteria for the diagnosis received by the client. In contrast to the previous studies, we directly manipulated causal coherence by informing participants about the absence of one of the three symptoms. Thus, we presented participants with three hypothetical clients that either lacked S1, S2, or S3. According to the causal status hypothesis, S1 is the causally most central symptom, because it affects S2 and S3, followed by S2 and S3. Therefore, participants should give the highest diagnostic ratings to clients missing S3, and the lowest diagnostic ratings to clients without S1. By contrast, the causal model theory of categorization predicts that the lowest diagnostic rating should be given to clients missing symptom S2. The reason is that the absence of S2 would be inconsistent with the presence of S1 and with the presence of S3 according to the S1 → S2 → S3 causal model for the disorder. Thus there are two causal inconsistencies. Conversely, the lack of S1 or S3 would entail only one causal inconsistency. Therefore, clients lacking S1 or S3 symptoms should be more likely to be diagnosed with the mental disorder stated at the beginning of the clinical report. More precisely, clients missing S1 should receive somewhat lower ratings than clients missing S3, because it is more likely that a person suffering from the disorder lacks S3 than S1 (see Rehder & Kim, 2006, for a detailed explanation why this is the case).
The clinical reports used in our experiment were constructed in a special way so we could analyze RTs to track participants’ intuitive causal reasoning. The second sentence of the report gave the preliminary diagnosis received by a hypothetical client. Then, participants read a sentence informing about a first symptom S1 followed by another sentence informing about a second symptom S2, followed by one more sentence informing about a third symptom S3. Only one of the three symptoms was explicitly pointed out as being absent. Times spent reading each sentence were measured unbeknownst to participants. Diagnostic judgments had to be made after reading each clinical report. As only one symptom in the causal chain was absent in each condition, we were able to compare RTs for three different types of sentences: (a) sentences stating the presence of a symptom after the previous sentence stated the absence of the preceding symptom, (b) sentences stating the absence of a symptom after the previous sentence stated the presence of the preceding symptom, and (c) sentences stating the presence of a symptom after the previous sentence stated the presence of the preceding symptom. We hypothesize that participants engage in online, Type 1 reasoning processes to compute causal coherence. If this hypothesis is true, then we should expect to find the following RTs. RTs should be longer for the first two types of sentences (causally incoherent sentences) than for the third type (causally coherent sentences). We also expected longer RTs for sentences stating the absence of the first symptom, because the absence of S1 stated in the third sentence is causally incoherent with the diagnosis stated in the second sentence of the clinical report.
Table 1 details the specific predictions regarding RTs and diagnostic judgments. Predictions for RTs are derived from our coherence computation hypothesis. Predictions for final diagnostic judgments are derived from the generative model, which is a quantitative extension of the causal model theory of categorization (Rehder & Kim, 2010).
Predicted Pattern of Reading Times (RT) for Target Sentences Drawn From the Assumption That Participants Engage in a Fast, Online Activation of Causal Knowledge, Leading to Fast Detection of Inconsistencies, Resulting in Longer Reading Times
Note: Predictions for diagnostic judgments are derived from the generative model. See the text for further explanations.
Method
Participants and design
A total of 61 undergraduate students from the School of Psychology at the University of Göttingen (Germany) volunteered in our experiment. All of them had been previously trained on the disorders used in our clinical reports in at least two courses of psychology, where they had learned about the theories used for the clinical cases. Although clinicians may be seen as a more appropriate target population for our study, we decided to focus on psychology students trained in clinical psychology. The reason is that our previous studies found evidence for Type 1 causal reasoning in clinicians but not in psychology students. One possible explanation is that, for unknown reasons, students are less prone to engage in causal reasoning. Another explanation is that the temporal causal connections between the symptoms used in our clinical reports were less obvious for students than for clinicians. According to this latter explanation, we should be able to find evidence for Type 1 causal reasoning processes in students if the causal relationships between symptoms are rather obvious. To explore this possibility, we decided to use a sample of trained psychology students.
The only factor manipulated was the type of client, defined by the missing diagnostic criterion within the causal chain model of the disorder (not S1, not S2, or not S3). As three types of disorders were used (specific phobia, obsessive-compulsive disorder, and depression), participants received a total of nine clinical reports in random order.
Materials
Depression, obsessive-compulsive disorder of cleaning, and specific phobia to dogs were used to create the clinical reports. The theories on which the causal chain models were based were the following: Beck’s (1967) cognitive theory of depression, which proposes that symptoms such as sadness or apathy are the result of an inadequate and biased processing of information; Salkovskis’s (1985) cognitive-behavioral model of obsessive-compulsive disorder, according to which the client tries to reduce her or his anxiety and unease produced by her or his obsessions by doing compulsive rituals; and Mowrer’s (1947) two-factor model of specific phobia, which states that, initially, an individual acquires an aversion to a stimulus, and then tries to avoid it to reduce the anxiety. The symptoms selected to play the role of S1, S2, and S3 were, in the case of depression, S1: “to think that bad things always may happen to oneself everywhere,” S2: “not to feel like going out with friends,” S3: “to be socially isolated and to have a lot of social problems,” respectively; in the case of obsessive-compulsive disorder, S1: “to feel anxious about getting a bacterial infection,” S2: “to wash hands around 40 times per day,” and S3: “to have strong problems in the workplace because of lack of time”; in the case of specific phobia, S1: “to have suffered from bad experiences with dogs during childhood,” S2: “to feel bad when passing close to dogs,” and S3: “to avoid going to pet shops or parks,” respectively. The symptoms were supposed to strongly activate previous causal knowledge according to which S1 should be the cause of S2, which, in turn, should be the cause of S3.
Every clinical report consisted of six sentences and was structured in the following way. The first sentence was an irrelevant sentence introducing the client. It was followed by a sentence informing about the diagnosis given by a professional. The third, fourth, and fifth sentences informed about the presence or absence of S1, S2, and S3, respectively, without providing any information about causal relationships between them. Every hypothetical client presented with two of the three symptoms (see the appendix). The absence of a symptom was made explicit by referring to an opposite state or behavior. For example, if the symptom was “she/he never feels like going out with friends,” its absence was made explicit by saying that “she/he always feels like going out with friends”; or if the symptom was “she/he washes her/his hands 40 times per day,” the corresponding sentence for stating its absence was “she/he washes her/his hands 4 times per day.” This way, the sentences referring to the presence and the absence of a specific symptom were almost identical regarding length, wording, structure, and number of syllables (in German). Finally, the clinical report ended with a final sentence that was held constant across the different clinical reports for the same disorder.
Procedure
The task was performed in a laboratory with six PCs equipped with home-built software programmed in Visual Basic 2005 and PowerPoint. Participants started by reading the instructions on the computer screen. They were informed that they were required to read a series of clinical reports about hypothetical clients who had been diagnosed with a mental disorder by a clinical psychologist. Participants were instructed to click on the screen to proceed from one sentence to the next and to read the material attentively and, at the same time, fluently. After reading the instructions, they were presented with an example of a clinical report based on a disorder (anorexia) that was different from those used in the actual experimental task. The example text had the same structure as the experimental clinical reports.
Sentences of each individual report were presented sequentially. Initially, every letter of the text was substituted by a mask consisting of a forward slash. Each click made all of the letters of a sentence visible while hiding the slashes. A second click had the reverse effect on the read sentence and rendered the following sentence visible. The RT for each sentence was the time that elapsed between the two consecutive clicks. As usual in self-paced reading tasks, readers were not allowed to go back during reading. Clicking after reading the final sentence allowed participants to proceed to the diagnostic judgment task. For this purpose, a new display was shown on the screen which included a message asking participants to rate the extent to which they agreed on the diagnosis mentioned in the clinical report, a rating scale consisting of a scroll bar, and the clinical report. The message was the following: “The diagnosis received by the client was [Disorder X]. Please rate the extent to which you agree on the diagnosis using the scale below” (translated from German). The ratings could range from 0 (completely sure that the correct diagnosis is another one) to 10 (completely sure that the diagnosis is correct). No feedback was provided.
After the diagnostic judgment task, participants had to judge the causal relationship for all possible pairs of events mentioned in the clinical report. This task was designed to check whether the participants’ causal beliefs conformed to the causal chain model (i.e., S1 → S2 → S3) on which we based our manipulation and predictions. After the causal judgment task, participants performed further clinical tasks for a different study. Once these tasks were finished, participants proceeded to the next clinical report.
Results
All of the statistical analyses reported here adopted an alpha of .05. Results about the causal-link judgments are reported in a separate section (see the Supplemental Material available online). These results clearly show that participants’ causal beliefs conformed to a causal chain model (i.e., S1 → S2 → S3) for every disorder. They also demonstrate that our manipulation of causes being absent versus present was successful.
Reading times
Statistical analyses were conducted to test whether RTs were longer for the inconsistent sentences than for the consistent sentences. Inconsistent sentences included those sentences stating the absence of a symptom given the presence of a previous symptom in the causal chain as well as those sentences stating the presence of a symptom given the absence of the previous symptom in the causal chain.
Before the analysis, RTs were normalized by the number of letters of the corresponding target sentence. Then RTs were subjected to a filtering process for statistical outliers. Specifically, for each clinical report, we removed those RTs that differed from the sample mean more than three standard deviations. As a consequence, 22 out of 1,647 RTs (3 disorders × 3 types of client × 3 symptoms × 61 participants) were removed. Then, RTs across the different mental disorders were collapsed into a single average RT per client and symptom condition. Table 2 shows the RTs for every target sentence in each of the different client conditions collapsed across the different disorders. Thus, the first column shows the RTs for sentences referring to Symptoms S1, S2, and S3 in the client condition in which the absent symptom was S1. The second and third columns show the corresponding RTs in those client conditions in which the absent symptoms were S2 and S3, respectively. As can be seen, the RTs for the inconsistent sentences were, in general, longer than for the consistent sentences. In the case of Symptom 1, the RTs in Client 1 condition were longer than in the remaining conditions. In the case of Symptom 2, the RTs in Client 1 and Client 2 conditions were longer than in Client 3 condition. Finally, in the case of Symptom 3, the RTs in Client 2 and Client 3 conditions were longer than in Client 1 condition. This pattern of results fits the predictions shown in Table 1. To confirm these impressions, we performed a repeated measures ANOVA 3 (client condition: absence of S1, S2, and S3) × 3 (symptom: Symptom 1, Symptom 2, Symptom 3), which yielded a significant main effect of client condition, F(2, 120) = 8.46, MSE = 110.85, p < .001, η2 = .12, a significant effect of symptom, F(2, 120) = 49.60, MSE = 133.19, p < .001, η2 = .45, and the significant interaction client condition × symptom, F(4, 240) = 33.25, MSE = 95.09, p < .001, η2 = .36.
Means and Standard Deviations of Reading Times (RT) for the Target Sentences (in Milliseconds per Letter) and Mean Final Diagnostic Judgments (From 0 to 10)
As a follow-up of the significant interaction, we conducted planned comparisons for all symptom sentences between conditions. Note that the slowing down of reading for sentences stating the absence of a symptom could be interpreted as the effect of an inconsistency between such sentences and the previously stated diagnosis, which would not require any causal coherence computation process. Therefore, inconsistencies specifically produced by the computation of causal coherence are better assessed by comparing RTs for present-symptom sentences preceded by absent-symptom sentences with present-symptom sentences preceded by present-symptom sentences. For Symptom 1, a t test for paired comparisons yielded a significant difference between Client 1 and Client 2 conditions, t(60) = 8.95, p < .001, η2 = .57, and between Client 1 and Client 3 conditions, t(60) = 5.62, p < .001, η2 = .34. Differences between Client 2 and Client 3 conditions were marginal (t = 3.10, p = .059) when the Bonferroni approach was taken to protect the statistical analyses against the accumulation of Type 1 error (i.e., with α = .016). For Symptom 2, a significant difference between Client 2 and Client 3 conditions was found, t(60) = 7.38, p < .001, η2 = .48. Furthermore, although Symptom 2 was present in Client 1 and Client 3, and its presence was consistent with the diagnosis stated in the clinical report, different RTs resulted between both conditions: t(60) = 6.59, p < .001, η2 = .42. If Symptom 2 was present despite its cause Symptom 1 being absent (Client 1) participants should have experienced a causal inconsistency. This was not the case when both Symptom 1 and Symptom 2 were present (Client 3), which is in accordance with a causal chain model. Therefore, no differences between Client 1 and Client 2 were found for Symptom 2 (t = 0.037, p = .971). For Symptom 3, a significant difference between Client 1 and Client 3 conditions was found, t(60) = 4.85, p < .001, η2 = .28, and also between Client 1 and Client 2 conditions, t(60) = 3.60, p = .001, η2 = .18. Note that these differences would still be significant even if we take the Bonferroni approach (i.e., with α = .016). As expected, no differences between Client 2 and Client 3 were found (t = 1.06, p = .293). These findings regarding Symptom 3 replicated the finding regarding Symptom 2. Again, Symptom 3 was present for Client 1 and Client 2, but it occurred despite Symptom 2 being absent in Client 2, which created a causal inconsistency according to the assumed causal chain model. Therefore, participants were apparently sensitive to causal inconsistencies, which seemed to have been computed on the basis of the causal chain model.
However, one may argue that the longer RTs spent on sentences informing about present symptoms given the absence of the antecedent symptom could be the result of a carryover effect rather than the expression of an online computation of causal coherence. To assess the plausibility of this account, we analyzed the RTs for the final sentence, which was an irrelevant sentence conveying no information about symptoms. If the reading of an inconsistent sentence stating the absence of a symptom produces a carryover effect on the following sentence, the RTs for the final sentence should be longer in the Client 3 than in the Client 1 and Client 2 conditions. Note that only in the Client 3 condition was the final sentence preceded by a sentence stating the absence of a symptom. The mean normalized RTs for the final sentence (again by the number of letters of the corresponding target sentence) were 35.61 ms, 32.27 ms, and 33.52 ms, corresponding to the Client 1, Client 2, and Client 3 conditions, respectively. It is not surprising at all that these RTs were shorter than the RTs for target sentences as the final sentence was irrelevant for the subsequent diagnostic judgment. There is a large body of evidence showing that readers spend less RT in irrelevant than in relevant sentences (see, for example, Goetz, Schallert, Reynolds, & Radin, 1983; Kaakinen & Hyönä, 2007; Kaakinen, Hyönä, & Keenan, 2002, 2003). Also, as can be appreciated, the RTs found for final sentences were clearly at odds with a carryover effect as the RTs in the Client 3 condition were not longer than in the remaining conditions. Therefore, the longer RTs found for Symptom 2 when comparing Client 1 to Client 3 and for Symptom 3 when comparing Client 2 to Client 1 are most likely due to the operation of online computation of causal coherence. To summarize, the results found in the participants’ RTs suggest the engagement of Type 1 processes responsible for the computation of causal coherence during reading of clinical reports.
Diagnostic judgments
Judgments for the different disorders were collapsed into a single mean per client condition for each participant. The analyses reported were conducted on these resulting means. Table 2 shows participants’ mean judgments in each condition. Participants agreed with the diagnosis stated in the text to a greater extent in the Client 3 than in the Client 1 condition. In turn, agreement ratings were higher in the Client 1 than in the Client 2 condition. These impressions are supported by a repeated measures ANOVA (client condition: Client 1 vs. Client 2 vs. Client 3) on participants’ judgments, which yielded the significant main effect of client condition, F(2, 120) = 102.96, MSE = 1.15, p < .001, η2 = .63. T tests for paired comparisons revealed significant differences between the Client 1 and Client 2 conditions, t(60) = 5.77, p < .001, η2 = .36, between the Client 1 and Client 3 conditions, t(60) = 8.24, p < .001, η2 = .53, and between the Client 2 and Client 3 conditions, t(60) = 14.94, p < .001, η2 = .79. Note that these differences would still be significant even if we take the Bonferroni approach (i.e., with α = .016).
These findings clearly show that assumptions about the causal relations among symptoms affected diagnostic judgments. Specifically, the pattern of diagnostic ratings supports the generative causal model of categorization (Rehder & Kim, 2010). As outlined earlier, this theoretical model predicts that reports in the Client 2 condition would be the least coherent, whereas reports in the Client 3 condition would be the most coherent (assuming probabilistic causal relations). Thus, the lowest and the highest agreement ratings should be found in the Client 2 and the Client 3 condition, respectively, which was in fact the case.
Discussion
The present study is part of an endeavor to find out why and how causal assumptions held by clinicians affect their diagnostic judgments. The aim of the present study was to investigate whether decision makers intuitively compare the observed symptoms with expectations that result from a causal model of the disorder the client is supposed to have and use the coherence between observations and expectation to make a diagnostic judgment. More precisely, our objective was to provide more compelling evidence that the diagnosis of mental disorders involves the computation of causal coherence through Type 1 processes. To achieve this goal, we manipulated the coherence of the observed pattern of symptoms with the causal model of the disorders presented to participants. We did so by changing which of three symptoms connected through a causal chain was absent. Our approach was similar to that taken by Rehder (2003) to study the influence of causal model assumptions on categorization. In his experiments, Rehder found a causal coherence effect. He showed that categorization was affected by the number of causal mechanisms violated by the observations made (see also Rehder & Kim, 2006, 2010). In our experiment, we went beyond these studies that investigated only the effect of causal beliefs and causal coherence on final judgments. We also collected an online measure of intuitive causal reasoning, that is RTs. RTs indicate whether the underlying Type 1 process is also affected by causal beliefs and coherence.
We found a causal coherence effect in diagnostic judgments in line with Rehder’s demonstrations of this effect. Our results are also consistent with Rehder and Kim’s (2006, 2010) findings on causal coherence effects in diagnostic decision making. In addition, we showed a causal inconsistency effect on RTs. This effect suggests (a) that participants’ causal beliefs affect online, fast, and efficient Type 1 processes during reading and (b) that these Type 1 processes involve the computation of causal coherence. Our findings complete previous findings by Flores, Cobos, López, Godoy, and González-Martín (2014). This study showed that RTs and diagnostic judgments are affected by the temporal coherence between the observed sequence of symptoms and the temporal sequence implicated by causal assumptions of the participating clinicians. The present findings indicate that there is a causal coherence effect in addition to a temporal coherence effect.
Maybe the most remarkable finding of our study is the high consistency between the results found through the online measure based on RTs and those found through the offline measure based on diagnostic judgments. The participants’ diagnostic judgments were consistent with the causal model theory of categorization, which states that categorization is the result of the computation of causal coherence. At the same time, the RTs results suggest that the computation of causal coherence forms part of the online processes at work at the very moment in which participants receive relevant information for the diagnosis during fluent reading. A sensible interpretation of this consistency is that Type 1 processes responsible for the online and fast computation of causal coherence had an impact on the later diagnostic judgments. Note that this interpretation does not necessarily entail that diagnostic judgments were the direct output of Type 1 processes or that Type 2 processes did not play any role at all in the diagnostic judgments. Type 2 processes might have contributed to the diagnostic judgment results by, for example, confirming the first impression resulting from Type 1 processes through deliberate causal reasoning or just by accepting totally or partially such first impression, which, in turn, may have had an anchor effect on judgments. Our results are compatible with several ways in which Type 1 and Type 2 processes may interact.
Limitations
One limitation is that we conducted our study with advanced psychology students who received training in clinical psychology, which may be considered as a limitation when it comes to generalize our results to experienced clinicians. One may question whether clinicians are as vulnerable to causal reasoning as students are. However, as said in the introduction, there is evidence showing that experienced clinicians are also biased by causal reasoning. In their studies on the effect of causal knowledge on diagnostic judgments, Kim and Ahn (2002) found an effect for both students and clinicians. Flores, Cobos, López, Godoy, and González-Martín (2014) and Flores, Cobos, López, and Godoy (2014) used samples of experienced clinicians and students. In general, they found stronger evidence of causal reasoning in experienced clinicians than in students. Thus, we would expect to find the causal coherence effect reported here in experienced mental health clinicians as well.
Another potential limitation is that participating students may have relied on their common causal beliefs about daily live rather than specific knowledge about the causal mechanisms underlying particular mental disorders. In the case of specific phobia, for instance, we do not need a psychology course to causally link “bad experience with dogs,” “feeling very anxious when being close to a dog,” and “very rarely spending time near pet shops or parks.” This may explain why our students’ RTs were sensitive to causally relevant manipulations whereas students’ RTs in Flores, Cobos, López, Godoy, and González-Martín’s (2014) study were not. In these experiments, the temporal order of symptoms affected clinicians’ RTs, with longer times for temporal orders inconsistent with the causal links between symptoms compared with temporal orders consistent with such causal links. Diagnostic ratings by clinicians and students were affected by temporal order of symptoms. However, students’ RTs were unaffected by temporal order of symptoms. These divergent findings indicate that we need to further explore which type of causal knowledge people use in clinical judgment and decision making and when they use it. But the possibility that participants may have relied on everyday causal knowledge rather than clinical knowledge in the present experiment does not invalidate our findings. The results still show that causal considerations affected Type 1 processing of the given information and diagnostic judgments.
Type 1 Causal Reasoning Processes in Diagnostic Decision Making in Clinical Practice
One may think that more experienced clinicians would be less affected by causal considerations in diagnostic decision making as they receive extensive training on the correct use of the DSM. The experimental findings cited earlier prove otherwise (Flores, Cobos, López, & Godoy, 2014; Flores, Cobos, López, Godoy, & González-Martín, 2014; Kim & Ahn, 2002). Conversely, it might be argued that experienced clinicians are especially vulnerable to influences of causal knowledge because Type 1 processes play an important role in experts’ reasoning. According to some researchers (see, for example, Charlin, Boshuizen, Custers, & Feltovich, 2007; Charlin, Tardif, & Boshuizen, 2000; Schmidt, Norman, & Boshuizen, 1990; Smith, 1989), expert clinicians’ knowledge is represented via structurally organized units (scripts) that allow automatic and efficient access from memory through fast activation processes. Parts of these scripts are the causal mechanisms and the typical causal-temporal development of the disease or disorder over time. Once a script is activated, it would participate in the production of fast inferences as well as in the effective integration of incoming information through dynamic top-down and bottom-up processes. According to this, diagnostic biases due to causal reasoning operated through Type 1 reasoning processes could be more likely in clinicians than in students. In support of this idea, previous studies with clinicians have provided evidence for the lack of adherence to previous versions of the DSM as a result of heuristic processes typically attributed to Type 1 processes (Maj, 2011; Westen, 2012; Westen & Shedler, 2000).
Implications for Clinicians’ Training
The persistent tendency of clinicians to use causal reasoning in diagnosis conflicts with DSM–IV recommendations and with how clinicians are trained to use this resource. It therefore appears that clinicians’ initial training regarding DSM–IV prescriptions is not sufficient to prevent them from using causal reasoning when making diagnoses. Our results suggest that one explanation for this difficulty may be that clinicians are not aware of their use of causal theories. The involvement of Type 1 causal reasoning processes that appears to have occurred during the reading task suggests that very rapid, efficient, and automatic processes may have been at work. If such is the case, clinicians’ training should be supplemented by training in causal reasoning that is aimed at describing the different and (occasionally) subtle ways in which it can influence judgments and decisions in the clinical context, especially in cases of patients who are potentially suffering from DSM–IV disorders. Consequently, training in causal reasoning should help clinicians gain further control of their reasoning and decision making. Furthermore, according to Kahneman (2011), the more we know about the activities and biases of Type 1 processes, the more aware we will be of how they work and how they influence and mislead Type 2 processes. In addition, the latter processes can be trained to improve (e.g., calculating probabilities and using statistics).
Our results also appear to have interesting implications for evidence-based clinical practice, specifically for the application of empirically supported treatments (ESTs). According to the American Psychological Association, ESTs are currently considered to be the best methods for addressing the treatment of mental disorders and patients’ behavioral problems. Although ESTs are quite standardized, there is evidence demonstrating that clinicians have difficulties in following the indications that are prescribed in textbooks (Waller, 2009) and tend to adapt the treatments to either the patients’ individual characteristics (McHugh, Murray, & Barlow, 2009) or to the clinicians’ case formulation, even when such formulations are not explicit or structured (Pain, Chadwick, & Abba, 2008; Persons, 2006). Moreover, this tendency has been considered to be inevitable by other clinicians (Persons, 2005). Our results suggest that causal theories, which appear to be readily available in the clinician’s mind and used through Type 1 processes, may play an important role in clinical case formulations (Eells, 2007). Such clinical case formulations would in turn be responsible for the difficulties that are experienced by clinicians when attempting to strictly follow the treatment protocol, especially when the theory on which the EST is based differs from the clinician’s causal theory (Anderson & Strupp, 1996; Beutler, 1999). Thus, clinicians’ application of ESTs may benefit from a certain degree of training in causal reasoning that is aimed to make clinicians aware of the different and subtle ways in which it can affect treatment decision making and treatment application.
Footnotes
Appendix
Target Sentences of Reports in Client 1, Client 2, and Client 3
| Target sentences | |||
|---|---|---|---|
| Diagnosis | Client 1 | Client 2 | Client 3 |
| Depression | “ “She never feels like going out with friends” “She has developed problems in her social relationships” |
“She feels that all bad things happen to her wherever” “ “She has developed problems in her social relationships” |
“She feels that all bad things happen to her wherever” “She never feels like going out with friends” “ |
| Specific phobia | “ “He feels very anxious when he is close to dogs” “He is never in commercial centers with pet shops or in parks” |
“He had bad experiences with dogs when he was a child” “ “He is never in commercial centers with pet shops or in parks” |
“He had bad experiences with dogs when he was a child” “He feels very anxious when he is close to dogs” “ |
| Obsessive-compulsive disorder | “ “He washes his hands about 40 times per day” “He has problems at job because of his bad management of time” |
“He is very worried about the possibility of getting contaminated with microbes” “ “He has problems at job because of his bad management of time” |
“He is very worried about the possibility of getting contaminated with microbes” “He washes his hands about 40 times per day” “ |
Note: Sentences were translated from German. The bold sentences indicate the absence of a symptom for that disorder.
Acknowledgements
All procedures performed in studies involving human participants were in accordance with the ethical standards of the institutional and/or national research committee and with the 1964 Helsinki Declaration and its later amendments or comparable ethical standards. Informed consent was obtained from all individual participants included in the study.
Author Contributions
Y. Hagmayer developed the study concept. All authors contributed to the study design. Testing and data collection were performed by Y. Hagmayer and A. Flores. A. Flores performed the data analysis and interpretation under the supervision of Y. Hagmayer and P. L. Cobos. A. Flores drafted the manuscript, and Y. Hagmayer and P. L. Cobos provided critical revisions. All authors approved the final version of the manuscript for submission.
Declaration of Conflicting Interests
The author(s) declared that there were no conflicts of interest with respect to the authorship or the publication of this article.
Funding
A. Flores and P. L. Cobos were supported by Grant 2008-SEJ-03586 from Junta de Andalucía and Grants 2007-63691/PSIC and PSI2011-24662 from the Spanish Ministry of Science and Innovation. A. Flores had an FPI PhD scholarship that was awarded by Junta de Andalucía.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
