Abstract
Previous evidence has shown that, in a recognition memory task, emotion leads participants to make more false alarms and decreases response times (RTs) for false alarm responses. This pattern could arise because participants adopt more liberal responding for emotional stimuli and/or because emotional lures are more likely than neutral lures to produce misleading memory retrieval. Recently, Starns et al. designed a new recognition memory paradigm and found that the speed of memory errors shows the influence of misleading information resulting in unavoidable memory errors. This study investigates the basis of false alarms to emotional lures by testing predictions of the diffusion model for a recognition paradigm similar to that by Starns et al. Participants studied lists of emotional words and then completed an old–new recognition memory test. After each old–new decision, participants were asked to make a forced-choice recognition decision that provided a chance to correct possible errors on the preceding old–new decision. Under the assumption that emotion promotes misremembering, the diffusion model predicts that forced-choice accuracy should be lower for pairs with emotional versus neutral lures and that faster old–new errors should be associated with lower forced-choice accuracy. This study tested these predictions, providing theoretical insights into how emotion affects memory retrieval and further developing a new methodology for measuring recognition performance.
Emotional events are often experienced in a more detailed and vivid manner than more common neutral events. This sense of vividness can promote a high degree of confidence in memories for emotional events, even if those events have never been experienced (Loftus & Bernstein, 2005). Consistent with this pattern, recognition memory experiments consistently show that people are more likely to say that a non-studied word was seen on the study list if this word is emotional (e.g., “cancer”) versus neutral (e.g., “toaster”; Boduroglu & Kapucu, 2018). However, these “false alarms” can be driven by decision biases in addition to memory processes, and both of these mechanisms could contribute to emotion effects. This study explores the basis of the emotion effect on memory errors by testing predictions of the diffusion model (Ratcliff, 1978) for a recognition paradigm in which participants had an opportunity to correct their recognition errors.
Emotion effects on memory
Extensive research has demonstrated that emotion has significant effects on both quantitative and qualitative aspects of memory (Kensinger & Schacter, 2008). In some scenarios, memory accuracy differs between emotional and neutral stimuli due to stronger encoding via increased attentional allocation (Cahill et al., 1996; Kensinger & Schacter, 2006; Talmi et al., 2007), the engagement of more effective consolidation (Hamann, 2001; Sharot & Phelps, 2004), or the availability of information at the time of retrieval (Bowen et al., 2016; Dolcos et al., 2005). Emotion can increase the likelihood of remembering the stimulus itself, as well as provide more detailed and vivid contextual memories for the stimulus (Kensinger & Corkin, 2003, 2004; Ochsner, 2000).
Some previous studies have shown that positive and negative emotional stimuli have comparable effects on memory (Bradley et al., 1992; Dolcos & Cabeza, 2002; Kensinger et al., 2002), while many others have drawn a somewhat different conclusion such that negative stimuli (especially those with high arousal) increase the likelihood of having more accurate and vivid memories than positive and neutral ones (Charles et al., 2003; Dewhurst & Parry, 2000; Kapucu, 2010; Talmi et al., 2007; Yüvrük & Kapucu, 2022). These memory benefits for negative emotional stimuli may be explained by the greater use of a recollection process, which is often assessed by the Remember/Know paradigm (Tulving, 1985). As claimed by the dual-process models of recognition memory (see the study by Yonelinas, 2002), “remember” responses reflect the contribution of a recollective process that brings back episodic contextual details of that memory, while “know” responses reflect the contribution of a familiarity process, which gives a sense that the memory has been merely experienced before without retrieving specific episodic details. Participants are more likely to give “remember” responses to negative emotional stimuli compared with positive or neutral stimuli (Dewhurst & Parry, 2000; Kensinger & Corkin, 2003; Ochsner, 2000). In contrast, proportion of “know” responses are at comparable levels for differently valenced stimuli. These findings have supported the claim that memory enhancement for negative emotional stimuli depends on a recollective memory process.
Model-based research has questioned whether emotion effects on “remember” responses should be attributed to a unique effect on recollection or a general increase in memory strength. Dougal and Rotello (2007) have argued that emotion only increases subjective reports of “remembering” through misleading familiarity information rather than actual use of the recollective process or retrieval of episodic contextual details. They have concluded that emotion creates an illusion of recollection and leads participants to adopt a more liberal bias towards the “old” response even if those stimuli are actually new. Namely, response bias for negative emotional stimuli stems from participants’ “increased sense of recollection” for those stimuli. Similarly, Kapucu et al. (2008) have extended the effects of emotion on response bias to older adults. They asked younger and older adults to make recognition memory decisions for neutral, positive, and negative words using remember-know paradigm. Their receiver-operating characteristic (ROC) data revealed no accuracy differences between emotion conditions in both age groups despite inflated “remember” judgements for negatively valenced words. Yet, robust emotion effects on response bias were observed in the study by Kapucu et al. (2008). Their modelling work has provided evidence that “remember” judgements arise from increased misleading familiarity information for negative emotional stimuli rather than true recollection, and this conclusion was supported by data from both age groups. Indeed, extensive research has demonstrated robust shifts in response bias to emotional compared with neutral stimuli, regardless of differences in memory accuracy (Dougal & Rotello, 2007; Kapucu et al., 2008; Thapar & Rouder, 2009; White et al., 2014; Yüvrük & Kapucu, 2022).
Recenlty, Bowen et al. (2018) have proposed a convincing neurocognitive model that may explain why negative items increase liberal bias more than positive or neutral ones. According to their NEVER (Negative Emotional Valence Enhances Recapitulation) model, sensory details of negative events are incorporated into memory traces more often than details of positive events during both encoding and retrieval. These details allow negative memories to be more likely to be reactivated and recapitulated at retrieval. These emotion effects on recapitulation result in the qualitative changes in memory, even when these changes are not accompanied by higher sensitivity to discriminate studied items from non-studied ones. Recapitulated negative memories may intensify the confidence and subjective vividness of memory and increase familiarity for all negative test items due to shared semantic similarity between studied items and unstudied lures. This heightened sense of familiarity would cause a more liberal bias to make an “old” decision with persuasive misleading information at retrieval (White et al., 2014).
Little is known, however, about how much this increased misleading information and liberal bias for negative emotional stimuli may be linked to the speed of retrieval processes for those emotional stimuli. In fact, limited previous evidence has shown that the emotional content of the stimuli influences the speed of recognition memory decisions. Reaction times to false alarms are faster for negative words than for neutral words (Maratos et al., 2000), while slower correct rejections are observed for negative than for neutral words (Jaeger et al., 2017; Maratos et al., 2000). In Deese-Roediger and McDermott (DRM) paradigm, participants tend to make faster “old” responses to negative than positive lures regardless of whether or not the lure items are associated with the studied words (Norris et al., 2019). Similarly, Yüvrük and Kapucu (2022) have found that “old” judgements were expedited to emotional (both negative and positive) words regardless of whether they were actually old or new. Considering the findings by Dougal and Rotello (2007) on “increased sense of recollection,” emotional (especially negative) items seem to expedite “old” judgements due to increased persuasive misleading familiarity information for lures. Still, it has not been directly examined whether or not faster reaction times for emotional lures are actually based on misremembering them as studied items. The diffusion model (Ratcliff, 1978) might be the best candidate to clarify this issue by virtue of better accounting for response time (RT) data.
The diffusion model
The diffusion model is a sequential sampling model accounting for the underlying cognitive processes of simple two-choice decisions (Ratcliff, 1978; Ratcliff & McKoon, 2008). The model provides a better understanding of recognition memory performance by considering the proportion of correct and error responses as well as the full reaction time distribution for both. The model allows for the individual estimation of each participant’s retrieval process by separating it into different components that represent the ability to take in information from the stimulus, the decision biases and speed–accuracy tradeoff adopted by the decision maker, and non-decision delays such as the time required to press a response key after a decision is made. Thus, the model allows researchers to uncover the individual cognitive mechanisms underlying performance.
The model assumes that the decisional component of retrieval is a noisy process and that evidence is accumulated over time from a starting point (z) towards one of the two decision boundaries (conventionally positioned at 0 and a) during this process. Evidence supporting an “Old” response accumulates towards the top response boundary (a), whereas evidence supporting a “New” response accumulates towards the bottom response boundary (0). Once the evidence accumulation reaches one of the boundaries, the decision process ends, and a response is executed. The distance between the two boundaries is assumed to reflect the speed–accuracy tradeoff engaged in by participants when making a decision. Greater distance indicates a more cautious response style in which the decision maker requires stronger evidence to terminate the decision process, thereby leading to slower but more accurate responses. If the starting point, z, is closer to one of the boundaries, than the decision is more likely to terminate at that boundary. For example, in a recognition memory task, if z is closer to the top boundary than the bottom one, the participant is biased to respond “Old” more often than “New.” Therefore, in the diffusion model, z is characterised as an a priori response bias that the participants have before the evidence accumulation process begins.
According to the model, the rate of evidence accumulation (i.e., drift rate, v) represents the quality of information extracted from the stimulus. For a recognition memory task, drift rate reflects the memory strength of the test item (Ratcliff et al., 2004). If the memory strength is high, then information accumulates quickly towards the decision boundary and drift rate will be high. Thus, any factor increasing the quality of memory strength (such as presentation duration, word frequency, or distinctiveness) will increase the drift rate, as well. Still, within-trial memory strength varies from moment to moment due to the noisy decision process. This variability indicates that processes with the same mean drift rate might not terminate at the same time (resulting in RT distributions) and also not at the same boundary (resulting in a mixture of correct responses are errors). We note that, although the basic diffusion model was introduced as a four-parameter (v, a, z, to) model described above (Stone, 1960), the current model includes several other parameters representing across-trial variability in the “main” parameters (see the studies by Ratcliff & McKoon, 2008; Wagenmakers, 2009 for further information).
A diffusion model-based approach to emotional memory
Despite the significant advantages of the model, only a few studies have examined the effects of emotional content of the stimulus on recognition memory decisions using diffusion modelling (Bowen et al., 2016; Kapucu, 2010; Spaniol et al., 2008). For example, Bowen et al. (2016) questioned what kind of cognitive mechanism underlay the negativity bias in memory in younger adults. In particular, they were interested in two distinct cognitive mechanisms, which could be distinguished thanks to the diffusion model: response bias and memory bias. Response bias was conceptualised as a motivational preference for one of the two memory decisions (i.e., old or new), and memory bias as a general tendency to extract mnemonic information in favour of either an “old” or a “new” response. In terms of diffusion model parameters, response bias was characterised by the relative position of the starting point between the two decision boundaries (i.e., z/a), while memory bias was characterised by the sum of the drift rates across target and lure items (i.e., vold + vnew). Bowen et al. (2016) showed that participants generally adopted more liberal responding for emotional items than for neutral ones, while liberal response bias for highly arousing items was observed only after a short retention interval. The more striking results from this study regarded memory bias. In line with the misleading retrieval account, negative or highly arousing items remarkably facilitated the information accumulation favouring the “Old” decision compared with positive or low-arousing items, respectively. That is, negative or highly arousing items gave rise to a familiarity memory bias in younger adults independent of actual discriminability. Bowen et al. also questioned at what stage of the memory process emotion modulation manifested by testing the memory in 1-day and 7-day retention intervals and showed familiarity memory bias for those items in both retention intervals. In support of the findings by Bowen et al., Kapucu (2010) also found enhanced familiarity memory bias (i.e., higher vold + vnew) for negative words compared with neutral ones in both immediate and delayed testing, which indicates that negative-related memory bias might manifest both before and after consolidation. Moreover, using a similar approach, Spaniol et al. (2008) showed that younger adults experience greater novelty memory bias (i.e., lower vold + vnew values) for positive items than older adults, which indicates a greater tendency to extract mnemonic information in favour of a “new” response for those items in young adults. In addition to existing findings on response bias, diffusion model findings have provided new evidence for the idea that the increase false-alarm rate for emotional lures is driven by misleading retrieval.
Drift rates and bias
The diffusion model results indicating an effect of emotion on memory bias are consistent with the claim that participants tend to retrieve stronger misleading memory evidence for emotional lures than neutral lures; in other words, emotional lures seem more like they might have been one of the studied items. However, there are alternative explanations for an effect of emotion on drift rates for lures. First, it is possible that emotion has no effect on the amount of evidence retrieved for lure items and instead affects the drift criterion, which is the standard for how strong the match to memory must be to count as evidence for an “old” response (Starns et al., 2012). Moving the drift criterion produces an equivalent shift in drift rates across stimulus classes, and applying a lower (more liberal) drift criterion for emotional than neutral words would shift drift rates to higher values for emotional stimuli. In other words, the drift criterion can produce a memory bias, to use the terminology of Bowen et al. (2016). Second, it is possible that higher false-alarm rates for emotional lures reflect an emotion-biased guessing process in response to failed retrieval. This explanation is consistent with a two-high-threshold (2HT) model of recognition memory (Snodgrass & Corwin, 1988), in which participants either detect whether the current stimulus is old or new (successful retrieval) or fail to detect and make a guess that is independent of the item’s presentation history (failed retrieval). If guessing is more biased towards an “old” response for emotional items, then these items would have a higher hit rate and false-alarm rate than neutral items. These accuracy effects could drive the drift rate differences found in previous studies.
This study
Our goal was to use a new recognition memory paradigm to investigate the possibility that higher false-alarm rates for emotional lures are driven by misleading memory retrieval. We investigated this claim by testing diffusion model predictions in a combined old–new and forced-choice recognition task. According to the model, when memory errors are triggered by misleading information, evidence accumulation tends to quickly reach the wrong boundary. Namely, the model assumes that fast errors indicate more persuasive misleading information than slow errors, which results in unavoidable errors (Starns et al., 2018). For example, new test words carrying greater misleading familiarity information would force evidence to accumulate inaccurately and quickly towards the top “Old” boundary, thus would tend to result in fast false alarms. In a similar vein, old test words with greater misleading novelty information would force evidence to accumulate inaccurately and quickly towards the bottom “New” boundary, and thus would tend to result in fast misses. In both cases, the decision process ends up with high reversed drift rate; that is, a drift rate that quickly approaches the boundary for the incorrect response. Based on these diffusion model predictions, we aimed to test three hypotheses about the processes that drive higher false alarm rates for emotional stimuli. Before explaining each specific hypothesis below, we will first explain how we measured misleading memory retrieval to make our hypotheses clearer.
Recently, Starns et al. (2018) experimentally demonstrated that the speed of memory errors showed the influence of misleading information using diffusion model-based analysis. In their study, participants were asked to memorise lists of neutral words and then to complete a recognition test with both old–new and forced-choice trials. The test was arranged in blocks that began with old–new trials and ended with forced-choice trials. The two words in a given forced-choice trial always got the same response on the earlier old–new trials even though one word was a target, and one was a lure. Thus, each forced-choice trial always had one word with a previous correct response and one with a previous error, and participants were asked to correct their previous errors by deciding which of the two words on the screen was actually studied. Model-based analysis revealed that the accuracy of forced-choice trials was related to the RT for the previous error response on the old–new recognition test block. Compared with slow errors, fast old–new errors were associated with lower accuracy in a subsequent forced-choice test block. This was consistent with a diffusion model analysis showing that fast errors were more likely to be based on persuasive misleading information from memory compared with slow errors. That is, fast errors more often had drift rates that approached the incorrect boundary. Starns et al. (2018) noted that the effect was inconsistent with a high-threshold model that assumes that participants make false alarms because they sometimes guess “old” from a failed retrieval state, not because they sometimes retrieve misleading evidence. Voormann et al. (2021) conducted a pre-registered replication of this effect and showed that it generalises to a procedure designed to rule out an alternative explanation of the effect in terms of decision strategies.
The combined old–new and forced-choice paradigm described above provides an excellent opportunity to test the claim that emotional lures trigger misleading memory retrieval to a greater extent than neutral lures. In this study, we used a test procedure that was closely based on the original design by Starns et al. (2018). Each trial began with the presentation of a single word for an old–new decision. When the response was made, a second word appeared alone on the screen for 1 s, and then both words appeared together for a forced-choice response (i.e., the participant selected the word that they thought was presented in the study phase). Unlike the blocked design by Starns et al., there were no intervening trials between old–new and forced-choice decisions and words were not paired based on participant’s old–new responses; instead, we paired a target and a lure for each test trial before the recognition test began. This procedure ensured that participants made an equal number of forced-choice decisions for each emotion class. So, we had two trial types based on whether the target or the lure would be presented first for the old–new decision (i.e., target-first and lure-first trials). Participants were told that the forced-choice decisions would serve as a second chance to correct possible errors in the preceding old–new decision. We used all possible combinations of trial type, target type (emotional or neutral), and lure type (see Table 1).
Trial types used in this study.
ET: emotional target; EL: emotional lure; NT: neutral target; NL: neutral lure. First words were presented for the old–new decision followed by the second for the forced-choice decision. n values represent the total number of trials for each condition, except for filler and practice trials.
We tested three main hypotheses. First, we expected to find a higher false alarm rate for emotional lures than neutral lures in old–new recognition (Hypothesis 1). In other words, we expected to replicate the emotion effect on false-alarm rate reported in previous studies (Bowen et al., 2016; Dougal & Rotello, 2007; Kapucu et al., 2008). The goal of including the forced-choice task was to help reveal the processes that drive the increase in false alarms by permitting a more stringent test of the misleading retrieval account compared with past studies. As we demonstrate with model simulations in the “Predictions and analysis plan” section below, implementing the misleading retrieval mechanism in the diffusion model motivates two additional hypotheses for the forced-choice data and the relationship between old–new and forced-choice performance.
Hypothesis 2 states that forced-choice accuracy should be lower for trials that include an emotional lure versus a neutral lure when the target type is held constant; that is, accuracy should be lower for neutral-target-emotional-lure trials than for neutral-target-neutral-lure trials and also for emotional-target-emotional-lure trials than for emotional-target-neutral-lure trials. The misleading retrieval account assumes continuous distributions of memory evidence, a basic property of the diffusion model. Thus, a natural way to make forced-choice decisions is to compare the memory strength of the two items and choose the one with a higher value. Although people do sometimes make absolute judgements in forced-choice tasks without considering both items (Starns et al., 2017), our specific paradigm was designed to encourage relative judgements. Namely, one item must be considered for an old–new decision, and the other was then presented alone to give participants a chance to assess memory for this item before they responded for the forced-choice decision when both items reappeared on the screen. Thus, we are confident that participants considered both items and chose the one with a stronger match to memory. If emotional lures tend to produce higher levels of misleading retrieval than neutral lures, then it is more likely that an emotional lure would exceed the memory strength of the target that it was paired with, producing an error. Thus, all other things equal, forced-choice percentage correct should be lower for trials with an emotional versus a neutral lure.
Hypothesis 3 states that forced-choice accuracy should be lower following fast versus slow errors for the old–new judgement. As discussed by Starns et al. (2018), this pattern is a signature of misleading retrieval. That is, items with strong misleading retrieval should tend to get fast error responses because they have strong “reversed” drift rates (e.g., a high positive drift rate for a lure), and strong misleading retrieval should also impair forced-choice accuracy (e.g., a lure with strong misleading evidence has a good chance of exceeding the memory strength of the target that it is paired with).
If all three hypotheses of interest are supported by the data, then results will provide strong evidence in favour of the misleading retrieval hypothesis. Critically, the anticipated results would also provide specific evidence for misleading retrieval, as decision-based accounts cannot explain a data pattern that confirms all three hypotheses. To explain why, we considered two versions of a decision-process account. One version holds that emotion affects the criterion for the amount of memory strength needed to support an “old” response, represented by the drift criterion parameter in the diffusion model (Starns et al., 2012). If the emotion effect on old–new recognition is produced by a shift in the drift criterion, then the effect should be eliminated in forced-choice trials, as they involve comparing two items instead of evaluating a single item against a criterion. Given that a strength criterion plays no role in the relative judgement, drift criterion effects are irrelevant for the forced choice task, and percentage correct should be equal for trials with emotional versus neutral lures. In other words, shifting the drift criterion could not account for results that confirm Hypothesis 2; that is, results in which forced-choice accuracy differs based on the type of lure (emotional or neutral) included in a test pair.
The second version of the decision-based account holds that emotion effects on false-alarm rate are based on changes in guessing tendencies in the face of failed retrieval. This account is based on a high-threshold characterisation of recognition memory (Snodgrass & Corwin, 1988), according to which retrieval either succeeds or fails, with failure resulting in a “guess” state that does not distinguish between targets and lures. Such a model can produce higher false-alarm rates for emotional versus neutral lures by assuming that participants are more willing to guess that an emotional lure was studied. Critically, any effect of emotion on guessing would also apply to forced-choice trials, because sometimes the participant would fail to retrieve information for both test items and would have to select a response based on guessing. Thus, in contrast to the drift-criterion explanation, a guessing account could also accommodate results that verify Hypothesis 2 by showing lower forced-choice accuracy for trials with an emotional versus a neutral lure. Luckily, a guessing account could not explain results that verify Hypothesis 3 by showing lower forced-choice accuracy following fast versus slow old–new errors. As demonstrated by Starns et al. (2018), the error-speed effect is uniquely predicted by models that interpret memory errors as the product of misleading retrieval as opposed to uninformed guessing. If errors are based on guessing in a state of failed retrieval, then there is no systematic difference between errors that are made quickly and slowly. In other words, there is no equivalent to the diffusion model pattern in which fast errors signify stronger misleading retrieval, because none of the errors are based on misleading retrieval in this account.
In summary, our main analyses tested three hypotheses inspired by the misleading retrieval account; that is, the claim that emotional lures tend to retrieve stronger misleading evidence of being studied than neutral lures. Hypothesis 1 states that old–new false-alarm rates will be higher for emotional than neutral lures. Hypothesis 2 states that forced-choice accuracy will be lower on trials that include an emotional versus a neutral lure given the same type of target item. Hypothesis 3 states that forced-choice accuracy will be lower following a fast old–new error than a slow old–new error; in other words, it will be more difficult for participants to recover from fast errors than slow errors. The misleading retrieval account predicts that all three hypotheses will be confirmed. Alternative decision-based accounts cannot explain results that verify all three hypotheses.
Hypothesis 1: FAR difference
Hypothesis 2: Forced-choice accuracy
Hypothesis 3: RT bias effect
Methods
Participants
Power analyses using the effect sizes predicted by the diffusion model were performed and reported in the pre-registration article, as explained in detail in the online Supplementary Material A. Based on these analyses, 70 participants with usable data were adequate to consistently detect effects predicted by the diffusion model under the assumption that emotional lures produce stronger misleading retrieval than neutral lures. As we pre-registered, we excluded participants with near-chance old–new recognition performance, defined as a difference less than 0.1 between the hit rate and the false alarm rate, and those with more than 10% of their RTs below 400 ms (Starns et al., 2018; Voormann et al., 2021). Considering possible technical difficulties as well, we first planned to recruit 20% more participants than the minimum number of valid datasets required. Actual recruitment went a little over this target. In total, we recruited 93 undergraduates aged between 18 and 35 years from the University of Massachusetts Amherst, with all participants compensated with extra course credit. The proportion of participants excluded by the criteria above was much lower than we anticipated; indeed, we excluded only one participant’s data by those standards. We also lost one participant’s data due to technical difficulties during data collection. Therefore, all planned and exploratory analyses were conducted on data that included 91 participants in total. No other exclusion or inclusion criteria were used for participation.
Materials
All materials, analysis codes and data files are publicly available on Open Science Framework (OSF) website (https://osf.io/f7a4j/?view_only=c5a7b71a11da4358899d1af1c4a0d3aa).
Emotional words
The words were selected from the English Lexicon Project (ELP, Balota et al., 2007), a large database that provides standardised descriptive data for variety of word characteristics such as frequency and concreteness for over 40,000 words. One-hundred sixty-eight words were selected for each emotion condition from this database. One-hundred forty-four words from each condition were used as old or new words in the test phases, while the remaining ones were used in trials probed for recall in the study phases. Also, an additional 75 neutral words were selected to use in filler trials or in practice trials. In the most recent version of the ELP, valence and arousal values were gathered from the norming study by Warriner et al. (2013). In the norms by Warriner et al., valence values were ranging from unhappy to happy, while arousal values ranged from calm to excited on a 9-point Likert-type scale. Thus, lower ratings mean more negative or less-arousing words, while higher ratings mean more positive or more-arousing words in this database. We selected the words based on these norm values so that the emotional test words and recall words were more negative (for test words, t (286) = 34.56, 95% confidence interval [CI] = [2.39, 2.68], for recall words t (46) = 17.87, 95% CI = [2.65, 3.32]) and more arousing (for test words, t (286) = 27.54, 95% CI = [–2.29, –1.98], for recall words t (46) = 9.08, 95% CI = [–2.38, –1.51]) than their neutral counterparts. To prevent any confounding effects stemming from word characteristics, emotional and neutral test words as well as recall words were matched on word frequency (Lund & Burgess, 1996), word length, and concreteness based on descriptive ELP data, all t < 1.96. Also, to make sure that semantic similarity within list words would not lead any confounding effects on recognition accuracy (Dougal & Rotello, 2007), we calculated semantic interrelatedness of lists by conducting latent semantic analysis with the LSAfun R package (Günther et al., 2015; Landauer et al., 1998), thereby emotional and neutral test words were matched on semantic interrelatedness as well, (t (286) = 0.21, 95% CI = [−0.01, 0.02]. The mean and standard deviation values of word characteristics for each emotion condition are presented in Table 2.
The mean and standard deviation values of word characteristics for each emotion condition.
SD: standard deviation; N/A: not available.
Emotional test words as well as recall words were matched on word frequency, word length, and concreteness with their neutral counterparts. Only emotional and neutral test words were matched on semantic interrelatedness, all t < 1.96.
Log-transformed estimates of word frequency were used as the raw frequency values were highly skewed (see the study by Balota et al., 2007).
Recognition memory task
In this experiment, participants studied a list of words and then completed a recognition test. A modified version of the original recognition memory test by Starns et al. (2018) was used to ensure that participants made equal numbers of forced-choice decisions for each emotion class. Like Starns et al., the test had both old–new trials and forced-choice trials, and forced-choice trials provided the opportunity to correct potential errors on the old–new trials. Unlike Starns et al., each old–new decision was followed immediately by the corresponding forced-choice decision (Starns et al. used blocks of multiple old–new decisions followed by multiple forced-choice decisions). Also, unlike Starns et al., we did not ensure that every word on the forced-choice test had the same previous old–new response. Following this procedure could result in different number of forced-choice decisions for each emotion class, and we instead had an equal number of forced-choice trials for all the trial types listed in Table 1. Initially, one of these words appeared alone for an old–new decision. After the participant responded, the second word appeared alone for 1 s, and then both appeared side-by-side for a forced-choice decision.
Negative and neutral words were used in this study. The experimental session included one practice cycle and three main study-test cycles. The practice cycle was composed of only neutral words not to be used in the further main study-test cycles, while each study-test cycle included equal numbers of negative and neutral words, except that all filler words were neutral.
On each study-test cycle, participants first studied a total of 72 words presented one-at-a-time in random order for 2,000 ms each in a block of four words with a 100 ms inter-stimulus interval. The first and the last blocks contained only filler items to control for primacy and recency effects. To ensure the memorisation of the study words, participants were asked to recall one of the randomly selected words from the last studied block of four words (first, second, third, or fourth word). Half of the words to be probed for recall in the main study blocks were negative, and the other half were neutral. If participants could not recall the selected word correctly, they saw an error message. Words probed for recall were not included in the later test phases.
Immediately after the study phase, participants completed a recognition memory test including 24 studied targets and 24 non-studied lures for each emotion condition (negative and neutral). Before the test began, targets and lures were pre-matched randomly to create word pairs with the following constraints: In half of the pairs, both words came from the same emotion conditions (i.e., negative–negative or neutral–neutral target–lure pairs), whereas those in the remaining pairs came from different emotion conditions (i.e., negative–neutral, neutral–negative target–lure pairs). So, a total of 48 word pairs were created. On each test trial, one of the words from each pair appeared alone on the left side of screen and was presented until the participant made an old–new recognition decision for that word. After the response, the word disappeared and the other word from the same pair was presented for 1,000 ms on the right side of the screen. Then both words were presented together, and participants were asked to choose which of the two words was actually an old word. Participants were not allowed to respond until both words appear together on the screen. Thus, participants made two types of recognition decisions on each trial: first, an old–new decision and second, a forced-choice decision. For half of the trials, the target word was presented first for the old–new decision followed by the lure for the forced-choice decision. For the remaining trials, the lure came first followed by the target. So, target-first and lure-first trials were constructed (see Table 1). Because participants made a forced-choice decision immediately after the old–new decision on the same test trial, possible delay-related memory changes should be less of a concern than that in the study by Starns et al. (2018). Also, considering the participants’ faster and more intuitive decision tendencies later in a task (Dutilh et al., 2009), making consecutive old–new and forced-choice decisions allowed us to prevent any confounding effects based on trial position.
Using this design, participants completed a recognition test of 48 trials in random order on each cycle, which included equal number of trials for each trial type (see Table 1). At the beginning of each test cycle, there were two additional filler trials comprising of two targets from the filler trials of the study phase and two lures that were not presented before in the experiment. Responses to these filler items were not included in the analyses. Thus, the recognition test included a total of 50 trials per cycle.
Procedure
Participants completed the experimental session individually in a dimly lit room. They were informed that they would participate in a memory task comprising of four study-test cycles in which they would study a list of words and then complete a recognition memory test for those words. The experiment was programmed using E-Prime 2.0 software.
On each cycle, participants studied lists of negative and neutral words in blocks of four words, as described in the previous section. After each block, a response screen appeared, and the participant was instructed to recall and type in one of the words from the latest study block by position (e.g., “Recall the second word”). If participants’ responses did not match the word prompted for recall, an ERROR message was displayed on the screen for 1,000 ms. The next block proceeded after a blank screen for 500 ms.
Following the study phase, participants completed a recognition test including both old–new and forced-choice decisions in each test trial. They were presented with one word and asked to decide whether the presented test word was “Old” or “New” by pressing the “z” or “/” keys on the keyboard. They were instructed to be as accurate as they could, but to respond as soon as they knew the answer. We decided to use this instruction since Starns et al. (2018) has previously showed that predicted error RT effects on forced-choice decisions are larger when participants have more cautious responding style. Importantly, participants were informed that they would have a chance to correct their possible errors that they might have made in their most recent old–new decision. They were told that one more word would be presented along with the word for which they have made the old–new decision, but only one of these two words was actually studied. So, they were asked to select the word that was actually presented on the study list using the “z” or “/” keys to indicate the left or right word, respectively. Participants were told to select whichever word they thought was most likely to be on the study list even if the forced-choice response contradicted their previous old–new response.
Throughout the experimental session, participants completed three main study-test cycles including a total of 216 study trials and 150 recognition test trials (6 of them were fillers) along with a practice cycle. The practice cycle was identical to the main cycles except for the total number of trials. Participants studied a list of 28 words in blocks of 4 words with 4 fillers at the beginning and 4 at the end. Next, they completed 17 recognition test trials.
Results
Predictions and analysis plan
All codes that implement the simulations and the primary and exploratory analyses are available on the project’s OSF website (https://osf.io/f7a4j/). Before the first-stage submission of this registered report, we defined predicted patterns of results with diffusion model simulations using the rtdists R package (Singmann et al., 2019).
The online supplementary Material A provides details on the prediction simulation methods and results as reported in our pre-registration article. Here, we will just provide a summary of the distribution of critical effects across the 500 simulation runs. For the effect of emotion on old–new false-alarm rates (Hypothesis 1), the distribution of effect sizes was well described by a Gaussian distribution with a mean of 0.06 and a standard deviation of 0.013, where effect size is defined as the emotional false-alarm rate minus the neutral false-alarm rate. For the effect of lure emotion class on forced-choice percentage correct (Hypothesis 2), the distribution of effect size was well described by a Gaussian distribution with a mean of 0.04 and a standard deviation of 0.01, where effect size is defined as forced-choice percentage correct for pairs with a neutral lure minus forced-choice percentage correct for pairs with an emotional lure. For the relationship between old–new error speed and forced-choice accuracy (Hypothesis 3), the distribution of effect size closely followed a Gaussian distribution with a mean of 0.07 and a standard deviation of 0.018, where effect size is defined as forced-choice percentage correct for trials with a slow (above the median) error RT minus trials with a fast (below the median) error RT.
Descriptive statistics
Before presenting the findings from primary and exploratory analyses, we briefly reported some descriptive statistics of the empirical data. Participants accurately recalled 85% of the emotional and 84% of the neutral study words prompted for recall in the study phase, indicating that they complied with the instructions to memorise the study words. In the test phase, neutral targets had a hit rate of 0.68 compared with 0.77 for emotional targets, while neutral lures had a false alarm rate of 0.17 compared with 0.23 for emotional lures. Thus, the empirical data were in line with the previous findings showing that emotion increased the likelihood of recognising the studied as well as unstudied items (Kensinger & Schacter, 2008). The average RT median was 1,198 ms for hits, 1,379 ms for misses, 1,246 ms for correct rejections, and 1,545 ms for false alarms. The average RT median values for emotional and neutral test items can also be seen in Table 3. The overall pace of responding was in line with the recognition tasks without time pressure (Starns et al., 2018; Voormann et al., 2021), and results replicated the common finding that error responses tend to have longer latencies than correct responses (Ratcliff, 1978; Starns, 2014).
Mean (M) and range in RT medians for each condition in the empirical data.
The values expressed in bold represent the mean values and corresponding ranges obtained from the cutoff points used to group participants’ errors into fast and slow.
Primary inferential analyses
Following our primary analysis plan, our first critical comparison was to test Hypothesis 1, which concerns old–new false alarm rate difference between emotional versus neutral lures. In the empirical dataset, neutral lures had a false-alarm rate of 0.17 compared with 0.23 for emotional lures, t(90) = 6.28, 95% CI = [0.04, 0.07]. Thus, the empirical data confirmed the prediction from our diffusion simulations, as participants were more likely to claim that they studied emotional lures versus neutral lures.
To evaluate our Hypothesis 2 that concerns the emotion effects on forced-choice percentage correct, we conducted a 2 (target emotion class) × 2 (lure emotion class) repeated measure analysis of variance (ANOVA). Table 4 shows mean forced-choice percentage correct values for different target and lure emotion classes. Note that the outcome of the preceding old–new decision was not a variable of interest in this analysis, thus we included all forced-choice trials regardless of participants’ previous old–new decision (see Table 1). A significant main effect of lure type confirmed the predictions obtained from our simulations by showing that performance was worse for pairs with an emotional lure (0.83) than pairs with a neutral lure (0.86), F(1, 90) = 20.02, 95% CI = [0.02, 0.04]. There was also a main effect of target type, as the percentage correct was higher for pairs with an emotional target (0.86) than pairs with a neutral target (0.82), F(1, 90) = 33.18, 95% CI = [−0.06, −0.03]. The results showed no evidence of an interaction between target and lure emotion class, F(1, 90) = 0.19, 95% CI = [−0.03, 0.02], which replicated our simulated results and indicated that the size of the lure-type effect was very similar for pairs with emotional and neutral targets as well as for pairs that had items from the same emotion condition versus pairs with items from different emotion conditions.
Mean forced-choice percentage correct values for each condition in the empirical data.
As a last critical planned comparison, we tested Hypothesis 3 concerning the relationship between old–new error RT and forced-choice accuracy with a paired-samples t-test. To do that, we first grouped old–new error RTs into fast and slow errors for each subject following the same rule that we applied for our simulated data. Namely, participants’ error RT medians for targets called “new” and lures called “old” were used as cutoff points to group their errors into fast and slow. Table 3 shows the mean values of and range in these cutoff points. Note that this analysis only included trials for which participants made an incorrect old–new decision. The average numbers of forced-choice trials preceded by incorrect old–new decisions were represented in Table 5. Our empirical data demonstrated that forced-choice percentage correct was lower following fast (0.58) than slow (0.60) old–new errors. Yet, this difference was smaller than we expected based on the prediction simulations and was nonsignificant with a paired-samples t-test, t(90) = 1.60, 95% CI = [−0.01, 0.06]. Thus, results were inconclusive, as the confidence interval included null or slightly reversed effects on the low side and also effects that were fairly consistent with the average effect size in our prediction simulations (0.07) on the high side. Recognition accuracy was higher overall in this study than in the study that provided parameter values for the prediction simulations (see the online Supplementary Material A), and the predicted size of the error RT effect decreases when accuracy increases (Starns et al., 2018). This could be a potential factor in the mismatch between the predicted and observed effect size.
Mean number of forced-choice trials preceded by incorrect old–new decisions in the empirical data.
The values expressed in bold represents the average number of forced-choice trials preceded by incorrect old–new decisions that were used to test Hypothesis 3. We note that emotion condition of neither target nor lure was a variable of interest in these analyses testing Hypothesis 3.
Exploratory analyses
As we previously stated in Footnote 1 of the pre-registration article, we also conducted 2 (trial type) × 2 (target emotion class) × 2 (lure emotion class) three-way repeated-measure ANOVA on forced-choice percentage correct (i.e., Hypothesis 2) to check for a possible main effect of trial type or an interaction of trial type with other variables (see Table 4). This analysis revealed that neither main effect of trial type, F(1, 90) = 0.15, 95% CI = [−0.01, 0.02], nor its interactions with other variables were significant, all F < 1.96, confirming our prediction that trial type (target-first or lure- first) would have no effects on forced-choice percentage correct, as the participant always had ample time to consider both items in a trial before making the forced-choice response.
We also investigated possible trial type effects on the relationship between old–new error RT and forced-choice accuracy to see if the error RT effect (i.e., lower forced-choice accuracy following fast old–new error response, Hypothesis 3) would emerge in a specific trial type in the empirical data (see Table 6). Exploratory paired sample t-tests showed that forced-choice percentage correct was lower following fast (0.54) than slow (0.61) old–new errors only for target-first trials, t(90) = 2.67, 95% CI = [0.02, 0.12], but not for lure-first trials, t(90) = 0.07, 95% CI = [−0.06, 0.06]. These findings indicated that our initial expectation that recognition errors are driven by misleading retrieval was partially supported by the empirical data. Although we planned to evaluate the RT effect by combining across trial types, we can check the pre-registered prediction simulations to determine whether the diffusion model predicts a larger effect for target-first trials, as observed. In fact, this pattern was consistent with the model predictions. Across all simulation runs, the RT effect was 0.091 for target-first trials versus 0.036 for lure-first trials.
Mean forced-choice percentage correct values following fast and slow incorrect old–new responses for the empirical data.
Hierarchical Bayesian logistic regression models
We also applied hierarchical Bayesian logistic regression models similar to the ones reported by Starns et al. (2018). We aimed to provide a Bayesian equivalent of our critical analyses, providing a test of whether our conclusions are robust across different analysis approaches.
For all the models reported below, each participant was allowed unique logistic parameters; for example, each participant in the false-alarm rate model had parameters defining the emotion effect. The logistic regression parameters were assumed to follow Gaussian distributions across participants with unique means (μ) and standard deviations (σ), which resulted in two across-participant hyperparameters for each logistic parameter. The online Supplementary Material B provides details on the logistic regression parameters that we used to estimate probability of a recognition response. We used loosely informed priors for the across-participant parameters for all models, as reported in the analysis code files that are publicly available on the project’s OSF website (https://osf.io/f7a4j/). We used JAGS to sample from posterior distributions (Kruschke, 2015), and for each model we ran four chains of 5,000 samples each with a thinning value of 50 and a burn-in period of 1,000 samples. Thus, a total of 20,000 samples were created for each model. We ensured that the chains converged by visually comparing posterior distributions across chains and computing the Gelman–Rubin statistic (Gelman & Rubin, 1992). This statistic was lower than 1.01 for all parameters in all models, which demonstrates good convergence between chains (Brooks & Gelman, 1997). We assessed credible intervals (CIs) as a way of summarising uncertainty of the posterior samples (Kruschke, 2015). CI indicates which points of a distribution are most credible by specifying an interval covering 95% of the distribution, and any point that falls within 95% CI has higher credibility than any point outside the interval. Thus, in this study, any effect of interest that does not include zero in 95% CI was interpreted as strong evidence for an effect in the specified direction. Table 7 shows median values and CIs for the posterior distribution for the effect of interests in the Bayesian models.
Median (M) values and 95% credible intervals (CIs) for the posterior distribution of the across-participant effects.
I: intercept; E: emotion; TT: trial type; TE: target emotion; LE: lure emotion; RT: response time group.
Our first critical comparison was to test the emotion effect on false alarm responses (Hypothesis 1). The posterior distribution for the average intercept value across participants (μi) was –1.50 (CI = [–1.62, –1.36]), corresponding to an overall false alarm rate of about 0.18 (see Model 1 in the online Supplementary Material B). The posterior distribution for the average emotion effect (μel) had a median of 0.34 (CI = [0.21, 0.34]), which confirmed the frequentist results above and showed that false alarm rates were higher for emotional lures (CI = [0.19, 0.23]) than neutral ones (CI = [0.14, 0.18]).
To test our Hypothesis 2 with a Bayesian approach (see Model 2 in the online Supplementary Material B), we did not concern the preceding old–new decision outcome and included all trials in this logistic model that corresponds to our previous 2 (trial type) × 2 (target emotion class) × 2 (lure emotion class) three-way repeated-measure ANOVA. The posterior median for the average intercept value across participants (μ i ) was 1.81 (CI = [1.70, 1.93]), which indicates a proportion correct of 0.86 for the overall forced-choice decisions. Consistent with the frequentist analyses, the posterior distribution of the across-participant average lure emotion effect had a median of −0.26 (CI = [−0.38, −0.16]), and demonstrated strong evidence of lower forced-choice accuracies for emotional lures (CI = [0.82, 0.86]) compared with neutral lures (CI = [0.86, 0.89]), as none of the 20,000 posterior samples were below zero (see Table 7). We also observed strong evidence that forced-choice accuracies were higher for pairs with an emotional target (CI = [0.86, 0.89]) than pairs with a neutral target (CI = [0.82, 0.85]), and the posterior for the across-participant mean target emotion effect had a median of 0.35 (CI = [0.24, 0.47]). As we expected, there was no clear indication for the interaction of target and lure emotion classes (−0.10, CI = [−0.31, 0.09]), which confirmed our frequentist results showing that the size of the lure-type effect was very similar for pairs with emotional and neutral targets as well as for pairs that had items from the same emotion condition versus pairs with items from different emotion conditions. The posterior distribution for the across-participant average trial type effect had a median of 0.002 (CI = [−0.11, 0.17]), which indicates no clear difference between target-first and lure-first trials. All remaining interaction effects did not rule out a null effect, as the CIs included zero (see Table 7). Therefore, the Bayesian logistic regression results reach the same conclusion with frequentist results for the forced-choice accuracy.
Finally, as we observed that the error RT effect only emerged in a specific type of trial, to test the Hypothesis 3, we preferred to set up a Bayesian logistic regression model (see Model 3 in the online Supplementary Material B) such that it corresponds to a 2 (trial type) × 2 (RT group) two-way repeated-measure ANOVA. 1 Similar to the standard frequentist analysis, we included only trials in which participants made an error old–new decision and defined fast and slow responses as those above and below the RT median, respectively. The posterior for the average intercept value across participants (μi) was 0.34 (CI = [0.22, 0.47]), which corresponds to a poor but above-chance performance of 0.58 for the forced-choice decisions after old–new errors. The posterior for the mean error RT effect had a median of 0.11 (CI = [−0.05, 0.26]), which indicates participants’ forced-choice accuracies tended to be higher after a slow old–new error compared with a fast one, as we hypothesised, but the results did not strongly rule out no overall effect or even a small effect in the other direction. Consistent with the frequentist analyses, the Bayesian model showed clear evidence for an RT effect on target-first trials (CI = [0.01, 11]) but not lure-first trials (CI = [−0.06, 0.05]). Results showed moderate support for a larger RT effect in target-first than lure-first trials, as the posterior distribution for the average intj had a median of 0.26, and 95% of the posterior samples were above zero. However, we could not strongly rule out a null interaction effect, as the CI for the interaction effect included zero, [−0.04, 0.54]. There was no indication of a trial-type effect (−0.004, CI = [−0.19, 0.19], see also Table 7). Thus, the Bayesian logistic regression results reached the same conclusion with frequentist results for the error RT effect on forced-choice accuracy and confirmed that error RT effect emerged only in target-first trials for the empirical data.
Discussion
In this registered-report, we explored the basis of the emotion effect on memory errors by testing predictions of the diffusion model (Ratcliff, 1978) for a recognition paradigm in which participants had an opportunity to correct their recognition errors. Previous studies testing emotion effects on memory processes have rarely based their predictions on diffusion modelling, despite its potential to help researchers to understand how memory errors are produced and what kind of memory and/or decision mechanisms underlie those errors. The diffusion model predicts that fast errors are more likely to be based on persuasive misleading information from memory compared with slow errors (Starns et al., 2018; Voormann et al., 2021). To test this prediction, we used a novel recognition memory paradigm combining old–new and forced-choice recognition tasks. In each trial of this paradigm, participants made an old–new decision followed by a forced-choice decision that serves as a second chance to correct possible errors in the preceding old–new decision.
After defining predicted patterns of results via diffusion model simulations, we collected data from participants and employed both standard frequentist and Bayesian analysis approaches to test whether the predicted patterns of results were confirmed. For Hypothesis 1, the results confirmed our predictions by showing a higher false-alarm rate for emotional than neutral lures. For Hypothesis 2, results again matched the prediction simulations by showing that forced-choice accuracy was lower for pairs with emotional lures compared with pairs with neural lures. Results were less clear for Hypothesis 3. The overall error RT effect was in the same direction as the prediction simulations, but the effect was smaller than predicted and could not be distinguished from a null effect with high certainty. Follow-up analyses showed an effect for target-first but not lure-first trials, and a larger effect for the former was consistent with the prediction simulations. Thus, there was some evidence for an error RT effect, but the observed effect was smaller than the predicted effect and possibly limited to target-first trials.
Our findings (Hypothesis 1) were consistent with previous ones that emotion has been shown to substantially increase the false alarm rates. Nevertheless, to what extent this increase in false alarm rates can be driven by decision biases versus memory processes had yet to be well investigated. Indeed, some previous ROC-based research has questioned these mechanisms and suggested that increased false alarms can be linked to the increased liberal response bias for emotional stimuli (Dougal & Rotello, 2007; Kapucu et al., 2008; Yüvrük & Kapucu, 2022). Still, although ROC analysis can help researchers measure memory discriminability independently of response bias, it also fails to separate response bias from decision bias due to measuring bias after decision process ends and a response has been made (White & Poldrack, 2014). On the contrary, only limited number of studies have employed diffusion model that allows researchers to investigate decision mechanisms underlying these bias effects, and these studies found evidence that participants tend to retrieve familiarity information (i.e., memory bias) for emotional items (Bowen et al., 2016; Kapucu, 2010; Spaniol et al., 2008). Although this shift in memory bias is consistent with the claim that participants tend to retrieve stronger misleading memory evidence (e.g., a high positive drift rate for a lure) for emotional lures, alternative decision mechanism could also account for this shift. Therefore, despite the use of diffusion model, those studies also fall shorts of explaining to what extent emotion effects on false alarms stems from the misleading evidence or other decision-based accounts.
Diffusion model predicts two outcomes that supports stronger misleading retrieval account for increased false alarms to emotional lures (Starns et al., 2018; Voormann et al., 2021). First, using the trial-by-trial design, we provided strong evidence for the predicted outcome (Hypothesis 2) that emotional lures are more likely to be selected and therefore to produce a forced-choice error regardless of target emotion. Moreover, our exploratory analyses did not find a difference between target-first and lure-first trials, confirming that participants had ample time to consider both options. On the contrary, although we expected participants to have similar levels of forced-choice accuracy for pairs with emotional versus neutral targets based on our simulated data, for the empirical data, both analysis approaches provided strong evidence that pairs with emotional targets had higher forced-choice accuracy, again regardless of lure emotion class, indicating that emotional targets tended to produce higher levels of accurate retrieval (i.e., higher drift rates) than neutral targets, and thus were more likely to exceed the memory strength of the lure that they were paired with, producing a correct response (see also the online Supplementary Material C for drift rates obtained from the empirical data). These findings suggest that emotion increases the availability of information at the time of retrieval thereby producing higher memory performance for targets (Bowen et al., 2016; Dolcos et al., 2005), while it also misleads participants for lures.
For the second predicted outcome of diffusion model (Hypothesis 3), we found mixed evidence, as we observed the error RT effect only for target-first trials but not for lure-first trials, as revealed in both frequentist and Bayesian analyses. Note that although we framed our original predictions in terms of overall error RT effect regardless of trial type, the error RT effect was larger for target-first trials in our simulated data as well. This asymmetry in effect size is driven by a pattern that is common across a number of previous recognition studies; namely, the proportion of targets with negative drift rates tends to be higher than the proportion of lures with positive drift rates (Starns et al., 2012, 2018; Starns & Ratcliff, 2014), which suggests that the proportion of errors that are avoidable is lower for targets than lures. Still, previous studies using original block design by Starns et al. (2018) failed to find clear evidence of a larger RT effect for misses than for false alarms (Starns et al., 2018; Voormann et al., 2021), potentially due to the methodological constraints that we discussed above (e.g., changes in memory between the single-item and forced-choice trials). Our results suggest that the trial-by-trial design might overcome these methodological constraints and be more suitable to reveal the pattern in error RT effect for targets relative to lures. On the contrary, unlike the simulated data, our empirical data did not show an error RT effect for lure-first trials, contradicting the diffusion model prediction of a smaller, but still present effect. This might represent a Type 2 error, or it might represent a shortcoming of the model.
One potential explanation for our smaller-than-expected error RT effects is relatively high overall performance. The observed false alarm rate was lower than the prediction simulation false alarm rate by about 0.1. Starns et al. (2018) showed that the size of the error RT effect predicted by the diffusion model depends on the proportion of errors that are avoidable with additional consideration time, with larger effects for lower proportions of avoidable errors (see also the online Supplementary Material C for proportions of avoidable errors obtained from our empirical data). High performance is associated with high proportions of avoidable errors. That is, if evidence from the stimulus is clear and drift rates approach the correct boundary on most trials, then most errors come from within-trial noise that caused the accumulation process to cross the boundary opposite the average direction of drift. Thus, the increase in memory performance in the empirical relative to the simulated data might have resulted in losing the RT effect for lure-first trials and observing a smaller-than predicted effect for target-first trials. Therefore, we suggest that future studies investigating the error RT effect should aim for low to moderate memory performance.
Although we did not find clear evidence for every hypothesis, our results support the misleading retrieval account over decision-based accounts of elevated false alarm rates for emotional lures. For instance, moving the drift criterion (i.e., how strong the match to memory must be to count as evidence for an “old” response) might produce an equivalent shift in memory bias for emotional stimuli as increased misleading evidence does (Starns et al., 2012), meaning that the increased familiarity memory bias for emotional stimuli observed in previous studies (Bowen et al., 2016; Kapucu, 2010; Spaniol et al., 2008) could be accommodated with the drift criterion account as well. The drift criterion, however, is irrelevant for the outcome of forced-choice decision as it plays no role in the relative judgements in which the memory strengths of two items needs to be compared. Contrary to our findings, the drift criterion account anticipates similar levels of forced-choice accuracy for emotional vs. neutral lures. As such, our results regarding the difference in forced-choice accuracy between emotional and neutral lures cannot be accommodated with changes in drift criterion for emotional stimuli, but they are consistent with a misleading retrieval account whereby the higher memory strength for emotional lures makes them more likely to exceed the strength of a paired target.
Moreover, one might expect that emotion-biased guessing in response to failed retrieval also contributes to increased false alarm rates for emotional lures and drives the shift in memory bias (Snodgrass & Corwin, 1988). Importantly, unlike the drift criterion account, the emotion-biased guessing can account for the forced-choice accuracy difference between emotion conditions because it assumes that participants would be more willing to guess an emotional lure being studied when they failed to retrieve both forced-choice items. This study was designed to differentiate the contribution of misleading evidence and emotion-biased guessing to emotional false alarms. As the error RT effect can only be accommodated with models that assume that recognition memory errors result primarily from misleading evidence, and as guessing account anticipates no error RT effect, the error RT-effect results are critical for distinguishing these mechanisms. Unfortunately, the RT-effect results were unclear, with an effect arising for target-first trials but not lure-first trials, and no clear overall effect. However, given that the RT bias effect was observed in other similar studies (Starns et al., 2018; Voormann et al., 2021), we think the misleading evidence explanation is the best-supported account given the current state of the data.
There are also some aspects of this study that can be improved in future studies. First, investigating the asymmetry in error RT effect for target-first versus lure-first trials might clarify whether the lack of an effect for lure-first trials reflects just a Type 2 error or reflects a failed prediction of the diffusion model. Even though our simulated findings suggested that the trial-by-trial design was suitable to investigate this, our sample size and number of trials per condition might have been too low to reliably detect this asymmetry, as we designed our study to detect the overall RT effect across all trials. Future studies investigating this asymmetry with a larger sample size and/or more trials per participant may improve our understanding of the mechanism underlying recognition memory decisions.
Another limitation is that this study utilised only negatively valenced words as emotional materials, therefore we were not able to distinguish which dimension of emotion (valence or arousal) drives these effects. As we primarily aimed to investigate the basis of emotion effects on false alarms, which have been more consistent in comparisons of neutral and negative words in previous studies (Bowen et al., 2016; Dougal & Rotello, 2007; Kapucu et al., 2008), we opted to include only negatively valenced materials for the sake of parsimony. Moreover, focusing only on negative and neutral words also allowed us to increase experimental power by increasing the total number of trials for each condition. Thus, we believe that the present design was best suited to our primary goal of discovering whether the increase in false alarm rate that is consistently observed for negative words is based on decision biases or the retrieval of misleading memory evidence. Still, examining the negative versus positive or high versus low-arousing emotional materials to understand to what degree emotional valence or arousal drive an error RT effect on forced-choice performance would be a reasonable future direction. Moreover, recent studies have suggested that specific negative emotions have differential impacts on recognition memory processes. For instance, although disgust and fear are both negative and highly arousing emotions (Russell, 1980), their impacts on recognition memory can substantially differ from each other. Disgust was shown to increase memory discriminability relative to fear, while it also encouraged a more liberal response bias relative to fear (Boğa et al., 2021; Chapman et al., 2013). An ongoing study in our lab using diffusing modelling also demonstrated that familiarity memory bias for disgust-related emotional materials tends to be larger relative to those for fear-related ones (Yüvrük & Kapucu, 2021). Altogether, these findings suggest that error RT effect might emerge in different magnitudes for specific negative emotions if the bias for disgust-related material is driven by misleading memory evidence, which is worth testing in future studies.
To conclude, in this registered report, we first defined a predicted pattern of results by simulating datasets following diffusion model predictions, and then empirically tested the predicted pattern by collecting data from real participants. Conducting both standard frequentist and Bayesian analyses allowed us to see how robust our conclusions were across different analysis approaches. The approaches converged to show strong evidence of an emotion effect on false alarm rate, strong evidence of a lure-emotion effect of forced-choice accuracy, and moderate evidence of the error RT effect. Overall, these findings are consistent with the claim that misleading evidence drives false alarms for emotional lures compared with neutral ones. Future studies need to investigate which aspects of emotion guide these effects by incorporating different emotional materials into their designs.
Supplemental Material
sj-docx-1-qjp-10.1177_17470218221137347 – Supplemental material for Does misremembering drive false alarms for emotional lures? A diffusion model investigation
Supplemental material, sj-docx-1-qjp-10.1177_17470218221137347 for Does misremembering drive false alarms for emotional lures? A diffusion model investigation by Elif Yüvrük, Jeffrey Starns and Aycan Kapucu in Quarterly Journal of Experimental Psychology
Footnotes
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by The Scientific and Technological Research Council of Turkey (TUBITAK) within the scope of 2214-A International Research Fellowship Programme for PhD Students [grant number 53325897-115.02-121630].
Data accessibility statement
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
