Abstract
Although long-term memory and Theory of Mind (ToM) are closely related across the whole lifespan, little is known about the relationship between ToM and semantic memory. Clinical studies have documented the co-occurrence of ToM impairments and semantic memory abnormalities in individuals with autism or semantic dementia. However, to date, no study has directly investigated the existence of a relationship between ToM and semantic memory in the typical population. We addressed this gap on a sample of 103 healthy adults (M age = 22.96 years; age range = 19–35 years). Participants completed a classical false memory task tapping on semantic processes, the Deese-Roediger-McDermott (DRM) task, and two ToM tasks, the Triangles and the Reading the Mind in the Eyes task. They also completed the vocabulary scale from the Wechsler Adult Intelligence Scale. Results showed that participants’ semantic performance in the DRM task was significantly related to that in the Triangles task. Specifically, the higher participants’ ToM in the Triangles task, the higher participants’ reliance on semantic memory while making false memories in the DRM task. Our findings are consistent with the Fuzzy Trace Theory and the Weak Central Coherence account and suggest that a (partially) common cognitive process responsible for global versus detailed-focus information processing could underlie these two abilities.
Introduction
Human long-term memory and the ability to understand/infer other’s mental states and intentions (generally labelled as Theory of Mind, ToM; Premack & Woodruff, 1978; Wimmer & Perner, 1983 and for a recent overview see: Devine & Lecce, 2021) are tightly intertwined across the whole lifespan. The ability to mentalise represents a key factor in memory development (Perner, 1991; Perner et al., 2007). This is particularly evident in explicit episodic memory, where ToM is involved in understanding the distinction between one’s memory (i.e., the mental representation of the reality) and reality (i.e., what has really occurred) (Hoerl, 2018) and in the existence of a relationship between ToM and destination memory, defined as the ability to remember to whom a piece of information was previously told (El Haj et al., 2017). The existence of a link between ToM and episodic memory has also been corroborated by neuroimaging studies showing that the brain areas active during tasks that require the attribution of mental states largely overlap with areas involved in autobiographical memory (Buckner & Carroll, 2007).
Insofar, however, little is known about the interplay between ToM and other types of memory systems. The present study addresses this issue focusing on semantic memory, a neurocognitive system deputed to store general knowledge about the world, including concepts, facts, and beliefs (Binder & Desai, 2011). There are compelling grounds for investigating associations between ToM and semantic memory.
A possible link between ToM and semantic memory has been outlined by works on clinical populations such as individuals with autism and semantic dementia, who are primarily characterised by deficits in ToM (the former) and in semantic memory (the latter). In addition to a deficit in ToM, individuals with autism typically display difficulties in retrieving semantically related information (Tager-Flusberg, 1991) and diminished capacity to use semantic relations to maximise memory performance (Bowler et al., 2008). These findings have been explained considering that individuals with autism are characterised by an efficient memory for low-level organisation (e.g., random organised items) but show impaired memory for complex levels of organisation (e.g., semantically organised items; Minshew & Goldstein, 2001). The higher memory performance for low, rather than high, levels of organisation fits well with the general view that individuals with autism are likely to focus on the local level of information while showing a difficulty in organising and integrating input/stimuli at a more general and complex level (Frith & Happé, 1994). Individuals with semantic dementia are also interesting with respect of the association between ToM and semantic memory. They are characterised by impaired semantic but intact episodic memory, usually due to focal atrophy in temporopolar and perirhinal cortices (for reviews: Hodges & Patterson, 2007; Klimova et al., 2018). As a consequence of this pattern, patients with semantic dementia show specific impairments in processing and recognising semantic associations (Simons et al., 2005). Interestingly, alongside with the semantic impairment, these individuals also display social dysfunctions (e.g., lack of empathy, loss of insight, deficits in complex social scenario’s understanding and in processing rule violation) that are related to ToM difficulties (for a review see: Kipps & Hodges, 2006; for behavioural evidence, see: Duval et al., 2012; for neuroimaging evidence see: Bejanin et al., 2017; Ross & Olson, 2010). Taken together, studies on autism and semantic dementia point to a co-occurrence of impaired ToM and semantic memory perturbations, suggesting the existence of a possible relationship between semantic memory processes and ToM abilities. In addition to the literature on clinical conditions, further evidence on the existence of a link between semantic processes and ToM can be found considering children’s language acquisition. With this respect, a meta-analysis has investigated the relationship between different language components and ToM development in 8,891 participants (Milligan et al., 2007). Results showed a significant association between language abilities (including semantic abilities) and ToM performance. In particular, semantics accounted for the 23% of the variance in ToM (measured using the classic first-order false-belief understanding task). Slade and Ruffman (2005) showed a bidirectional association between semantics and ToM in typically developing preschool children, supporting the idea that changes and improvements in mental state understanding (ToM) are connected to those in semantic memory. Further evidence comes from a work by Biscevic and colleagues (2018) on 38- to 72-month-old pre-schoolers in which ToM abilities were found to be related to the development of semantic fluency (i.e., in tasks involving the production of words belonging to the same category). Finally, from a neural standpoint, a recent study indicates that the same white matter tracts (specifically, the superior longitudinal fasciculus III) contribute to both semantic processes and ToM (Zekelman et al., 2022).
Building upon this set of evidence, the present study was designed to directly investigate the existence of a relationship between individual differences in ToM and in semantic memory in typical individuals. To reach this goal, we employed the Deese-Roediger-McDermott task (DRM; Deese, 1959; Roediger & McDermott, 1995) to assess semantic memory. The DRM is one of the most used tasks to induce semantic associations as it is extremely successful in producing false memories whose occurrence can be traced back to semantic processing (Brainerd, Yang et al., 2008; Cann et al., 2011; Gatti, Rinaldi et al., 2022; Roediger et al., 2001). In the DRM task, participants are first presented once with several lists of words that have to be memorised (within each list, the words are semantically/associatively related to a non-shown target word, named critical lure; for example, door, glass, pan, shade, ledge—critical lure: window); after a brief distracting task, participants are then asked to perform a recognition task in which they have to indicate whether a given word was part of the memorised lists or not. Interestingly, during this last phase, participants tend to falsely recognise the critical lures (i.e., they recognise new words as if they were part of the memorised lists), although these words were never presented in the encoding phase (for a review, see Gallo, 2010). Specifically, for the DRM task, we applied an already established method to compute the semantic similarity between the words presented in the recognition phase and the ones previously memorised (for a complete discussion, see: Gatti, Marelli, et al., 2022; Gatti, Rinaldi et al., 2022; Gatti et al., 2021), by using indexes extracted from distributional semantic models (DSMs). DSMs induce word meanings from large databases of natural language data, representing them as high-dimensional numerical vectors: These models indeed are thought to well capture the structure of semantic memory (Günther et al., 2019; Jones, Willits, et al., 2015).
In addition to the DRM task, participants were asked to complete two advanced ToM tasks: the Triangles task (Abell et al., 2000; Castelli et al., 2000, 2002) and the Reading the Mind in the Eyes task (RMET; Baron-Cohen et al., 2001). While the Triangles task evaluates the extent to which people spontaneously attribute mental states, the RMET assesses the ability to infer another individual’s mental state. In the Triangles task, participants are asked to watch short video clips of two triangles moving around and to describe what happened in the clips. Descriptions are then coded for the grade in which the participant attributes mental states to the triangles. In the RMET, participants are shown a number of images depicting the eye region of different individuals, and they are required to match these images with the term that best describes what the person in the photo is thinking or feeling choosing one of four given alternatives. Both tasks have been found to reliably differentiate between high-functioning autism spectrum disorder groups and verbal ability-matched control groups (Abell et al., 2000; Murray et al., 2017) and have been used in studies with typical adults including college students (e.g., Devine & Hughes, 2019; Eddy & Hansen, 2020; Kynast et al., 2021; Olderbak et al., 2015). For example, a recent study on 180 undergraduates showed that no ceiling effect on any of the RMET items and the existence of substantial variability across participants (Eddy & Hansen, 2020).
We chose to assess ToM using these two tasks because the present study takes as one of its premises that ToM is a complex and multidimensional construct that encompasses a wide range of different abilities (Osterhaus et al., 2016; Rusch et al., 2020). While the Triangles task evaluates the extent to which people spontaneously attribute mental states to geometric shapes based on their physical movements, the RMET assesses the ability to attribute the correct mental state to a corresponding image of the eye region. Existing studies in developmental (e.g., Bottiroli et al., 2016; Wellman et al., 2001), clinical (e.g., Shamay-Tsoory et al., 2007, 2010), lesional (e.g., Corradi-Dell’Acqua et al., 2020), and neuroimaging (Schlaffke et al., 2015) fields show that these two tasks are not significantly related to one another, supporting the view considering them as reflecting two (partially) independent components of the more general ToM skills. These two tasks differ on many dimensions, and some of these differences are crucial in predicting which should be associated with performance in the DRM task. First, they rely on different stimuli: moving geometric shapes for the Triangles task and images of the eye region for the RMET. Second, they have a different response format: While the Triangles task requires participants to freely describe the animation, in the RMET participants are asked to select the mental-state term that best describes the presented eye region by selecting one of four options. Third, while the Triangles task involves a variety of complex mental states (including, intentions and beliefs), the RMET mainly focuses on attributing and recognising emotions. Perhaps more importantly, while the RMET explicitly requires participants to recognise mental states focusing on local details, the Triangles task measures the ability to infer the triangles’ mental states to make a general sense of the viewed clips. This is, we believe, crucial for the purpose of the present study because the DRM also prompts participants to integrate and associate different information. Therefore, we expect to find a stronger association between the DRM and the Triangles task than between the DRM and the RMET.
To investigate the relationships between ToM and semantic memory, we also controlled for general verbal ability using the Vocabulary subtest of the Wechsler Adult Intelligence Scale (Wechsler, 2008). Following the preliminary evidence available in the literature, we expect participants’ reliance on semantic memory in the DRM task to be associated with participants’ ToM. That is, participants whose false memories rely more on semantic memory processes should also display higher ToM abilities.
Method
Participants
The sample size was determined a priori through G*Power (Faul et al., 2007). The minimum sample size for a linear multiple regression including four predictors was estimated using effect size f2 = .2, α = .05, 1 – β = .95. The minimum sample size was 98. One hundred three Italian students (30 males, with one participant reported being not binary; M age = 22.96 years, SD = 3.39; participants’ age ranged between 19 and 35) participated in the study. All participants were native Italian speakers, had normal or corrected to normal vision, and were naïve to the purpose of the study. Informed consent was obtained from all participants before the experiment. The protocol was approved by the psychological ethical committee of the University of Pavia, and participants were treated in accordance with the Declaration of Helsinki. Sixty-eight participants were tested individually during an online video call due to Covid-19 restrictions and asked to both share their computer screen and turn their camera and microphone on, while 35 participants were tested in-person in group sessions when social restrictions were lifted. 1
Cognitive and ToM measures: stimuli and procedure
After giving informed consent and filling in a first brief questionnaire assessing sociodemographic information (i.e., age, gender), participants performed the DRM task; then, in counterbalanced order, they completed the other tasks investigating ToM (e.g., RMET and Triangles task) and verbal intelligence. The sociodemographic information questionnaire and these last three tasks were collected online using Google Forms.
Reading the mind in the eyes
Participants were presented with the RMET proposed by Baron-Cohen and colleagues (Baron-Cohen et al., 2001, 2003), in the Italian version adapted by Serafin and Surian (2004). The RMET is composed of 37 black and white photos (i.e., the first photo is presented as a practice trial and the remaining 36 photos as stimuli). Each image depicts the eye region of different individuals of different ages and of both genders (Baron-Cohen et al., 2003). See Figure 1a for an example of one trial.

Example of the ToM tasks. (a) The practice trial of the RMET retrieved from https://www.autismresearchcentre.com/. (b) Five screenshots taken from a video of the Triangles task.
Each image is surrounded by four different adjectives or terms regarding emotional states/expressions, placed one at each corner of the picture, as in the original English version (Baron-Cohen et al., 2003). Participants were required to look at the image, read each of the four terms presented with that picture and select the one that best described what the person in the photo was thinking or feeling (i.e., a forced-choice task between four response options). Instead of the classical pen-and-paper form, we adapted the test to an online modality, where participants were presented with the image surrounded by the four terms at the corners (as the original task) and, below it, a multiple-choice format, where participants had to select among the four options.
To minimise vocabulary differences among participants, they were told to ask the examiner for the vocabulary definition whenever they had uncertainties regarding the meaning of the proposed terms: In such cases, the examiner would read the definitions proposed in the glossary (Serafin & Surian, 2004).
Following the scoring guidelines provided by the authors (Baron-Cohen et al., 2001), each response was scored as 1 or 0 (i.e., the value 1 indicates a correct response). The total score of each participant is given by the sum of the scores obtained in all 36 test items.
Triangles task
Participants were shown nine silent animations, lasting 35 to 45 s each, depicting two triangles: a big-blue triangle and a small-orange one moving in a framed white landscape (Abell et al., 2000; Castelli et al., 2000, 2002). The animations were of two types: (1) Action clips (three videos), in which triangles moved in a goal-directed fashion (e.g., chasing and fighting) and (2) ToM clips (six videos), in which triangles moved interactively with implied intentions (e.g., coaxing and tricking). See Figure 1b for an example of one trial.
Participants were instructed to watch each video just once and, after it, to answer the following question (i.e., the same for all the videos): “In your opinion, what happened in this video?.” Participants’ answers were scored based on the index of intentionality, which reflects the degree to which the subject attributes intentional mental states to the moving triangles. The intentionality score for each description ranged from 0 (no deliberate action, e.g., “bouncing,” “rotating”) to 5 (deliberate action aimed at affecting another’s mental state, e.g., “persuading,” “pretending,” “deceiving”). Two scores of intentionality, one for the Action clips and one for the ToM clips, were computed by summing scores in each answer. Action clips were used as a control measure, as they are expected to obtain low (or null) intentionality scores. No ToM involvement is in fact required to make sense of the triangles’ movements: Accordingly, individuals with autism get lower scores in ToM clips as compared with control individuals, but no difference is observed in the Action clips (e.g., Livingston et al., 2021; White et al., 2011). The interrater agreement (based on double-coding of 25% of the responses) was good (Cohen’s K = .83, p < .001). The summed score could range from 0 to 30 for the ToM clips and from 0 to 15 for the Action clips.
Vocabulary
To measure a proxy for participant’s verbal intelligence abilities, we used the Vocabulary subtest of the Wechsler Adult Intelligence Scale–Fourth Edition (Wechsler, 2008), in the Italian standardised version proposed by Orsini and Pezzuti (2013). Participants were presented with 27 words (4–30 item, Wechsler, 2008) and were asked to provide the meaning of each word. We adapted the original oral modality to an online form, in which participants were presented with the written words one by one and were asked to read each of them and write down a definition for each of them. We adopted verbal ability as a control measure, in line with previous studies on ToM (e.g., in Triangles task; Bianco et al., 2019).
Each item was scored by using a value from 0 to 2. A score of 2 indicates a good comprehension of the word, in which the participant provided an exhaustive definition, with one or more main or definitive features. A score of 1 indicates a correct answer but poor in content (e.g., providing a vague or irrelevant synonym or just an example of the term). Finally, a score of 0 indicates clearly wrong or absent answers that do not show any correct comprehension of the proposed word. The total score of each participant is given by the sum of the scores obtained in all 27 items.
DRM task: stimuli and procedure
Participants performed the DRM task (Deese, 1959; Roediger & McDermott, 1995), a typical false memories paradigm, in which they are instructed to remember several lists of words and then, after a brief distracting task, to perform a recognition task. The words that compose each list are associatively/semantically related to a non-shown word (called critical lure), and this association is thought to be responsible for participants’ false memory (Roediger et al., 2001; but see also: Gatti, Rinaldi et al., 2022).
The task is composed of two phases: an encoding phase and a recognition phase. For the encoding phase, we selected 12 lists of words out of 24 from the normative data for the Italian DRM test (Iacullo & Marucci, 2016). Each list was originally composed of 15 words: We selected the first 12 words (144 words in total), while 2 of the 3 remaining words were used as weakly related lures (see below).
The recognition phase was composed of 96 words, 48 of which had been presented in the previous phase (i.e., studied words) and 48 of which had not been previously presented (i.e., new words). The 48 studied words presented in this experimental phase were those in serial positions 1, 4, 7, and 10 in the studied lists. Of the 48 new words, 12 were the critical lures from the studied lists (i.e., the non-shown words mostly associated with the words composing each list), 24 were weakly related lures, and 12 were unrelated words. The weakly related lures were 2 of the 3 words of the studied lists that were not presented in the list, those in positions 13 and 14. The unrelated words were chosen randomly among the words of the excluded lists; this criterion was established arbitrarily.
During the encoding phase, participants were required to study 12 lists of words. Participants were shown the 12 words that composed each of the 12 lists in descending forward associative strength (i.e., the association strength from the critical lure to the words that compose the list). The order by which the lists were presented was random, while the order of the words within each list was fixed (see: Roediger & McDermott, 1995). Each trial started with a central fixation cross (presented for 500 ms) followed by a word (presented for 1,500 ms) and a blank screen (presented for 300 ms), then the script moved automatically to the next fixation cross. At the end of the 12th list, participants were requested to perform a distracting task (i.e., to solve as many arithmetical operations as they could) for 2 min.
Then participants were asked to perform the recognition task. Participants were instructed to make old/new judgements and to respond as fast and accurately as possible by pressing the left/right key (A and L) using both hands; the response keys were counterbalanced across participants. After the old/new judgement, the script moved to the next trial. The trials were shown in random order.
Each recognition trial started with a central fixation cross (presented for 1,000 ms) followed by a word (presented for 2,500 ms), participants’ judgement ended the recognition trial and moved to a confidence judgement. 2 The confidence judgement ended the trial, and the fixation cross of the next trial was presented.
The DRM task was administered online using Psychopy (Peirce, 2007, 2009; Peirce et al., 2019) through the online platform Pavlovia (https://pavlovia.org/).
Word-embeddings
In the present study, we applied a method capitalising on seminal studies on distributional semantics (for a review see: Günther et al., 2019) to compute semantic similarities values between the words in the recognition phase and those in the encoding phase of the DRM task (for an extended discussion, see Gatti, Rinaldi et al., 2022; see also below Computation of semantic similarity values). This method allows for a quantification, at the item level, of the degree of semantic similarity between the to be recognised words and the encoded (i.e., memorised) ones by building word representations from language usage (i.e., predicting a target word from the linguistic context in which it typically appears). We note that previous studies predicted the occurrence of false memories using human-based indexes, such as the backward associative strength (BAS; e.g., Roediger et al., 2001). This index quantifies the association strength from the words that compose each list to the critical lure. In the present study, we rather opted for the adoption of an independent index derived from natural language use. The decision to adopt this index and not the BAS is grounded in the need to employ predictors built from sources independent from human ratings (Westbury, 2016; for an example of this effect on semantic fluency, see Jones, Hills et al., 2015).
Vector representations for the words used in the DRM task were extracted from a semantic space obtained by inducing word embeddings using the Continuous Bag of Words (CBOW) method, an approach originally proposed by Mikolov and colleagues (2013). The model, released by Marelli (2017), was trained on itWaC, a free Italian text corpus based on web-collected data and consisting of about 1.9 billion tokens. The model used is set on the following parameters: 9-word co-occurrence window, 400-dimension vectors, negative sampling with k = 10, and subsampling with t = 1e–5. This set of parameters defines the learning procedure used to induce word vectors (Mikolov et al., 2013). CBOW indicates the applied learning procedure: When using CBOW, the obtained vector dimensions capture the extent to which a target word is reliably predicted by the contexts in which it appears. Co-occurrence window size indicates how large the considered lexical contexts are; in our case, a nine-word window indicates that we estimated predictions concerning four words on the left and four words on the right of the target word. The number of vector dimensions indicates how many nodes are included in the hidden layer, representing the result of the dimensionality reduction process implicitly applied by the network. Negative sampling estimates the probability of a target word by learning to distinguish it from draws from a noise distribution; the parameter k specifies the amount of these draws. The subsampling parameter t specifies a threshold-based procedure that limits the impact of very frequent, uninformative words.
From this semantic space, we extracted vector representations for the words used in this study. Specifically, for each word pair, it is possible to obtain a semantic-similarity index (hence SSim) based on the cosine of the angle formed by vectors representing the meanings of these words. In particular, the higher the SSim value, the more semantically similar the words should be as estimated by the model. A heatmap matrix of the semantic similarity structure among the words composing a DRM list is represented in Figure 2.

A heatmap matrix of the cosine values among the words composing the list doctor (i.e., critical lure; words list taken from Roediger & McDermott, 1995).
Computation of semantic similarity values
For each new word (12 critical lures, 24 weakly related lures and 12 unrelated words), we computed an SSim. That is, SSim was computed as the frequency-weighted average SSim (for a similar approach see: Gatti, Rinaldi et al., 2022; Marelli & Amenta, 2018) between each new word in the recognition phase and each of the 12 words that composed its list. For unrelated words, we computed the index randomly matching each word with a list. The formula used was:
where nw is a new word shown during the recognition task,
Data analysis
All the analyses were performed using R-Studio (RStudio Team, 2020). Generalised linear mixed models (GLMMs) were run using the lme4 R package (Bates et al., 2015), while linear models were run using the stats R package (R Core Team, 2019). The graphs reported were obtained using the effects R package (Fox, 2003; Fox & Weisberg, 2019).
First, before testing the main study hypothesis, we aimed to replicate the positive effect of SSim on participants’ old responses for new words (i.e., false memories; for a complete discussion of this analysis see also: Gatti, Rinaldi et al., 2022). We, thus, estimated a GLMM having participants’ responses (“old” responses scored as 1 and “new” responses scored as 0) for new words (i.e., critical lures, weakly related lures, and unrelated words) as dependent variable and SSim as a continuous predictor; subjects and items were included as random intercepts and the effect of SSim was included as random slope across participants. Then, we aimed to quantify for each participant an estimate of the effect of the SSim predictor on new words while performing the DRM task. We thus extracted the conditional modes of the random effects estimated for each subject in the GLMM, namely, the individual-level random slopes. Extracting the individual-level slopes for subjects allows us to observe the interparticipant variability in the sensitivity to the semantic similarity predictor. This index (labelled hereafter as SSim sensitivity) can be conceptualised as the SSim effect on participants’ performance as extracted by the ranef R function, which provides the conditional modes of the random effects from a fitted model object, that is, the set of differences between the population-level average predicted response for a given set of fixed-effect values (SSim in our case) and the response predicted for each participant (i.e., thus each value describes how much each participant’s slope differ from the slope of the total sample; see: Bates et al., 2015).
Next, to examine how individual differences related to false memories’ reliance on semantic processes are associated with Vocabulary and ToM variables, we estimated a linear model including the previously extracted Ssim sensitivity index as a dependent variable and additively the three ToM measures along with the verbal intelligence score (i.e., Triangles task ToM clips, Triangles task Action clips, RMET, and Vocabulary) as continuous predictors.
Results
Replication of the effect of semantic similarity on false memories
Trials in which overall RTs were faster than 300 ms or slower than 5,000 ms (1% of the trials) were excluded from the analysis. Participants’ proportion of “old” responses for studied words was .70, for critical lures .58, for weakly related lures .20, and for unrelated words.07.
As expected, the effect of SSim on the recognition of new words was significant, z = 5.75, p < .001, b = 9.94, Pseudo-R2 (total) = .44; Pseudo-R2 (marginal) = .16 (Figure 3), indicating that the higher the SSim (i.e., the higher the semantic similarity between the to be recognised new word and the studied words of each list), the higher the chances of recognising it as “old” (i.e., the higher the chance of making a false recognition). From this model, we thus extracted the participants’ random slopes, namely, SSim sensitivity.

Results from the GLMM were estimated using SSim as a continuous predictor, illustrating the positive relationship between SSim and false recognition.
Descriptive statistics for the ToM and vocabulary tasks
Descriptive statistics for the four cognitive and ToM measures are reported in Table 1.
Descriptive statistics for the Triangles task (ToM and Action clips), Reading the Mind in the Eyes (RMET) task, and the Vocabulary ability.
SD: standard deviation.
The correlation matrix between the four measures and the dependent variable (SSim sensitivity) is reported in Table 2. Overall, the correlations between the four predictors ranged from very low, as the one between Triangles ToM clips and RMET (r = –.02), to low, as the one between RMET and Vocabulary (r = .26), which was the only significant one, p = .007. In addition, here we found a significant correlation (r = .28) between SSim sensitivity and Triangles ToM clips. We thus next tested the relative contribution of each of these variables to false memories in a multiple regression analysis approach.
Correlation between four predictors and the dependent variable included in the current study.
RMET: Reading the Mind in the Eyes; SSim: semantic-similarity index.
p values < .05.
Do ToM abilities relate to participants’ semantic performance in the DRM task?
The effects of the linear model having SSim sensitivity as dependent variable and the three ToM variables (i.e., Triangles task ToM clips, Triangles task Action clips and RMET) along with the vocabulary task as continuous predictors are reported in Table 3 and Figure 4. Globally the model explained the 8% of the variance, R2 = .08. Results showed that only the effect of Triangles task ToM clips was significant, indicating that the higher participants’ score in the Triangles task ToM clips, the higher participants’ reliance on semantic memory while performing the DRM task. No other significant effect was found. 3
Results of the linear model on participants’ semantic performance in the DRM task including ToM and Vocabulary predictors.
RMET: Reading the Mind in the Eyes.
Bold font indicates p values < .05.

Results from the linear model including the effect of Triangles task ToM clips (a), Triangles task Action clips (b), RMET (c) and Vocabulary (d) on participants’ reliance on semantic memory while performing the DRM task. Only the effect of Triangles TOM clips was found to be significant.
Discussion
In the present study, we investigated, for the first time in the literature, the relationship between ToM and semantic processes subserving false memories in a typical population. By taking advantage of recent machine-learning techniques from computational linguistics, we extracted an index informative about the individual reliance on semantic memory processes while falsely recognising new words in the DRM task, and then investigated the relationship between this index and two tasks of ToM. Our findings showed a significant link between participants’ reliance on semantic memory while falsely recognising new words in the DRM task and their performance in the Triangles task, over and above verbal abilities. On the contrary, we did not find a significant relationship between the DRM task and the RMET.
Before commenting on these main findings, it is important to acknowledge that the semantic index computed using the distributional semantic model was extracted from participants’ responses for new words. This choice was made because, according to seminal theories, semantic processing is involved in both false and veridical recognition but would be more crucial for the former (Brainerd & Reyna, 1998, 2002; Reyna & Brainerd, 1995; Roediger et al., 2001). Accordingly, previous studies have shown that this semantic index is reliable in predicting participants’ performance for both studied and new words, but in the latter case, the effect is higher (Gatti, Rinaldi et al., 2022). Thus, this index provides a general measure of how much each participant’s false recognition can be predicted by the semantic similarity between the new words shown in the recognition phase and the studied ones.
The key finding of the present study is the significant association between individual differences in semantic memory as evaluated via the DRM task and Triangles task but not the RMET. This dissociation in findings between the RMET and the Triangles task was expected and can be explained by the differences between these two ToM tasks outlined in the introduction. We propose that the Triangles task could be seen as a higher level ToM task, requiring participants to integrate single movements to create a global and coherent representation of the scene. On the contrary, the RMET task would represent a lower level ToM task, which relies on participants’ ability to focus on subtle and local details of a restricted face region, involving basic abilities of emotion recognition (Osterhaus & Bosacki, 2022). This pattern of results suggests that the Triangles task and the DRM share a common ground, and we argue that it can be traced back to domain-general abilities allowing individuals to integrate and structure information in different contexts, whether social or linguistic. The Triangles task allows measurement of ToM as the ability to adopt a general and global picture of the surrounding (social and mental) environment, “[a] cohesive interpretative device par excellence: it forces together complex information from totally disparate sources” (Frith, 1989, p. 174). On the contrary, the use of semantic memory in the DRM task is related to participants’ ability to build upon the global meaning of the words studied and to use this information in the subsequent recognition phase (Brainerd & Reyna, 1998, 2002; Reyna & Brainerd, 1995). As such, the positive relationship between the Triangles task and semantic memory reported here can be traced back to the activation of a partially similar cognitive process, that is, one’s ability to integrate a variety of information into a common meaning, which allows a coherent and global vision in various domains (whether social or conceptual).
This line of interpretation fits with data on individuals with autism who, according to the Weak Central Coherence account (Frith, 1989; Frith & Happé, 1994; Happé & Frith, 2006), tend to focus their attention on the local level of the incoming information, without integrating it in a general framework or at a global level (Frith, 1989; Frith & Happé, 1994). Consistent with this, it has been proposed that a deficit in central coherence induces individuals with autism to focus on detail, rather than on the global level, in a variety of domains such as memory performance (Smith et al., 2007). Accordingly, previous research conducted using the DRM task showed that individuals with autism are less susceptible to false memories as induced by the DRM task (Beversdorf et al., 2000; Wojcik et al., 2018; for evidence on the visual version of the task, see Hillier et al., 2007). This higher resistance to the classical false memory effect of the DRM task (Beversdorf et al., 2000; Hillier et al., 2007; Wojcik et al., 2018) can be due to a bias towards a detailed-focused processing (i.e., hyper encoding of the verbatim trace) and an impaired global processing (i.e., hypo encoding of the gist trace) (Happé & Frith, 2006). This interpretation is also in line with the Fuzzy Trace Theory (FTT; Reyna & Brainerd, 1995, but cfr also: Roediger et al., 2001), a key theoretical frameworks accounting for the performance in the DRM task. According to the FTT individuals’ memory performance is based on two different memory traces: a verbatim one, linked to specific perceptual details of the encoded information, and a gist one, linked to the semantic themes of the information (Brainerd & Reyna, 1992). This distinction would be reflected in the development of the memory system, with children’s memory that shifts from being mainly verbatim-based to gradually integrate semantic and associative information (Brainerd & Reyna, 1992; Reyna & Brainerd, 1991a, 1991b). Following this perspective, the diminished false memory occurrence in autistic individuals (e.g., Griego et al., 2019) can be traced back to an unbalanced reliance on the verbatim trace at the expenses of the gist trace. That is, individuals with autism would tend to largely focus on specific details of events while acquiring new information, thus reducing their tendency to rely on conceptual knowledge (for an in-depth discussion see: Miller et al., 2014).
The positive relationship reported here between false memory and ToM, an adaptive social ability enabling individuals to interact with each other, may suggest that memory distortions could also be adaptive. DRM errors can be the result of an underlying adaptive system, as converging evidence indicates that the human memory system is far from being a perfect recorder as its processes often involve active rather than passive reproduction. Indeed, many studies reported evidence consistent with the idea that false memories can have an adaptive value or being considered a normal by-product of an adaptive memory system (e.g., Howe & Derbish, 2010; Howe et al., 2010, 2015; Thakral et al., 2019; and for an in-depth discussion see: Klein, 2013; Schacter et al., 2011; Vecchi & Gatti, 2020).
However, empirical works on the occurrence of false memories across the life cycle depict what seemingly is a more complex scenario about the possible functional role of false memories. That is, while studies in aging demonstrate that older adults produce higher rates of DRM false memories compared with young adults (e.g., Balota et al., 1999), developmental evidence during childhood indicates that the suggestibility to the false memory illusion increases during child development (i.e., older children make more false DRM memories compared with the younger ones; Brainer, Reyna, & Ceci, 2008; but see also Howe, 2008). Hence, a higher number of false memories may at the same time represent a sign of a declining memory (i.e., as the one observed in aging) but also of developmental improvement (i.e., as the one observed in childhood).
We suggest here that false memories are neither adaptive nor maladaptive by themselves, but they likely are the consequences of normal memory processes. In particular, the occurrence of false memories in the DRM is likely caused by the fact that the associative/semantic links between the words in the list and the lure produce an automatic over-activation of the latter (Roediger et al., 2001). Yet, these memory distortions do not necessarily reflect a “bad” memory performance. Indeed, a similar mechanism (i.e., the activation of semantically related words) is at play in priming tasks, in which the chronometric performance relies on the semantic similarity between the two words: the more the prime word pre-activates the meaning of the target word, the faster the performance (e.g., Hutchison et al., 2013, and for recent evidence using DSMs see: Gatti, Marelli et al., 2022). Hence, while the output of participants’ performance is task-dependent, the cognitive process behind both outputs is likely the same. Taken together, false memories would be a “normal” by-product of a flexible cognitive system and cannot be considered necessarily positive or negative manifestations. This is also consistent with what has been suggested by recent studies maintaining that producing a certain amount of false memories is normal (although an excess of false memory could be, e.g., indicative of source-memory failure, like in older adults, e.g., Balota et al., 1999).
Several limitations should be acknowledged when interpreting the results of this study. Among these, first, we note that in studying individual differences linked to false memory production, more ecological paradigms could be adopted, to provide direct evidence from the practical point of view of justice trials and witness evaluation (Schacter & Loftus, 2013). Second, as measures of ToM, participants were required to complete two advanced tasks that differ on a number of dimensions (stimuli, response format, ToM component, level of integration required), and we cannot clearly identify which of them is responsible for the pattern of results found.
In conclusion, the present study investigated the role of ToM in predicting the semantic performance in the DRM task. Results showed that the higher participants’ ToM abilities in the Triangles task, the higher participants’ reliance on semantic memory while falsely recognising new words the DRM task. These findings are consistent with theories that explained ToM deficits and memory performance across clinical populations, such as individuals with autism, and indicate that a (partially) common cognitive process could underlie these two abilities.
Research Data
sj-csv-1-qjp-10.1177_17470218221135178 – Research Data for Individual differences in theory of mind correlate with the occurrence of false memory: A study with the DRM task
Research Data, sj-csv-1-qjp-10.1177_17470218221135178 for Individual differences in theory of mind correlate with the occurrence of false memory: A study with the DRM task by Daniele Gatti, Serena Maria Stagnitto, Chiara Basile, Giuliana Mazzoni, Tomaso Vecchi, Luca Rinaldi and Serena Lecce in Quarterly Journal of Experimental Psychology
Research Data
sj-csv-2-qjp-10.1177_17470218221135178 – Research Data for Individual differences in theory of mind correlate with the occurrence of false memory: A study with the DRM task
Research Data, sj-csv-2-qjp-10.1177_17470218221135178 for Individual differences in theory of mind correlate with the occurrence of false memory: A study with the DRM task by Daniele Gatti, Serena Maria Stagnitto, Chiara Basile, Giuliana Mazzoni, Tomaso Vecchi, Luca Rinaldi and Serena Lecce in Quarterly Journal of Experimental Psychology
Footnotes
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by funding from the Italian Ministry of University and Research (PRIN 2017 no. 201755TKFE) to TV and from the Italian Ministry of Health (Ricerca Corrente 2022) to LR and TV.
Data availability statement
The data used in this study are reported as Supplementary Materials.
Supplementary material
The supplementary material is available at: qjep.sagepub.com
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
