Abstract
Semantic context effects are well established using both words and pictures as stimuli. One such effect, semantic interference, is observed in naming latencies when a categorically related distractor word or picture is presented together with a target picture (e.g., dog-LION). Recently, this effect has also been shown to occur when an environmental sound (e.g., a dog barking) is presented as an auditory distractor during picture naming and when a distractor picture is presented with a target sound for naming. The purpose of the current study was twofold: (1) to replicate the semantic interference effect in the picture–sound interference (PSI) paradigm and (2) determine whether a semantic interference effect is also observable when distractor words are presented with environmental sounds as target auditory objects for naming, using a novel sound–word interference (SWI) paradigm. We replicated the semantic interference effect in Experiment 1 with environmental sound distractors. Experiment 2 demonstrated significant semantic interference during an SWI paradigm for the first time. We discuss the implications of these results for our understanding of the origin and locus of the semantic interference effect according to current theories of lexical selection.
Introduction
There is general consensus among models of spoken language production that activation spreads between conceptual and lexical stages of processing (de Zubicaray et al., 2006; Dell & Sullivan, 2004; Levelt et al., 1999). Using picture naming as an exemplar production task, the first stage following object perception is typically assumed to involve access to conceptual representations followed by retrieval of the lexical representation and finally articulation (e.g., Glaser, 1992; Humphreys et al., 1988; Roelofs, 1992). According to most accounts, multiple related concepts are activated simultaneously (e.g., when naming DOG, other animals are activated such as wolf and fox), and spread their activation to lemmas, a type of mediating representation in declarative memory that codes conceptual and syntactic (e.g., lexical category and grammatical gender) representations of a word (see Roelofs & Ferreira, 2019). However, production accounts diverge in their explanations as to how the target is selected from among the multiple activated candidates. Non-competitive selection models propose the one with the highest level of activation is selected, regardless of the level of activation of other activated candidates, whereas competitive selection models take into account their activation levels (Dell & Sullivan, 2004; Levelt et al., 1999; Spalek et al., 2013).
A considerable amount of research over the past 40 years has been dedicated to investigating how picture naming latencies are affected by semantic contexts. The picture–word interference (PWI) paradigm has been particularly informative in this respect (e.g., Glaser & Düngelhoff, 1984; Lupker, 1979; Roelofs, 1992; Schriefers et al., 1990; Starreveld & La Heij, 1995). The PWI paradigm requires participants to name a target object while ignoring a spoken or written distractor word that is related, unrelated, or congruent (i.e., identical) with the target picture. Naming latencies are affected by the relationship between the target and distractor. Facilitated naming occurs when a target (e.g., CAT) is presented with a congruent distractor (cat) compared with neutral (e.g., a row of Xs) or unrelated (violin) distractors. Slower naming (i.e., semantic interference) occurs when a target is presented with a categorically related (wolf) versus unrelated or neutral distractor. Congruency facilitation effects are largest (e.g., 60–100 ms; Gauvin et al., 2018; Glaser & Düngelhoff, 1984), while semantic interference tends to be observed reliably in a small range of distractor stimulus onset asynchronies (SOAs) around target picture presentation (−160 to 160 ms) with an overall magnitude of ~21 ms (see meta-analysis by Bürki et al., 2020).
The finding of semantic interference in PWI has been fundamental to the development of models describing the mechanisms of lexical selection. Competitive models of lexical selection propose that semantic interference arises because the related target and distractor become activated after stimulus presentation and through their shared coordinate category connection increase each others’ and semantically related words’ activation levels, making lexical selection more time-consuming compared with unrelated distractors (e.g., Glaser & Glaser, 1989; Levelt et al., 1999; Roelofs, 1992; Starreveld & La Heij, 1995). The Swinging Lexical Network variant of this account (Abdel Rahman & Melinger, 2009, 2019) assumes that related distractors produce priming during the conceptual activation stage, so that the combined effects of conceptual priming and lexical competition will be either facilitation or competition dominant according to whether there is sufficient converging activation to activate a cohort of multiple lexical candidates. Alternative non-competitive accounts propose a post-lexical selection process to explain semantic interference in PWI. According to this account, categorically related distractors invariably produce conceptual–lexical priming (i.e., facilitation). However, this priming is outweighed by a decision mechanism operating on distractor and target representations in a pre-articulatory buffer (Mahon et al., 2007). Specifically, the selection process is delayed because the categorically related distractor satisfies a response-relevant criterion (category membership) requiring more time for it to be excluded from the buffer. By contrast, all current production models assume congruent facilitation effects occur due to target activation levels being raised by overlapping target and distractor representations, with this overlap potentially occurring across multiple stages of production from conceptual preparation to phonological encoding, and so speeding selection.
Recently, semantic interference effects have also been demonstrated with auditory objects as distractors or targets to be named, indicating that they may occur irrespective of the distractor or target modality. For example, Mädebach et al. (2017) introduced the picture–sound interference (PSI) paradigm to investigate congruent facilitation and semantic interference effects produced by non-linguistic, environmental sound distractors during picture naming. Over a series of four experiments, they consistently observed facilitation with congruent sound distractors (e.g., a COW mooing) and interference with semantically related sounds (e.g., a COW with the sound of a horse neighing), analogous to PWI effects at an SOA of −200 ms (sound onset before picture onset). Mädebach et al. (2017) concluded that perceived concepts are assigned a lexical form, including those incidentally co-activated, regardless of their stimulus form (i.e., picture, word, or environmental sound) or modality of perception (i.e., visual or auditory).
These findings were confirmed by subsequent experiments by the same group (Mädebach et al., 2018), where they observed the semantic interference effect across different task SOAs in a psychological refractory period (PRP) paradigm (e.g., Ferreira & Pashler, 2002; Piai et al., 2015). The PRP paradigm involves participants being presented with two stimuli from different tasks requiring them to respond to Stimulus 1 before responding to Stimulus 2. To identify dual-task interference, the SOA between Stimulus 1 and Stimulus 2 (task-SOA) is varied. Mädebach et al. (2018) found that despite varying Task 2 (PWI) SOA between 0 or 500 ms, the magnitude of semantic interference was similar (Task 1 was shape classification). According to the PRP logic, this reflects additivity of task SOA and distractor type, suggesting the locus of the effect occurs during central (i.e., lexical or post-lexical selection) processes. The congruency effect was similarly unaffected by task SOA.
The above evidence suggests comparable effects of verbal and sound distractors on picture naming, yet surprisingly few studies have investigated similar effects when naming auditory objects. It seems reasonable to expect that naming an auditory object would, by virtue of lexicalisation, also be susceptible to semantic interference. For example, a recent study by Wöhner et al. (2020) adapted the picture–picture interference paradigm by having participants name environmental sounds as auditory target objects while ignoring distractor pictures. They observed congruent facilitation and semantic interference effects when using an SOA of −200 ms. In addition, they noted the magnitude of the effects were greater than those typically observed with picture naming, which they attributed to the overall longer latencies when naming sounds and/or the connections between sounds and their corresponding lexical–semantic representations being relatively weaker than those for pictures. Additional experiments with the PRP paradigm revealed the semantic interference effect was additive with dual task interference, suggesting its locus occurs during central (i.e., lexical or post-lexical selection) processes. However, the congruency effect was underadditive, suggesting it arises during pre-central (i.e., perceptual, or conceptual) processing.
The present study
While the above studies indicate that semantic interference effects may be observed when environmental sounds are presented as either distractors or targets for naming, the interference effect reported by Mädebach et al. (2017, 2018) with distractor sounds in German speakers is yet to be independently replicated, and no study to date has demonstrated a semantic interference effect in naming auditory objects in the presence of a related distractor word. Given picture distractors were able to elicit congruent facilitation and semantic interference effects during sound naming in Wöhner et al.’s (2020) study, it seems reasonable to assume that written word distractors would do likewise during sound naming, with the effect attributable to identical mechanisms to those invoked to explain written distractor effects in the PWI paradigm. The purpose of the current study was therefore twofold: in Experiment 1, we sought to replicate the semantic interference effect in the PSI paradigm in native English-speaking participants using a similar range of SOAs to Mädebach et al. (2017), whereas in Experiment 2, we investigated whether semantic interference is observable when distractor words are presented with environmental sounds as target objects for naming, using a novel sound–word interference (SWI) paradigm. In both experiments, we also tested for congruency facilitation effects. Wöhner et al.’s (2020) results demonstrate the feasibility of using sound naming to study semantic context effects and so provide complementary evidence to evaluate the task-specificity of critical findings in the production literature. However, as words are processed more quickly than picture distractors, it is likely that the time course of semantic interference will differ between the two when naming sounds. Given the longer perceptual identification and prolonged lexical selection for sounds (see Chen & Spence, 2013), we reasoned more positive distractor SOAs (i.e., distractor word presentation after target sound onset) would be needed in sound naming than in picture naming to observe semantic interference.
Experiment 1: picture naming with sound distractors
Method
Participants
Twenty-four native monolingual speakers of English participated (20 female, mean age = 21.9, range 17–36 years). All participants were right-handed, had normal or corrected-to-normal vision and normal hearing, and provided informed written consent.
Apparatus and stimuli
Target picture stimuli consisted of 32 simple black line drawings of the 32 target objects listed in Mädebach et al. (2017) (see the online Supplementary Material). All line drawings, except one, were obtained from various freely available online databases, were approximately 6.5 × 5 cm2 in size and presented on the same grey (RGB: 220 220 220) background as used in Mädebach et al. (2017). One picture (doorbell) was not available and so was manually drawn to the same style, size, and background as the other pictures. The distractor stimuli were environmental sounds obtained directly from Mädebach et al. (2017) and were either from the same category as the target picture (semantically related), unrelated or corresponding to a target item (congruent) (see Figure 1). These sounds were normalised in amplitude to a monophonic sampling rate of 48 kHz (mean length 1,246 ms, range 702–1,572 ms). Distractors were also members of the response set (i.e., also target names).

Examples of the target picture and distractor environmental sound pairings used in Experiment 1.
Visual stimuli were presented on a Dell Latitude 5580 laptop screen (1920 × 1080 pixels) using Cogent 2000 v125 (Cogent 2000, 2013), positioned at a comfortable viewing distance and angle as indicated by the participant. Distractor auditory stimuli were presented through Sennheiser HD6 Mix headphones. Picture naming latencies were recorded using a noise cancelling microphone and were measured automatically (in milliseconds) by Chronset (Roux et al., 2017). Errors were manually recorded online and checked offline, with any naming or technical errors (e.g., incorrect responses, verbal dysfluencies, voice-key malfunctions) being excluded.
Design
The first independent variable (IV) was the distractor type which comprised three levels: congruent, semantically related, and unrelated (see Figure 1). Unlike Mädebach et al. (2017; Experiments 1–3), we did not include a neutral condition comprising a 500-Hz sine tone (Mädebach et al. also omitted the neutral condition in their Experiment 4). The second IV was SOA. We adopted SOAs of −400,−200, and 0 ms as Mädebach et al. (2017) reported reliable semantic interference and congruency effects at −200 ms across their Experiments 1–3 that became less reliable at earlier (−500 ms) and later (i.e., 0 ms) SOAs. The dependent variable was the reaction time (RT), measured using the speech onset extracted from the audio recording using Chronset (Roux et al., 2017). The order of the blocks was counterbalanced to prevent practice, learning, and fatigue effects. Each subject was presented with each picture three times, each time with a different distractor. The stimuli were pseudo-randomised for each subject using Mix (van Casteren & Davis, 2006). The randomisation was constrained such that the same target picture would not appear within five trials, each condition would not appear more than twice consecutively, and distractor condition was not repeated in more than three consecutive trials.
Procedure
Prior to commencing the experiment, participants were familiarised with the pictures and sounds. The familiarisation phase consisted of three cycles: the first cycle presented the picture simultaneously to the congruent sound and written name of the target; the subsequent two cycles contained the target picture and congruent sound only. The target names were direct English translations of those used in Mädebach et al. (2017). Participants were asked to respond with the name of the picture.
Three experimental PSI blocks followed (96 trials per block), each block presented the distractor sounds at a different SOA (i.e., −400, −200, and 0 ms). Participants were required to name the target picture aloud as quickly and as accurately as possible while ignoring the distractor sound. Each trial began with a black fixation cross-presented on the laptop screen for 500 ms at an SOA of −500 ms. The environmental distractor sound was then presented (with an SOA of −400,−200, or 0 ms, depending on which SOA block was being presented). The target picture was presented from 0 to 800 ms before being replaced by a white screen for 2,200 ms. Each trial lasted for approximately 3.5 s, making each block approximately 5–6 min in length. Response latencies were calculated from picture onset to speech onset.
Results
Of the total dataset, 5.05% of trials (349) were excluded from analyses due to naming or technical errors (e.g., incorrect responses, verbal dysfluencies, voice-key malfunctions) and recorded responses below 300 ms. Naming latencies exceeding two standard deviations from a participant’s and an item’s mean, calculated within each experimental condition, were considered outliers and also removed (137 trials; 2.0%). In total, 7.0% of the dataset (486 trials) was removed using R (R Core Team, 2018). Table 1 presents mean naming latencies and percent errors as a function of condition and SOA.
Mean naming latencies (in ms) and error rates (in %) as a function of condition and SOA in Experiment 1.
SOA: stimulus onset asynchronies; RT: reaction time. Standard errors of the mean (SEMs) are in parentheses.
Naming latencies were analysed with linear mixed-effect (LME) models with crossed random effects for participants and items (Baayen et al., 2008) fitted using the lme4-package (version 1.1.21; Bates et al., 2015) in R (version 3.6.1, R Core Team, 2016), as per Mädebach et al. (2017). Q-Q plots of the residuals (Pinheiro & Bates, 2000) did not reveal any obvious deviations from normality.
A “parsimonious mixed-model method” was used (Bates et al., 2016). First a maximal model including random intercepts, random slopes, and correlations for all fixed effect terms, participants, and items, was fitted. The complexity of this random-effect structure was then iteratively reduced by stepwise removal of random correlations, random interactions, and random main effect slopes that were not involved in remaining random interactions (see Mädebach et al., 2017). Random effects were retained in the model with a liberal criterion of α = .20 (Barr et al., 2013; Matuschek et al., 2017). The resulting model was then used to evaluate the significance of fixed effects using model comparisons (likelihood ratio tests). Significance of fixed parameter estimates was determined using the Satterthwaite approximations of degrees of freedom provided by the lmerTest package (version 2.0.32; Kuznetsova et al., 2017). Error rates were analysed with generalised linear mixed-effect models’ binomial families following the same procedure as above.
This selection procedure resulted in models for the latency analyses including main effect of condition, SOA and interaction between the two, random slopes for condition by participant and item, random slope by SOA by participant, and random intercept per item (see Table 2 for estimated marginal means).
Estimates of naming latencies as a function of condition and SOA in Experiment 1 from the model fit.
SOA: stimulus onset asynchronies; RT: reaction time. SEMs in parentheses.
For naming latencies, there was a significant main effect of condition, X2(2) = 35.15, p < .001 and significant main effect of SOA, X2(2) = 14.46, p < .001. The interaction was not significant X2(4) = 3.44, p = .49. Compared with the unrelated distractor condition, naming latencies were significantly faster in the congruent condition, B = −38.14, SE = 5.60, t = −6.81, p < .001, and significantly slower in the semantic condition, B = 13.71, SE = 6.57, t = 2.09, p = .043. Although latencies were numerically slower at 0 ms compared with −200 and −400 ms SOAs, none of the post hoc comparisons were significant (all ps > .05).
In the analyses of error rates, there was no main effect of SOA, X2(2) = 1.79, p = .407 (cf. Mädebach et al., 2017), but there was a main effect of condition, X2(2) = 10.18, p = .006, but no interaction, X2(4) = 3.14, p = .534, with fewer errors in the congruent condition (B = −0.50, SE = 0.22, Wald’s Z =−2.30, p = .021), with no significant difference between the unrelated condition and semantic (B = 0.15, SE = 0.19, Wald’s Z = 0.82, p = .411).
Discussion
Experiment 1 largely replicated Mädebach et al.’s (2017) results, extending them to English language speakers. Specifically, we observed semantic interference and congruent facilitation effects. However, unlike Mädebach et al. (2017), we did not observe a significant interaction with SOA.
Experiment 2: sound naming with written distractor words
Experiment 2 introduced a novel SWI paradigm. We hypothesised that a semantic interference effect would be observable in auditory target naming with written distractor words, analogous to the effects observed in PWI and PSI paradigms. Like Wöhner et al. (2020), we suspected participants might try to perform the auditory naming task by looking away from the computer screen or closing their eyes to avoid the distractor words. To ensure participants processed the written distractors and kept their attention focused on the computer screen, we included a neutral condition comprising a string of five Xs and informed them that trials could include either real words or letter strings presented with the sounds, and that they should maintain fixation on the crosshair in the centre of the screen. We reasoned this manipulation would ensure the distractors would attend to and inform us about whether participants processed the words in the related, unrelated, and congruent conditions.
Method
Participants
Twenty-four native monolingual speakers of English participated (18 female, mean age = 22.4, SD = 5.3, range 18–41 years). Inclusion and exclusion criteria were identical to Experiment 1. None of the participants in Experiment 2 had performed Experiment 1.
Apparatus and stimuli
Target environmental sound stimuli consisted of the 32 auditory objects used as distractor stimuli in Experiment 1. The written distractor stimuli were either names of these environmental sounds (e.g., “dog”) from the same category as the target picture (semantically related), unrelated, corresponded to the target item (congruent) or a letter string “XXXXX” (neutral). Distractor words and letter strings were presented in black, size 20, Cambria font and were displayed in the centre of a white background.
For the familiarisation task, the 32 environmental sounds were presented with their congruent written names. Target auditory stimuli were presented through Sennheiser HD6 Mix headphones. Picture naming latencies were recorded/measured identically to Experiment 1. Distractor visual stimuli, font size 50, were presented on a Dell Latitude 5580 laptop screen (1920 × 1080 pixels) using Cogent 2000 v125 (Cogent 2000, 2013).
Design
The first IV was distractor type, comprising four levels, congruent, semantically related, semantically unrelated, and neutral (see Figure 2). In the congruent distractor condition, the agent responsible for the target environmental sound was referred to by the distractor word (e.g., the “neighing” of a horse paired with distractor word “horse”). In the semantically related distractor condition, the target environmental sound was presented with a distractor word semantically related to its referent (e.g., the “neighing” of a horse with the word “donkey”). In the unrelated distractor condition, the environmental sound was presented with a semantically unrelated word (e.g., “neighing” with word “drum”). Finally, the neutral condition presented a non-word with the target environmental sound (e.g., “neighing” with letter string “XXXXX”). The second IV was SOA (0, 200, and 400 ms). Positive SOAs were used to account for the slower perception and identification of sounds compared with written words (see Chen & Spence, 2010). Our use of positive SOAs reflected the uncertainty about when the environmental sound target would be identified lexically given the unfolding nature of auditory scenes (see Scott, 2005; Wöhner et al., 2020).

Examples of the target environmental sound and distractor written word pairings used in Experiment 2.
The two IVs were tested within items and participants. The dependent variable was naming latency. The order of the blocks was counterbalanced to prevent practice, learning, and fatigue effects. Each subject was presented with each environmental sound four times, each time with a different distractor word. The stimuli were pseudo-randomised for each subject using Mix (van Casteren & Davis, 2006). The randomisation was constrained such that the same target environmental sound would not appear within five trials, each target condition would not appear more than twice consecutively, and distractor condition was not repeated in more than three consecutive trials.
Procedure
Subjects were sat comfortably in a semi-dark room with a screen positioned ~30 cm away to limit their peripheral vision and so keep them focused on the screen (none of the participants indicated this distance was uncomfortable). Prior to commencing the experiment, participants were familiarised with the sounds and words. The familiarisation phase consisted of two stages, which were repeated until participants could correctly name all environmental sounds. During the first stage, each target environmental sound was presented simultaneously with the congruent written name of the target. In addition to the 32 target sounds, one trial displayed the letter string “xxxxx” with no sound present. In the second phase, each target environmental sound was presented in isolation. Participants were asked to produce the name of the sound. The order of trials was randomised. Naming errors were corrected immediately.
Three experimental blocks followed (128 trials per block), each block presented the distractor words at a different SOA (i.e., 0, 200, or 400 ms). Participants were informed the task involved naming sounds while ignoring distractors comprising real words or letter strings, and they were required to name the target sound aloud as quickly and as accurately as possible while ignoring the distractor, and to maintain their fixation on the crosshair in the middle of the screen. Each trial began with a black fixation cross-presented on the screen for 500 ms at an SOA of −500 ms. The distractor word remained on the screen for 800 ms and was replaced by a blank white screen until an SOA of 3,000 ms (allowing for a 3-s response window).
Results
The same criteria were applied to the raw data as in Experiment 1. After applying them, 6.33% of trials (639) were excluded from analyses. Applying the same trimming method resulted in 1.6% of trials (144) being identified as outliers. Table 3 shows mean naming latencies and error rates as a function of condition and SOA.
Mean naming latencies (in ms) and error rates (in %) as a function of condition and SOA in Experiment 2.
SOA: stimulus onset asynchronies; RT: reaction time. SEMs (in ms) are in parentheses.
The same analysis methods to Experiment 1 were conducted. Q-Q plots of the residuals did not reveal any obvious deviations from normality. A parsimonious mixed-model selection procedure resulted in models for the latencies including main effects for condition and SOA, and random slopes for condition, SOA per participant and per item (see Table 4).
Estimates of naming latencies as a function of condition and SOA in Experiment 2 from the model fit.
SOA: stimulus onset asynchronies; RT: reaction time. SEMs in parentheses.
For naming latencies, there was a significant main effect of condition, X2(3) = 215.06, p < .001, but the main effect of SOA, X2(2) = 0.64, p = .73, and interaction, X2(6) = 10.27, p = .11, was not significant. Compared with the unrelated condition, naming latencies were significantly faster in the neutral (B = −27.54, SE = 10.39, t = −2.65, p = .008) and congruent conditions (B = −96.99, SE = 10.28, t = −9.44, p < .001), and significantly slower in the semantic condition (B = 51.72, SE = 10.44, t = 4.96, p < .001).
In the analyses of error rates, there were main effects of both condition, X2(2) = 28.96, p < .001, and SOA, X2(2) = 7.34, p = .025. The interaction was not significant, X2(4) = 6.11, p = .19, with fewer errors in the unrelated condition (intercept) (B =−2.96, SE = 0.22, Wald’s Z = −13.33, p < .001), the congruent (B = −1.07, SE = 0.18, Wald’s Z = −6.12, p < .001), and the neutral (B = 0.34, SE = 0.18, Wald’s Z = −1.93, p = .052) conditions and no significant difference for the semantic condition (B = −0.02, SE = 0.13, Wald’s Z = −0.15, p = .88). More errors were produced during trials presented with an SOA of 400 ms compared with SOA 0 ms (B = 0.20, SE = 0.08, Wald’s Z = 2.42, p = .02) but no significant difference between SOA 200 ms and 0 (B = 0.06, SE = 0.08, Wald’s Z = 0.66, p = .51).
Discussion
Experiment 2 demonstrates, for the first time, semantic interference and congruency facilitation effects when naming a target environmental sound in the presence of a written word distractor. The observed effects provide clear evidence that the concurrent processing of a written distractor word affects naming latencies for target auditory objects as it does with target pictures (i.e., as in the PWI paradigm). The results also converge with Wöhner et al.’s (2020) findings for picture–sound distractor–target combinations.
General discussion
The two experiments presented here provide convincing evidence that congruency facilitation and semantic interference effects in spoken word production are not specific to verbal stimuli. In Experiment 1, we replicated Mädebach et al.’s (2017) findings with environmental distractor sounds and picture naming in English speakers. The results from Experiment 2 demonstrate, for the first time, similar effects during an SWI paradigm.
Although we replicated Mädebach et al.’s (2017) finding of semantic interference in Experiment 1, it is worth noting that we did not observe a similar significant interaction with SOA. This is possibly because their earliest SOA of −500 was 100 ms earlier than our own. Hence, our findings might indicate a lower bound of −400 ms for the time course of the effect. Numerically, we also observed a maximal effect at −200 ms. Mädebach et al. (2017, 2018) interpreted their findings as supporting a locus for semantic interference at the lexical–semantic (lemma) level, consistent with the lexical selection-by-competition account (e.g., Abdel Rahman & Melinger, 2009, 2019; Roelofs, 1992). According to this explanation, the conceptual activation from the sound distractor spreads to the lexical level, resulting in increased competition with the target picture’s lexical representation. They argued that a post-lexical account devised for PWI effects that relied upon more rapid written distractor processing could not provide a viable explanation for semantic interference from environmental sounds (e.g., Mahon et al., 2007), because this would require an additional ad hoc assumption that perception of an environmental sound automatically leads to retrieval of its referent’s production-ready (i.e., articulatory) representation in the same manner as a written distractor. In addition, such a process would be too slow to influence post-lexical processes, that is, an articulatory response to the target picture would be prepared faster than an articulatory response to the sound distractor due to the relatively longer time needed to identify auditory objects (see Chen & Spence, 2013).
We also replicated Mädebach et al.’s (2017) finding of significant congruent facilitation. Together, these findings are consistent with Chen and Spence’s (2010, 2011, 2013, 2018) work on cross-modal priming showing that target picture processing can be rapidly primed by environmental sounds via cross-modal conceptual associations and, moreover, this effect is relatively short-lived. They interpreted this as evidence that environmental sounds access their conceptual representations directly, like visual depictions of objects, whereas spoken words access their meanings more slowly via lexical representations. However, to maintain ongoing representations of sounds, either auditory imagery or lexicalisation mechanisms are needed which entail additional processing (see Chen & Spence, 2018). Thus, for semantic interference to be elicited, auditory and visual target and distractor objects need to be presented within a relatively short time-window (~200 ms) so that their conceptual co-activation can influence lexical-level processes in picture naming.
The novel results from Experiment 2 demonstrate the congruent facilitation and semantic interference effects can also be elicited by written word distractors when naming target sounds. All production models attribute congruent facilitation effects to target activation levels being raised by overlapping target and distractor representations, with this overlap potentially occurring across multiple stages of production. Our results are consistent with and extend those of Wöhner et al.’s (2020) recent study in which participants named environmental target sounds and ignored distractor pictures. They too observed facilitated environmental sound naming for congruent distractors and a semantic interference effect. However, Wöhner et al. (2020) reported effects with picture distractors at SOAs of −200 and 0 ms. The different onsets of the effects between studies are attributable to the use of word and picture stimuli as distractors, as the former are processed more quickly (e.g., Glaser & Düngelhoff, 1984). Interestingly, the ranges of SOAs (i.e., durations) of the effects in both studies are comparable to those reported for PWI (see Bürki et al., 2020). As Wöhner et al. (2020) proposed, this may be because both pictures and written words “unintentionally activate their lexical representations while they remain on the screen” (Damian & Martin, 1999, p. 351), and so may be relatively robust to variations in relative processing speed.
The finding that written distractors produce semantic interference during sound naming in Experiment 2 is explicable by all current models of PWI effects. Hence, it could either reflect conceptual activation from the written distractor spreading to the lexical level, resulting in increased competition with the target sound’s lexical representation (e.g., Abdel Rahman & Melinger, 2009, 2019; Roelofs, 1992), or a response-ready representation of the distractor word arriving in an output buffer before that of the target sound (e.g., Mahon et al., 2007). However, it is possible the effect we observed is limited to distractors that are also members of the response set (i.e., also target names) as employed in the current experiments, if the connections between sounds and their corresponding lexical-semantic representations are relatively weaker than those for pictures (e.g., Wöhner et al., 2020). This would be interesting to test in future studies. It is now well established that semantic interference can be elicited in PWI without response set membership (e.g., Caramazza & Costa, 2000; Gauvin et al., 2020). However, Mädebach et al. (2017; Experiment 4) did not observe semantic interference when distractor sounds were not part of the response set. Low-level visual features (i.e., form similarity) are also able to induce interference in picture naming, but only when members of the response set (de Zubicaray et al., 2018).
In summary, we replicated the PSI results of Mädebach et al. (2017, 2018) and introduced our own novel findings from an SWI paradigm. In both studies, we observed congruent facilitation and semantic interference effects using environmental sounds as distractors or targets for naming. Hence, these effects are not limited to verbal stimuli and occur regardless of the stimuli form (i.e., picture, word, or environmental sound), or modality of perception (i.e., visual or auditory).
Supplemental Material
sj-docx-1-qjp-10.1177_17470218221137007 – Supplemental material for Neighing dogs: Semantic context effects of environmental sounds in spoken word production - a replication and extension
Supplemental material, sj-docx-1-qjp-10.1177_17470218221137007 for Neighing dogs: Semantic context effects of environmental sounds in spoken word production - a replication and extension by Sonia LE Brownsett, Matteo Mascelloni, Georgia Gowlett, Katie L McMahon and Greig I de Zubicaray in Quarterly Journal of Experimental Psychology
Footnotes
Acknowledgements
The authors are grateful to an anonymous reviewer and Andreas Mädebach for their helpful comments on an earlier version of this paper. This research was supported by Australian Research Council Discovery Project Grant DP200100127 & DP150103997. This experiment was realised using Cogent 2000 developed by the Cogent 2000 team at the FIL and the ICN and Cogent Graphics developed by John Romaya at the LON at the Wellcome Department of Imaging Neuroscience.
Author contributions
All authors included on this manuscript contributed significantly to its development.
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The authors were supported by Australian Research Council Discovery Grants (DP150103997 and DP200100127).
Data accessibility statement
Supplementary material
The supplementary material (list of target picture stimuli consisted of 32 simple black line drawings of the 32 target objects listed in Mädebach et al. (2017)) is available at ![]()
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
