Abstract
Onomatopoeia are words signifying sounds by phonetically imitating or suggesting them. Research into the use of onomatopoeia in literature, language and comics has been rich and varied while the study of the use of onomatopoeia in film, television and their functions has so far been limited. This article focuses on two aspects: (a) the relationship between onomatopoeia and sound effects in US animation and Japanese anime; and (b) audiovisual humour caused by the use of onomatopoeia and sound effects in US animation and Japanese anime. The author introduces classifications and usage patterns of onomatopoeia in English and Japanese, and analyses the different types of onomatopoeia and sound effects in recent popular US animation films and TV, and in Japanese anime and their contributions to creating humour by applying humour theory (incongruity theory in particular). This study combines film analysis, soundtrack analysis, linguistic analysis and theories of humour, and will be of interest to students and scholars of film, sound, linguistics, psychology and other subjects interested in the study of onomatopoeia and audiovisual humour.
Keywords
In comics, including manga, 1 onomatopoeia – that is, words that signify sounds by phonetically imitating or suggesting them – are ubiquitous. But what happens if the still images of comics start moving? What roles do onomatopoeia play in animated films and television, and how do they interact with other elements of the sound and image tracks? That is the primary focus of this study which, in order to answer this question, looks at (and listens to) examples from recent US animation films and TV, and Japanese anime, ‘the two predominant forms of animation garnering substantial popularity in the present era’. 2 The second focal point of this article is humour. While not all animation is comic in intent and effect, much of it is, in both US and Japanese cultures (though more so in the US than in Japan), and the contribution of soundtrack elements other than dialogue to the humour of animated films and TV shows has been sporadically studied. 3 Onomatopoeia can contribute to a sense of dynamism and immediacy in a comic or manga, and to some extent also in animation/anime but, together with other visual and sonic elements, it can also contribute to their creation and humour, and this study shows a range of options for that.
Literature on onomatopoeia in animation is limited. Fukusato and Morishima (2014) have designed a computer system to estimate, depict and quantify onomatopoeia based on physical parameters. Naomi (2007, 2019) analyses the adaption of onomatopoeia from manga to anime in several manga-based anime and Putri et al. (2020) examine emotions connected to onomatopoeia in the first season of Haikyuu!! (2014, a comedic anime for young teen boys). The examination of onomatopoeia within the context of comics and manga has yielded a comparatively substantial body of research, spanning diverse facets of analysis. For instance, Xie (2002) and Li (2007) analyse the use of onomatopoeia and mimetic words in Shōjo manga (i.e. manga targeting female teens). Petersen (2009) explores the rich aesthetic possibilities of sonic and visual spectacle in manga. Forceville et al. (2014) explore the relationship of onomatopoeia and speech/thought balloons in comics, Guynes (2014) provides a Peircean analysis of mainstream US comic book onomatopoeia. Pratha et al. (2016) compare English and Japanese onomatopoeia in comics and manga.
However, there is no study comparing the different types of onomatopoeia and sound effects and their relationships in anime and US animation, let alone an examination of the contribution onomatopoeia make to the audiovisual humour of animation. This article revisits US animation and Japanese anime with a focus on onomatopoeia and humour, on the basis of humour theories – (incongruity theory and relief theory), using examples from selected popular US comedy animation and Japanese anime of the last 20 years.
Onomatopoeia in language
The uses of onomatopoeia in anime and in US animated films and TV shows are largely influenced by the traditions of using onomatopoeia in daily life and language in both cultures, and to understand their functions in animation, it is necessary to look first into the respective linguistic and cultural backgrounds. Onomatopoeia in English literature are words ‘imitating the sound associated with the object or action designated’ or a combination of words to ‘evoke a certain image or mood by sound’ (Birch, 2009). The term can be traced back to Greek – ‘onoma’ (meaning name) and ‘poiein’ (meaning to make). Although a specific sound is heard similarly by people from different cultures, there are considerable differences between different cultures regarding the use of onomatopoeia.
Scholars started to study and classify English onomatopoeia in comics already in the 1940s (Hill, 1943a, 1943b). Bredin (1996: 559) argues that the use of onomatopoeia (in English) is not merely governed by ‘the nature of acoustic resemblance’, but also determined by convention. He classifies them into three categories (pp. 558–566):
1. Direct onomatopoeia: ‘the denotation of a word is a class of sounds’ and ‘the sound of the word resembles a member of the class’ (p. 558). The degree of the onomatopoeic resemblance can be either higher (e.g. buzz) or lower (e.g. rustle).
2. Associative onomatopoeia: the sound does not resemble the object or action that it denotes. It is the ‘acoustic resemblance’ (p. 560) that is embodied in the sound of the word and whatever associated sound the word denotes (e.g. bubble – like the sound of a bubbling liquid).
3. Exemplary onomatopoeia: ‘its foundation rests upon the amount and character of the physical work used by a speaker in uttering a word’ (p. 563). Namely, the way in which the word is spoken suggests the quality of things when speakers pronounce a word using different amounts and kinds of physical work. It thus exploits the innumerable associations between sounds and things (e.g. sluggish and sloth require more physical effort and time to say while nimble and dart require less physical effort and time to say, which corresponds to the qualities the words denote).
Attridge’s (2004[1988]) category system is wider and classifies English onomatopoeia into lexical and non-lexical ones. Lexical onomatopoeia pertains to well-established linguistic expressions whose phonetic form symbolically imitates the entities or actions they represent in the external world beyond language (pp. 145–150), including all the types of onomatopoeia in Bredin’s system. Non-lexical onomatopoeia covers ‘the use of the phonetic characteristics of the language to imitate a sound without attempting to produce recognizable verbal structures’ (p. 136) (e.g. utterances like vroom or brrrrm, imitating the sound we hear when a car revs up). Except for exemplary onomatopoeia, which are more abstract and do not aim to imitate a sound directly, the other types of onomatopoeia can all be found in (Western) comics.
The earliest recorded Japanese onomatopoeia can be dated back to the 8th century (Millington, 1993: 14). As Petersen (2009) suggests, Japan has a long history of storytelling that explores a considerable wealth of possibilities of sounds (including spoken and written onomatopoeia) and visuals, which greatly influenced the use of onomatopoeia and sounds in manga, although contemporary manga are largely influenced by US newspaper comics in the 20th century (pp. 167–168). 4 According to Ono’s (2007) dictionary, Japanese has around 4,500 onomatopoeia and mimetic words. They are generally regarded as onomatopoeia, though some of them are actually mimetic words, that is, words that imitate other things which can but do not always make sounds (Hamano, 1986). These types of sound symbolism are fairly widely used in contemporary Japanese life by people of all ages (Akita, 2009: 9–10). They fill everyday conversation, literature and media due to their expressiveness and the amount of information they convey. It has been the use of such sound symbolism in manga and anime, however, that has made Japanese onomatopoeia become world-famous.
According to Kojien (widely regarded as the most authoritative dictionary of Japanese), there are five categories of such words (Niimura, 1969: 2390):
1. Giseigo (擬声語): Words that imitate human and animal voices (e.g. nyaa [ニャア]: the SFX (sound effect) for the yowling of a cat).
2. Giongo (擬音語): Words that imitate sounds from non-living things (e.g. poo- [ポォーッ]: SFX for the whistling of a machine, such as a train whistle).
3. Gitaigo (擬態語): Words that describe visual, tactile, and other non-auditory sensitive impressions, especially conditions and states (e.g. nyari [ニャリ]: leering or grinning in a derogatory, teasing, or evil manner).
4. Giyougo (擬容語): Words that describe movements and motions (e.g. pyo-n [ピョーン]: SFX for a long jump).
5. Gijougo (擬情語): Words that describe feelings (e.g. waku [ワク]: a state of happy, cute anticipation).
Giseigo and giongo are what would be classed as onomatopoeia in English, that is, words imitating a voice and other sounds. Gitaigo, giyougo and gijougo belong to the category of mimetic words (which is important in Japanese, but which English and many other languages do not have in a similarly well-defined sense), i.e. words that phonetically express states that do not produce sounds, including emotions, movements and the state of a character/object. Many Japanese onomatopoeia contain more than one meaning, and some of them can be used either as mimetic words or onomatopoeia indicating real sounds, like gangan (ガンガン) which means: (a) ‘bang bang’ – something knocking ferociously (giongo); (b) loud sounds (giongo); (c) nagging (giseigo); (d) throbbing (gijougo); (e) a great deal of momentum, doing something powerfully (giyougo). The specific meaning depends on the kind of context it is used in.
Susan Millington (1993) thinks English onomatopoeia is more straightforward and overall less poetic than Japanese onomatopoeia, as Japanese encompasses plenty of onomatopoeia to describe emotions, feelings and subtle nuances of sounds while English onomatopoeia is mainly used for real sounds caused by people/things or actions. One reason for this, Millington proposes, is that the Japanese language relies on various onomatopoeia to modify verbs with supplementary information because the verbs themselves are much less varied than in English (pp. 12–13).
Onomatopoeia in manga/comics
The historical relationship between animation, comics and manga is multifaceted and intriguing. The rise of early US animation was deeply influenced by comic strips (e.g. George Herriman’s Krazy Kat, 1913–1944) and US newspaper comic strips featured in 231 films from 1916 to 1940), although today’s US animation has largely distanced itself from its comic roots. In contrast, as Novielli (2018: 1–2) claims, anime’s evolution has involved assimilating US animation experiences while preserving and enriching Japanese cultural characteristics. Emakimono (picture scrolls) 5 and kage (shadow puppetry) 6 are considered anime precursors, while manga, originating from Emakimono, shares significant parallels with anime, as numerous anime adaptations stem from manga sources. Given this historical connection, exploring the functions of onomatopoeia in these still-image traditions, especially in manga, becomes a meaningful endeavour.
Onomatopoeia are an important element of Western comics, but indispensable in Japanese manga. Japanese manga use much more onomatopoeia (in quantity, range, variety, etc.) than US comics. Onomatopoeic sounds can be found in the majority of manga while a number of classic Western comics do not use onomatopoeia at all (e.g. Hal Foster’s Tarzan, 1929, and Prince Valiant, 1937).
Contemporary US comics can be broadly divided into mainstream comics, mostly in the superhero genre, including DC and Marvel Comics, Image, Dark Horse, IDW, America’s Best Comics, Big Dog Ink, Boom! Studios, Dynamite Entertainment and independent comics that are typically more flexible in style (Guynes, 2014: 58–59). Pratha et al. (2016) conclude that independent comics often contain both onomatopoeia (e.g. ‘pow’) and descriptive words (words that ‘describe’ the action or sound being portrayed instead of merely mimicking the sound it produces, e.g. ‘punch’) to capture sounds and actions, while mainstream comics hardly use descriptive words (pp. 96–99). Onomatopoeia in mainstream comics are usually written in a highly conventionalized style – colourful hollow words with or without balloons in different shapes (wavy, jagged, square, or other shaped borders) and colours around them, ‘combined with its panel border-exceeding nature’, making the emphasized sounds look loud (Forceville, 2013: 259) (see Figure 1 for an example).

An example of onomatopoeia in the style of mainstream comics from Shutterstox website. It is a hollow word in red with jagged balloon and clouds as decorations, indicating the harshness and power of the sound.
Japanese uses different writing systems: kanji, logographic characters derived from Chinese, and Kana, consisting of a pair of syllabaries – hiragana and katakana. They function very differently. Onomatopoeia in manga are usually written in hiragana (e.g.あ) or katakana (e.g. ア). Hiragana is the first form of writing that Japanese children learn, and it is the most widely used standard form of Japanese writing. Katakana is typically used to write foreign words or loan words (also scientific words, names of plants and animals, etc.), to add emphasis or to represent prominent externally audible sounds. Its characters look more abstract (they are derived from small fragments of complex kanji) and look sharper and more irregular.
According to the statistical analysis by Pratha et al. (2016), shōjo manga (mainly aimed at a teenage female readership) ‘used hiragana and katakana in roughly equal proportion in the writing of sound effects’, while shōnen manga (mainly marketed towards a teenage male readership) ‘used katakana far more than hiragana’. ‘Shōjo manga used substantially more gitaigo than shōnen manga’ to describe the emotional or physical state of animate objects and characters, featuring in genres like romantic manga, stories about magic girls, etc. Shōnen manga ‘tended to use more giongo than shōjo manga’ to describe the sounds of inanimate objects, featuring in genres such as fantasy, action, adventure, sports, martial arts, etc. (p. 102).
Onomatopoeia bring vividness to the description of sounds, movements, states and emotions without using tedious words to explain the details. Comics, especially manga make extensive use of all of these types of onomatopoeia to make the silent images become ‘audible’ by proxy and to explore the intrinsic aesthetic possibilities of sounds in the panels (an individual frame depicting a frozen moment in the multiple-panel sequence of a comic/manga).
Onomatopoeia in manga (and also in Western comics) can both mean a word delivering information and a component of visual decoration for the panel, integrated with the images (Fusanosuke, 2012: 109; Pratha et al., 2016). There is a wide variety of character options, and the ‘size, boldness, tilt, exclamation points’, ‘font styles, shapes’, typesetting, use of shadows or glow effects, placement, spatial orientation, ‘line and colour’ (if any) etc. are varied and can be either written in different types of balloons or directly in the panel, indicating the different ‘loudness’,‘quality, roughness, waviness, sharpness, fuzziness’, ‘source’, pitch and duration of the sounds (McCloud, 2006: 146–147) (see Figure 2 for an example).

An example of giant onomatopoeia in yellow with jagged balloon and dense speed lines (SFX for an explosion) in manga from Shutterstox website, indicating the harshness and lethality of the sound.
Onomatopoeia enrich the panels and also influence readers’ perceptions of the virtual world of sounds or emotional states in a medium-specific form of synaesthesia, a perceptual phenomenon in which stimulation of one sensory or cognitive pathway (i.e. the visual sense) leads to involuntary experiences in a second sensory or cognitive pathway (i.e. the auditory sense) (Salgueiro, 2008). Onomatopoeia as a type of varied pattern are even comprehensible to an illiterate child or to people who cannot read Japanese.
Meanwhile, the pronunciation of onomatopoeia also influences readers’ perception of the images through synaesthesia. A classic experiment examined the relationship between the pronunciation of onomatopoeia and the visual realm (Köhler, 1929; Ramachandran and Edward, 2005). It proves that words spoken in different forms and with different positions of lips and tongue have connections, to a certain extent, with the shape of the icons (e.g. people tend to match words spoken with round lips, like ‘Bouba’, with round images and vice versa). Likewise, onomatopoeia and balloons depicted in different shapes and colours also influence people’s perceptions of the onomatopoeia themselves.
In addition, readers understand and imagine written onomatopoeia through what Petersen (2009) calls ‘subvocalization’. They perform the internal sounds for themselves, based on personal experiences and habits. Therefore, they are both the audiences and participators of the performance ‘performing the comic for themselves, just as a ventriloquist might bring a voice to a puppet while acting as witness to the character he/she manipulates’ (p. 164).These internal sounds are unique for each reader and help to bring comics or manga to life for them.
Translating Japanese onomatopoeia into English inevitably leads to problems, whether in literature (Bartashova and Sichinskiy, 2014; Inose, 2007; Iwasaki, 2013) or in manga (Jüngst, 2014; Sell, 2011). It is common for manga that are translated into other languages to leave these sound words untranslated, for three reasons. Firstly, it is ‘costly and time-consuming’ to edit the sound effects and replace them with English for this ‘retouching the whole picture’ (Jüngst, 2014: 64–65), while the budget for such translations is usually limited. Secondly, English onomatopoeic expressions in most cases are considered childish and informal. A literal translation could not retain the register of the original Japanese text and thus messes up the flow and impact of the story. Last but not least, English lacks most of the sound effect words found in Japanese, so English cannot convey the subtle nuances of the diverse onomatopoeia (e.g. ‘gossu’ and ‘gosu’ in the example of Gin Tama in this article) and mimetic words. Manga readers who do not speak Japanese have to learn these words (which means that, at least in this respect, manga readers have an edge over traditional book readers). There are free online Japanese SFX translation websites (e.g. <http://thejadednetwork.com/sfx/>; <https://onomatopedia.jp/>) which include general explanations of the majority of common Japanese onomatopoeia that appear in manga and so can help non-Japanese readers to grasp the main idea of these words.
Humour theory
Before analysing examples of onomatopoeia in animation and their comical effect, it is necessary to introduce humour theories first. The question of what makes something funny has occupied scholars for centuries. Multiple humour theories have been put forward by scholars from different fields of study, including psychology, linguistics, philosophy, etc. This article introduces two of the classic psychological humour theories – incongruity theory (and its developments) and relief theory, both of which provide theoretical perspectives to examine and explain the mechanics of onomatopoeia’s contribution to audiovisual humour in animations and anime, an interdisciplinary approach that has not been systematically undertaken in the existing literature. I will also use Spencer’s idea of the ‘descending incongruity’ to explain the humour caused by incongruities involving a shift in the perceived value or importance of the interpretation of a stimulus; and I will use Freud’s version of relief theory to explain humour playing with social taboos.
Incongruity theory is widely acknowledged as the most comprehensive theory to explain humour, and it is the theory applied to the majority of humorous examples in this article. 7 Incongruity theory pays close attention to the cognitive aspects of humour and posits that humour is the result of perceiving and interpreting incongruity between two distinct explanations for the same item or fact. Martin and Ford (2018) put forward their explanation for humour by combining incongruity theory with the idea of a playful mindset as a necessary condition for an incongruity they perceive to be humorous. They propose that there are two cognitive–perceptual processes in an individual’s mind activated at the same time when humour is to work: ‘(1) the perception of incongruity and (2) appraisal of incongruity in a nonserious humor mindset’ (p. 4). Namely, there are two opposing ways of making sense of a stimulus (the incongruity), and that the humour stimulus must give cues to the receivers to appraise the stimulus playfully and non-seriously so that ‘people temporarily abandon rules of logic and expectations of common sense and congruity’, as they describe the necessary condition of humour.
In relief (or release) theory, humour is generated by the build-up of psychological tension and its sudden release in an individual. Spencer (1860), one of the most prominent proponents of relief theory, introduces an idea that can become an addition to incongruity theory – the concept of the ‘descending incongruity’ to describe a process in which the consciousness of the recipient of a joke or humorous event: ‘laughter naturally results only when the consciousness is unawares transferred from great things to small – only if there is what we may call a descending incongruity’ (p. 400). Freud (1960[1905], 1928), another proponent of a relief theory of humour, believes that we are all filled with excess energy all the time because, in our everyday lives, we have to suppress urges that are taboo (e.g. aggression, violence, sexuality): ‘suppressed purpose can . . . gain sufficient strength to overcome the inhibition, which would otherwise be stronger than it’ (1960[1905]: 187). Once the veiled insult or sexual reference is made in a socially acceptable manner (e.g. by allowing people to talk about taboo subjects in unserious ways), the repression is thus overcome, the ‘incomparably greater’ amount of humour and pleasure can be released, carrying therapeutic value. 8
Adapting onomatopoeia from manga to anime/comics to animation
The use of sound effects in anime/animation is largely directed by the onomatopoeia in the manga or comics they have been adapted from. This section focuses on the relation between onomatopoeia in manga and sound effects in anime because a great many anime are adapted from manga, which facilitates comparisons while there have been far fewer popular comics-based US animations in recent years. There are three main approaches for anime to represent the sounds written in manga: (a) directly playing sounds that onomatopoeia imitate or tend to express; (b) using onscreen written onomatopoeia; and (c) characters or voiceovers verbalize the onomatopoeia (this will be discussed in the next section). These three modes can be used simultaneously. Whether such adaption supports or creates humour depends on the right kinds of incongruities between words, sound effects and images.
Contemporary anime still keep a large amount of the features of ‘limited animation’, i.e. a type of animation that reduces the total number of frames to save time and budget. The projection rate of traditional live-action film is 24 frames per second. The frame number in full animation is no less than 12 frames per second to guarantee fluency and a sense of realism, while in limited animation, the rate of frames is only around 8 per second (Lamarre, 2009: 187–189). In order to keep a sense of dynamism and continuity, despite the fact that the low frame rate makes individual images visible for longer, anime is often cut quite rapidly. Thomas Lamarre believes ‘cutting from image to image increases in importance, as do the rhythm and speed of cuts. Cutting between static drawings tends to work well with scenes of characters talking . . . to become more important in introducing a sense of continuity across cuts’ (p. 191). But this still lacks ‘in-between’ frames that fill in the missing trajectories of movements. Moving onscreen onomatopoeia as a type of visual pattern can be used to compensate to some extent for this lack of fluid movement.
When sounds effects are played and the corresponding onomatopoetic words appear onscreen simultaneously, Furuhata (2012) uses the term ‘audiovisual redundancy’ – a technique that strengthens the sense of intermediality and attracts attention to the similarity and dissonance between anime and manga. Petersen (2009: 166) writes that this is not a real redundancy but that it serves ‘to slow the reader down and create greater visual depth and texture to the scene’. The increase in the size and frequency of the onomatopoeia to fully represent dramatic sounds for key moments ‘provides time for the intensity of the action to develop’ and ‘gives force and dimension to the dramatic action’. This ‘redundancy’ also helps to express and transmit emotions. Smith (2003) calls perceptual stimuli transmitting emotional signals ‘emotion cues’. The process can be provoked by both audio (dialogue, intonation, sound effects, music) and visual output. Each single cue is supposed to induce a kind of emotion. In consideration of the range of experiences and differing perceptiveness of different audience members, filmmakers prefer to mobilize multidimensional filmic cues to underline specific motivations and emotions. Such ‘redundant emotion cues’ (Smith, 2003: 43) guarantee that the majority of viewers will be able to pick up the signal. Onscreen written onomatopoeia in animation are such ‘redundant cues’ that arouse the audiences’ attentions and deliver the information and emotions through both the audio and visual channels, which enables viewers of different ages and levels of media experience and literacy to grasp the main emotional tendency of a moment or scene.
A great many onomatopoeia in manga, especially giseigo and giongo, disappear in anime adaptations because they can be replaced by the corresponding realistic sounds or unrealistic cartoon sound effects. One way for anime to create humour based on manga is to create sound metaphors – the incongruity between the actual sound ‘X’ that would be heard in a realistic setting and a different sound ‘Y’ with some metaphoric connection to the action. For example, in Haven’t You Heard? I’m Sakamoto (2016, 1: 5, ‘Charisma Yankee Senior 8823’, 09.00–09.20), Sakamoto’s shoes are stolen after he has used the toilet in school, but he is unfazed and as poised as usual. He draws the shape of a slipper and writes his name on his white socks. Then he opens his arms, slides forward swiftly on one leg with the other leg hanging high in the air, making the shape of an aircraft. ‘Fu-’ (フー) written in the original manga is the SFX for a hiss, indicating the sound of the sock rubbing on the floor. But we can hear the swish and a sharp grating that imitates the sound of an aircraft gliding across the airfield, which echoes the pose of the character and the speed-lines in the frame. Both the sounds and the visuals contain the metaphor that compares a human to an airplane. The incongruity causing humour in this scene is that: (1) Sakamoto pretends to be stylish and fast and to appear like an airplane, for which he receives the appreciation of his classmates; and (2) whatever the reason, the location, the props and the gestures, he looks like a fatuous, overly self-confident show-off rather than a superhero.
Written giseigo and giongo sometimes appear in images to enhance the impact of the sounds that they represent and which have been put into the anime. As with onomatopoeia in manga, the font, size, shape, typesetting, etc. of the onomatopoeia in anime are various, without a fixed style and standard, in order to tailor their impact to the requirements of the scene.
US animations, especially those adapted from comics, also occasionally add written onomatopoeia (mostly direct and associative onomatopoeia) to the images to pay homage to classic comics. The 1960s Batman TV series that is based on the DC comic book character of the same name is best known for adding obtrusive and decorative written onomatopoeia to images of the characters during comic-book fight scenes, synchronized with dynamic sounds played by loud trumpets in a high range. For one thing, the onomatopoeia and the sounds underline the dramatic effect of each movement in a playful form. On the other hand, the onomatopoeia help to avoid having to show a potentially uncomfortably violent scene.
Although few of such scenes in the original Batman films and TV series are played for laughs, many later animations have paid homage to or parodied Batman and the fighting scenes for comic effect. Incongruity can be created through the similarities and differences in the looks and behaviour of the original Batman and the new ‘Batman’ (who may, for example, have the classic mask and cape, and speak in a deep raspy voice but behave ridiculously). Instead of using the proper onomatopoeia to represent different fighting noises, such parodic animations tend to use or create different onscreen onomatopoeia or other words in a similar pattern with similar sounds and music during fighting scenes, but the words are absurd, ironic or even taboo. Characters in Futurama (2003, 4: 58, ‘Less Than Hero’, 11.20–11.30) transform into superheroes to fight crimes. The onscreen word used is ‘01001010!!!’ (that is, binary computer code) when Bender, a robot, is beaten by a kangaroo. The incongruity is: (1) the word is expected to describe the noise of hitting metal objects and the glitch; (2) the numbers are not onomatopoeia but reflect Bender’s disordered system caused by the strike that has hit him. (One could see this as a case of using the on-screen text for internal focalization: it represents Bender’s inner state, his robot experience of the punch rather than the action itself.) In The Simpsons (1995, 7: 2, ‘Radioactive Man’, 02.40–02.50), the word ‘SNUH!’ appears on screen during a parodic fight scene. It is the acronym of an organization (Springfieldians for Nonviolence, Understanding, and Helping) against cartoon violence from an earlier episode (1990, 2: 9, ‘Itchy & Scratchy & Marge’), but ironically used as onomatopoeia relating to violence in this scene.
Mimetic words in manga that do not involve real sounds are also often attached to sound effects in anime to underline their effects. Sound effects of different types can be applied to the same or to similar elements in the images to bring out their funny side through the implicit comparison. In The Disastrous Life of Saiki K. (1: 3, ‘How Shady! Dark Reunion’, 12.30–12.50), Kaidou is an arrogant and innocent middle-school student who always believes he is the chosen person to save the world. In the episode in question, a classmate makes a fool of Kaidou by asking him to copy his two movements and slogans to put up a strong barrier to prevent being annihilated by the devil who of course does not actually exist. Throughout the scene, we hear intense electronic background music that seems to be taking the defence-against-the-devil conceit seriously, while the sound effects create the comedy. The scene uses two different sound effects: a harsh bang for Kaidou’s classmate followed by a comical cartoon squeak sound 9 for Kaidou himself, repeated in each pose. The contrast between the sound effects satirizes Kaidou’s gullibility and cluelessness, and creates a status difference between the movements his classmate and he himself make. Here, the anime goes beyond what the manga it is based on can do: In the manga, the scene uses two words (‘キュ’) (squeezing one’s hand in frustration or anger) and (‘カッ’) (footsteps; bang) describing their two body movements, and only ‘(カッ)’ contains real sounds. The sound effects in the anime do not present the real sounds of (‘キュ’) and (‘カッ’) and their differences, but make a comparison between the two boys’ movements for comic effect.
Anime sometimes use written mimetic words in the images to repeatedly emphasize certain movements/emotions. In the second episode of Back Street Girls (2018, 1: 2, 16.30–17.30), the three boorish yakuzas are forced to get sex reassignment surgery and are trained to become ‘female’ idols. 10 Clueless about anything to do with girls, they mistakenly think that ‘massage’ is a type of communication between girls. They kneel on the ground and rub each other’s breasts but do not feel anything pleasant. The incongruity is between the following perspectives on the scene: (1) girls rubbing their breasts is a stereotypically male fantasy, and it is reasonable for the yakuzas to do so because their gender identity is still male; (2) their sex becomes that of girls, and they look and sound like girls, but the idea that girls communicate with each other by massaging their breasts is a kind of sexist misunderstanding, used for an incongruity that plays with the mismatch between gender and sexual identity. This could be a very indecent scene if it were played by real-life adult actors and if it used real rubbing sounds. But, in the anime, it becomes a comic scene because both images and sound effects convince the audience that what happens is ridiculous. We see a series of the giyougo words ‘もみ (groping)’, coloured in poppy pink and bouncing around the characters, synchronized with their movements, and we hear five stereotypical plucked cartoon sound effects, widely used for sundry unserious movements in anime, to provide a sense of briskness, as if the yakuza are pinching elastic toys. All of this diminishes the sense of salacity and provides a sense of unseriousness that makes fun of the sexist misunderstanding.
Onomatopoeia can not only occur as on-screen characters or be replaced by the corresponding sound effects in anime and animations, but can also be spoken by voiceovers or by characters, fictional or extrafictional, for different purposes.
In both English and Japanese, people, especially children, more frequently use spoken onomatopoeia in informal conversation than in formal written language. Japanese adults use far more onomatopoeia and mimetic words than English adults in both conversation and writing. English adults tend to avoid the use of onomatopoeia in formal conversation and writing because they are regarded as too emotive, childish and unserious (Schourup, 1993: 51). Characters in anime and animations use verbal onomatopoeia far more than characters in live-action films in order to show the characters’ cuteness, liveliness, childishness, etc.
Characters in US animations often speak non-lexical onomatopoeia (or direct onomatopoeia containing a high degree of onomatopoeic resemblance) in a way similar to an interjection (an utterance on its own and expresses a spontaneous feeling or reaction). For example: saying ‘ta-da’ to imitate a fanfare, strengthening an impressive entrance or a dramatic announcement; saying ‘dun-dun-dunn’ as a dramatic pause or to emphasize that something frightening or thrilling; and saying ‘nom-nom-nom’ 11 when eating something delicious.
Sometimes, new onomatopoeia are made up for a certain character or motion. The most typical one is the catchphrase – a repeatedly used phrase or expression that also mimics a kind of sound or contains onomatopoeia. For example, Lighting McQueen’s catchphrase in Cars (2006) is ‘Ka-chow!’ It is a made-up word that seems to mimic the sound of lightning. Lightning does this when he poses and brightens his lightning bolt stickers to express his excitement and happiness, similar to shouting ‘hooray’. His antagonist Chick ‘Thunder’ Hicks always wants to win in the racing competition but never succeeds (because thunder comes after lightning!). He imitates Lightning and says ‘Ka-chigga’ as his own catchphrase (and the fact that, unlike ‘Ka-chow!’, it ends on an unstressed syllable already indicates its inferiority).
Onomatopoeia in adult animations can be used as a pun related to social taboos and to create both incongruity and relief humour in a seemingly innocent way. The cheerful song ‘Poo-Choo Train’ in South Park (2002, 6: 17, ‘Red Sleigh Down’, 05.00–05.40) seems to innocently sing about a running train made of poo, and the lyrics ‘Poo-Choo’ represent the steam train whistle sounds. In reality, the song refers to tormenting an unsuspecting victim and causing them to poo their pants. The children in the episode also sing ‘Poo choo train is my favourite thing, spreading Christmas joy as we ride and sing! Christmas time would not be the same without hugs and kisses and a poo choo train!’ in a seemingly innocent manner even though they are well aware of the indecent implication.
In English, some verbs and nouns used to describe processes of sound-making more or less resemble the sounds they represent, a phenomenon on the edge of lexical onomatopoeia. These words mostly exist in written English (rather than in spoken dialogue) to replace the corresponding real sounds. There is an example of the use of ‘yawn’ in The SpongeBob SquarePants Movie (2004, 19.00–19.10): Mr. Krabs, the boss, asks his new manager Squidward to keep a sharp eye out for paying customers. Squidward listlessly responds ‘Yawn’ while putting a hand over his mouth. ‘Yawn’ signifies a motion and state but also imitates the real yawn sound. Instead of having a real yawn (that is, a sound like ‘ahhhh-hhaaaaaa’), this spoken ‘yawn’ sounds unnecessary and unnatural. But it gives a knowing, self-conscious quality to the reaction. Squidward seems to use this word consciously and deliberately to describe both his real physical state and his reluctant psychological state.
Since spoken onomatopoeia in US animations tend to provide a general sense of innocence and unseriousness, incongruity arises when it is applied to more serious situations. In both The LEGO Movie (2014) and The Lego Movie 2: The Second Part (2019), numerous diegetic sounds are replaced by spoken onomatopoeia in fighting scenes. The escaped pigs are dubbed ‘oink’ by a voice actor in different tones, a cat is dubbed with an emotionless ‘meow’ when the Lego world is in turmoil, we can also hear ‘pew pew’ that replace the gun sound effects, and the fatal shark weapons make ‘nom nom’ sounds. Sound effects like a voiced bilabial trill (which can be regarded as a spoken non-lexical onomatopoeion) can be heard in the stop-motion scenes to replace the sounds of a ship or aircraft engine during escapes shown in distant shots. The animators try to echo and hark back to how a child might make a film. In effect, they ‘alternate between thinking like responsible filmmakers working on a large-budget Warner Brothers animated film, and then suddenly approach a scene like a kid animating in their basement’ (Freckelton, 2017). To be exact, the incongruity is: (1) the fighting/quarrel scenes shown in CGI with harsh sound effects and intense music in high quality convince audiences of the seriousness of the scene; (2) the distant shot reveals the unseriousness when the escape is shown stop-motion with the vocal imitation sounds. These examples are the conceit of the film: what we see and hear is the fantasy of Finn (Jadon Sand) who is playing with his father’s elaborate Lego set and using onomatopoeic sounds as kids usually do when they play with toys.
There are also ‘superfluous’ spoken onomatopoeia that are invented to represent unrealistic cartoon sound effects. For instance, the neologism ‘yoink’, famous in The Simpsons, is an onomatopoeion of a fictional comical cartoon sound effect, spoken by characters to emphasize the playful theft of an item in front of others.
Animations occasionally deliberately cause a misunderstanding of the sources of the spoken onomatopoeia to create humour, usually achieved through different angles of shots. In Dan Vs. (2011, 1: 5, ‘The Animal Shelter’, 07.30–07.40), Chris gets poisoned and is treated in hospital. We can hear two different intermittent beep sounds during a high-angle shot, as if both sounds are coming from medical apparatus, hinting at the seriousness of the condition although one of them is low in sound fidelity. Then a medium shot reveals that Chris’s best friend, Dan is saying ‘beep’ following each real beep sounds for fun. His frivolous behaviour amuses the audiences by releasing the previous tension, a classic descending incongruity. In The Lego Movie 2: The Second Part (2019, 09.30–10.30), a diegetic chart-topping pop song played by a beachgoer’s radio ends with an extremely long ‘boo’ in rising pitch that indicates contempt for Lloyd, the protagonist. Then the camera pans out, and we see millions of enemy aircraft flying toward the city centre. The onomatopoeic lyrics ‘boo’ shift attachment to represent the sound of the engines (and are remixed with real engine noise later). The shifting attachment of the ‘boo’ sound (as part of the song, as a comment on Lloyd and as the engine sound) shows the playful nature of the make-believe reality of the Lego world embedded in the first-level fictional reality of the film.
Nevertheless, there are far fewer limitations on characters speaking any type of onomatopoeia in everyday Japanese dialogue. For example, it is not strange to find characters of all ages saying unheard mimetic words like Gitaigo ‘jii-’ (ジィーッ: SFX for gazing fixedly at something or someone) or ‘doki doki’ (ドキドキ: SFX for love situations; a scared or anxious heart thump) in many comedy manga and anime with/without onscreen written onomatopoeia and with/without different sound effects underlining them. Experienced anime audiences may get used to this, but it still sounds and looks strange for English speakers to see a grown man do this when Japanese spoken onomatopoeia in dialogues and onscreen written onomatopoeia are translated into the English version.
This also applies to character names that contain onomatopoeia. Many of the Devil Fruits (mystical and mysterious fruits) in One Piece (1999–present) are named after sound effects (Japanese onomatopoeia) related to their respective abilities which endow the eaters with a specific super-human power. For example, Goro Goro no Mi (ゴロゴロの実) is named after the sound effect of thunder rumbling. It allows the user to turn into ‘Lightning Human’ to create, control, and transform into electricity at will. The double form is usually used as an adjective, expressing a continuing state of the sound or feeling. Its official English name is Rumble-Rumble Fruit. The doubling of the onomatopoeia in English does not mean the same thing as in Japanese and makes it sound childish, making the powerful warriors become somewhat silly every time they are mentioned (at least for English-speaking audiences).
Onomatopoeia can also be applied as lyrics in pre-existing or impromptu songs or other music, like ‘bang’, ‘dada’, ‘lalala’, animal sounds, etc., an effect more like casual humming, usually sung in light mood. Incongruity arises when the music or song contains a serious meaning undermined by the unmatched lyrics. For example, the Japanese word ‘nyan’ (にゃん) is onomatopoeic, imitating the call of a cat, an equivalent to English ‘meow’. A girl humming a song with ‘nyan’ is considered cute. 12 But humming a splendid classical symphony containing great emotional depth with ‘nyan’ creates a crass incongruity. In Nodame Cantabile: Finale OVA ‘Lesson Semiquaver: Mine and Kiyora’s Reunion’ (2010, 07.40–08.40), the music college students in a Vienna tavern are upset because they do not have musicians performing there during the winter season. To cheer the group up, Nodame ‘conducts’ the first movement of Beethoven’s Symphony No. 3 ‘Eroica’ 13 with a knife while singing the tune with an innocent ‘nyan’, along with the first violin Kiyora who also sings ‘nyan’ to accompany her, in front of many other restaurant guests, rendering the impromptu concert absurd and parodic.
In a few unique cases in anime, spoken onomatopoeia read by voiceovers or other extrafictional or fictional people are used as sound effects with or without onscreen written onomatopoeia to replace the corresponding (cartoon) sound effects. The general incongruity is: (1) the written onomatopoeic words in manga are replacement sounds, states or emotions that cannot be easily shown in manga, but which one would expect to be restored into the real sounds and moving images in anime; (2) sounds that it would be possible to make are replaced by their replacements – by the conventionalized onomatopoeia. It is an efficient way to save budget and also create a comic effect, though there is a risk that viewers get tired of this trick if it is overused. A typical example occurs in the anime Zan Sayonara Zetsubō Sensei Bangaichi (2009). Many sound effects in this anime are written onscreen in various exaggerated ways, voiced by a sweet young girl in a consistent intonation, which causes various additional incongruities that cause humour (see Table 1).
Examples of spoken onomatopoeia that create incongruities and humour in different aspects in Zan Sayonara Zetsubō Sensei Bangaichi (2009).
Source: Lei Ye.
Materialized and personified onomatopoeia
In this category, onscreen onomatopoeia are not ‘nondiegetic’ (i.e. sounds that are real only for the audience and cannot be perceived by the fictional characters), but exist in the animated world and participate in the narration, causing multiple types of incongruities.
In Gin Tama (2015, ep 267, 14.10–19.00), the world stops turning and everything is still except for Gintoki, Kagura and Shinpachi. The trio have to search for a spare battery in the Universal Clock used to control time to resume the flow of time and bring back balance to the universe. They move the clock’s hands to travel into the future and start to play with the materialized and personified onomatopoeia in the following scenes:
1. At some point, the battery is held in a giant rocket fist and stops milliseconds away from hitting the forehead of Otae (Kagura’s sister) with a giant red 3D SFX – ‘ゴッス’(gossu: strong impact) hanging in the air accompanied by a short muffled explosion sound.
2. Different from other SFX appearing in the scene, the trio do not just realize it, but also change and play with it. To relieve Otae’s pain, Gintoki cuts off ‘ッ’ and pushes ‘ゴ’ and ‘ス’ together accompanied with a grating sound as if the words have to be moved against resistance in the air (‘ゴス’ (gosu: strike, smash) sounds less painful than ‘ゴッス’). He writes an annotation of the new SFX in the air and we can hear the sound of a marking pen rubbing on glass.
3. Gintoki further modifies and anthropomorphizes the new word into a souled person who even attempts to stop the rocket to protect Otae. Different comical sound effects can be heard when Mr. SFX moves like other characters.
4. The discarded ‘ッ’ dropping to the floor is later modified by Gintoki into one long line and two balls like the shape of a phallus for Kyuubei.
5. The trio move time forward and come to Otae’s wedding. She is getting married to her saviour – Mr. SFX in a black suit. She lost her memory because of the hit except for the kind voice saying ‘gosu’ by Mr. SFX she heard at that moment, which she could never forget.
The farce ends with the fact that Ham-san takes out the proper battery and turns the clock back in time to the point where their troubles began. There are three successive different incongruities during the transformations of the SFX:
Firstly, the trio break the fourth wall to realize the existence of the onscreen onomatopoeia, which should only be accessible to the audience, not to the characters. The fourth wall is a convention of fictional narratives on stage and screen so that there is an invisible, imagined fourth wall that separates actors from the audience. Normally, only the audience can see through this wall while the characters cannot see it or realize that they are fictional. Breaking the fourth wall means that characters momentarily change this ontological difference and relationship, and directly address the audiences/directors/animators to acknowledge the mechanics of their medium: that this is just a show. It is usually used as an ironic device for comedic purposes because, as the characters ignore the barrier by jumping down from the screen and make comments as casual observers, they close the distance to their audiences but also lose their roles’ specific authority and mystique. Thus, self-conscious characters can more or less create a descending incongruity and produce humour during the process of traversing the borderline between the fiction and the world of the audience. Similarly, Batman in The Lego Batman Movie (2017, 1.27.10–1.27.30) once breaks the fourth wall and comments as self-mockery: ‘We are gonna punch these guys so hard, words describing the impact are gonna spontaneously materialize out of thin air!’
Secondly, when the trio start to modify the toy-like SFX, they are no longer normal onomatopoeic words but become onscreen objects to create visual humour. Similar examples can be: onomatopoeia shouted by characters become solid forms flying in accordance with the shouter’s direction and can be ridden on like aircraft in Doraemon (2010, ep 370); and Dirk in Tuca & Bertie literally eats the written onomatopoetic word ‘Boioioioing’ (which would have been an offensive comment on Tuca’s big breast) that is released from his mouth, a play on the expression ‘eating one’s own words’ to indicate someone retracting a statement after having been humiliated in some way.
Finally, Mr. SFX in Gin Tama is anthropomorphized and endowed with a new identity: from a visual pattern or invisible sound effect, he is successfully transformed into a sentient character. But he still keeps several features of the SFX (his head, name and experience), which make him look incongruous and funny.
Conclusion
This article explains the two main reasons causing the difference in the use of onomatopoeia in anime and US animations – the different linguistic cultures and the different traditions of manga and comics, respectively. Spoken and written onomatopoeia as well as sound effects can be either concomitant or interchangeable. They can either provide a sense of unseriousness to underline and frame funny images or create audiovisual humour themselves. As the above examples show, humour is the result of various incongruities that can be caused by the translations of different language, sound metaphors, superfluous or improper spoken onomatopoeia, parodying the classic onomatopoeia, misunderstanding the sources of the sounds, breaking the fourth wall, materialization, personification, etc.
Haverkamp (2012) expands the range of onomatopoeia which ‘also encompasses the imitation of tones with musical means’ (p. 234) to ‘increase the emotionality of the expression’ (p. 235). He considers that techniques in rock music played in a high register to ‘imitate screeching or squealing sounds’ belong to the realm onomatopoeia (p. 235). This also makes sense in anime or animations. In Gin Tama (2006, 1: 10, 21.00–21.20), when Kagura exhausts her superpower to save her best pet that has been kidnapped by robbers, her lethally supersonic yell overturns their car. The incongruity lies in the fact that Kagura’s loud yell and the nondiegetic electric guitar have the same pitch. They are from different levels of narration – diegetic yell and nondiegetic guitar – but their pitch is the same, and they cooperate well. But this does not cause humour because the purpose and the effect of the voice and the music is coincident. Music prolongs her vocal utterance and helps to reinforce and extend its power. What is more, at least in this scene, she is seriously fighting with the robbers and the situation is urgent (although overall, this is a comedy anime, and she often does something absurd, so we may perceive a comic undertone in this moment as well).
It makes sense to explore more diverse means that imitate sounds or voices as a type of onomatopoeia in animations, anime or other media. Their functions and effects are also definitely not restricted to comic effects, which were the focus of this article. There are more possibilities of the mutual transformation and interactions of onomatopoeia and sounds in animation, and they deserve consideration and future study.
Footnotes
Acknowledgements
The author owes a huge debt of gratitude to Guido Heldt who gave her immense support when writing this article. She also sincerely thanks the peer reviewers for their helpful suggestions and support.
Correction (December 2025):
Article updated online to correct to remove one phrase in the section “Adapting onomatopoeia from manga to anime/comics to animation” and deleted the 9th Notes section.
Declaration of conflicting interests
The author declared no potential conflicts of interest with respect to the research, authorship and publication of this article.
Funding
This work was supported by Key Project of Natural Science Research of University in Anhui Provence [grant number 2023AH051263] and Hefei Normal University High-Level Talent Research Funding Program [grant number 2023rcjj37].
