Abstract
In primary education, children frequently combine drawings with alphabetic writing in their compositions, and research has shown that drawing can support early writing development. Building on existing research on multimodality and intermodality in early writing, this study examines how drawing and writing function together in children’s multimodal compositions produced for diverse social purposes. The study analyses multimodal texts written by Swedish learners aged 7–8 years and explores intermodal relations between writing and drawing using a systemic functional, transitivity-based analytical approach. The findings indicate a predominance of word-led and combinatory intermodality, with alphabetic writing generally assuming a dominant role. At the same time, children strategically integrate drawing and writing in ways that vary according to social purpose, task theme, and the specific affordances of each mode. In several cases, drawing enables the expression of meanings that are less readily conveyed through writing alone, underscoring its distinct representational potential. The study contributes to ongoing work on intermodality in early school writing by demonstrating the analytical affordances of transitivity for examining multimodal meaning-making, and it highlights the need for pedagogical practices and teacher metalanguage that support flexible movement across semiotic modes in early literacy education.
Introduction
Young children’s compositions frequently combine drawing and writing in rich and purposeful ways as they navigate academic meaning-making. As research has shown, drawing not only supports early writing development but also enables children to participate in classroom discourse, represent ideas, and express identities (e.g., Christianakis, 2011; Mackenzie, 2011; Moses & Serafini, 2022). In this study, we approach children’s multimodal compositions as utterances in Bakhtin’s (1986) sense, that is, as situated, socially responsive products that enter into dialogic relations with teachers, peers, and institutional contexts. This perspective differs from approaches that, for example, prioritize children’s own accounts of intention, since our focus is on the products that circulate and are acted upon in classroom practice. While children’s intentions are never fully retrievable from the artifact itself, the dialogic framework foregrounds how meaning is constituted in and through the product as it takes on social purpose (e.g., as a narrative, an argument, or a report). Teachers’ assessments, interpretations, and pedagogical responses are, more often than not, directed toward these textual products rather than children’s internal reflections, particularly in contexts where teachers are pressed for time or working under staffing constraints. By analysing the texts as utterances, we underscore how drawing and writing function as communicative acts with affordances and interrelations that shape how children are recognized as meaning-makers in school.
Moreover, although the significance of multimodality is widely acknowledged, we contend that a crucial dimension remains undertheorized: intermodality, understood as the relational dynamics among modes within a single composition, which constitutes a fundamental aspect of multimodality. Existing scholarship has made strong cases for attending to multimodality broadly (Kress, 2003; Unsworth, 2008) but has offered less guidance for systematically describing or analysing how modes like writing and drawing co-function to produce meaning, particularly in analogue (non-digital) texts produced by young learners. This lack of shared terminology and analytic precision creates both research and pedagogical challenges. For example, educators often lack the metalanguage needed to assess how drawings and writing operate together, which in turn may reinforce the marginalization of visual modes in classroom practice (Macken-Horarik, 2016; Shanahan, 2013). Young learners in the initial stages of formal education (corresponding to current Swedish school years F–3, ages 6–9 years) occupy a transitional developmental phase wherein both drawing and writing function as legitimate semiotic resources for meaning-making. However, within prevailing educational practices, drawing tends to be afforded limited recognition as a valued communicative mode compared to writing (e.g., Stagg Peterson & Friedrich, 2022). Such an asymmetry risks obscuring the fact that drawing has, throughout children’s earlier years, functioned as a primary means of communication and representation (cf. MacKenzie, 2024).
In the present study, drawing is foregrounded as a valuable communicative mode and examined in relation to print writing in order to understand how the two modes jointly construct meaning. Drawing on a dataset of 148 multimodal texts produced by Swedish students aged 7–8 years, we analyse how writing and drawing interact across five distinct writing assignments, each tied to different social purposes. The study is informed by a framework for categorizing intermodal relationships (Björk et al., 2025) that provides a means for analysing how children mobilize drawing and writing not as parallel or discrete modes, but as interconnected resources for integrated meaning-making. By applying this framework, we aim to (1) advance a more systematic approach to analysing intermodality in early writing, and (2) support the development of pedagogical strategies that reflect the complex semiotic work involved in young children’s multimodal compositions for different social purposes. These aims are specified in the following research questions:
Research Question 1: What characterises intermodal relations in children’s compositions made for different social purposes?
Research Question 2: How can possible variations of intermodality, depending on writing assignment, be understood?
The subsequent sections outline prior research and theoretical underpinnings of the study, focusing on children's multimodal compositions and the concept of intermodality in educational contexts. Next, we detail the study design and analytical procedures. Finally, we present the findings of the study and conclude with a discussion of its implications.
A Social Semiotic, Multimodal, and Dialogic Framework for Early Writing
This study adopts a social semiotic (Halliday, 1978; Halliday & Matthiessen, 2014), dialogic (Bakhtin, 1986), and multimodal (Kress, 2003, 2010) framework for understanding writing. These three strands are conceptually distinct yet complementary. Social semiotics provides the overarching orientation: meaning-making is understood as a socially situated, motivated, and value-saturated process through which individuals design utterances using the semiotic resources available to them (Halliday, 1978; Kress, 2010). From this perspective, all forms of communication, such as verbal, visual, or bodily, are modes of social action shaped by cultural, institutional, and ideological conditions. Within this broader semiotic view, the dialogic strand (Bakhtin, 1986) is employed to account for the fundamentally relational and responsive nature of children’s multimodal meaning-making in school. Further, research on multimodality in early education positions multimodal resources as central to children’s contextually embedded meaning-making. This work, among other aspects, illustrates how multimodal practices operate as mediational means that bridge school-based learning with children’s everyday experiences, encompassing the dynamic negotiation of peer cultures, institutional norms, and wider cultural repertoires (cf. Anderson, 2013; Bezemer & Kress, 2008; Compton-Lilly, 2014).
The following sections present an integrated overview of the theoretical and empirical foundations of the study. Drawing on social semiotics, multimodal theory and dialogic perspectives, we review and synthesize research on drawing and writing in primary school, with particular attention to how children’s meaning-making unfolds across modes. Within this body of scholarship, we situate educational research on intermodality, focusing on how relations between modes have been conceptualized and empirically examined in early years contexts. These perspectives jointly inform the analytic focus of the study, and the section concludes by outlining the analytical considerations and procedures guiding the present investigation.
Multimodality in Early School Writing and Drawing
The multimodal strand extends beyond the linguistic mode to include a fuller range of semiotic modes through which meaning is designed and interpreted (Bezemer & Kress, 2008; Kress, 1997). Multimodality refers to the understanding that meaning is made through the coordinated use of multiple semiotic modes, such as writing, image, layout, gesture, and colour, each offering distinct affordances for representation and communication (Kress, 1997; Kress & van Leeuwen, 2001). Kress (1997) emphasises that children’s early inscriptions are purposeful acts of meaning-making, in which choices of mode reflect both available resources and communicative intent. Similarly, Kress and van Leeuwen (2001) conceptualise multimodal texts as designed ensembles, where meaning emerges in the orchestration of modes rather than in any single mode in isolation. In early school writing, this implies that drawing and writing function together as complementary semiotic resources through which children explore, structure, and communicate experience, rendering multimodality not an exception to writing development but a foundational condition of early textual practice. Against this backdrop, the present study focuses specifically on the interplay between written text and drawing, understood here as a central visual mode, in order to examine how these two modes operate together in children’s early school compositions.
Research that examines how early drawing and writing function as complementary resources is particularly relevant to this study. Mackenzie (2011, 2014, 2022) demonstrates that integrating drawing into writing instruction fosters more complex texts and strengthens writer identities in the early years. Similarly, Moses and Serafini (2022) show how one first-grade student distributes meaning across modes in nonfiction compositions, highlighting the value of moving beyond strictly verbal texts to support communication and authorial identity (p. 351). Stagg Peterson and Friedrich (2022) indicate that children aged 4–6 years demonstrate substantial knowledge of print literacy through multiple modes, including drawing. Using video recordings of 64 children, field notes, and their written and illustrated texts, Stagg Peterson and Friedrich (2022) examine oral strategies for discussing multimodal compositions. The study demonstrate that children’s verbal reflections provide crucial insight into how writing unfolds across communicative resources, underscoring the need to attend both to texts and to children’s oral reflections (cf. Friedrich et al., 2021). The study further identifies an imbalance in adult feedback, with substantially greater attention directed to print than to drawing, a disparity reflecting the broader privilege of written over visual modes in classroom assessment. Friedrich et al. (2021) also focus on the relationship between drawing and writing among 3- to 6-year-olds. Their analysis identified six multimodal patterns and revealed significant correlations between drawing and writing performance (p. 12). The study concludes that drawing can serve as a developmental foundation for expression within the conventionalized sign system of writing, a view also supported by Mackenzie and Veresov (2013, p. 69).
Dialogism and Children’s Texts as Utterances
Bakhtin’s dialogic perspective (1986) is mobilized to theorize the inherently relational and responsive nature of children’s multimodal meaning-making in school contexts. From a dialogic standpoint, texts are understood as utterances: concrete, situated communicative acts that gain meaning through their responsiveness to prior discourse and their orientation toward anticipated responses. Meaning is thus not located solely in formal textual features or in the individual intentions of an author but emerges through participation in ongoing chains of communication shaped by social purposes, genres, and institutional norms. In educational contexts, dialogism has been productively taken up to examine how students’ texts respond to and are shaped by classroom interactions, pedagogical framing, and circulating discourses (e.g., Nelson et al., 2008; Ranker, 2008). While these studies are not primarily concerned with Bakhtin’s theory per se, they are particularly relevant for the present study because they foreground how multimodal compositions function as socially situated responses within classroom practices. Nelson et al. (2008), for instance, demonstrate how students’ multimodal self-representations are shaped by expectations of audiences and institutional values, while Ranker (2008) shows how children’s multimodal narratives draw dialogically on peer cultures, popular media, and school-sanctioned literacy practices. These studies, therefore, illustrate dialogism at work in multimodal classroom texts, aligning closely with the analytical concerns of the present study. Adopting a dialogic approach is especially important given our focus on children’s use of writing and drawing rather than on elicited accounts of authorial intention. Classroom literacy practices are largely organised around the circulation, assessment, and interpretation of textual products, and it is these products that mediate pedagogical decisions and (social) recognition. Viewing children’s compositions as utterances allows the analysis to foreground how drawings and writings respond to teacher prompts, genre expectations, and prior texts, while simultaneously projecting anticipated readers and evaluative stances. The boundaries of an utterance are thus defined not only by formal or modal features, but by the social purposes and value positions the text comes to enact.
From a Bakhtinian perspective, every utterance is oriented toward other voices, genres, and discourses, many of which remain outside the conscious awareness of the speaker or writer. A child’s utterance may therefore draw dialogically on narrative conventions from picture books, popular culture, or scientific discourse, while also responding to explicit instructional demands such as task framing, assessment criteria, or expectations regarding handwriting and illustration. These dialogic relations shape multimodal compositions in ways that exceed individual intention, situating even seemingly personal texts within broader communicative practices. In the present study, dialogism provides a crucial complement to social semiotic and SFL-based approaches by foregrounding responsiveness, positioning, and social purpose in children’s multimodal utterances. It frames intermodality not merely as a structural relationship between modes, but as a socially and historically situated interaction of semiotic resources through which children participate in classroom discourse. This orientation underpins our analytic focus on how drawing and writing operate together as dialogically responsive acts of meaning-making in early school contexts.
From Multimodality to Intermodality
From a social semiotic perspective, each mode offers distinct affordances for representing experience and enacting social relations, yet these modes rarely operate in isolation; multimodality refers to the presence and use of multiple semiotic modes within a text or communicative event (e.g., image, writing, and gesture), highlighting that meaning is made through more than one mode (Kress, 2010). Barthes (1977) laid important groundwork for later multimodal theory by his analysis of image–text relations, demonstrating how different semiotic systems mutually shape and activate one another in an interplay of meaning-making. While not writing from a multimodal perspective, Kristeva’s (1986) notion of intertextuality offers a closely related account of meaning as relational and transformative, conceptualising texts as a mosaic of texts that continuously draw on and transform each other. We argue that this understanding of meaning as emergent through semiotic relations resonates productively with subsequent work on multimodal composition. Hull and Nelson (2005), for example, extend this relational logic to multimodal composition, showing that their expressive power lies precisely in the interactions among co-present modes. Meaning in such compositions is thus not additive, but emergent, arising in the interplay—or orchestration (Kress, 2010, p. 161; cf. Bateman & Wildfeurer, 2014; O’Halloran, 2008; Unsworth, 2008)—of resources such as word, image, and sound: a mosaic of interrelated modes. Taken together, these perspectives anticipate what later scholarship has conceptualised as intermodality (e.g., Siefkes, 2015, 2018). Siefkes defines an intermodal relation as “present when one mode has a definable influence on the expression, semantic, and/or stylistic properties of another mode in a specific text” (Siefkes, 2015, p. 115). Our understanding of intermodality reflects this definition. It emphasizes the relational and meaning-making dynamics among modes, focusing on how different modes interact, complement, elaborate upon, or transform one another in the production of meaning. In this article, our analytic focus is therefore intermodal rather than merely multimodal. While the analysed students’ texts are multimodal in that they include both drawing and alphabetic writing, we examine how these modes function together as a single communicative unit or utterance; drawing and writing are not treated as parallel or additive resources, but as interdependent elements within the same utterance, where meaning emerges in and through their interrelation.
Within social semiotic and Systemic Functional Linguistics (SFL)-oriented scholarship, intermodality has been conceptualized as the semantic and functional relations that arise between co-present semiotic modes. Siefkes (2015) advances a broad view of intermodality as meaning-based relations across modes that may operate beyond ideational meaning, whereas Martinec’s account (Martinec & Salway, 2005), structured through logico-semantic and status relations, has been characterized as largely confined to the ideational metafunction. Building on this work, Unsworth (2008) foregrounds the pedagogical relevance of intermodality in educational contexts, showing how relations between image and writing shape meaning-making and learning in school genres. Taken together, these perspectives position intermodality as the dynamic organization of meaning across semiotic resources within specific social purposes. While the present study analytically foregrounds the ideational metafunction through transitivity analysis, this focus is not intended as a representation of ideational meaning alone. Rather, by examining intermodal realizations of processes across drawing and writing, the analysis demonstrates how intermodality also mobilizes structuring, relational, and attitudinal resources, highlighting the analytical affordances of transitivity for multimodal analysis of children’s compositions. Anchoring the analysis in SFL is valuable for interpreting children’s multimodal compositions, as it provides analytic tools to examine how meaning is distributed and realized across modes (Mills & Unsworth, 2017; Painter et al., 2014; Reid & Moses, 2022).
Examinations of intermodality in educational contexts have encompassed both analogue and digital settings. Reid and Moses (2022) address intermodality in 4th grade through the metaphor of “orchestration” to conceptualize the writer's agency and the complex layering of decisions involved in producing a multimodal text (p. 726). These researchers argue that the constituent elements of a composition do not operate in isolation but are “interwoven” with meaning, forming integrated ensembles that constitute more than the sum of their parts. On a less metaphorical level, Flewitt et al. (2009) emphasize more broadly the necessity of examining how different modes interact within a single text, advocating for analytic approaches that attend to the relational dynamics of such multimodal communication (see also Hull & Nelson, 2005). Research on multimodal composing also highlights how intermodality shapes meaning-making and classroom practice. Shanahan’s (2013) study of 10–11-year-olds’ digital compositions shows that visual–verbal relations position students within classroom discourse, but that limited teacher knowledge of SFL metafunctions can restrict learners’ semiotic choices, suggesting that merely acknowledging multiple modes is insufficient without a critical understanding of how they interact. Mills and colleagues extend this line of work: Mills et al. (2020) demonstrate how 9–11-year-olds communicate attitudinal meanings multimodally through comics, while Mills and Exley (2014) identify tensions between mandated testing and the assessment of multimodal text production, arguing for reconceptualising writing epistemologies to accommodate digital practices. Although such studies show how linguistic and visual modes jointly realise meaning, intermodality between drawing and writing in the early years need further scrutiny. Recent work answers this call: Pacheco-Costa and Guzmán-Simón (2020) use SFL to show how 7–8-year-olds express ideational, interpersonal, and textual meanings across drawing and writing, advocating assessment models that recognise children’s lived, multimodal practices. A related study (Björk et al., 2025), with learners of the same age, proposes an intermodal transitivity framework, also used in the present study, to analyse how semiotic resources are coordinated across modes in early composing.
In sum, research consistently shows that the integration of multiple modes enriches children’s compositions and expands their meaning-making processes (e.g., Friedrich et al., 2021; Mackenzie, 2011; Moses & Serafini, 2022; Stagg Peterson & Friedrich, 2022). Research has also begun to address intermodality between drawing and writing in early school compositions through more systematic analytical approaches. There is, however, a need for more fine-grained analyses of how modes interact intermodally in analogue classroom contexts, particularly regarding how meaning is distributed across drawing and writing. Emerging frameworks (e.g., Björk et al., 2025; Pacheco-Costa & Guzmán-Simón, 2020) underscore the value of systematic pedagogical and analytical approaches to multimodal composition. Although drawing has been examined as a communicative mode in early years education, existing research has primarily focused on particular linguistic and educational contexts, most notably English-language settings. Björk et al. (2025) propose a framework for categorising intermodal relations in children’s descriptive texts, demonstrating how drawing and writing distribute meaning-making responsibilities in diverse ways in Grade 1 classrooms in Sweden. While this work establishes a metalanguage for identifying patterns of intermodality, it also opens further analytical questions regarding how meaning is construed within these intermodal relations, and how such relations vary across social purposes and classroom tasks. The present study builds directly on this work by extending the analytic focus from categorical relations to the intermodal realisation of ideational processes, using transitivity as a means of examining how children express experience across writing and drawing in multiple genres and instructional contexts.
Analysing Intermodality: A Systemic Functional Linguistic Approach
This study draws on Systemic Functional Linguistics (SFL), which complements social semiotic and dialogic theories by offering a systematic account of how language and other semiotic resources function in context (Anderson, 2013; Halliday & Matthiessen, 2014). SFL-based approaches to analysis of multimodal discourse (e.g., Martinec & Salway, 2005; Unsworth, 2006) provide frameworks for examining intermodal relations – how linguistic and visual modes interact to realize meaning within specific social contexts. Anchoring the analysis in SFL is valuable for interpreting children’s multimodal compositions, as it offers analytic tools to explore how meaning is distributed, realized, and negotiated across semiotic modes (Mills & Unsworth, 2017; Painter et al., 2014; Reid & Moses, 2022). While SFL foregrounds systemic structures and metafunctions rather than the agency of the sign-maker, combining SFL with dialogic and social semiotic perspectives ensures that children’s multimodal utterances are interpreted as socially responsive and situated acts of design.
While the preceding discussion established the theoretical foundations of a social semiotic, dialogic, and multimodal understanding of writing, the following section outlines how these principles are operationalized in analysis. Specifically, this section details how SFL’s functional grammar is used to interpret children’s texts in terms of social purpose, communicative intent, and the representation of experience across linguistic and visual resources. The study draws on research that applies SFL to multimodal text analysis, including Martinec and Salway (2005) and Unsworth (2006), who investigate intermodal relations between image and language. As Shanahan (2013) and Anderson (2013) note, such frameworks foreground systemic structures and metafunctions rather than the agency of the sign-maker, marking a distinction from other social semiotic traditions. Nonetheless, anchoring the analysis in SFL is valuable for interpreting children’s multimodal compositions, as it provides analytical tools to examine how meaning is distributed and realized across modes (Mills & Unsworth, 2017; Painter et al., 2014; Reid & Moses, 2022). From an SFL perspective, categorizing children’s texts based on the social purpose of the prompts they respond to enables a nuanced understanding of meaning-making patterns. Rather than applying arbitrary genre labels, this approach foregrounds communicative intent – such as narrating, explaining, or arguing – and reveals how children mobilize multimodal resources to meet such specific rhetorical demands (Honig, 2010). This categorization supports SFL’s emphasis on the dynamic relationship between text, context, and audience, and also facilitates systematic comparisons across tasks, shedding light on how children navigate multimodal design within school-sanctioned genres (Anderson, 2013; Reid & Moses, 2022). The understanding of genre and purpose is particularly important in light of research showing that young writers draw on both curricular expectations and their own social worlds when composing texts (Björk & Iyer, 2023; Compton-Lilly, 2014; Dyson, 2013; Mills et al., 2020).
Transitivity analysis provides a central analytical lens for examining how experience (ideational meaning) is represented in children’s multimodal utterances. In SFL, transitivity refers to the system through which clauses construe processes (what is happening), participants (who or what is involved), and circumstances (the conditions under which the process unfolds) (Halliday & Matthiessen, 2014). In other words, it captures how language encodes ways of perceiving and organizing the world. When extended to multimodal texts, transitivity analysis allows for the examination of how meanings of action, perception, and emotion are distributed and realized across linguistic and visual resources. For example, an action represented verbally in a caption (“I run fast”) may be simultaneously elaborated or reinterpreted through a drawing that foregrounds movement, spatial direction, or affective stance. Such instances highlight that meaning does not reside within any single mode but emerges in the intermodal orchestration of semiotic choices, a process that is both socially and dialogically situated.
This dialogic orientation is essential: each multimodal choice responds to and anticipates other communicative acts within the classroom and beyond. The child’s depiction of an event or process thus becomes an utterance within a larger chain of discourse, shaped by pedagogical expectations, genre conventions, and personal experience. Through transitivity analysis, these relations of agency, evaluation, and positioning can be traced across modes, revealing how children express experience and negotiate meaning through their multimodal designs. Hence, transitivity serves as a key analytical bridge within the study’s social semiotic–dialogic–multimodal framework. It links the functional structures of language and image with broader social processes of communication, allowing the analysis to account for both the systemic organization of semiotic resources and the responsive, value-laden ways in which young writers mobilize them.
Analytical procedure in three steps
The analytical procedure is performed in three steps: 1) transitivity analysis 2) analysis of intermodal relations and 3) analysis of intermodalities. These three steps are presented in more detail below.
Step 1. Transitivity analysis
The first step of the analysis draws specifically on the system of transitivity as outlined in SFL. Transitivity analysis focuses on how processes (actions, events, or states) and participants (people, objects, or entities) are represented in a text. These elements are not restricted to written language but are fundamental to meaning making in various semiotic modes, including visual representations like drawings. Halliday’s (Halliday & Matthiessen, 2014) system of transitivity, a key component of SFL, provides a framework for analysing how language, including early writing, represents experiences and events. At its core, transitivity examines the processes expressed in clauses, the participants involved, and the circumstances surrounding them. Halliday identifies three primary process types: material processes (MaP), which denote actions and events (e.g., “She built a sandcastle”); mental processes (MeP), which relate to thoughts, perceptions, and emotions (e.g., “He felt sad”); and relational processes (RP), which establish relationships between entities (e.g., “The sky is blue”). Additionally, he includes behavioural processes (actions tied to mental states), verbal processes (VP; acts of saying), and existential processes (EP; expressing existence). Drawing on Karlsson & Holmberg’s (2006) account of functional grammar in Swedish, we focus our analysis on how material, relational, verbal, and mental processes are realized in the children’s texts. Each process type is accompanied by specific participant roles, for example “actor” and “goal” in material processes or “sensor” and “phenomenon” in mental processes. Transitivity analysis thus allows researchers to explore how language choices reflect ideational meaning, revealing patterns of agency, focus, and perspective in texts.
In drawings, processes, whether they involve actions, verbal dialogue, or mental states, can still be expressed through visual cues. For example, an action process can be depicted by showing a figure performing a visible action, such as running or holding an object, which parallels the way such actions would be described in text. Similarly, mental or verbal processes can be indicated by facial expressions or dialogue bubbles, which reflect inner thoughts or spoken words. Thus, even though the semiotic resources differ between writing and drawing, the core elements of transitivity–the processes, participants, and circumstances–are applicable to both modes. This makes transitivity a flexible analytical tool for examining how children construct meaning in their drawings and how these visual processes can correspond to, or complement, the written processes in their multimodal compositions. Figure 1 displays how each visual process type has been analysed, exemplified by a drawing from the current dataset.

Transitivity analysis in drawings: Empirical examples of different process types.
Step 2. Analysis of intermodal relations
In the second step of the analysis, the processes identified in the compositions have been further elaborated according to the framework of intermodal relations developed in Björk et al. (2025). In this study, the researchers expand on the intermodal relations in young children's compositions. The framework is built on analyses of 100 multimodal compositions and proposes four types of intermodal relations: Concurrency, Complementarity, Divergence, and Inactivity, each representing different ways that drawing and writing interact to convey meaning. These main categories further include subcategories, represented in Table 1 below.
Framework for the Categorisation of Intermodal Relations Developed by Björk et al. (2025).
Step 3. Intermodal classification
The categories from the second step, have been classified into three overarching categories: word-led, image-led, and combinatory intermodality. These categories are defined as follows: (1) Word-led intermodality, where alphabetic writing carries the predominant meaning-making responsibility (corresponding to categories B2 and C2); (2) Image-led intermodality, where the drawing assumes a greater meaning-making load than alphabetic writing (corresponding to categories B1 and C1); and (3) Combinatory intermodality, where the responsibility for meaning-making is distributed and varies between modes in diverse ways (corresponding to categories A, B3, and B4). The category of D (inactivity) has been excluded from this classification (cf. Björk et al., 2025).
To ensure reliability and minimize subjective influence in the analytical procedure, the analysis was conducted independently by the two researchers. In Step 1, each researcher identified transitivity processes within the compositions. In Step 2, these processes were categorized, with both steps carried out separately to prevent bias. Following this initial phase, the researchers compared their results to assess the consistency of categorization. To enhance transparency and reliability, illustrative examples of transitivity in drawings are presented in Figure 1. In Step 3, the intermodal classification was conducted through interpretative dialogue between the researchers. Additionally, the data and analysis were presented in research seminars, allowing other researchers to review and critically engage with the findings. Alternative categorizations were discussed both in these seminars and in researcher dialogues, reinforcing the rigor and credibility of the analysis.
The present study
This study is part of the FEAST-project: an intervention study in grades 1-2, age 7-8, in primary school in Sweden 2020-2023. The project focused on a teaching intervention that addressed functional writing (af Geijerstam et al., 2024). In 15 intervention schools teachers were given 16 teaching assignments with associated assignments for students. Texts written and created by children in the intervention schools comprise the primary data set of this study. The intervention (and teaching) is not foregrounded per se in the study. The assignments in focus are described in detail below. The data has been collected in accordance with the research ethical guidelines in Sweden (Swedish Research Council, 2024). Consent has been given by both children and caregivers to participate in the research project. The FEAST-project has gained ethical approval for the project from the Swedish Ethical Review Authority (Dnr 2020–00738 and Dnr 2020-02990).
The selection of compositions in the study is considered representative of the broader population from which the data were drawn, encompassing children for whom Swedish may function as either a first or an additional language. Sweden is characterized by a multilingual landscape, with a diverse range of languages represented in classrooms. It is however important to note that the intervention was conducted within a Swedish-language subject context. Further, the multilingual dimension does not constitute an analytical focus of the present study. Accordingly, no further demographic information about the children is included, as our interest, in line with previous reasoning, lies in the multimodal properties of the texts themselves rather than in the linguistic or biographical backgrounds of their authors.
The assignments
Five assignments were selected for inclusion in the study (see Table 2). These assignments were chosen from the broader corpus of texts based on the extent to which the children’s responses incorporated a combination of drawings and alphabetic writing. The selection process prioritized variation, resulting in five assignments that were notably distinct in nature. This diversity was a key criterion, ensuring alignment with the research questions and the study's overarching analytical framework.
Included Assignments, Main Function of the Assignments, The Time When Written and Number of Text Included in this Study.
A total number of 148 children’s compositions are included in the analysis. Regarding assignment A1, the analysis is grounded on the work presented in Björk et al. (2025). In this assignment specifically, 100 texts were randomly selected out of a corpus of 271 multimodal texts in this category (i.e., descriptions). In assignments A2-A5 all of the multimodal texts that were collected in the project were included in the study. The small number of assignments in A2-A5 reflects the limited availability of multimodal data within these text samples, mirroring a decrease in multimodal texts in relation to time.
A1: Descriptive (inform)
Assignment A1 was conducted in the first semester of year 1. Teachers read a letter to the class asking students about their favourite recess activities, optionally using pictures for inspiration. After discussing recess activities, teachers explored the concept of description, the role of a scientist, and letter writing. Students chose their topics and wrote using pen/digital tools. The intermodal relations in these texts were analysed and communicated in Björk et al. (2025).
A2: Narrative (narrate)
In the second semester of year 1, the students continued a story the class had read, focusing on narrative cohesion and creative endings. The activity began with reading and discussing stories to understand plot structure, especially resolution and conclusion, using the "dramaturgical curve" as visual support. Students wrote their endings individually or in pairs and illustrated their stories.
A3: Argumentative (persuade)
During the first semester of year 2, students wrote job applications to a fictional employer (Santa Claus), focusing on persuasive writing and building a connection with the reader. The activity began with discussions about the employer, followed by reading a fictional job invitation. Students included personal details, qualifications, and reasons for being hired. Letters were "sent" to Santa, with a possible response that all were hired.
A4: Narrative (narrate)
In the second semester of year 2, students wrote and illustrated ghost stories, emphasizing suspense, tension, and cliffhanger endings. The teacher introduced the genre through ghost stories, films, or storytelling, discussing effective narrative elements. Students then brainstormed and wrote their stories, focusing on spooky events and character reactions, and shared their stories through readings or class book compilations.
A5: Descriptive (inform)
In the second semester of year 2, students wrote informational texts about the life cycle of the frog, integrating relevant images. After studying the animal through films and discussions, and possibly observing it in nature, students used a shared outline to write structured, descriptive texts. Their work was presented in booklets displayed at school or shared with others. The assignment concluded by comparing various life cycles.
Findings
The following section presents the distribution of intermodal relationships for each assignment, followed by an analysis of a specific example from each assignment. A summary of the results is provided at the end of the section. In the majority of examples analysed, the students’ page layout follows a two-part structure: the upper portion is intentionally left blank to accommodate an illustration, while the lower portion is ruled to provide space for alphabetic writing. This configuration is consistently observed in examples A2 through A5. However, variations occur across the sample. In contrast to this pattern, the descriptive texts generated in A1 are presented on fully ruled pages. Typically, as illustrated in example AI.a below, the alphabetic text occupies the initial page(s), followed by a subsequent page containing the drawing, which is also placed on ruled lines rather than on a blank section. Each example below is presented with an annotated representation that specifies the process types identified in the analysis, anchored in the central process core of each clause for writing (marked with squares) and in the visually discernible central processes within the drawings (marked with circles). Inferred processes are marked by dotted lines.
A1: Descriptive (inform)
The analysis shows that the A1 descriptive texts are primarily word-driven, as the most salient intermodal relation are word-enhanced complementarity (B2, 29%), and word-driven divergence (C2, 25%) (cf. Björk et al., 2025). Composition A1.a has been categorised as an example of word-enhanced complementarity (B2).
In Composition A1.a in Figure 2, the drawing depicts a material process (playing soccer) and a mental process (mood, two smiling faces). Both the material action of playing soccer and the positive mood is explicitly represented in the alphabetic writing through the sentence “I like to ” and the final relational process: “because football is fun”. While the processes represented in the drawing are present in the writing, the alphabetic text introduces additional (relational) processes (“they are really kind friends”), resulting in an intermodal relationship characterized by word-enhanced complementarity.

Composition A1.a.
A1 is predominantly classified as word-led, with a discernible proportion of combinatory intermodality, indicating that children chose to express their preferred recess activities in diverse ways. The drawings often illustrate one or several activities that are also described in the text. The transitive processes expressed through both drawing and writing often depict different aspects of the same activity. While these modes may complement one another, they frequently convey distinct types of content, even when representing the same process.
A2: Narrative (tell a story)
The analysis of the A2 narrative texts show that the most salient intermodal relation is distributive complementarity (B4, 47%). The intermodal relations are further distributed between word-driven divergence (C2, 13%), integrative complementarity (B3, 13%) and word-enhanced complementarity (B2, 27%). Composition A2.a has been categorized as an example of distributive complementarity (B4).
In Composition A2.a in Figure 3, the drawing conveys one verbal process (a speech bubble) and one mental process (a mood indicated by the smiling face). These processes are not reflected in the alphabetic writing, even though the writing reflects the same process types. The writing instead emphasizes other circumstances of the story. The written text highlights the relational process “she was white in the face,” which is absent from the drawing. Instead, the mother is depicted as dark or shaded in the doorway. Following this, the written text presents four verbal processes: the first and third processes (“asked” and “said”) are attributed to “mom” as the sayer, while the second and fourth processes are responses from the focalized “I,” suggesting inferred (inf.) process equivalent to “said” and “said.”.

Composition A2.a.
The primary overall intermodal relation expressed in A2 is combinatory, which may be expected since the assignment tasked the children with finishing a specific story, rendering the texts narratively “incoherent” and dependent on the foregoing storyline. The drawings, therefore, may be understood as an intermodal continuation or complement to the preceding story provided to the children, displaying how the children use intermodality as a narrative device and explaining the prevalence of distributive complementarity in A2.
A3: Argumentative (persuade)
The analysis of A3 argumentative texts show that the most frequent intermodal relationship in the compositions is word-driven divergence (C2, 44%). The category of distributive complementarity (B4), 44%) is the second most common intermodal relation in the assignment. The category of B3 represents 6% of the sample. Composition A3.a has been categorised as an example of Word-driven Divergence (C2).
In Composition A3.a in Figure 4, the alphabetic text predominantly features relational and mental processes, such as “I am eight” and “I want to,” while the drawing does not depict any discernible processes. More broadly, the texts in A3 align with A3.a in that most exhibit word-driven divergence: the alphabetic text conveys processes describing attributes and values associated with the “I,” whereas the drawings primarily depict various Christmas-related symbols (e.g., Christmas trees, gifts, ornamentation) without representing processes. However, some texts within A3 feature processes expressed both alphabetically and visually, demonstrating intermodal relationships of distributive complementarity (B4). In these instances, most of the drawings include verbal and material processes attributed to figures such as Santa or an elf, as illustrated in composition A3.b below. A3.b. is categorised as distributive complementarity as the written text expresses mental and relational processes whereas the drawing expresses verbal (“Ho Ho”) and material process (steering a sleigh), as seen in Figure 5.

Composition A3.a.

Composition A3.b.
Overall, the intermodal relations in the compositions A3 are classified as combinatory. These compositions are consistent with A3.a in that the majority display word-driven divergence through an alphabetical text construing a number of processes describing attributes and values associated to an “I”, while the drawings depict various symbols associated with Christmas (e.g., Christmas tree, gifts, ornamentations) albeit without processes. There are, however, texts in A3 which consist of processes expressed both alphabetically and visually, and where the intermodal relations show distributive complementarity. In these texts, most of the drawings constitute verbal or material process in relation to Santa or an elf, which is shown in example A3.b.
A4: Narrative (tell a story)
The most frequent intermodal relationships in the A4 narrative texts are integrative complementarity (B3, 40%) and distributive complementarity (B4, 40%). Word-led complementarity (B2) is represented in the sample with 20%. Composition A4.a has been categorised as an example of distributive complementarity (B4).
In Composition A4.a in Figure 6, the alphabetic text conveys a variety of processes and process types, primarily material and mental. In contrast, the drawing represents a mental process, evoking a mood of fear or fright. The drawing features a frightening depiction of a clown with sharp teeth, blood, a sinister grin, and intricate makeup, appearing to hover ominously above the alphabetic text. This pattern is consistent across most texts in A4, where drawings, rather than depicting explicit material processes from the story, emphasize the narrative’s horrific elements through symbolic representations, such as the depiction of a menacing clown.

Composition A4.a.
Overall, the intermodal relations in the sample are classified as complementary or distributive. The analysed sample is characterized by compositions that convey various processes through alphabetic writing, narrating a scary story. This is either illustrated with drawings that evoke scariness, rather than directly mirroring the processes described in the written story, or by drawings where a single material process, such as a specific event from the text is focalized.
A5: Descriptive (inform)
The most frequent intermodal relationship in the compositions is word-enhanced complementarity (B2, 71%). The intermodal relations are further equally distributed between word-driven divergence (C2, 14%) and integrative complementarity (B3, 14%). Composition A5.a has been categorised as an example of word-enhanced complementarity (B2).
In Composition A5.a in Figure 7, the alphabetic text conveys relational processes, such as “firstly, they are frog eggs,” and material processes, such as “they eat plants,” which together describe the developmental stages of frogs. Similarly, the drawing represents these stages but does so visually, using arrows and numbers to connect four illustrations of central stages of the frog’s life cycle (material processes). This visual representation can be interpreted as conveying an overarching material process, such as “the frog develops.” The composition is therefore classified as expressing word-enhanced complementarity (B2), as the alphabetic writing provides a more detailed representation of the processes involved in the stages of frog development compared to those conveyed in the drawing.

Composition A5.a.
Overall, the texts in A5 are classified as primarily word-led. The analysed sample is further distinctly characterized by a scientific mode of representation, or, in Westlund’s (2018) terms, portraying a scientific content in theoretical terms. In nearly all the analysed compositions, connections, such as those representing developmental processes, are depicted using arrows. Although this dataset is limited in size, it offers important insights into how intermodal relations can be expressed within this specific text type.
Summary of Results
Figure 8 shows the distribution of intermodal relations across the assignment groups chronologically, from A1 in spring of Year 1 to A5 in spring of Year 2. The analysis reveals that differing social purposes provide the foundation for varying intermodal relations.

Bar chart showing the distribution of intermodal relations in Samples A1–A5. The x axis shows percentage of intermodal relation in each of the assignments.
Further, the diagram in Figure 9 indicates that all the children's compositions in this sample are predominantly categorized into the intermodal relations of word-enhanced complementarity (B2), integrative complementarity (B3), distributive complementarity (B4), and word-driven divergence (C2).

Bar chart showing the distribution of intermodal relations and classification of all compositions. The y axis represents the number of compositions.
While A1 and A5 compositions exhibit word-led intermodality, A2, A3, and A4 are primarily characterized by combinatory intermodality. These assignments variously enable or invite content expression through multiple modes, including writing and drawing, pertaining to both the social purpose and the represented content. These findings are further discussed in the next section.
Discussion
The findings of this study underscore the significance and variability of intermodality in young children’s school compositions (cf. Hull & Nelson, 2005; Pacheco-Costa & Guzmán-Simón, 2020; Reid & Moses, 2022). While alphabetic writing is often privileged in school contexts, our analysis shows that children’s drawings play a central role in expanding the semiotic range through which ideas are expressed and negotiated. When children are enabled to draw as part of their writing, they distribute meaning across modes in ways that support conceptual exploration, identity work, and communicative precision (cf. Stagg Peterson & Friedrich, 2022). These findings reinforce the pedagogical importance of inviting and valuing multiple modes of representation, as well as supporting teachers’ metalinguistic awareness of how modes operate together in young children’s meaning-making.
In addressing RQ1, we identified three overarching intermodal patterns: word-led, image-led, and combinatory intermodality. These patterns describe how responsibility for meaning-making shifts between writing and drawing. Although word-led intermodality dominates the corpus overall, combinatory intermodality is also strongly represented, revealing that children routinely coordinate drawing and writing in complementary ways. This suggests that instructional design should support flexible movement across modes, allowing children to make strategic semiotic choices.
Responding to RQ2, we found systematic variation across assignments, indicating that the social purpose and thematic content of a task strongly shape the forms of intermodality children employ. While the two descriptive tasks (A1 and A5) differ in genre and complexity, both remained largely aligned with alphabetic writing. By contrast, the narrative and argumentative tasks (A2, A3, A4) invited richer and more diverse intermodal relations, especially distributive complementarity (B4), where visual and verbal modes carry different, but equally important, dimensions of meaning. These patterns likely reflect both the affordances of particular genres and children’s interdiscursive familiarity with visually rich cultural forms such as picturebooks or seasonal symbolism. In the argumentative task, for example, the theme of Christmas afforded symbolic visual representation, even though “argument” as a rhetorical form seem less visually conventionalised.
The findings suggest that intermodality is shaped not only by social purpose but also by the thematic and semiotic affordances of the content domain. Concrete and symbolically rich themes, for example, Christmas or horror, seemingly lend themselves more naturally to visual representation. More abstract or informational content, such as an animal's life cycle, tends to be expressed more systematically through writing, with drawings serving an illustrative function. This underscores the need for educators to consider how the design of writing tasks enables or constrains multimodal expression and to explicitly guide children in recognising the affordances of different genres and modes.
A key contribution of this study is the development of a refined account of intermodality in early school writing that advances current multimodal theory in two ways. First, by integrating transitivity analysis with a dialogic conceptualisation of children’s texts as utterances, we demonstrate that intermodal relations are not merely structural alignments between modes but socially situated responses shaped by genre, audience, and thematic framing. Existing frameworks (e.g., Martinec & Salway, 2005; Siefkes, 2015; Unsworth, 2008) tend to conceptualise image-text relations either ideationally or taxonomically. Our analysis extends these accounts by showing how children distribute agency, evaluation, and epistemic stance across modes in patterned ways corresponding to the social purposes of school writing tasks. Second, we propose an analytically operationalizable model of intermodal transitivity that captures how processes are realised visually and verbally within single compositions. This model offers a more granular description of multimodal meaning-making in the early years than do existing typologies and provides researchers and educators with a systematic tool for tracing how children mobilise semiotic resources across modes.
Finally, extending transitivity analysis to drawings highlights how children use visual modes to represent material actions, mental states, and relational configurations in ways that parallel but are not subordinate to writing. Although each semiotic mode has distinct affordances (Painter et al., 2014), our analysis demonstrates complementarity across semiotic systems when intermodality is examined at the level of process (or ideational) realisation. This perspective challenges literacy hierarchies that privilege alphabetic writing and underscores the importance of recognising the full semiotic repertoires that children bring to school. It also offers a concrete analytical resource for understanding and supporting multimodal literacy development in early school education.
Footnotes
Ethical Considerations
The data has been collected in accordance with the research ethical guidelines (Swedish Research Council, 2024). The FEAST-project has gained ethical approval for the project from the Swedish Ethical Review Authority (Dnr 2020–00738 and Dnr 2020-02990).
Consent to Participate
Consent has been given by both children and caregivers to participate in the research project.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This article is an output of the project Functional writing in Early School Years: Assessment, Teaching and Professional Development (FEAST), funded by the Swedish Institute for Educational Research, led by professor Åsa af Geijerstam at Uppsala University, Sweden. This work was financed by Swedish Institute for Educational Research 2020–2023 (head of the project: Åsa af Geijerstam, participating researchers: Oscar Björk, Charlotte Engblom, Jenny W Folkeryd, Sofia Hort, Caroline Liberg, Kimberly Norrman, Maria Rasmusson and Maria Westman).
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data Availability Statement
No data available
