Abstract
This paper introduces a novel method for investigating how public service, market-driven, and populist news outlets differ in their construction of ideological positions in audiovisual reporting. A semi-automatic framework of multimodal semiotic content analysis is presented that combines computational audiovisual feature detection with theory-driven manual annotation of discourse and narrative elements and shows early results concerning the construction of distinct positions on COVID-19. The method presented suggests a powerful reconceptualisation of what counts as “large-scale” analysis in multimodal video research, arguing that scale should be defined not merely by the number of videos but be explicitly related to analytical complexity, levels of abstraction, and the density of annotated units, such as shots and many other multimodal features. Based on a corpus of 107 German news reports (January–March 2022, containing around 5000 shots) from Tagesschau (public service), BILD TV (market-driven), and CompactTV (populist/right-wing), we systematically conduct shot-by-shot annotation for multimodal semiotic content analysis. The analysis identifies recurring narrative patterns across channels and uncovers how multimodal configurations materialise distinct ideological orientations. Statistical comparisons show clear cross-channel differences. Tagesschau foregrounds expert voices and largely avoids sen-sationalist framing. BILD TV emphasises emotionalised language associated with shock and “craziness”, while CompactTV amplifies oppositional actors and frequently frames COVID-19 measures in terms of “dictatorship” and “coercion”. These findings demonstrate that multimodal discourse patterns function as ideological performances, shaping whose voices are legitimised, how events are staged, and which emotions are mobilised. The study illustrates how semi-automatic multimodal analysis enables scalable, fine-grained investigation of narrative form and ideological signalling in contemporary news environments.
Keywords
Introduction
“Performing ideologies: To suggest that news is a performance is saying nothing new. Events, issues, and social tensions are performed daily in television news, brought to life through narratives, acted out through visuals, embroidered with emotions and so forth… TV news has operated by a standard set of performance codes or tropes, each used to establish the legitimacy and authority of those imparting such ‘truth’, and each central to its stated purpose of representing ‘reality’” (Jones, 2012: p181).
As the above quotation shows, scholars have long recognised that television news operates as a performative practice through which news events are edited, sequenced and constructed into engaging narratives. Rather than merely reporting realities, news media produce news videos that need to be seen as mediated stories, often constructed deploying established filmic narrative conventions which have the effect of legitimising particular ideologies of information and truth. Building on this understanding of news as narrative, this paper advances a framework that applies recent advances in multimodal analysis to show how the patterned orchestration of linguistic, visual, and auditory resources materialise the ideological orientations of different types of news outlets. By comparing and contrasting the three main types of outlets – public service media (here after PSM), populist channels, and market-driven organisations – we conceptualise audiovisual news as a site where ideologies are embedded in multimodal discourse forms, and where distinct “performances of truth” are negotiated through multiple aspects and modes of narrative features.
Methodologically, we use Multimodal Semiotic Content Analysis (hereafter, MSCA: Tseng et al., 2026a) to systematically annotate and analyse the multimodal choices deployed in news videos for determining what is shown and how, whose voices are heard, and which emotions and evaluations are elicited. MSCA builds on traditional content analysis (Krippendorff, 2004) and its multimodal extension (cf. Bell, 2001; Serafini and Reid, 2023), complementing detailed multimodal semiotic analysis with reliable quantitative approaches. The method allows us to capture multimodal discourse patterns from a large-scale body of news video data, showing the ways in which these videos integrate multiple expressive modes – spoken language, news imagery, infographics, editing, sound, and music – to structure events, evoke engagement, and guide interpretation.
Whereas close analysis of individual videos can usefully suggest the existence of such multimodal discourse patterns, determination of the patterns’ extent and potential, variation in use according to social factors demands a treatment at scale. This is itself a considerable challenge, both theoretically and practically. Consequently, we adopt a mixed-methods framework that combines automatic audiovisual feature detection with MSCA in order to support large-scale annotation across extensive news video corpora. Recent advances in computational audiovisual research enable the systematic identification of an ever broadening range of audiovisual features at a scale previously unattainable in film and video studies. Our annotation approach leverages techniques such as temporal video segmentation (dividing continuous video streams into discrete units such as shots, scenes, and speaker turns), face recognition, linguistic phrases and terms realising particular semantic groups, and many more. These automated annotations are subsequently refined manually and enriched with additional discourse and narrative-level annotations, such as narrative event types and cohesive links between people and locations.
This kind of semi-automatic framework for multimodal discourse analysis opens up new pathways for understanding how multimodal elements shape news interpretation and ideological signalling. By enabling systematic multimodal semiotic analysis of fine-grained audiovisual and discourse features and automatic annotations for large-scale quantitative analysis, the framework presented in this paper allows for robust comparisons of recurring narrative patterns across diverse types of news outlets.
The paper is structured as follows. Section 2 first reviews the conceptualisation of news as narrative and sets out the corresponding challenges facing the analysis of such audiovisual news narratives. Section 3 describes our methodology, first explaining how we combine tenets of content analysis with recent multimodal semiotic analysis, then discussing the required scale of broad film and video analysis, motivating in particular why we opt for a semi-automatic rather than fully automatic analytical method, and then finally outlining our methodological architecture, detailing how we interconnect the notions of multimodal semiotics, content analysis, discourse analysis and ideologies within the genre of audiovisual news. Section 4 presents our data and our social semiotics-based annotation scheme. Section 5 presents the automatic tools employed in the study and how we integrate automatic detection with manual refinement of multimodal features, for which Section 6 then reports the intercoder reliability scores achieved in that annotation. Section 7 turns to our contrastive findings across the three types of channels analysed, before Section 8 briefly concludes with a review and outlook.
Narrativisation and filmic discourse patterns in news videos
Narrative has long been established as a highly effective mode of communication across media, characterised by properties that foster engagement, most notably the build-up and release of tension, the presence of identifiable characters, and the perception of causal coherence among events (cf. Berning, 2011; Dahlstrom, 2014; Fludernik, 1996; Hogan, 2011; Ryan, 2004; Smith, 1995). Television news has consequently also long been recognised as a form of filmic narrative (Dunn, 2005; Ekstr¨om, 2000; Glasgow University Media Group, 1976; Philo, 1995; Sperry, 1981; Wahl-Jorgensen and Schmidt, 2020). By the late 1980s and early 1990s, TV news had increasingly adopted techniques familiar from narrative film to capture viewers’ attention and enhance comprehension (Baym, 2004).
In adopting such narrativised forms, however, audiovisual news has, to some extent, displaced the traditional ideal of objective detachment with what has been described as ‘engaged journalism’ (Dunn, 2005, 144). But research also suggests that audiences transfer interpretive habits from filmic storytelling to news contexts, even when viewing ostensibly non-narrative content (Cohen and Roeh, 1990). This means that narrativised news formats may well elicit interpretations concerning protagonists, causality, and moral evaluation that diverge from those intended by journalists.
The prevalence of narrative structures in contemporary news poses many analytical challenges for assessing the degree and impact of narrativisation across media outlets. As narrative traits now permeate all types of platforms – from PSM to market-driven broadcasters and social media channels – it becomes relevant to examine how far PSMs have moved toward market-oriented narrative conventions, or whether narrative patterns characteristic of populist news formats are sometimes integrated into PSM and commercial news alike.
The presence of narrativisation in a medium is itself a complex phenomenon, how-ever, involving configurations and discourse organisations that spread over quite diverse levels of abstraction. For example, on the one hand we have the established technical and formal features by which narratives are constructed filmically – including camera features, editing style, event types, protagonists, narrative arcs, music and sound, emotional and evaluative visual images and language, and so on – and, on the other hand, more abstract levels of organisation such as balancing (or not) authoritative detachment with affective engagement, blurring boundaries between journalism, journalistic practice, and dramatic storytelling, and deploying specific discourse patterns with the purpose of distinguishing themselves ideologically from others.
Addressing these questions requires a systematic framework encompassing technical and formal as well as semantic and discourse dimensions to capture the multiple dimensions of narrativity so that differences and similarities across commercial, market-oriented channels and public service media can be revealed. Only by contrasting these multiple layers of narrativisation can cross-channel differences be reliably revealed (Tseng et al., 2026a). The aim of the current paper is to contribute further to this line of research by proposing a comprehensive framework for analysing discourse and narrative patterns in news videos. By employing multi-dimensional annotation of audiovisual elements supported by semi-automatic methods that increase scale, our approach enables systematic cross-channel comparison and insight into how narrative forms and strategies mediate journalistic ideology and audience engagement.
Methodological framework
Multimodal semiotic content analysis as an extension of content analysis
Although content analysis is a long established and well defined method for revealing significant differences in media practices of the kind that we wish to address here (cf. Krippendorff, 2004; Neuendorf, 2002), certain problems and limitations have been noted in the literature (cf., e.g., Karlsson and Sjøvaag, 2016). Particularly important for our current concerns is the question of scale. Since most content analysis coding schemes rely on interpreting data, this usually requires human intervention. Current work employing large language models to improve on this situation show some promise (e.g., Boji´c et al., 2025), but only for specific properties and also at the cost of decision transparency. The need to rely on human coders thus continues to impose severe limits on either the range of analytic categories that can be applied, or the scale of data analysis, or both.
Our approach in the present paper explores a different path by which computational support can be provided for the coding task. Riffe, Lacy, Watson and Lovejoy define content analysis as: “the systematic assignment of communication content to categories according to rules specified in a coding protocol, and the analysis of relationships involving those categories using statistical methods” (Riffe et al., 2024, 4).
And so this means that one needs to specify above all just what coding categories are to be taken and how these are to be operationalised.
We suggest that a rich source of categories for coding protocols can be supplied from multimodal semiotic descriptions. Many of these have their origins in the critical social semiotics introduced by Hodge and Kress (1988), itself a forerunner of critical discourse studies, particularly multimodal critical discourse studies (Machin, 2013). Although sometimes overlapping with the concerns of content analysis when applied to media (cf., e.g., Bell, 2001; Serafini and Reid, 2023), significant differences in approach have until now led to these research traditions remaining distinct. Whereas both seek to engage with materials at an interpretative level, addressing content and social meaning, critical discourse studies have paid less attention to the methodological demands of achieving reliable coding schemes and, as a partial consequence, have also typically not engaged in sufficient statistical evaluation of any patterns claimed. Interpretation remains by and large discursive and informal.
More recent multimodal work of the kind we build on here (Bateman et al., 2017; Tseng et al., 2026a) has attempted to renew its empirical connection and so moves closer to the aims of content analysis. What is more, the descriptions offered within multimodal semiotics suggest that coding catalogues can be beneficially organised as being spread over several levels of abstraction, ranging from form-related classifications to socio-cultural interpretations (Bateman, 2022; Bateman and Tseng, 2023; Tseng, 2013). This rejects the dichotomous separation into ‘form’ and ‘content’ which has previously prevented content analysis and other descriptive approaches interacting.
The provision of multiple levels of abstraction then also supports a staged approach to computational support whereby lower levels of abstraction are already often amenable to computational analysis, less abstract levels might be best supported by semi-automatic annotation procedures, and highly abstract levels remain best handled by human coders. Critically, however, we have found that there are many intermediate levels of descriptions where lower level descriptions can be recoded as higher level descriptions, thus providing computational support for ever-increasing levels of abstract description. Due to their relation to lower level descriptions, such descriptions maintain a documentary trace to the formal evidence used in their application, thereby supporting more transparent coding results. In short, this enables a systematic analysis at formal and discourse levels which then supports social and ideological interpretations.
In the following, we fill in this conceptual framework with the concrete descriptions necessary for actual analysis of audiovisual news reports, showing how semi-automatic analysis steps provide the basis for abstract content analysis-like categories that can be used to isolate differences and similarities between news channels in the usual way, but now with respect to both larger scale data sets and broader collections of phenomena.
The scale and degree of automation of large-scale news video analysis
Given the re-conceptualisation of the kinds of coding catalogues being applied in the current work, we can also offer a productive re-consideration of the notion of ‘large-scale’ news video analysis. In this subsection, we consequently focus specifically on what constitutes a sufficiently ‘large’ dataset in relation to, on the one hand, the research questions and, on the other hand, the complexity of the analytical dimensions involved in audiovisual analysis. These facets combined allow motivated decisions to be made between fully automatic and semi-automatic approaches to audiovisual content analysis.
Today, with the availability of numerous open-source computational tools capable of annotating a wide range of linguistic and visual features, processing large datasets for individual features and correlating them with social or political categories has become highly accessible (Clever et al., 2023; Li et al., 2025). Substantial challenges remain, however, when expanding analytical dimensions to include discourse and narrative features that unfold across temporal sequences rather than the individual frames more commonly handled. In addition, current computational technologies are well suited for formal and affective analyses of text and images, but higher-level discourse phenomena – such as identifying recurring places, objects, and individuals, and tracking their actions and movements across shots and scenes to ‘tell a story’ – require a far more integrated combination of multimodal information extraction. Only when such complex integration becomes feasible would fully automated corpus analysis of narrative patterns in film and video be feasible. At present, however, the fully automatic tracking of narrative patterns is still highly challenging (Tseng et al., 2026b).
Such considerations have many consequences. Principal among these is a reassessment of just what ‘large’ corpora means in the context of highly complex, multi-layered, multimodal data that is itself receiving multiple descriptions at different levels of abstraction. It is no longer sufficient, or even particularly useful, to simply state that there are 60, 600, or six million ‘texts’ or ‘videos’, and so on because each item in the collection may itself be internally complex along many dimensions. Certainly when now reporting on any multimodal corpora, it is important to give indications of the particular kinds of units receiving analysis, their modalities, and levels of descriptive abstraction employed. ‘Large’ can then be seen with respect to adequacy in providing sufficient quantities of some phenomena of interest to support the discovery of patterns of use relevant for a research question. Even relatively small ‘absolute’ numbers, such as the quantity of videos in a corpus, may give rise to considerable bodies of data that, triangulated against more abstract discourse level categories, are more than sufficient for showing regularities. Quite concretely, in our present case, even when dealing with collections of news reports numbering only in the hundreds, these already include thousands of shots, one of our basic unit of analysis, which themselves each contain 10–50 audiovisual features, all or any of which may be contributing to discourse or narrative patterns.
This issue also relates to the choice to be made between fully automatic and semi-automatic methods. This choice depends largely on the analytical level and complexity required to address particular research questions. On the one hand, as previous studies have demonstrated, various formal and semantic features can exhibit significant correlations with social or political ideologies when examined at sufficient scale. In such cases, an automated approach is both necessary and productive for processing the required volume of data. On the other hand, when conducting corpus-based video studies aiming to uncover associations between narrative patterns and social ideological issues – particularly those aiming to uncover how viewers’ narrative interpretations are shaped – bottom-up large-scale analysis is best complemented by top-down constraints derived from discourse and semiotic frameworks.
This requires not only annotations of lower-level formal and semantic features but also methods for tracking their cohesion, coherent uses, and their roles in discourse and narrative structures across shots and scenes in extended video datasets. Achieving such multiple levels of analysis currently necessitates combining computational techniques with manual refinement, as the connection between automatically extractable features and higher-level discourse patterns raises significant theoretical issues that will continue to place fully automated reliable analysis beyond current capabilities for some time to come.
The inclusion of theoretically informed, top-down annotation schemes supported by manual annotations within a semi-automatic workflow can then in addition support the adoption of a more functional notion of ‘scale’. Given reliable corresponding ‘top-down’ guidance, the dataset size required does not need to be exceedingly large in some absolute sense. As we will now illustrate, more focused annotations support a more precise allocation of attention to properties of the data that are distinctive of particular strategy uses. In particular, we detail how we integrate such multiple levels of annotation for news videos data – spanning formal, semantic, and discourse levels – through a semi-automatic approach. We begin by setting out some core challenges of video narrative annotation, and then explain how bottom-up feature annotation is complemented with top-down constraints derived from semiotic theory in the overall design of our annotation scheme.
Analytical strata of the current study
Drawing on the description, Figure 1 provides a graphical overview of the underlying organisation and motivation of our methodological framework. We annotate audio, visual, and linguistic features that collectively shape recurring multimodal discourse patterns based on semiotic theory (characterised in Section 4.2) – patterns through which news topics are framed, evaluated, and protagonists represented. We then compare these discourse patterns across our three selected news channels to examine their similarities, differences, and the ways in which they reflect or reproduce the channels’ underlying ideological stances toward specific topics. Interlinking multimodal features, discourse patterns that combines multimodal features, and cross-channel comparisons of ideologies.
Data and multimodal annotation scheme of the study
Data, units of annotation and semiotic analysis
The dataset discussed in this paper comprises 107 news reports on COVID-19 drawn from three German outlets – Tagesschau (public-service), BILD TV (market-driven), and CompactTV (populist/right-wing) – broadcast between January and March 2022. Each report runs approximately 5–7 minutes, yielding a corpus of roughly 700 minutes of video material.
When conducting semiotically motivated annotation, one of the first decisions that needs to be made concerns the kinds of units of analysis that will be adopted. Units can be motivated on both practical and theoretical grounds, ideally both. Distinct kinds of patterns may involve quite different units, even for the same data, and so the decisions made in this respect must always be documented. Several distinct kinds of units play a role in our annotation. Most important are divisions into shots, normally tracked following the principal visual frame present in a video sequence, and divisions into speaker turns, in order to capture linguistic data. There is no claim that these are the only units that might be relevant. In our study, each video contains about 55 shots, resulting in a total of more than 5,000 shots analysed.
Even though the primary unit of visual annotation reported here is the shot, it is important to emphasise that our semiotic and discourse analysis extends beyond shot boundaries. In other words, while the shot provides one useful and systematic unit for quantitative analysis – particularly for measuring the frequency of observations – the semiotic categories employed in this study are not constrained to individual shots. As will become clear in the annotation categories discussed below, linguistic, visual, and auditory features may evolve or persist across multiple shots, and can also vary within single shots. This allows us to analyse patterns and structures at different levels of abstraction and detail, and also makes it more straightforward to consider extensions of the analysis to further features subsequently.
Our dataset was selected to include reports covering comparable topics (e.g., policy updates, vaccination, protests, restriction debates) within the shared time window in order to enable close comparisons of content and form. The three channels are also popularly well-known for their distinctive presentational styles and so offer a certain baseline for evaluation of our methods. The public service media Tagesschau encompasses concise bulletin packages, BILD TV is more personality-driven and talk-centered, while CompactTV focuses substantially on opinion-led reports, longer interviewee monologues, and oppositional, activist narratives.
While these general stylistic orientations are often assumed and can be documented on a case-by-case basis, the focus in our study is to articulate methods capable of revealing narrative features that contribute to a channel’s ‘identity’ overall. More specifically, we examine whether – and how – we can compare similarities and differences along multiple dimensions of multimodal features across our data from the three channels so as to uncover more nuanced characterisations of any contrasts exhibited.
Annotation scheme along three main communicative dimensions
This subsection outlines the annotation scheme employed in the study. As introduced above, the scheme is grounded in a semiotic framework that delineates the analytical dimensions relevant to communicative functions and strategies. By drawing on this framework, we are able to produce theoretically and practically robust annotations that nevertheless directly reflect the formation of discourse patterns. This then serves as a strong comparative scaffold both for contrasting lower-level features and for considering the interpretative consequences of those lower-level features.
The specific semiotic framework we employ in the study reported here draws on several strands of work deployed in the area of multimodality. One of these is Social Semiotics (Kress, 2010; Kress and Van Leeuwen, 2001), an approach building on a mul-timodal extension of Systemic Functional Linguistics (Halliday and Matthiessen, 2014) and in which the relationship with the social is considered central. In Social Semiotics, modes of communication are consequently understood as resources that materialise three generalised kinds of meanings, called ‘metafunctions’. The ideational metafunction characterises how a communicative mode represents experience, depicting participants, actions, and circumstances through resources such as images, text, speech, movement, spatial arrangement. The interpersonal metafunction describes how communicators establish relationships, attitudes, and social roles, whether through facial expression, gaze, positioning, or other interactive cues. And the textual metafunction describes how messages are organised into coherent wholes, encompassing sequencing, image layout, visual salience, rhythm, and other structural features that guide interpretation. These provide an overall organisation for our complementary sets of coding categories.
The classification of meanings according to metafunctions offers in addition a valuable guide for focusing analytic attention on how audiovisual news reports construct representations of reality, negotiate social relations, and shape messages into meaningful, contextually coherent narratives or multimodal compositions. Our annotation scheme reflects this organisation directly: (1) News events and contents (ideational function): annotation of different categories of action and speech to distinguish event types. (2) Evaluation and social role (interpersonal function): annotation of semantic groups of terms related to evaluation and emotion. (3) Structure and segmentation (textual function): shot and scene segmentation.
This combined scheme seeks to capture the varied multimodal resources that are characteristic of the news-video genre, articulating a principled framework supportive of empirical investigations ranging from the distributional differences we document below to further explorations of how the resources may combine for the purpose of directing viewers’ interpretations.
In addition, we specifically tailor the annotation categories of the framework and their operationalisations to our corpus of Covid-19 news reports. By these means we seek to mark out the most significant recurring elements in our data in a theoretically well motivated manner. This is one of the principles of multimodal semiotic content analysis set out in Tseng et al. (2026a) — the particular features of the coding scheme are designed in response to specific properties of the data and the research questions, as well as incorporating insights from pilot studies where the coding features were trialed prior to more extensive application. This allowed us to avoid prematurely importing predefined categories from earlier studies that may or may not have relevance for the data at hand. All proposed coding scheme categories were then subjected to reliability testing prior to analysis in order to ensure the robustness of the scheme.
In the following sections, we introduce the coding categories and criteria, and then describe how the annotations were semi-automatically implemented using computational tools for feature extraction.
Annotating categories of Covid-19 related news events and contents
Four main categories of news events and contents were distinguished in the dataset: speech processes, interactions between actors and objects, actors’ movements, and ana-lytical processes. These categories reflect the ideational dimension of meaning-making by capturing how Covid-19 news reports represent actions, actors, and situational con-texts. The coding criteria and operational definitions for each category are outlined below.
News speech process
Sequences in audiovisual news reports can be categorised by the types of talk presented. That is the news situation (Cheema et al., 2024) in which news is being talked about. We distinguish the following speech process types: • Interview: direct questioning and answering between journalists and news actors. • Talking head: on-camera reporting delivered either by on-site journalists or by anchors in the studio. • Commenting: extended explanatory or evaluative commentary, often produced by studio-based editors or expert contributors. • Voice-over: off-screen narration accompanying visuals without showing the speaker.
Differentiating among these talk types is essential for analysing their co-occurrence with other communicative features along other communicative dimensions, such as the social roles of actors, the distribution of evaluative language, or the staging of events. For example, distinguishing interview segments allows us to identify which social actors are given a platform to speak. And noting in which talk process evaluative expressions appear (e.g., in talking-head reports vs studio commentary) sheds light on institutional stance-taking practices. These distinctions also enable us to trace how verbal discourse interacts with visual cues to construct the overarching narrative of the Covid-19 crisis.
Actors’ interactions and actions on objects
The dataset contains several recurrent forms of interaction involving actors handling, or acting upon, objects – activities that visually foreground how individuals respond to or manage pandemic-related situations. The most common interaction types include: • administering or receiving vaccinations, • conducting or undergoing Covid-19 tests, • holding signs during demonstrations, • dining or serving in caf´es and restaurants, • caring for patients in clinical or hospital settings.
Annotating these interaction types enables cross-analysis of how specific actions are distributed across different news segments, channels, and speech-process types. It also allows us to investigate patterns such as the alignment between particular actor roles (e.g., public service workers, politicians, experts) and the actions they are depicted performing. Furthermore, these annotations help reveal how certain activities become symbolically associated with key pandemic narratives, such as public compliance, medical expertise, or political resistance.
Actors’ movements
In addition to object-oriented actions, the dataset features recurrent forms of bodily movement that structure how news actors are visually positioned within the unfolding events. The most salient movement types include: • participating in Covid-19 demonstrations, • walking in public spaces such as streets or transport hubs, • queuing for testing or vaccination services.
Annotating these movement patterns provides insight into how mobility, crowd behaviour, and spatial organisation are represented in Covid-19 news coverage. Such annotations support analyses of how news reports visually construct public responses to health regulations and how they depict everyday life under pandemic conditions.
Analytical processes
Beyond the three dynamic scene types described above, Covid-19 news reports frequently incorporate static, informational elements that are commonly assumed to support analytical reasoning and factual grounding. These segments typically present data, summarise policies, or foreground authoritative sources. To capture these forms of static, evidence-oriented communication, we define the overarching category of analytical process and distinguish the following primary forms of information presentation within that category: • infographic: visual or semi-visual representations that condense information for quick comprehension. This includes bullet-point lists, fact boxes, timelines, maps, flow charts, and schematic illustrations of procedures (e.g., how a virus spreads or how a policy is implemented). • graph: quantitative data visualisations that depict numerical trends or proportions, such as bar charts, line graphs, pie charts, scatter plots, etc. These segments often track case numbers, vaccination rates, transmission curves, or regional comparisons. • citation: explicit textual references used to anchor claims in authoritative discourse. This includes (a) on-screen quotation highlights, (b) excerpts from documents or articles, (c) spoken or captioned citations accompanied by the author’s photograph, name, or institutional affiliation, and (d) policy excerpts displayed as text overlays.
Annotating categories of Covid-19 related evaluation and social roles
To capture the interpersonal, evaluative and emotional dimensions of Covid-19 news discourse, a set of semantic annotation categories was developed (Martin and White, 2005). These categories represent recurrent linguistic patterns through which journalists sensationalise, emotionalise, and evaluate unfolding events and social actors. Within the Covid-19 news corpus, several dominant semantic groups emerge, each reflecting characteristic ways of emotionalising and evaluating Covid-19 related policies. • positive evaluation: 1. trust: words and phrases that refer to trust or distrust, e.g. “vertrauen, zweifeln, missvertrauen”, etc. 2. free, peace or hope: phrases, such as free, freedom, hope, hopeful, peace, peaceful. 3. togetherness and solidarity: phrases that suggests togetherness, e.g. “Zusammenheit, Zusammensein, Solidatit¨at”. • negative evaluation: 1. dictator and coercion: phrases related to dictatorship or coercion, e.g. “Diktatur, Zwang”, usually depicting forced obligation or reinforce-ment of vaccination and Covid policies. 2. uncertainty, confusion and skepsis: phrases related to uncertainty, unsure, unclear, confusing, e.g. “unklar, Klarheit, unsicher, skeptisch, zuru¨ckhalten, verwirren, verwirrt”, 3. craziness, nonsenseand insanity: e.g. “Wahnsinn, Irrsinn, Unsinn, sinnloss, Irre”. • emotionalisation: 1. shock, horror, terror: phrases related to shock, horror or terror, e.g. “Schock, schockierende, schockiert, Horror, Terror” 2. worry and panic: e.g. “Sorge, befu¨rchten, Panik, Angst, besorgt, sorglos”. • violence and conflict: word or phrases including violence/violent, e.g. “Gewalt/gewalt¨atig”
In addition to these semantic evaluative groups, the social roles involved in Covid-19 news videos were also annotated according to the categories presented in the taxonomy shown in Figure 2. This taxonomy organises the social roles set out for critical discourse analysis by Van Leeuwen (1996) into a hierarchical structure consisting of broad groups and their corresponding subroles as appropriate for our data. At the highest level, roles are distinguished into journalist, public service, layperson, and elite. Each group then contains more specific social roles that together span the actors most frequently depicted in our data thus: • journalists: encompassing anchor person or reporters, either on the site journalists or studio commentators. • Covid-19 related public services: police, who are often seen in demonstrations, and firefighters/paramedics, who are in or near ambulances. • layperson: layperson is further categorised into protesters and non-protesters, who represent the general public in public space, restaurants, offices, etc. Protesters are distinguished between anti Covid protesters or Covid policy supporters. • elites: this category includes experts (medical or others), politicians from non-German countries or from different local parties within Germany, and celebrities reported with regard to Covid-19, for instance, Djokovic, the tennis player who refused to receive vaccination. Social roles of actors appearing in Covid-19 news videos.

In news videos, the social roles of actors are often identified with verbal subtitles. By explicitly distinguishing the types of social roles involved, the taxonomy enables a clearer analysis of how different actors are depicted and evaluated within Covid-19 news reporting. It also provides a consistent analytical framework for comparing the representation patterns across the three channels.
Annotating structure and segmentation of shots
Each news report was segmented into shots, which served as the primary analytical unit of the current study. Shot properties – including segmentation boundaries and shot sizes – were automatically analysed using computational tools as detailed in the following section.
Automatic tools used and semi-automatic analytical process
This study employed two main computational tools for automatic analysis and data exploration: TIB AV-Analytics (TIB-A-VA: Springstein et al., 2023), developed at TIB Hanover, and Zoetrope (Tseng et al., 2023), developed at the University of Leipzig.
Automatic detection of shot and scene segmentation
TIB-A-VA is an open-source, web-based video analysis platform (publicly available at https://service.tib.eu/tibava) that integrates automatic solutions for information retrieval. It integrates state-of-the-art AI approaches in the fields of computer vision, audio analysis, and natural language processing for many relevant video analysis tasks including, but not limited to, shot boundary detection, shot size classification, person clustering, place classification, visual object detection, and automatic speech recognition. An overview can be found in Burghardt et al. (2024). The present study leveraged the automatic segmentation of shot boundaries based on TransNet V2 (Soucek and Lokoc, 2020), classification of shot sizes, and automatic detection of places (Zhou et al., 2018), to structure shots and scenes throughout the video (Tseng et al., 2026b).
Automatic detection of semantic groups in language, actions and movement, particular faces of social roles
The second tool this study employed is Zoetrope, a prototype combining a range of computational techniques, such as speech recognition, face detection, etc., into a flexible tool supporting the progressive annotation of larger data sets by means of a more interactive interface. This tool is capable of visualising automatic annotations and provides some basic functionalities for querying a video, for instance for specific key words or named entities. The interface of Zoetrope is shown in Figure 3. Prototype tool for visualizing and exploring news videos.
A particularly useful function of Zoetrope is the keyword search function (segment 1 in Figure 3). Any keyword can be used as a query, which is then searched for in both the spoken language (Mozilla DeepSpeech 1 with a German model by Agarwal and Zesch 2 ) and written language (scene text detection framework easyOCR 3 ) of the video (Tseng et al., 2023).
The keyword search function supported this study specifically by effectively locating the Covid-19 reports from the complete news video reports; each 15-min news video programme generally combines reports on different news topics making selection necessary. With keyword search, the sequences that mention a particular Covid-19 topic (e.g., Covid, vaccination, etc.) are automatically highlighted on the timeline. Furthermore, the tool also allows queries that draw on word embeddings 4 , so that words semantically related to the query can also be found. This function supported us in automatically detecting all semantic groups of phrases that we aimed to annotate.
As shown in Figure 3, the results of such a query are visualised in segment 2. Query results found in spoken language are shown in green, while results in written language are shown in red. Note that for the query “Krieg” (war), we also find semantically related concepts such as “Truppen” (troops), “Invasion”, or “Militär” (military), due to a flexibly adjustable semantic similarity threshold. This supported the detection of semantically related terms and phrases from our targeted set of terms. The automatic image description function also supports the annotation of particular actions and movements, such as holding signs, walking, or dining as required for our ideational annotation categories
In segment 3, we also see some exemplary automatically determined features, including face detection and audio analysis by means of a spectrogram. The face detection function tracks particular faces throughout the video – using this function, one can automatically detect the reoccurring anchors or politicians evidently most relevant to Covid-19 news in early 2022. Finally, segment 4 shows an interactive video player that can be navigated by means of the timeline or by clicking on specific results, such as a keyword or a face. Visible features, such as written keywords or faces, can also be rendered with their bounding boxes within the video.
Semi-automatic analysis
The automatic analysis described above is complemented by subsequent manual refinement as motivated in earlier sections. Figure 4 outlines the full workflow. First, we import the video corpus into both TIB-A-VA and Zoetrope to obtain automatic detections of shot boundaries, faces, written and spoken phrases belonging to our defined semantic groups, and relevant movement cues. We then export these automatically generated annotations into ELAN (Wittenburg et al., 2006), where they are manually refined, corrected, and extended. This manual stage enables the addition of discourse- and narrative-level annotations, namely, the annotation categories specified in Section 4.4 that are difficult to detect automatically – for example, distinguishing news speech types (e.g., interviews, talking-head segments, commentary, off-screen narration), identifying more fine-grained narrative events such as caring for patients, queuing, or taking COVID tests, and differentiating social roles. The multitier annotations are then exported to R (R Core Team, 2025) for correlation analyses and pattern identification. Semi-automatic annotation process.
Finally, the refined manual annotations may also be taken as ground truth labelling for further rounds of training for the computational models (Cheema et al., 2024), enabling more accurate and narrative-level automatic detection in subsequent annotation cycles and future studies.
Intercoder reliability tests
To assess the reliability of our annotation scheme, we conducted intercoder reliability tests for the annotation categories over a 30% sample of the entire dataset (30 out of 107 videos) using Krippendorff’s alpha (Krippendorff, 2011). These tests were conducted by two master students trained as annotators for this study.
As shown in Figure 5, most categories achieved high reliability, with alpha values exceeding the commonly accepted threshold indicating good levels of agreement of 0.67. Positive evaluations related to trust and togetherness, negative evaluations, and emotionalisation reached perfect agreement (alpha = 1.00), while social role (alpha = 0.94), conflict (alpha = 0.86), analytical strategies (alpha = 0.85) and speech types (alpha = 0.93) also demonstrated strong consistency between coders. Movement (alpha = 0.66) fell near acceptable reliability levels. Positive evaluation related to freedom (alpha = 0.74) and interaction (alpha = 0.70) showed moderate reliability. Overall, the results indicate that the coding scheme was robust and so could be applied consistently across annotators, providing a reliable foundation for the subsequent analyses. Results of intercoder reliability tests.
Analytical results: Significantly different discourse patterns across the three channels
We now report the statistically significant findings from our analysis. We conducted Fisher’s exact tests to examine the relationships between the three news channels and each of the annotated categories introduced in Section 4.4. This analysis allowed us to identify which of the categories showed a statistically significant dependence on channel, thereby providing evidence of differences in use across the three channels.
In addition to analysing each category separately, we also combined the annotations for the speech type interview, with each of the other categories – i.e., interview with social roles and interview with actors interactions or movement, to identify which types of social actors were most frequently invited to express opinions on each channel, in which event types interviews take place, and whether these differed significantly across channels.
Who is being interviewed?
Observations of the combined annotations of “interview” and “social roles” across the three channels.

Test result of correlation between channels and interviewed social roles; mosaic plot of residuals generated with the vcd R package (Meyer et al., 2022).
Bild TV can therefore be seen to exhibit a pronounced over-representation of anchors acting as interviewees — suggesting that its anchors engage in extended dialogic commentary rather than the conventional “talking head” format typical of most news broadcasts. The anchors frequently articulate their own opinions, functioning more like invited commentators. In Tagesschau’s COVID-19 coverage, medical experts are substantially interviewed, but the voice of anti- Covid policy protesters as well as the right-wing politicians are under-represented. In contrast, CompactTV shows an over-representation of right-wing politicians and anti–COVID-policy protesters, while relatively under-representing medical experts’ opinions.
Overall, therefore, the channels differ significantly in whose voices they foreground. These narrative choices reflect distinct ideological emphases in their respective reporting.
Which news narrative event is related to interviews
Observations of the combined annotations of “interview” and “event actions” across the three channels.

Test result of correlation between channels and event where interview takes place; mosaic plot of residuals generated with the vcd R package (Meyer et al., 2022).
How are Covid-19 policies evaluated
The counts of three main types of negative evaluations directed at Covid-19 policy, namely, portraying the government as dictatorial or coercive, expressing uncertainty, and framing policy as craziness and nonsense, are displayed in Table 3 Figure 8 displays the association between the three news channels and three forms of negative evaluation. The significant result of Fisher’s exact test (i.e. p-value = 2.22e-07, Cram´er’s V = 0.48, indicating a large effect size) indicates that the distribution of these evaluative frames differs systematically across channels. Test result of correlation between channels and different semantic groups of phrases of negative evaluation; mosaic plot of residuals generated with the vcd R package (Meyer et al., 2022). Observations of the annotations of the three main semantic groups of negative evaluations across the three channels.
The populist CompactTV shows strong overuse of “dictator/coercion” to depict the governmental policy. The sensational, market-driven channel Bild TV, by contrast, do not use much of these terms but uses phrases related to craziness and nonsense to describe the Covid policy.
The analysis shows that there is an underrepresentation of these terms in Tagesschau. However, the unexpected results is the overrepresentation of the phrases related to “uncertainty and skepsis” spoken in the PSM Tagesschau and the underrepresentation such phrases in CompactTV.
Overall, the systematic differences of use in negative evaluative terms uncovers the channels’ respective ideological, critical position with regard to Covid-19 policies.
How is Covid-19 emotionalised
Observations of the annotations of the two semantic groups of emotions across the three channels.

Test result of correlation between channels and different semantic groups of phrases of emotionalisation.
The plot shows that shock/horror is used most dominantly in Bild TV and CompactTV, as reflected in the blue shading and positive Pearson residuals. In contrast, Tagesschau underuses this framing, shown by the small pink tile. Overall, the figure suggests that sensational emotional language – particularly the evocation of shock or horror – is more characteristic of Bild TV and CompactTV, while Tagesschau avoids such heightened emotional descriptors in news reporting.
Conclusion
This paper has demonstrated how a semi-automatic multimodal analytical framework can meaningfully broaden the empirical scope of narrative and ideological research on audiovisual news. By integrating automatic video segmentation, keyword-based linguistic detection, and computational identification of faces, movement, and scene types with manual annotations under a semiotic theory-based annotation scheme, we show that fine-grained narrative and discourse features in large-scale news video corpora can be analysed both systematically and reliably. The multimodal annotation scheme – grounded in social semiotic theory and validated through strong intercoder reliability – provides a robust foundation for capturing how linguistic, visual, and auditory cues interact to construct news narratives and perform ideologies. As a result, the approach bridges high-level ideological interpretation with quantifiable audiovisual evidence, offering a methodological model for future studies of news storytelling.
Our comparative analysis of public service, market-driven, and populist news channels reveals clear and systematic differences in how Covid-19 was narrativised across outlets. Each channel draws on distinct multimodal discourse patterns – foregrounding different social actors, selecting specific event contexts for interviews, and deploying contrasting patterns of evaluation and emotionalisation. Tagesschau largely avoids sensational framings and gives prominence to expert voices; Bild TV relies heavily on people’s queuing imagery and uses emotionalised and evaluative language related to shock, terror and craziness; and CompactTV amplifies oppositional stances, featuring anti-policy actors and using strong negative evaluations related to “dictatorship” and “coercion”. Together, these findings support the conclusion that audiovisual discourse patterns are not merely stylistic choices but operate as ideological performances that may then well influence audience interpretation, although establishing this connection more firmly would naturally require complementary recipient studies as well. The study thus underscores the analytical importance of multimodal features for understanding contemporary news communication and highlights the value of semi-automatic techniques for scaling narrative research to increasingly complex audiovisual news environments.
Footnotes
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research was supported by grants from the German Ministry of Education and Research (BMBF; Award ID: 16KIS1515K).
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
