Abstract
The COVID-19 pandemic has made videoconferencing tools an essential part of our lives as these tools are what allowed us to keep in touch in a time of social distancing. Having said that, however, many people have found virtual interactions to be surprisingly exhausting. This has given rise to the concept of Zoom fatigue. The purpose of this article is to explore the dynamics that give rise to this peculiar phenomenon. The article first discusses the concept of Zoom fatigue and critiques the brain-centrism of current explanations. It then proposes a more embodied approach to interaction, discusses the mediating role of technology in videoconferencing, and proceeds to presents a list of five videoconferencing dynamics that may induce Zoom fatigue: Awkward turn-taking, inhibited spontaneity, restricted motility, lack of eye contact and increased self-awareness. Finally, it is argued that these dynamics should make us temper our collective expectations about the hybrid future.
Keywords
Highlights
1. The COVID-19 lockdown has prompted discussions about Zoom fatigue 2. The present article seeks to understand this peculiar new phenomenon 3. It is argued that Zoom fatigue can be attributed to changed social dynamics 4. These dynamics impede the rhythmic coordination of our social interactions 5. The present article thus seeks to deflate current hype about the hybrid future
Introduction
In 2020, social venues in the form of offices, schools, cafés and bars had to be closed down overnight to contain the spread of a new coronavirus sweeping the globe. In the wake of this worldwide lockdown, all our meetings, lessons and casual hangouts had to be moved online. Prior to the pandemic, most people probably thought of videoconferencing as a handy way of maintaining long-distance relationships. After the outbreak, however, videoconferencing became the primary way of maintaining relationships with anybody outside of one’s immediate household. As a result, videoconferencing instantly became a crucial social and societal lifeline. ‘Video conferencing has become an essential communication tool for businesses, education and personal connections during the pandemic’, as an article in the Wall Street Journal put it (Morris, 2020). As Guenther’s (2013) harrowing study of solitary confinement shows, humans are radically social beings who will experience a kind of ‘social death’ when depraved of interpersonal contact, so we should certainly be thankful that videoconferencing enabled us to keep in touch in a time of social distancing. Having said that, however, many people have found virtual interactions to be surprisingly exhausting. This has given rise to the novel concept of Zoom fatigue.
In a time when more than six million people have lost their lives to the coronavirus, it may seem borderline sacrilegious to complain about Zoom fatigue, yet it remains relevant to devote scholarly attention to the phenomenon for three reasons: First, there is the sheer scale of the issue. While only 10 million people had attended Zoom meetings at the end of 2019, usage had already exploded to 300 million by April 2020 when lockdowns were effectuated (Grant, 2020). Second, while coronavirus may be worse than Zoom fatigue, it does not follow that Zoom fatigue is somehow trivial or irrelevant. Indeed, the experience is ostensibly bad enough for people to have coined a new term to describe it. The final reason concerns powerful actors’ discourses about the future of videoconferencing: Business leaders, for example, have claimed that remote working will eliminate commuting, boost productivity and save operational costs (BBC, 2020), UNESCO has stated that investments in remote learning will help us develop ‘more open and flexible education systems for the future’ (UNESCO, 2020), and the Zoom company itself has chimed in to argue that brick-and-mortar universities will soon be facing digital disruption (Lee, 2021). The glaring contrast between these optimist discourses of videoconferencing and the very notion of Zoom fatigue calls for critical analysis. Not to deny the (real) silver linings outlined above, but to shine a light on the (equally real) downsides. The rest of the article is devoted to doing so. The article first introduces Zoom fatigue and critiques the pronounced brain-centrism that undergirds current explanations of the phenomenon. It then proposes a more embodied approach to interaction, discusses the mediating role of videoconferencing technology and proceeds to presents a preliminary list of five dynamics that may induce Zoom fatigue: Awkward turn-taking, inhibited spontaneity, restricted motility, lack of eye contact and increased self-awareness. In the final section, the article compares these proposed dynamics to a recent account by Bailenson (2021) and discusses their practical implications. It is ultimately argued that the five dynamics identified here should make us temper our expectations about the hybrid future.
On the concept of Zoom fatigue
For scholars interested in the psychology of technology use, it is imperative to pay attention whenever new words like Zoom fatigue pop up into our everyday language to help us describe and make sense of our technological experiences. Such linguistic inventions often offer us important clues about the psychological significance of everyday technology use, and we therefore have to acknowledge, investigate and analyze the underlying lived experiences that lend credence to them. As an example, I have previously studied the contemporary phenomenon known as phubbing, which is a portmanteau of the words ‘phone’ and ‘snubbing’ that refers to the act of snubbing conversational partners in favor of one’s phone (Aagaard et al., 2021). A more recent example of this type of neologism is doomscrolling, the endless consumption of negative online news, which is a concept that also gained notoriety during the COVID-19 pandemic (Ytre-Arne and Moe, 2021). Like doomscrolling (and perhaps even because of doomscrolling), the concept of Zoom fatigue entered our collective vocabulary during the pandemic lockdown of 2020 due to its rapid spread in a spate of news articles (Google Trends, 2022).
The first part of the word, Zoom, designates a specific videoconferencing software that has indeed become emblematic of our lives during COVID-19, but the dynamics giving rise to Zoom fatigue are presumably shared across all videoconferencing tools (including Microsoft Teams and Skype), so the concept is best understood as an umbrella term that points beyond itself. The second part of the word, fatigue, has explicitly negative connotations and Wikipedia (2022) currently defines Zoom fatigue as ‘tiredness, worry or burnout associated with the overuse of virtual platforms of communication, particularly videoconferencing’. (In addition to this focus on quantity of use, however, this article will highlight the quality of interaction). While it is true that the concept of Zoom fatigue gained notoriety during the pandemic lockdown, we should not assume that all fatigue experienced in this period stems directly from videoconferencing, since much of this hardship can likely be attributed to the stressful experience of being isolated in one’s home during an unprecedented health crisis. For many of us, this situation corroded any semblance of work/life balance, and it felt less like ‘working from home’ than ‘living at work’ (Bagger and Lomborg, 2021). Homeschooling, health scares and job insecurities only exacerbated these issues. With these caveats in mind, however, it is still worth exploring the notion that videoconferencing itself can be exhausting, that participating in these online interactions somehow feels different. It is this idea that the rest of this article seeks to develop.
Brain-based accounts of Zoom fatigue
Zoom fatigue is still such a new phenomenon that it has primarily been discussed in news articles and blog posts, where it is mainly explained in neuropsychological terms (for an exception, see Sacasas, 2020). ‘Zoom fatigue is taxing the brain’, says an article in the National Geographic (Sklar, 2020). ‘The brain gets fatigued from overexertion’, says another article in Wired (Jackson-Wright, 2020). Not all these brain-based accounts are identical, however, and explanations of Zoom fatigue can be divided into two distinct camps: A deficit model and a transformational model. As the name implies, the deficit model argues that the problem with videoconferencing is that something vital is missing from these online interactions and that that ‘something’ is information. An article in The Conversation poignantly encapsulates this idea when it claims that the problem with videoconferencing is that ‘Our brain needs to do extra work to fill in the gaps’ (Hines and Sun, 2020). Appropriating a phrase from Chomsky (1980), we might say that the deficit model stipulates a ‘poverty of the stimulus’ in which a loss of information forces our brains to work overtime. Unfortunately, the deficit model never quite specifies exactly what information is missing from videoconferencing (is Zoom fatigue caused by the loss of smell, for instance?) or how the brain manages to fill in these ostensible gaps. Furthermore, the deficit model simply takes for granted that videoconferencing feels bad because it is exhausting (i.e., because it is mentally overwhelming), yet it may also be relevant to explore the inverse idea, namely that videoconferencing may be exhausting because it feels so bad (i.e., because it is socially underwhelming).
This brings us to the transformational model. Writing in Psychiatric Times, Lee (2020) implicitly draws on the transformational model when she asserts that fatigue is usually prevented by activation of dopaminergic pathways in the brain associated with reward, but ‘if the audio delays inherent in technology are associated with more negative perceptions and distrust between people, there is likely decreased reward perceived when those people are videoconferencing’ (np.). As is evident here, the transformational model of Zoom fatigue does not argue that the problem with videoconferencing is that it lacks something (like information), but that it changes something, namely the dynamics of everyday interaction. In making this argument, however, the transformational model effectively demotes brain activity to an epiphenomenon that is at least one step removed from the decisive question: Why is it that videoconferencing leads to these impaired dynamics in the first place? Ultimately, an account that is overly focused on the brain may be a hindrance to answering this question. 1 In what follows, I will therefore replace an intracranial focus with a more embodied approach to interaction.
From inner minds to rhythmic bodies
One promising venue for a more embodied approach to interaction is the phenomenological tradition. Merleau-Ponty’s (2002) phenomenology showed us that the body is not just a cylinder for housing the brain, but a living breathing entity that pulsates with life. Based on this idea, Merleau-Ponty dissolved the age-old problem of other minds by emphasizing that we usually have a direct understanding of other people based on their bodily behavior and expressions. ‘I do not see anger or a threatening attitude as a psychic fact hidden behind the gesture, I read anger in it. The gesture does not make me think of anger, it is anger itself’ (p. 214). 2 According to this framework, then, behavior is not just symptomatic of other people’s inner mental states but meaningful in itself. When understood accordingly, interaction starts at the level of mutually interacting bodies, or what Merleau-Ponty (1964) calls intercorporeality. Phenomenological scholars have since latched onto the concept of mirror neurons to explain intercorporeality (e.g., Carman, 2008; Dreyfus, 2012), so this approach in no way dismisses the importance of the brain. While mirror neurons may serve as the neurological basis for intercorporeality, however, they do not tell us much about the dynamics involved in this process. To get a closer look at such dynamics, we instead turn to the work of Daniel Stern.
Stern (2010) offers us the helpful vocabulary of vitality, which is a dynamic unit that arises from a combination of movement and its ‘four daughters’ of force, time, space and intentionality. When describing the vitality of an event, Stern argues, one is not describing its content or purpose, but its style, which is best captured through adverbs or adjectives like surging, gliding, tense, gentle and fleeting. As an example, think of the difference between tense and gentle laughter. In everyday speech, changes in vitality enacted through timing, pitch and stress procures the experience of talking to an actual, living person (as opposed to the flat and monotonous intonation of robots in old sci-fi movies). Furthermore, during everyday interaction, we often match and share vitality across sense modalities, or what is also known as affect attunement: We express the vitality of another person’s actions without imitating their exact behavioral expression. By doing so, we are effectively brought into synchronization with each other. At this basic level of intercorporeality, no mind reading is needed. Building on these insights, scholars have therefore begun to replace the ocularcentric metaphor of mind reading with more embodied metaphors like rhythms of dialogue (Jaffe et al., 2001) and communicative musicality (Malloch and Trevarthen, 2009). Instead of cognitive ‘in-sight’, these metaphors stress the fundamental importance of coordinating bodily movements over time, of getting ‘in-tune’ and ‘in-sync’ with each other.
The question concerning technology
Having shifted focus from inner minds to rhythmic bodies, we now need to take a look at the role of technology in videoconferencing. Although videoconferencing is temporally synchronous, it not immediate in the perceptual sense outlined above. Instead, videoconferencing can best be characterized as a form of ‘mediated immediacy’ (Lindemann and Schünemann, 2020). Traditionally, studies on mediated immediacy have focused on mediated presence, which refers to situations in which technologies render physically absent partners experientially copresent. The higher degree of mediated presence, the more technology recedes into the background, and vice versa. An absolute degree of mediated presence is defined as ‘a psychological state in which the virtuality of experience is unnoticed’ (Lee, 2004: 32) or as ‘the perceptual illusion of nonmediation’ (Lombard and Ditton, 1997:np). Essentially, then, most studies on mediated presence have focused on the extent to which technologies can become invisible. Here, the concept of Zoom fatigue invokes an important gestalt shift: Rather than asking whether technologies can become neutral conduits for human interactions, the question now becomes how technologies actively affect our interactions. This move toward viewing technologies as full-fledged difference-makers is indebted to theories of technological mediation (e.g., Latour, 2005; Verbeek, 2005). A basic assumption in these theories is that technologies do not just carry meaning from A to B but influence or ‘mediate’ how the world is present to us and how we are present in the world. When technologies are viewed from this mediational perspective, they cannot be taken as neutral tools, but must be regarded as active shapers of human perception and action (Verbeek, 2005). In other words, we want to examine how videoconferencing mediates our sense of immediacy.
While exploring technological mediation from the explicitly negative perspective of Zoom fatigue inevitably puts us on the path to technology critique, we will try to avoid the popular phenomenological strategy of juxtaposing the virtual domain with a supposedly ‘realer’ and more authentic domain. As an example of this foundationalist strategy, Borgmann (1999) contrasts natural, cultural and technological information and portrays the last of the three as a dangerous new phenomenon that threatens to overflow and suffocate ‘actual reality’. Similarly, Dreyfus (2009) contrasts fully embodied presence with what he calls ‘disembodied’ telepresence and disparages the latter. Although leading to thought-provoking critiques, this analytical strategy neglects the fundamental intertwinement of body, world and technology, and therefore risks regressing into digital dualism, which is the fallacy of perceiving the real/virtual as separate and distinct realities (Jurgenson, 2012). While we do want to acknowledge that unmediated interactions constitute our everyday baseline, we do not want to elevate these interactions to a higher ontological status. However exhausting they may be, virtual interactions are very real indeed. Eschewing digital dualism is no easy feat, however. When we colloquially refer to unmediated interactions as ‘real life’, for example, we might take a page from the deconstructionist playbook and ask, ‘As opposed to what?’ In other words, this designation effectively positions virtual interactions as less real or even unreal, which is exactly what we want to avoid. Other popular terms like ‘physical’ and ‘face-to-face’ interactions run into similar problems by implying that virtual interactions somehow lack these qualities. Perhaps the best we can do is thus to simply state our basic assumptions openly and explicitly: We do not mean to imply that videoconferencing suffers from inherent ontological lacks (i.e., that it is ‘unreal’ or ‘disembodied’), we simply wish to explore the empirical dynamics that might give rise to Zoom fatigue. Let us now proceed to do so.
Five possible dynamics of Zoom fatigue
Five proposed dynamics of Zoom fatigue.
This list is not intended to be exhaustive, and it is very plausible that future research may identify additional dynamics. While the list will be developed in a theoretical manner, I should say a few things about its underlying methodology: First, although it is preliminary and explorative, the list is not strictly inductive and the basic approach taken here, which is to explore the technological mediation of intercorporeal dynamics, is inspired by a previous study on phubbing (Aagaard, 2016). Furthermore, the list deliberately excludes two types of problems: Technological issues like breakdowns and human issues like people not muting themselves, not unmuting themselves, or not turning on their webcams. The list also does not touch upon the issue of distraction, which is of course a real temptation (Fosslien and Duffy, 2021), but one that is by no means isolated to videoconferencing (Aagaard, 2021). While all these issues are surely frustrating (after all, who wants to crack a joke in front of blank screens and dead silence?), the list solely targets dynamics intrinsic to fully functional videoconferencing. Working from a principle of charity, the goal is to analyze videoconferencing at its best. Second, although the list is partly based on self-observations and everyday conversations, the article has a theoretical rather than empirical ambition, which means that I will not be delving into my own lived experiences. Instead, for each individual dynamic, I will be consulting relevant theoretical literature as well as empirical studies from the field of human-computer interaction (HCI).
Awkward turn-taking
The first proposed dynamic of Zoom fatigue is awkward turn-taking. Normally, adult human conversations consist of dynamic interchanges of ‘turn-constructional units’ like sentences, clauses, phrases and words (Sacks et al., 1974). Each conversational turn is punctuated by a brief switching pause that begins when the speaker falls silent and ends when another person begins to speak. This rhythmic coordination of turn-taking constitutes the fundamental temporal structure of dialogue (Jaffe et al., 2001). Human beings are highly adept at turn-taking, and empirical studies show that the average gap between conversational turns is between 0.1 and 0.3 s (Holler et al., 2016). To achieve and maintain a conversational rhythm, however, turn-taking not only has to be fast, it also has to be smooth. If not, our conversations will either have too long gaps or constant interruptions. We therefore have certain devices for dealing with violations like overlap, the most common of which is for one of the speakers to simply ‘drop out’ and give up their turn (Schegloff, 2000). Of course, for this device to work, the speakers must first figure out who that person is. So how are these concerns relevant to the technological mediation of intercorporeal dynamics that occurs in videoconferencing?
One of the distinct characteristics of videoconferencing is that it involves transmission latencies or ‘lags’ that consist of delays between the production and perception of information. This brief excerpt from a 1993 study vividly illustrates how lag affects three people engaged in conversation: ‘When he hears A finish, C assumes that he can take the channel. However, because of the transmission lag, he is unaware that B has already begun to speak. Both B and C then drop out to allow the other to speak’ (O’Conaill et al., 1993:412). Even today, we continue to struggle with lags, and according to a recent study of videoconferencing, lags of about 0.7 s are enough to cause frictions that have negative effects on turn-taking (Seuren et al., 2021). The study also demonstrates how lags frustrate our overlap-resolution devices. By prompting false starts and overlaps, lags hamper smooth turn-taking and coordination of turn-allocation, and we thereby get stuck in a vicious circle of false starts and interruptions. The satirical newspaper The Onion thus hit the nail on the head when they published an article entitled ‘Coworkers on Zoom Trapped In Infinite Loop of Telling Each other ‘Oh Sorry, No, Go Ahead’. If the popular metaphorical understanding of interaction as a kind of dance is true, it seems fair to say that interacting on Zoom is like dancing to a broken record. To summarize, we can say that videoconferencing induces false starts, broken rhythms and failed overlap resolutions, and this dynamic makes such interactions feel strangely erratic.
Inhibited spontaneity
The second proposed dynamic of Zoom fatigue is inhibited spontaneity. This dynamic is closely connected to the awkward turn-taking described above. According to Schegloff (2000), a fundamental precondition for what he calls viable social organization is the basic conversational norm of one-speaker-at-a-time with conversational overlaps being one of two departures from this pattern (the other of course being silence in which nobody speaks at all). Importantly, however, this conversational norm only applies within individual conversations, which means that, in everyday meetings and gatherings, the main conversation will often break off into several parallel conversations or ‘side-conversations’ that all follow the basic conversational norm of one-speaker-at-a-time. Accordingly, it is normally possible to have a room full of what Schegloff calls ‘separate but ecologically near’ conversations (p. 4). So how are these concerns relevant to the technological mediation of intercorporeal dynamics that occurs in videoconferencing?
One of the distinct characteristics of videoconferencing is that uses a joint audio stream and ‘nondirectional sound’ (O’Conaill et al., 1993). Because all sounds of the virtual meeting are in focus at once, you cannot initiate a quiet side-conversation with the person next to you. In fact, you cannot address any one specific individual but are forced to address the entire assembly at once. Accordingly, videoconferencing can only encompass one main conversation, and if more than one person in this main conversation speaks, the joint audio stream will quickly devolve into auditory chaos. As a result, any informal communication ends up being centered on the present speaker, who is thrust into the spotlight and forced to deliver a sort of monologue. This issue is compounded by the awkward turn-taking, which can make other participants hesitant to speak for fear of interrupting. On top of that, the software itself will often place the speaker in visual focus on the screen. Ultimately, this interrogation-like setup leaves little room for casual banter and ‘watercooler chat’, but instead encourages formalized and sequential turn-taking. Indeed, O’Conaill and collegues (1993) found that the absence of turn-taking cues in videoconferencing leads speakers to use more explicit handovers in which one speaker express names the next. They also found that listeners showed a reduced ability to spontaneously take the conversational floor and describe a ‘lecture-like style of interaction that lacked spontaneity’ (p. 421). There is thus a certain levity and spontaneity to everyday interactions that gets bogged down in videoconferences. To summarize, we can say that videoconferencing encourages explicit and sequential turn-taking, and this dynamic makes such interactions feel strangely formal.
Restricted motility
The third proposed dynamic of Zoom fatigue is restricted motility. Movement arguably constitutes our most primitive and fundamental experience. As Pascal (1995) once argued: ‘Our nature consists in movement; absolute rest is death’ (p. 641). Although we seldom give it much thought, movement is crucial to the vitality of human interaction. In Tronick and collegues (1978) ‘still-face’ experiment, mothers were instructed to remain impassive and expressionless whenever their babies attempted to interact with them, and results showed that infants quickly became upset when interacting with still-faced mothers. In a different context, however, Foucault (1991) demonstrated that our movements are not completely free and self-determined but heavily regulated by what he called disciplinary power. In school, for example, students are constantly exposed to behavioral regulations like sitting still and raising one’s hand before speaking. Foucault compared such training to the dressage of horses: While spontaneity and unruliness may be the outset, orderliness is the outcome. He also stressed that disciplinary power is not just employed by humans but can literally be built into our material surroundings (as demonstrated in his famous example of the Panopticon). So how are these concerns relevant to the technological mediation of intercorporeal dynamics that occurs in videoconferencing?
One of the distinct characteristics of videoconferencing is that the webcam has a limited visual range. A webcam is remarkably fixed and inflexible in its demands on the comportment of its user, which entails being hands-on with the keyboard and face-to-face with the screen (Friesen, 2011). As such, the webcam instantiates an ongoing disciplining of the body in which we must sit physically still to remain within its frame. This inhibits, constrains and limits our motility. There is a certain unruliness to our everyday interaction that get eliminated. Such restraint may lead to physical fatigue on behalf of the speaker, but it may also render interactional understanding less intuitive to listeners. A further dimension of this problematic is that it leads to only our heads and shoulders being included in the frame. According to Heath and Luff (1992), it is crucial for the smoothness of interaction that the other person’s body language is allowed to unfold subtly on the periphery of our visual field, but this kind of perception is made difficult in videoconferencing. ‘Only occasionally are relatively gross movements, such as the other standing or blocking the screen, noticed and noticeable’ (p. 336). Ultimately, the empirical configuration of bodies and technology in videoconferencing results in people being transformed to mere talking heads. To summarize, we can say that videoconferencing restricts our freedom of movement, and this dynamic makes such interactions feel strangely static.
Lack of eye contact
The fourth proposed dynamic of Zoom fatigue is lack of eye contact. It is hard to overstate the importance of eye contact, which serves a host of social functions from tracking the behavior of others to assisting in turn-taking and expressing emotions (Kendon, 1967). Etymologically speaking, the word ‘eye contact’ means eyes touching (tangere) each other (com). Eye contact, in other words, does not signify distant visual impressions, but a felt sense of closeness. Indeed, it is often said that the eyes are windows to another person’s soul (Angus et al., 1991). While our auditory system is physically divided into mouths for speaking and ears for hearing, our eyes thus serve a powerful ‘dual function’ in that they both perceive and express information simultaneously (Gobel et al., 2015). What this means is that, of all the fascinating dimensions of human interaction, eye contact is one of the few cases in which we have direct reciprocal interaction between two qualitatively similar processes (Heron, 1970). As Simmel (1921) poetically put it, ‘The eye cannot take unless at the same time it gives’ (p. 358). However, Simmel also argued that eye contact is a brittle relation in which the smallest deviation ‘destroys the unique character of this union’ (p. 358) and presciently went on to argue that, if this were to happen, social relations would be changed in unpredictable ways. So how are these concerns relevant to the technological mediation of intercorporeal dynamics that occurs in videoconferencing?
One of the distinct characteristics of videoconferencing is the offset placement of camera and screen. The ‘geometry of video-conferencing’ thwarts direct eye contact (Bohannon et al., 2013). Put in another way, during videoconferencing, we cannot obtain true eye contact with our conversational partners because the webcam is placed above our visual focal point. What this means is that when you actually look at your conversational partner, you seem to be looking downward. Conversely, you can give the appearance of making eye contact, but doing so requires you to look away from that person and into the camera (Friesen, 2014). Dreyfus (2009) summarizes this infamous curse of the webcam accordingly: ‘You can look into the camera or look at the screen, but you can’t do both’ (p. 16). The effect of this changed dynamic is that something feels ‘off’ about videoconferencing, and we get out of touch with our conversational partners. Not in the literal haptic sense of the word, but in the intimate, interactional sense that eye contact is otherwise so remarkable at facilitating. To summarize, we can say that videoconferencing prohibits direct eye contact, and that this dynamic makes such interactions feel strangely disconnected.
Increased self-awareness
The fifth and final proposed dynamic of Zoom fatigue is increased self-awareness. Normally, our experience is characterized by self-forgetting in the sense that we are completely absorbed by the world. Sartre (2011) describes how, when I am running to catch the bus, there is no reflective self-awareness at stake, I am simply plunged into a world of objects with attractive and repellant qualities, ‘but as for me, I have disappeared’ (p. 13). However, Sartre (2003) also describes how the experience of being looked at by another person can effectuate a radical change in this everyday awareness. In his famous description of ‘the look’, Sartre describes an imagined instance of spying on a couple in another room by listening at the door and looking through the keyhole. While peeping voyeuristically through the keyhole, I am caught in what Sartre calls a pure mode of losing myself in the world, but when another person suddenly catches me in the act, I am abruptly torn away from this absorbed mode of being and instead become intensely aware of myself as an object to the other person’s gaze. I start seeing myself ‘from the outside’, as it were. The look is objectifying in the sense it changes me from an active subject to a passive object. While it is debatable whether Sartre was correct in portraying this antagonistic experience as fundamental to all intersubjectivity (‘hell is other people’), it is doubtlessly an eloquent illustration of self-awareness. So how are these concerns relevant to the technological mediation of intercorporeal dynamics that occurs in videoconferencing?
One of the distinct characteristics of videoconferencing is that the software projects an image of your own feed onto the screen. While there is an option to turn off this so-called ‘self-view’ in most videoconferencing programs, it does not appear to have been widely employed during the pandemic lockdown. At a time when our haircuts were arguably at their all-time worst, we were effectively pushed into an unprecedented focus on our own appearances. This is distracting in the sense that it draws (tracts) attention away (dis-) from the ongoing conversation: Instead of being absorbed by the interaction, we begin to pay attention to matters of self-presentation as described by Goffman (1990). We are moved from an everyday self-forgetfulness to an intense (and sometimes embarrassing) self-awareness that prevents us from becoming fully immersed in the ongoing conversation. Indeed, research suggests that seeing one’s own image on the screen during videoconferences can exacerbate negative emotions and interfere with one’s engagement in the primary task (Wegge, 2006). To summarize, we can say that videoconferencing is like carrying out a conversation in front of a gigantic mirror, and that this dynamic makes such interactions feel strangely objectifying.
Discussion
We have now looked at five taxing dynamics of videoconferencing: Awkward turn-taking, inhibited spontaneity, restricted motility, lack of eye contact and increased self-awareness. To summarize and synthesize, we can say that the combination of these five dynamics sets the stage for some downright exhausting interactions. The problems described here stem from a variety of sources including hardware issues (webcam placement), software issues (self-view) and connectivity issues (lag). Whether or not these issues will be fixed by future technological developments is an open question. As it stands, however, the issue with videoconferencing is not that it lacks something (such as information), but that it changes the dynamics of everyday interaction, which challenges interactional synchrony. As Gumbrecht (2004) writes about the concept of presence: ‘And if it became clear again that sitting together at a table for dinner (or making love, for that matter) is not only about communication, not only about ‘exchange of information’, then it might indeed become important and helpful […] to have concepts that would allow us to point to what is irreversibly nonconceptual in our lives’ (p. 140). Indeed, the framework presented here focuses on dynamics that are logically prior to concerns about information, meaning and interpretation. By taking these obvious but unnoticed dynamics and putting them into words, this article hopes to provide a framework for analyzing, discussing and making sense of videoconferencing. Hopefully, this helps us deflate overly grandiose notions of remote teaching, learning and working. Ultimately, we might do well to temper our expectations about the hybrid future.
Having said that, however, certain limitations to the study do follow. First and foremost, at this juncture in the article, it is worth reiterating that humans are radically social animals, and that videoconferencing remains preferable to being cut off from social contact entirely. Second, while we have purposefully focused on the limitations of videoconferencing, it does not follow that virtual interactions are inevitably terrible (for an excellent defense of virtual interactions, see Osler, 2021). Conversely, it also does not follow that unmediated interactions are always blissful and fulfilling. Indeed, many scholars can probably attest to the fact that regular meetings can feel highly draining and exhausting, too. The goal of this article was neither to vilify videoconferencing nor to glorify unmediated interactions, but to delve into the empirical existence of Zoom fatigue and try to understand why this peculiar phenomenon exists. In attempting to do so, however, we have only taken the first logical step, and the ideas proposed here need to be put to the test: Do people actually recognize the dynamics suggested here? Are they experientially resonant (Aagaard, 2018)? To answer such questions, we need empirical research that delves into the lived experiences of videoconferencing.
Perhaps, such empirical research could help us settle a burgeoning theoretical dispute: Since this article was initially submitted on 11 February 2021, a host of Zoom fatigue studies have been published, and while it is beyond the scope of this article to review and discuss these studies, there is one development that has to be addressed: On 21 February 2021, Bailenson (2021) published an article in which he gave four explanations for Zoom fatigue. While some of our arguments overlap (reduced mobility, the mirror effect), Bailenson notably posits that one of the main causes of Zoom fatigue is extreme amounts of intense eye-contact. ‘On Zoom’, he argues, ‘behavior ordinarily reserved for close relationships – such as long stretches of direct eye gaze and faces seen close up – has suddenly become the way we interact with casual acquaintances, coworkers, and even strangers’ (p. 2). Being the first of its kind, this account quickly rose to prominence in the news (e.g., Busby, 2021; Well, 2021), went on to become one of the American Psychological Association’s 10 most downloaded articles of 2021 (Palmer, 2022) and is now regarded as the go-to explanation of Zoom fatigue. As a result, it is worth pointing out that the conclusions drawn in this article directly oppose Bailenson’s interpretation: While he worries about excessive eye contact, I have argued that one of the major problems with videoconferencing is a distinct lack of eye contact. Which one of these accounts that gets the phenomenon right remains an open question, but the stark disagreement clearly indicates that important work on the subject remains to be done. In this endeavor, the article hopes to have charted a helpful course for future research.
