Abstract
In this paper we describe expertise as a way of seeing. We use match analysis `punditry’ as a setting to show how professional vision is interactionally achieved in TV sport broadcasts through environmentally coupled gestures enhanced by camera actions and a new technology of vision called telestrator. The paper is based on data from video sequences of (English) football TV broadcasts where the pundit shows to the TV host in the studio and to the non-expert audience at home what happened during a football match. We argue that the transparency of seeing is the product of an artfully instructed process whereby the pundit shows what should be seen, how it should be made accountable, and what the audience should expect in order to fully appreciate what they see. The paper shows how broadcasted match analysis expertise interactionally achieves this through the time-critical linking of talk, gesture, and technological environment.
Introduction: Match analysis and professional vision
This paper is about expertise as a way of showing how to see. In what follows we describe how match analyst expertise is displayed and the ways in which its public orientation is manifested to a television audience. We present match analyst expertise as a broadcasted multimodal activity involving talk, gesture and an optical technology called telestrator or video marker.
Unlike the objects of endoscopic (Mondada, 2014) or telescopic scrutiny (Garfinkel et al., 1981), the object of match analysis optical technologies is not seen for the first time. Match analysts use footages from televised football match that have already been watched. The work of the match analyst is that of rendering recognisably meaningful what audiences have already seen without noticing.
In this paper we argue that the work of match analysis in sport broadcasts is (1) based on an array of devices enabling the commentator to show the viewers patterns and connections in the spatial and temporal unfolding of the game, and that (2) the construction of this discourse of visibility is part of a demonstration of ‘professional vision’ (Goodwin, 1994). TV pundits enact professional vision by making visible and pointing to correspondences between an intended object of vision (an action in a match) and its technical or tactical meaning.
Match analysis is our case study. Match analysts show to each other, the TV host and the general audience how to see relevant actions in a football match. Broadcasted match analysis is the way in which a lay viewer is taken to appreciate the competent understanding of football actions in their constituent details. In our analysis we show how the match analyst imposes an organization to the constituent details of a visual pattern so that an order emerges e.g. from the distribution of the players on the football pitch.
This article is a contribution to the study of expertise in that it examines the work of experts in action. Contrary to the received view in sociology (Parsons, 1953) where expertise is broadly discussed in terms of how social groups come to define boundaries around their profession, we show expertise in how it is concretely enacted as a way to analyse and dissect visual elements.
The paper is structured as follows. In the next section, we review studies of the public dimension of expertise as well as studies of video practices and show how our case connects the two. We then articulate the conceptual relationships between video practice and expertise. In the empirical section of the paper, we analyze how TV pundits use the telestrator in sports broadcasts to diagram and analyze football tactics. A discussion section will conclude the paper.
Video practices and broadcasted expertise
There are two bodies of literature that are relevant for our study of match analysis expertise in sport broadcasts (Perry et a., 2019). On the one side there are studies of the public (i.e. broadcasted) dimension of expertise (Clayman, 2008) - see section “Expertise in public” below. On the other there is literature on video practices (see Broth et al., 2014), and video practices in sport in particular (Perry et a., 2019) - see next section “From instant replay to instant match analysis”. Our work expands the focus on broadcasted expertise by attending to match analysis expertise as video practice.
Thinking about the audiences for whom the expert is legitimated, Turner (2001) divides experts up in two main categories according to the way they obtain legitimacy. There are experts whose authority is legitimated only by restricted and/or pre-established audiences, like the physicist who are experts only for the community of physicists or the theologian who is legitimated by a restricted group of sectarian believers. There is another category of experts whose audiences are not restricted or pre-determined. These are experts who create their own audience and that they have to publicly prove themselves to these audiences by their actions. Moreover, they are not so different from the public with respect to their actual source of information. Among this category of experts are TV experts such as those studied by Hutchby (1995) and Raymond (2000).
Scannel (1991) studied forms of expertise in this category by looking at the management of expertise in broadcast talk and Hutchby (1995) in advice-giving in the radio. They did so by analytically addressing the ways in which that public orientation in the display of expertise is manifest in the organizational details of interaction. Similarly in our study we want to extend the approach to televised sport where talk combines with video practices.
Given the speed of so many of the actions that happen within a football match, televised sport has become a locus for the routine replaying of events and actions, to an extent that action-replay has become thoroughly interwoven with the coverage of major sports events. While the action replay is so commonplace, the systematic use in modern match analysis of electronic aids to back up informed tactical opinions stands out for its particular properties. The electronic aids commonly used in modern football match analysis consist of a video mark-up tool that allows the TV commentator to superimpose color lines indicating movement or direction onto the video footage. The tool clicker also allows the commentator to rapidly show a play, stop the action, back it up, and show it again.
From instant replay to instant match analysis
When looking at scholarship interested in sport broadcasting we recognize a primary focus of the expertise on talk (Delin, 2000; Ferguson, 1983; Kuiper, 1996; Kuiper et al., 2017). Ferguson (1983) describes sportscasting as ‘the oral reporting of an ongoing activity, combined with provision of background information and interpretation’ (Ferguson, 1983: 155–156). The skill exhibited by sportcasters is that they are smooth talkers (Kuiper, 1996). Fluency is regarded as an important aspect of the expertise of the sportcaster in that the talk and the accuracy of the depiction of events is produced together with the speed of the actions:
‘The task of a commentator of fast sports is to be fluent enough to keep up with the pace of what is happening in the visual field as otherwise the commentator would miss some episodes of the game. So the faster the sport, the more difficult it is for the commentator to speak and report immediately what is taking place’ (Kuiper, 1996).
An interesting body of interactional studies is emerging also in relation to sport broadcasts (Camus, 2015, 2017; Perry et al., 2019). Literature on video replay for example examines the threading together of visual image streams that are temporally separated (live and non-live video footage) under real-time conditions and how image work is coordinated within members of a television crew. Given the heightened temporal constraints under which live replay is coordinated, relevant sequences mainly cover events such as umpiring decisions or displays of skills and emotional involvement of individual players.
The growing demand in the audience for tactically informed commentaries combined with the orientation by major data companies to push tactical data from the realm of the esoteric to the everyday has produced a new breed of TV pundits that are keen to address the game in more detail. TV formats such as Match of the Day or Goals on Sunday are for discussions to take a different slant toward more engaging and deeper tactical analyses. Match analysts in this category of sport broadcasts combine the smooth talk expertise of sportcasters with the studio director’s ability to thread together talk and replay images under real-time conditions.
Expertise in public
In his study of sport refereeing, Collins (2010) describes the introduction of sport-decision technology and its implications for spectators, television viewers, and commentators. With being broadcasted, the display of referee expertise shifts from presumptive to ‘transparent’ (Collins, 2010). Presumptive is when one has good reasons to assume that expertise could be made visible if only one was in the position to see it (see also Raymond, 2000). When the display of expertise becomes televised, the mechanisms through which experts establish their authority from being implied become open to scrutiny. A shift from presumptive to transparent expertise applies to football analysis as well. After having practiced it for decades in the locker room or in the backroom of elite coaching staff, transparent expertise means that high status football professionals turned analysts are in fact seen to show match analysis expertise in front of a camera.
To demonstrate match analysis expertise to an unspecified audience outside the professional field (Lymer, 2009) is a challenge to professional vision (Goodwin, 1994). In match analysis, this requires not only that the analyst displays the specialized technique of tactical video analysis. Analysts attending sport broadcasts should also make accountably visible (Lynch, 2006: 97–98) the ability to pair images and other instructions in a sequence that makes tactical nuances accessible to the generic viewer. In doing so, the scopic system that affords assessment of a football match (i.e. the telestrator) is re-oriented and its qualities as a mean of instructing vision are topicalized.
In the remainder of the paper, we show how commentators shift between different modes of scoping access to the details of play in televised match analysis. We document for example how switches between the video clips operated by the analyst and the camera mounted in the TV studio correspond to a shift between different projections of the commentator stance towards the event e.g. from descriptions to evaluations. Through a detailed analysis of interactional sequences in which the mundane object of vision is transformed in an instructably observable arrangement (Garfinkel, 2002: 211) that is openly inspectable by others in its constitutive details, we show how match analysts help the audience tease out ‘the animal from the foliage’ (Garfinkel et al., 1981: 132). In so doing we extend the notion of ‘instructed vision’ (Licoppe and Tuncer, 2020; Majlesi, 2018; Tuncer and Haddington, 2020) to the domain of broadcast football analysis.
Data and method
The empirical part of the paper analyses how TV pundits use the telestrator in sports broadcasts to diagram and analyze football tactics. A telestrator (or video marker) is a device that allows its operator to draw a sketch or overlay the moving or still video image with shapes (ovals, arrows, halos).
Our dataset initially consisted of 30 minutes of match analysis video clips from TV sport broadcasts covering various games and involving various analysts.
In this paper we have selected just one case, taken from Sky Sport Monday Night Football. We focus on two different kinds of expertise on display: ability in the use of the equipment and knowledge of the game. In this case the match analyst is James (‘Jamie’) Carragher, a retired English footballer turned TV commentator. Carragher is analyzing one episode of a match between Manchester United and Chelsea played on the 26th of October 2014. Chelsea was playing an away game and was winning 0–1 till 4 minutes into extra time. At this point Manchester U. won a free kick on the right side of the goal defended by Chelsea. As the BBC sport writer Phil McNulty described what happened then: ‘Di Maria’s resulting free-kick saw Chelsea goalkeeper Thibaut Courtois save brilliantly from Marouane Fellaini but Van Persie was on hand to thrash home the rebound and spark wild celebrations around Old Trafford’. 1 Manchester U. was able to equalize the game well beyond the regular time.
The whole episode (from the free kick to the goal) lasted about 12 seconds. In our analysis we will show how Carragher – the match analyst – is going to dissect these 12 seconds in order to explain to the viewers (and to the TV host and the other analyst in the studio) what happened on the pitch: how was it possible that a goal in this occasion has been conceded? Is there anything in the positions of the players that would allow the viewer to better understand the dynamics of the events? Is there any tactical football feature to be disclosed so that the viewer can better appreciate the details of the action? Is there any detail in the footage that if shown would make what appears – even after repeated views – a chaotic assemblage of an organized performance?
Analysis
In this section we focus on three modes of doing descriptions: (1) comparing images, (2) framing a visual Gestalt - when relevant details are put in context as a reflexive pattern of figure and background - and (3) prospective vision i.e. when an image is described by anticipating the future course of action.
Comparing images
In this example, the match analyst shows how a video clip can be compared to another. The two clips being compared are the one leading to the goal and a similar episode earlier in the game. In both occasions an indirect free kick has been awarded to the attacking team in red shirt (Manchester United). The other team in blue shirt (Chelsea) is defending the goal. In order to compare the position of the players the screen has been split in two parts. Each part has been frozen at the exact moment in which the player in red shirt (bottom of the screen) is preparing to take the free kick (Figure 1).

The split screen.
The left part of the screen shows an earlier action in the game, while the right part shows the sequence leading to the goal. The analyst wants to show that the positions of the defending players in blue shirts in the occasion of the goal (right half of the screen) put the attacking players in red shirts in a better position compared to what happened earlier. In the right half of the screen one player in red shirt takes advantage of the position of the defending team and scores a goal.
In our analysis, we will show how the analyst actively demonstrates similarities and differences in the player positions in the two occasions through a combination of talk, oriented gaze, gesture, sound, and image – including the split screen image. In the Excerpt 1.1 below the analyst uses the split screen to compare images:
Ex. 1.1.
2


The camera immediately switches from full screen (#1) to a format where the telestrator’s toolbar is visible (#2). In this way the audience is given visual access to the actions that the analyst takes on the telestrator.
The analyst starts to demonstrate similarities and differences between the two footages by looking at the position of two players in blue shirt (Chelsea, the defending team) (l. 29, #1). He points the touch pen to the player to the bottom of the left half of the screen, and then to the player to the bottom of the right half of the screen (#2), naming the player ‘Fabregas’. The name of the player is equated to a position (‘role’) on the pitch that is shown to be the same in the two situations. The touch pen generates a halo effect, circling the player with a persisting round mark 3 .
Then (l. 31) the analyst points to two other players: one to the left half of the screen (#3) and names him ‘Willian’ and the other to the right. He notes that while the player to the right is different (‘Mikel’), it appears that he is occupying the same position (‘role’) as Willian in the earlier free-kick (l. 31, #4).
Note that while marking the similarities between the cases, the analyst uses the term ‘role’ to refer to the player positions (l. 29 and 31). ‘Role’ in this case is to be understood as a relational term, referring to the position a player occupies on the pitch at a given time and in relation to the other players. We argue that differently from what could be called ‘coding’ - that is, organizing the world into categories (Goodwin, 1994: 668), here the analyst avoids using tactical categories (for instance, he does not use the terms ‘defender’ or ‘striker’). Sense-making in this case does not require the viewer to know tactical categories. It immediate perceptual apprehension: seeing what is happening on the field in certain ways is understanding what is happening. The analyst uses a pointing device to unveil positions on the pitch, exhibiting positions in the visual field. He attracts the attention of the audience toward specific features of the visual field, rendering them relevant in order for the viewer to see the two pictures as displaying similar organized configurations of players
In the following excerpt we analyze the interplay of visual display and verbal description.
Ex. 1.2.


In the above excerpt, the analyst starts drawing an arrow around a group of blue players lined up along the edge of the small box. He draws an arrow around the players saying ‘these six giants here for Chelsea’ (l. 32) and then a second drawing saying ‘the six giants there for Chelsea’ (l. 33). It is important to note here how the organization of the verbal description matches exactly the visual display (‘here’, ‘there’; l. 32–33, picture #7).
In the figure below, we magnify clip in #8 to show that while the arrow to the right comprises all players on the edge of the box, the first player in the line of blue players is left out from the arrow drawn to the left (In figure 2).

A close up to show that there is one man missing in the picture to the right.
It is apparent in this excerpt how the unveiling of certain similarities and differences in the visual field is achieved through the simultaneous use of language, gesture, and the features of the video marker which mutually elaborate each other. Deictic terms such as ‘here’ and ‘there’ could not be worked out without this multimodal package of complementary meaning-making practices – an example of what Charles Goodwin calls ‘environmentally coupled gestures’ (Goodwin, 2007: 55; see also Arminen and Auvinen, 2013).
At this point the commentator starts talking about similarities and differences in the players’ positions. Cleaning all markers on the screen using the erase button on the telestrator toolbar (figure #9 below), the pundit asks what the difference between the two images is (l. 34). He then points the pen to the first player in the wall to the left screen and says: ‘the difference is they’ve got no one in that role that Oscar is taking up there, in this position’ (l. 35–36).


What is important to note here for the purpose of our analysis is that the pundit does not just display knowledge regarding the name of the players and their role in the game. He also shows the ability to spot differences between game situations. He does so first by pointing the pen to a player to the left half of the screen (drawing a ring around him; see figure #12). He then points to a void in the right half (#13). In this way the pundit is making concretely visible an absence that finds its meaning through reference to a previous situation.
Comparison is the way expert vision is demonstrated in publicly accountable ways. By spotting a difference and by marking it with an evaluation (i.e. “that’s the big difference”), the analyst shows what to see and why what is seen matters for the understanding of the game. With his expert ways of looking, the analyst brings distinction between two apparently identical crowds of bodies in colored shirts, perceptually orienting the audience towards features that are relevant for the understanding of the situation.
To summarize, so far we have shown that making images comparable through selecting relevant features is one element of what makes match analysis a form of visual expertise. In the next section we will discuss the ability of the analyst to put relevant details in context as a reflexive pattern of figure and background.
Analyzing details in a visual Gestalt
Another mode of expert vision is that of identifying details and refer to them as an organized Gestalt. Expert vision in this case consists in showing how an organization emerges from the constituent details of an image.
In the following three excerpts the analyst proceeds to examine the development of the situation seen earlier. When the ball is being crossed in the box, nearly all the players on the pitch are crowded in the penalty area. The players in blue shirts (Chelsea team) are defending their goal. The players in red shirts (Manchester U. team) are attacking. The footage has been frozen when the ball has just been kicked.
Ex. 2.1.


The analyst starts by focusing on just two players: Rojo (in red shirt) and John Terry (in blue shirt) (l. 112). The analyst uses the touch pen to orient the audience toward the two players. A ring appears around them that turns out to be a magnifying bubble (#15). The appearance of the magnifying bubble over the frame is precisely coordinated with the words ‘I want to highlight’ and a switch of the camera from full-screen game footage to the telestrator dashboard. Figure 3 shows from close-up what’s inside the bubble.

A close up showing Rojo (lighter gray) ‘blocking’ John Terry (darker gray).
The magnifying bubble makes the two players stand out from the rest and a closer scrutiny is possible.”Also, the image is available to the audience in full screen without the toolbar to frame it (#18).
The analyst is now on camera talking to the TV host (#17). He is going to tell why he used the magnifier effect. He previously mentioned the ‘six giants’ of Chelsea: the six players organized in a defensive line along the edge of their six-yard box. Here the focus is just on one of them (John Terry). According to the analyst, Terry is one of the two most important defensive players in the blue team when it comes to prevent the red from scoring a goal when the ball is in the air (the other one would be Gary Cahill).
The point here is that showing an image bigger is not in itself what it makes the audience see it better. It is the coupling of the magnifying effect with the reasons as to why the particular effect has been used that makes the picture meaningful. First, by highlighting the two players (John Terry and his direct opponent, Rojo) the commentator separates them from the rest of the crowd. By considering the duel between the highlighted players, he then prepares the audience to appreciate the relevance of what is being described. Finally, the analyst instructs the audience to see that the magnifying lens is oriented to appreciate how well the specific player (i.e. Rojo) is doing in this duel.
The analyst then describes and evaluates the action of John Terry direct opponent in lighter gray shirt: ‘Rojo does fantastic, blocks him [John Terry, in dark grey shirt], stops him winning the ball’ (l. 115–116). Line 1.115 contains an overt evaluation of Rojo’s action (‘fantastic’). It is important to note here how the camera actions accompany the switch from descriptions to evaluations. As soon as the analyst expresses an evaluation, that is, ‘Rojo does fantastic’ the body of the commentator becomes visible from the waist up. Once the conditional relevance is set for what has to be seen next, the analyst looks again to the screen (#17) and the camera switches back to the game footage (#18). This is an example of how camera actions contribute to make visually relevant a selected detail of an image by reflexively tying it to a verbal description (the activity of ‘blocking’, ‘stopping’) and to an evaluation of that action (‘fantastic’).
In the next section, we complete our account of the analyst’s ability to produce a visual Gestalt by looking at how individual player’s physical posture is topicalized in relation to the game situation. In Excerpt 2.2., the pundit turns to examine the position of another key player in blue: Gary Cahill.
Ex. 2.2.


The analyst invites the audience to ‘have a look’ (l. 117) and then moves the magnifier lens toward the other end of the six-yard box (see Fig. 4).

Detailed view of Cahill (blue) ‘on stilts’.
This time the audience is not asked to focus just on one player inside the bubble but to make a comparison with other players around him. The invitation to make a comparison is accompanied by a movement of the magnifying bubble over other players in the box. The analyst verbally focuses on a feature of the player under the lens: ‘he looks like he is on stilts’ (l. 118).
The pundit is acting as an expert who is teaching how to look: with the help of a lens that magnifies details that would have been otherwise overlooked and by offering a metaphorical (‘he looks like he is on stilts’, l. 118) and then literal (‘he so far higher than the other players’, l. 119) description of the player’s posture, he is instructing the audience to see that one player is taller than the others. The player’s physical feature of being taller is made accountable through constant reference to the local context and the other players.
Once the point has been made completely transparent in terms of its actual perceptual understanding (through enhanced visual access coupled with descriptions) the analyst provides reasons for this apparently `weird’ physical appearance. As seen earlier, here again the camera marks a move to the more evaluative mode by turning to the analyst in the studio. Offering reasons doesn’t need visual access to the phenomenon.
Ex. 2.3.
What is interesting to note here is how the player’s weird physical appearance is made accountable as a plausible body shape. The analyst first identifies some unusual visual features of the image, allowing the viewer to single out one player among the others for his appearance. Then he explains that this is not a physical attribute (i.e. ‘being on stilts’). Accompanying it with a sort of re-enactment (Sidnell, 2006), he concedes that the player appears so much taller than others because of the ‘the bouncing’ (l. 121: ‘bouncing up and down’, l. 123: the player ‘is caught in the air’). The analyst subsequently introduces some reasons to make this ‘thin’ description of the player behavior fully accountable. The reason is that the player is clearing his view because from his position he cannot see the ball (l. 123: ‘there is that many bodies’). With a ‘thick’ description, the analyst renders the action of the player that of an act of ‘jumping’ (l. 130).
Let’s have a look to the epistemic basis of the description offered by the analyst (‘I have been there myself’, l. 124). It is interesting to reconsider here what Collins (2010) says about the `epistemological privilege’ of decision-aid technology in sport. One of the sources of the epistemological privilege that confers authority to whoever uses them for Collins is, quite literally, the superior view of the camera: an elevated position that provides a better view of the action on the field. Rather differently, in our case the analyst’s epistemological privilege is construed in an interplay between the ‘superior’ view of the camera - that is, when the match is seen from the vantage point of the skycam above the pitch with no obstacles to impede the panoramic view of the play - and the `inferior’ view linked with direct experience: ‘I have been there myself’ (l. 124). It is only through first-hand experience of the player’s view from the ground (‘you can’t see the ball’, l.125) that ‘being on stilts’ can be understood as a result of ‘bouncing’. What might appear to be a weird posture (i.e. ‘being on stilts’) can now be seen as an ordinary action in the game of football (‘he’s obviously jumping’, l. 130).
To summarize, in this second section on expert vision we have seen how relevant details of an image are put into context. We have shown how the bodily appearance of a player is explained by putting it in relation to the surrounding players (the player is taller than the others) and to the game situation (the player is jumping to see the ball). The expert description allows viewers to understand a potentially problematic visual feature in and as a reflexive pattern of figure and background: ‘being taller’ is made accountable by the fact that ‘he is jumping’, and the ‘jumping’ explains and makes the apparently weird posture obvious. Expert vision in this case consists in showing how an organization (i.e. fighting to reach the best position for winning the ball in the air) emerges from the constituent details of an image, where at the same time these details make the whole picture apparent.
Prospective vision
Next we describe a third mode of expert vision: anticipating actions and moves on the pitch so that the viewer already knows what is there to be seen when the clip is played.
In Excerpt 3, the analyst is still focusing on the stoppage time goal scored by Man United. This time described is the trouble the defending team (blue shirts) encounters when a free kick is taken with the technique of the ‘outswing’. The analyst focuses on the trajectory of the ball has been kicked. He starts by referring to one player of the defending team (‘he’, l. 167, is the same Gary Cahill we have encountered earlier) and ton his ‘problem’. The problem is that the cross is an ‘outswing’: a cross taken with the internal part of the left foot so that the ball’s trajectory goes toward the goal before arching back away from it.
Ex. 3

The commentator stops the clip when the ball is mid-air and draws a white arrow with the video marker to show the trajectory of the ball. The analyst utterance of the word ‘outswinging’ is finely attuned to the completion of the drawing: the analyst drags to make the word ‘outswinging’ end exactly at the same time as the drawing.
The arrow makes apparent the full trajectory of the ball from where the free kick is taken to the head of the player where the ball will eventually land. Obviously, this is a replay of the action: this particular clip has already been watched several times. The viewers are knowledgeable regarding what it is to be expected next. But this is the first time that this particular visual feature of the action (the particular trajectory of the ball) and its practical consequences for the game are brought to the viewers’ attention. In this way, the analyst is making visible a ‘seen but unnoticed feature’ of the scene under scrutiny. Making the end result visually and verbally available to the viewers prior to being shown in the clip (169: ‘so it’s coming to Fellaini’), the analyst connects discrete elements of the picture in order to render understandable the temporal development of the action. It is the drawing of the outswinging trajectory of the ball that shows that the defenders find themselves in a weaker position compared to that of the opponent. The attacking player (i.e. Fellaini) is shown to be eventually able to jump and make a successful header before it actually happens.
The expert description and drawing here are a kind of foresight of a future state of affair, the prediction of an expected result achieved by introducing a stable course of action where the contingencies could have produced very different outcomes. The expert description consists here in providing a visual foresight whose outcome will be subsequently confirmed in the video clip.
The environmental coupling of visual sport punditry
In this section we come back to how the epistemological conditions afforded by the introduction of optical technologies in sport broadcasts affect the display of match analysis expertise.
With sport punditry, a pattern of play in a football match emerges as an object of vision through discursive and camera practices. Unlike coding practices in scientific disciplines when talk is addressed to an apprentice, the encounter between talk and image in televised match analysis is not organized by any system of inscription of the kind of a Munsell Chart (Goodwin, 1994: 609). As shown in our case, perception is organized by mundane acts of simple visual comparison between similar footages as made available by the split screen image. This is because the unit within which the intersubjectivity of football analysis is lodged does not only include advanced trainees but also members of the audience of a sport broadcast that are not necessarily trained as football analysts. The TV commentator is indeed expected to also ‘perceive the perceptions’ (Goodwin, 1994: 619) of an audience of lay people and be able to describe for them what is happening at a football match.
Another difference from similar studies on environmental coupling of camera and talk such as Mondada’s study of surgical work (Mondada, 2014)is that the camera is operated directly by the speaker. As a consequence, instructing vision in the case of match analysis does not take the form of directives and requests as in Mondada’s case where the surgeons’ talk directs an endoscopic camera operated by an assistant in the operation room (Mondada, 2014). To this respect, a great deal of background work in the TV studio backroom remained invisible to our analysis. We do not know who chose and edited the video clips made available, how the material has been selected and prepared, how decisions were made as to what material was deemed relevant for the audience and who made them.
Specifically, our data show three different modes that together describe how match analysts interactionally enact expertise through technologically-enhanced environmentally coupled gestures (Goodwin, 2007: 55).
The first mode of scoping access is the commentary of the game, where the video clip of the match is normal speed and full screen. The role of the analyst in this first mode is not remarkably different from that of an ordinary sportcaster (Delin, 2000; Ferguson, 1983; Kuiper, 1996). Talk accompanies events as they unfold and narration is composed of time critical utterances, which occur at the time of play and serve to describe it (Delin, 2000).
The second is the demonstration of expert conduct where match analyst expertise is seen to be shown as the specialist skill of coupling words, images, and video markers in a time-critical fashion. In this second mode, the video clip is shown within the telestrator dashboard frame. The hand of the commentator appears gesturing on it with a touch pen.
The time-critical coordination of talk and screen touches is key to achieve meaning-making in this mode. While the flow of the commentary is slowed down together with the video image arguably making the temporality of the action a more ‘docile object’ (Lynch, 1985: 43–44), the sound of the crowd is still audible in the background. The computer-generated crowd noise soundtrack stands to emphasize the time-critical coordination of talk and gesture on the video marker and that there is an element of multimodal sequencing that makes somatic skills central to the display of expertise in this mode.
For what concerns superior view, sometimes neither the view from above of the video images nor the telestrator tools are sufficient to account for how a player sees the game. It is when the camera switches to the wider angle of the TV studio and the body of the commentator becomes visible from the waist up that the crowd noise fades and we enter a different mode. That’s the ground level view that players have when for example they look for the ball and are covered by other players: a view that one can have only from being there. The switch of the camera marks the end of the environmentally coupled description in real time and the beginning of the specialist evaluation of what has just been described. It is only in this more explanatory mode that the analyst mobilize evidence also from offline sources including his credentials of former professional footballer.
Conclusions: Instructed vision in the case of match analysis
This study examined a case in which the taken-for-grantedness of visual data is achieved as a result of artful practices of instructed viewing. What is there to be seen is the result of a highly specialized technical eye that scans the visual phenomena, describes the bodily configurations in the visual field and illustrates them to a general audience through discursive and technical apparatuses. Ours is a case where a professional sport commentator instructs a lay audience on how to see visual configurations in their relevant details. Our analysis is reminiscent of Garfinkel’s idea that the recognizability of actions relies on these situated instructed actions being made openly inspectable by others:
‘The idea is this: worldly objects, as of the cogency and the cohesion of details, are available in the looks of organizational Things. If not, then where else in the world are you going to find them? Ethnomethodologically, they are available in an instructably observable arrangement, of apparent details – of details in and as their coherence producedly provided for’. (Garfinkel, 2002: 211; italics not in the original).
The instructed vision in the case of match analysis is organized through the construction of a spatial layout in which the elements of the game (the players, their actions and their reciprocal interactions) are made immediately visible and apparent, revealing aspects of the ‘endogenously produced coherent appearances of Things’ (Garfinkel, 2002: 211).
As it is for all practical productions, the elements of the game ‘are all highly observable ones (“practical” here having to do with accountably observable)’ (Baccus, 1986: 3).
By showing clear spatial relationship between relevant players, so that they are easier to identify, the commentator does impose a visual order to the scene.
In professional fields like surgery, the instructed action is aimed at teaching novices how to perform very skilful actions on the body of a patient (Mondada, 2014). In our case expert vision become apparent in that it offers the audience an instructed way of seeing the game, in the same way natural objects are made visible and analyzable in scientific research (Lynch, 1985) . Simultaneously, the way transparency of vision is achieved through the digitally enhanced manipulation of video clips in real time contributes to define the expertise in the new field of video analysis in sport broadcasts. This finding resonates with Goodwin when he says that ‘Discursive practices are used by members of a profession to shape events in the domains subject to their professional scrutiny. The shaping process creates the objects of knowledge that become the insignia of a profession’s craft: the theories, artefacts, and bodies of expertise that distinguish it from other professions’ (Goodwin, 1994: 606).
Footnotes
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
