Abstract
Research has shown that people’s gaze is biased away from faces in the real world but towards them when they are viewed onscreen. Non-equivalent stimulus conditions may have represented a confound in this research, however, as participants viewed onscreen stimuli as pre-recordings where interaction was not possible compared with real-world stimuli which were viewed in real time where interaction was possible. We assessed the independent contributions of online social presence and ability for interaction on social gaze by developing the “live lab” paradigm. Participants in three groups (N = 132) viewed a confederate as (1) a live webcam stream where interaction was not possible (one-way), (2) a live webcam stream where an interaction was possible (two-way), or (3) a pre-recording. Potential for interaction, rather than online social presence, was the primary influence on gaze behaviour: participants in the pre-recorded and one-way conditions looked more to the face than those in the two-way condition, particularly, when the confederate made “eye contact.” Fixation durations to the face were shorter when the scene was viewed live, particularly, during a bid for eye contact. Our findings support the dual function of gaze but suggest that online social presence alone is not sufficient to activate social norms of civil inattention. Implications for the reinterpretation of previous research are discussed.
Humans cannot help but look at faces, so the majority of the research into social attention has suggested. This work has largely come out of traditional cognitive psychology paradigms adapted to study social processes (Birmingham, Bischof, & Kingstone, 2008; Birmingham & Kingstone, 2009; Fletcher-Watson, Findlay, Leekam, & Benson, 2008). Although these studies have informed our understanding of the perception of social stimuli as largely inanimate objects, in recent years, some researchers have begun to question whether such frameworks are really able to tap into genuinely social processes (Nasiopoulos, Risko, & Kingstone, 2015; Risko, Laidlaw, Freeth, Foulsham, & Kingstone, 2012; Risko, Richardson, & Kingstone, 2016). The fundamental problem is that everyday social behaviour does not occur in situations where a “lone observer” views others without even the possibility of an exchange with those individuals (Risko et al., 2016). Yet, this critical function of social behaviour—to reciprocate with others—is simply not possible in traditional cognitive tasks. If paradigms wherein participants free-view photographs of others are really examining social cognitive processes, then viewing Angelina Jolie in Hello! magazine while you are alone in your bedroom should represent the same experience as if she were standing in front of you. Intuitively, we know that this is not the case and there is mounting evidence from cognitive and neuroscientific research to support this notion (Cavallo et al., 2015; Hietanen, Myllyneva, Helminen, & Lyyra, 2016; Myllyneva & Hietanen, 2015; Pönkänen, Alhoniemi, Leppänen, & Hietanen, 2011; Redcay et al., 2010).
Differences in social gaze between the lab and real world
Recent research has suggested that viewing others in real-world scenarios may alter the manner in which individuals deploy their attention to one another compared with when a lone observer simply views a recording of the same individual. For example, Foulsham, Walker, and Kingstone (2011) showed that in real-world scenarios, people look less at others’ faces than when viewing them as pre-recorded videos on computer screens. This has recently been corroborated by the work of Kuhn, Teszka, Tenaw, & Kingstone (2016). Similarly, Laidlaw, Foulsham, Kuhn, and Kingstone (2011) found participants were far less inclined to look at a real person sat in a waiting room than when that person was displayed on a screen in an otherwise identical setup.
The disinclination of people to look at strangers for extended periods is not a new suggestion. The theory of “civil inattention,” originally proposed by sociologist Goffman (1963), describes the amount of attention considered appropriate to show to strangers when encountered in public spaces; enough to acknowledge their presence (e.g., a brief glance) but not so much as to indicate that they are of special interest (e.g., not staring). This social norm of not showing excessive interest in others is tacitly adhered to in public spaces and violations of it are viewed negatively (Ocejo & Tonnelat, 2014; Zuckerman, Miserandino, & Bernieri, 1983). What is implicated in Goffman’s original proposal is the assumption that when encountered in authentic social situations, gaze is a powerful social signal.
The idea that human gaze serves multiple functions is not new (see Kleinke, 1986 for a review). One theory which has recently been revived and applied to social attention literature is the dual function of gaze (Argyle & Cook, 1976; Gobel, Kim, & Richardson, 2015; Nasiopoulos, Risko, & Kingstone, 2015; Risko et al., 2016). The theory posits that gaze serves two main functions: to perceive information and signal to others. This concept may explain the discrepant findings in social gaze between lab and real life. The suggestion is that when participants view a photograph of a fellow human their gaze can only fulfil the first of these functions as there is no one behind that image for their gaze to act as a signal to. Hence, the dual function theory suggests that participants have the freedom to view this highly rewarding stimulus as they wish without fear of their gaze being observed and therefore their interest being communicated (Gobel et al., 2015). This view would account for the face bias reported in countless social attention studies. In a genuine social scenario however, where one’s gaze is capable of both perceiving and signalling, people tend to avoid gazing at others, as they are reluctant to signal that their attention is directed at the other: instead, they adhere to “civil inattention”(Risko et al., 2016).
Limitations of previous research
The recent work on real-world social attention has added immeasurably to our understanding of human social behaviour in the wild. However, some very important issues have been overlooked in discussions of this topic. On closer inspection, it appears previous studies may have been confounded by comparing gaze behaviour towards pre-recorded scenes shown on computer monitors without the means for social interaction, to that occurring during real-time face to face scenarios where interaction is possible. Specifically, in the screen conditions in Laidlaw et al. (2011) and Foulsham et al. (2011) the stimulus was also viewed via a different medium to that in the live condition. Therefore, not only has the potential for social reciprocity changed from the equivalent real-world scenario, so too has the medium through which that interaction might occur if it were even possible. Viewing an image on a screen may be enough to increase looks to the face: on a small display even a complex stimulus is nevertheless still a pre-selected complex stimulus. This effect may be further enhanced by the very purpose of a screen: it is designed to be viewed and may attract attention simply because it serves no other purpose. Therefore, in order to be confident that these gaze effects are genuinely due to social influences and not simply differences in stimuli, the stimuli themselves must be kept constant across conditions.
A further previously overlooked issue is the reason that the participants cannot interact with a stimulus onscreen is due to two, entirely dissociable factors, either of which could result in increased bias to gaze at the face and which will be described in turn below.
The first issue to consider is the physicality of the screen. In previous research (Foulsham et al., 2011; Laidlaw et al., 2011) there have been no means of communication between stimuli and participant because the stimuli have been viewed through a spatial barrier of the monitor without an audio-visual link. However, it is entirely possible to view someone onscreen and also have the ability to interact with them, as anyone who has used video conferencing will know. Enabling this ability may profoundly affect gaze behaviour. In an attempt to directly assess the impact of potential social interaction on social viewing behaviour, a classic study by Argyle and colleagues (Argyle, Ingham, Alkema, & McCallin, 1973) showed that participants were more inclined to look at an unacquainted other when viewed through a one-way mirror, than when viewed face to face. The authors suggested that social gaze acts to control the level of intimacy between two individuals. In a situation where the observed cannot see the observer, there is no need for the observer to inhibit intimacy, hence the increased gaze to the others’ face. However it is still not clear from this study what role the physicality of the screen plays in increasing social attention as the presence of the screen and the degree of interaction were changed concurrently between the two conditions.
The only study to control for the presence of a screen while altering some degree of potential interaction between stimulus and participant was conducted by Gobel and colleagues (2015). These authors manipulated participants’ beliefs about whether their viewing behaviour would later been seen by the target in the video, but did not actually involve interaction at the time of testing. They found that when participants believed their responses would later be observed, they spent 5% less time gazing at the eyes of high ranking targets (to avoid challenging their dominance) relative to when they believed their responses would not be seen, although the trend in the opposite direction for low ranking targets was not significant. These findings suggest that social signalling has a role in determining social gaze behaviour, adding support to the dual function theory. Although Gobel et al.’s study was a welcome contribution, it was not without limitations which may have masked potentially interesting effects. Most importantly, the scenario did not involve the potential for real-time interaction, which, as already discussed, may have resulted in gaze behaviour unlike that which would have otherwise emerged. It might be argued that the technological advances in eye tracking since Argyle’s (Argyle & Cook, 1976; Argyle et al., 1973) research have actually led to less flexibility in terms of experimental setups which may in turn have dissuaded researchers from addressing the issues raised here. Requirements for participants to sit still and tight control of experimental stimuli together with difficulties analysing dynamic eye tracking data may have inhibited the creativity of contemporary researchers in a way that did not affect their predecessors. However, clearly demonstrating that a compromise in ecological validity for the sake of precision is not a necessity in twenty first century research, Hessels, Cornelissen, Hooge, and Kemner (2017) used a setup where participants, seated in the same room, could only view one another on screen but could verbally communicate directly, while both participants had their eye movements recorded at a rate of 120Hz. Although not designed with the intention of manipulating the degree of interaction possible in the dyad, it is clear from this study that experimental designs can be developed to overcome this particular limitation without a reduction in data quality being inevitable.
Gobel et al.’s (2015) study had a further limitation in that targets were not present at the time of data collection. This brings the discussion to the second reason that participants in past studies have been unable to interact with the stimuli: they have been pre-recorded. Social presence, either actual or implied, has been shown to be important in modulating social gaze (Nasiopoulos, Risko, Foulsham, & Kingstone, 2015; Nasiopoulos, Risko, & Kingstone, 2015; Risko & Kingstone, 2011). In line with numerous studies reporting “mere presence” effects, whereby the presence of another individual is sufficient to influence participant performance in a range of tasks (e.g., Markus, 1978; Platania & Moran, 2001; Rajecki, Ickes, Corcoran, & Lenerz, 1977; Ukezono, Nakashima, Sudo, Yamazaki, & Takano, 2015; Zajonc, 1965). Gregory et al. (2015) showed that if an onscreen social scene was viewed in real time via a webcam, participants looked less at the faces within it than when the same scene was viewed as a pre-recording. It made no difference whether participants thought they would or would not meet the people in the scene after the experiment: the critical factor was whether the actors were perceived to be temporally (albeit not spatially) present at that moment or not. These results suggest that viewing others in real time onscreen, even if not physically present with them, could be sufficient to activate social norms of not staring, even without interaction being possible or imminent. This could explain the increased gaze to the face often reported when viewing onscreen faces compared with those viewed in real life (Foulsham et al., 2011; Laidlaw et al., 2011).
The current study
In light of the limitations of previous research, the first aim of this study was to determine the relative roles of potential social interaction and online social presence on social gaze using a novel lab-based paradigm. To this end, we compared participants’ eye movements when they passively viewed an unacquainted confederate as (1) a pre-recording, (2) a live stream but where interaction was not possible (“one-way”), or (3) a live stream where interaction was possible (“two-way”).
Our second aim was to explore the effect of an overt attempt at interaction by the confederate in the form of a bid for eye contact across these different social viewing contexts. An attempt at direct gaze between unacquainted individuals has been shown to increase the likelihood of a subsequent conversation (Cary, 1978) and is thought to increase intimacy between social partners (Argyle & Cook, 1976; Argyle, Lalljee, & Cook, 1968). In addition, recent neuroscientific evidence has demonstrated that mutual direct gaze in a live setup activates not only cortical regions associated with social cognition but also those involved in language processing. Critically this did not occur when an attempt at eye contact was one-sided or when the stimulus was a photograph (Cavallo et al., 2015). These findings suggest mutual gaze between two co-present individuals may facilitate social communication between them, supporting the dual function of gaze theory. We were concerned with how a bid for interaction from the confederate would be responded to by the participant under our different viewing conditions.
We predicted that if social norms of looking behaviour, characterised by “civil inattention,” occur only when reciprocity is possible as the dual function of gaze theory would suggest, gaze behaviour in the pre-recorded and one-way condition ought to be very similar, that is, characterised by increased looking towards the face reflected in increased total dwell time, longer and more numerous fixations to the face, and consequently reduced attention to other parts of the scene in comparison with when the confederate is believed to be able to see and hear the participant, and where interaction is possible. Reduced looking to the face in the two-way condition might be particularly pronounced when the confederate attempted to make eye contact, if gaze avoidance functions to inhibit intimacy between partners. However, if online social presence (i.e., the belief that one is viewing people in real time) is the driver for social norms of not staring, gaze to the face in the one-way and two-way conditions should be similar to one another, but reduced, relative to that of the pre-recorded condition. This reduction may be particularly pronounced during the eye-contact period but would not be expected to occur in the pre-recorded condition.
Methods
Participants
There were no exclusion criteria for this study, except that participants should have good vision (with or without glasses) and be free from neurological disorder. Students and participant pool members from Bournemouth University volunteered to take part in exchange for £5 or course credit. In total, 132 participants took part in the study (mean age [Mage]: 23.29 years, standard deviation [SD]: 7.33; 101 females). The post-experiment manipulation check which is detailed in the “Procedure” section resulted in the exclusion of 42 participants who did not believe our experimental manipulations. Of those, one participant was excluded due to poor calibration of the eye tracker. The final sample size was 91 with 28, 29, and 34 participants in the pre-recorded, one-way, and two-way groups, respectively (Mage: 22.72 years, SD: 6.57; 64 females).
Data collection was conducted at Bournemouth University, and the study was approved by the Ethics Committee of the Faculty of Science & Technology, Bournemouth University (reference 8960).
Stimulus, materials and apparatus
The stimulus was a 1 min 5 s video of a young Caucasian woman, waiting in a testing lab within the Department of Psychology at Bournemouth University. The video was filmed using a webcam placed on top of the monitor of the computer located in the lab. The confederate sat side on to the camera, so that, to look directly at it, she had to turn her head 90° towards the screen. After 20 s of the scene, the confederate turned to look directly at the camera, giving the appearance of “making eye contact” with the participant. Although the tendency is to gaze at the screen during video mediated interactions when attempting to engage in mutual gaze, this gives the impression of the other averting their gaze downwards due to the misalignment of the screen and the webcam (Bohannon, Herbert, Pelz, & Rantanen, 2013). Mindful of this, we explicitly instructed the confederate to gaze directly at the webcam on top of the screen during the eye–contact period. To the viewer, this gave the appearance of a bid for eye contact initiated by the confederate as she appeared to gaze directly at the participant. This gaze shift, from the time she began to turn to the camera to the point where she was again looking down at the clipboard was 4 s. For the remainder of the scene, the confederate completed paperwork on a clipboard. The confederate did not speak, but the audio stream was included to improve the authenticity of the situation from the participants’ perspective. Screenshots from the non-eye-contact and eye-contact phases can be seen in Figure 1.

Video scene during the (a) no eye-contact and (b) eye-contact periods.
Eye movements were recorded using the EyeLink 1000 desk-mounted eye tracker (SR Research, Canada). Participants sat 60 cm from the display screen, a 22″ ProNitron 21/750 CRT monitor, connected to a HP Compaq dc7800 display computer which was connected to a Dell OptiPlex 760 host computer. Participants’ faces were stabilised by a chin rest. Pupil and corneal reflection were recorded monocularly at a rate of 1,000 Hz.
A webcam was placed on top of the monitor of the eye tracking computer and a computer microphone was placed on the desk next to the participant in the two-way condition to improve the authenticity of the supposed interactive nature of the experimental setup. For the same reason, the webcam and microphone were removed in the one-way and pre-recorded (non-interactive) conditions.
Post-study manipulation check
In the pre-recorded group, 29 participants believed our manipulation that they could not interact with the confederate, although one of those also did not believe that the scene was pre-recorded, so this participant was excluded. In the one-way group, 32 participants believed our manipulation that they could not interact with the confederate (scoring 4 or above on the “interaction belief” 7-point Likert-type scale), but of those, 3 participants did not believe the scene was live (scoring less than 4 on the 7-point “live belief” scale) and were excluded. In the two-way group, 38 participants believed they could interact with the confederate, but of those, 3 did not believe that the scene was live suggesting some confusion about the manipulation and as such these 3 were excluded.
Procedure
Prior to the testing session, participants were randomly allocated to one of the three conditions: pre-recorded, one-way, or two-way. On arrival, participants gave written informed consent to participate and provided basic demographic information. It was explained to the participants in the one-way and two-way conditions that they would be watching another experimental participant in a nearby lab to the eye tracking lab while their eye movements were recorded. Participants in the two-way condition were told that the confederate would also be able see and hear them through the webcam and microphone in the eye tracking lab, whereas participants in the one-way condition were told that the confederate could not see or hear them. Both groups were then shown the lab along the corridor where the confederate would later be seated (the same lab as the stimulus recording took place), which contained an empty chair, a desk with a clipboard containing a consent form, and a computer with a monitor, on top of which a webcam was placed. Figure 2 shows the layout of the experimental suite at Bournemouth University, where the eye tracking lab is situated and where the confederate was assumed to be sitting. In the two-way condition, the screen on the confederate’s computer contained a screenshot of the eye tracking lab as seen from the webcam atop the eye tracking computer monitor to improve authenticity of the manipulation. In the one-way condition, the screen in the confederate’s lab was left blank.

Diagram of the testing suite where data collection took place where (a) in the pre-recorded condition, only the eye tracking lab was employed for the study, but (b) in the one-way condition and (c) two-way condition, a second lab was set up as the confederate’s lab, which participants were shown prior to data collection.
Participants were then escorted to the eye tracking lab where the monitor already displayed a screenshot of the confederate’s lab, as seen from the webcam atop the confederate’s monitor. Only in the two-way condition, the webcam and microphone were present. Figure 3 shows the participants’ view when seated in front of the eye tracker in each of the conditions as well as the confederate’s lab setup, as seen by participants.

View of the eye tracking lab desk setup in the (a) pre-recorded and one-way conditions (with absence of microphone and webcam) and (b) two-way condition (with microphone and webcam circled), together with the (c) view of the confederate’s lab as seen by participants in the one-way and two-way conditions (with microphone and webcam highlighted). Note that in (c), the screen in the confederate’s lab displays a screenshot of the view from the webcam atop the eye tracking computer showing the empty seat and chin rest as seen in the two-way condition, but in one-way condition, the confederate’s screen was left blank.
Participants in the pre-recorded condition were told explicitly that they would watch a pre-recording of another psychology participant. They were not shown the second lab, the screenshot of the second lab was not displayed on the eye tracking monitor, and no microphone or webcam was present.
All participants were seated in front of the eye tracker, where a 9-point calibration procedure was conducted. In the two “live” conditions, the experimenter instructed the participant to remain still while they left the lab for a few seconds to pretend to check that the second participant was ready. Participants were then informed that the live stream/recorded scene would be displayed on the screen and that they should watch this until told to stop by the experimenter, without any specific viewing instructions.
To improve the authenticity of the two “live” conditions, a message appeared on the computer screen indicating that the computer was attempting to connect to the webcam in the second lab. A further message appeared stating, “Ready? Press Y to record.” The experimenter pressed the Y key on the host keyboard which initiated a final drift correct procedure; a single dot displayed in the centre of the screen which the participant was asked to fixate. The video was then presented at 800 × 600 pixels resolution and was displayed at 30 frames per second. After the video had terminated, a message appeared stating, “Connection to the webcam lost; Retry Cancel Abort?” which the experimenter responded to by pressing R on the host keyboard which terminated the experiment.
A post- study questionnaire was completed by all participants to ascertain their belief in the experimental manipulation. Those in the pre-recorded condition were asked: “Whilst you were watching the video on the screen, to what extent did you believe that the stream was pre-recorded?” and answered on a 7-point Likert-type scale with 1 being did not believe and 7 being believed entirely. A follow-up question stated, “Whilst you were watching the video on the screen, to what extent did you believe the person could not see and hear you?,” and participants gave responses on a similar scale.
Meanwhile, participants in the one-way and two-way conditions were asked slightly differently worded questions, but on a similar scale: “Whilst you were watching the video on the screen, to what extent did you believe that the stream was live” with a follow-up of: “Whilst you were watching the video on the screen, to what extent did you believe the person could see and hear you?,” and participants gave a responses on a similar scale.
Participants were then verbally debriefed, and those in the “live” conditions were informed about the necessity for deception.
Results
Data handling and eye movement measures
Freehand dynamic interest areas (IAs) were drawn around the face, body, and background of the scene using Data Viewer v.2.6.1 (SR Research, Canada). The IAs moved with the confederate’s own movements. The background IA constituted the whole video area excluding the head and body of the actor. The eye movement data, explained in detail below, were then averaged across two interest periods: (1) a period where the confederate made “eye contact” with the participant by looking directly at the webcam (“eye-contact period” or EC period) and (2) the period where she did not make eye contact (“no eye-contact period” or No EC period). The latter was calculated by averaging the data from the periods before and after the eye-contact phase.
We explored several eye movement parameters in our analyses. In line with the majority of social attention research, our primary dependent variable of interest was the total mean dwell time to each IA across the three groups and two eye-contact conditions. Total dwell time, which sums all the samples recorded in each IA and averages those over each condition, provides a measure of the amount of attention different regions attract over the whole trial duration. However, it is important to note that other eye movement parameters may change without impacting on total dwell time. For example, several small fixations may total the same length of time as one long fixation, yet relying on total dwell time alone would not permit this more subtle difference in viewing behaviour to be highlighted. As such, we also analysed and reported two further measures which contribute to total dwell time: mean proportion of fixations and mean fixation duration. Proportion of fixations refers to the number of individual fixations executed within an IA as a proportion of the total number of fixations made on the scene as a whole. Higher numbers of fixations have been suggested to reflect increased processing of that area which may arise when encountering processing difficulties, complexity, or lack of expertise with the stimulus (Holmqvist et al., 2011). Fixation duration refers to the mean length of each fixation to each IA averaged over the trial. Fixations with shorter durations are suggested to be the result of decreased cognitive processing (Henderson, 2003), but when viewing, social stimuli, specifically, may reflect an increased level of social anxiety (Horley, Williams, Gonsalvez, & Gordon, 2003).
The majority of the eye movement measures reported here were not normally distributed. However, as analysis of variance (ANOVA) is robust to such violations of normality (Mayers, 2013) and to aid easier comparison between these results and those published elsewhere in the field, we preformed analyses on non-transformed data. Cell means for all dependent measures across all conditions are shown in Table 1.
Mean eye movement data.
Summed face interest area (eyes, lower face, outer face) data for percentage fixations and dwell time may not sum to 100% or exactly equal face total value due to slight overlaps of the dynamic interest areas.
M: mean; SE: standard error.
Mean fixation durations are presented for all cells, together with the sample this is based on.
Scene analyses
Total dwell time
A mixed ANOVA on mean proportion of dwell time to the different IAs across the trial, with the between-subjects factor of group (pre-recorded, one-way, two-way) and the within-subjects factors of period (EC, No EC) and IA (face, body, background) was conducted.
Critical to the study’s main hypothesis, the IA × group interaction was significant, F(4, 176) = 3.51, p = .009,
There was a significant interaction between period and IA, F(4, 176) = 263.422, p < .001,

Mean proportion of dwell time to (a) head, body, and background across the three groups; (b) face, overall, between eye contact conditions and groups; (c) the individual facial IAs in the two eye-contact periods. Error bars represent standard error of the mean. Brackets denote a significant difference at the p = .05 level.
Proportion of fixations
A further mixed ANOVA was conducted to assess differences in the proportion of fixations allocated to the scene’s areas of interest across conditions. The main effect of period was highly significant, F(1, 88) = 58.65, p < .001,
Fixation duration
Because the face IA was the only one to attract fixations from every participant in every condition and, therefore, because all other conditions contained a substantial proportion of “missing data,” it was only possible to conduct meaningful analyses on fixation duration differences between groups and conditions to the face IA. For completeness, Table 1 shows the mean fixation duration for all levels of all conditions.
There was a main effect of period, F(1, 88) = 25.94 p < .001,
Face analyses
To explore precisely where in the face participants were looking, we further analysed our data by dividing the face up into individual IAs which included the eye region, lower face region, and outer face/head.
Dwell time
A mixed ANOVA revealed a main effect of period, as found in the whole scene analysis, with longer dwell time to the face in the EC period, F(1, 88) = 102.19, p < .001,
Proportion of fixations
A further ANOVA revealed a main effect of period, F(1, 88) = 446.59, p < .001,
Fixation duration
As described earlier, due to the fact that not all participants looked at all IAs during each period, particularly, during the EC period, it was not possible to conduct meaningful analyses on fixation duration data which included every face IA. Specifically, only 32 participants’ data existed for the outer face IA in both EC and No EC conditions, with group sizes varying between 6 and 18 participants. Therefore, although not reflecting the whole sample (pre-recorded: N = 13; one-way: N = 19; two-way: N = 15), an exploratory mixed ANOVA was conducted on fixation duration data for the eyes and the lower face IAs, with period as a further within-subjects factor (No EC, EC) and group as the between-subjects factor. Participants’ fixations were longer in the EC period than in the No EC period, F(1, 44) = 13.43, p = .001,
Discussion
Previous research has suggested that people avoid looking at others when physically present with them, but direct their attention towards them when they are viewed onscreen. However, recent work has shown that even when viewed onscreen, faces are avoided if participants view the stimulus as a live stream compared with when it is pre-recorded. We aimed to determine whether this real-time gaze avoidance, which may be driven by social norms of “civil inattention,” occurs when participants believe they can interact with the online target or whether simply being temporally present with the target in real time, regardless of the ability for interaction, is enough to activate this avoidance response. In addition, we wanted to ascertain the impact of an isolated bid for mutual gaze by the target on the participants’ attention. We achieved this by showing the same recorded stimulus to three groups of participants under different viewing conditions: as a pre-recording, as a live stream but without the ability to interact (“one-way”), and as a live stream with the ability to interact (“two-way”).
Our results support the interaction explanation. Participants who believed they were watching a pre-recording looked more at the face of the confederate than when they believed the scene was a live stream with an audio-visual link. Importantly, participants looked as much to the face when they believed the scene was live, but without the audio-visual link, as they did when it was pre-recorded. There was also a trend for participants to look more to the face during the one-way condition compared with the two-way condition and a significant increase in dwell time to the background in the two-way condition compared with the others, all of these results demonstrating a medium effect size. Thus, the belief in the ability to socially interact with the confederate appeared to be causing a reduction in gaze to the face in favour of increased looks to the background, whereas viewing a confederate without the means to interact with them, whether in real time or as a pre-recording, resulted in increased gaze to the face and reduced looks to the background.
This effect appeared to be driven by differences in social gaze during the EC period. When the confederate gazed directly at the camera, a large increase in dwell time and proportion of fixations to the face were found for all participants. However, those in the two-way condition looked significantly less at her face than those in the other condition. Again, this supports the interpretation that the ability for reciprocity in the two-way condition was causing a relative avoidance of direct gaze in this group. In contrast, where participants knew they could not interact with the confederate, either because there was no audio-visual link or because the confederate was not temporally present (i.e., because they were pre-recorded), participants looked more towards the face. Importantly, there was no difference between the pre-recorded and one-way conditions in dwell time to the face during the EC period. Conversely, there were no differences in gaze towards the face between groups during the No EC period, suggesting that social norms of looking behaviour influence gaze particularly when attention is overtly directed towards the observer and therefore, where an interaction may be immediately imminent. These findings support the dual function of gaze theory, as the ability of gaze to act as a signal in the two-way condition caused the reduction in social attention. In contrast, our findings did not support the idea that mere online social presence (viewing another onscreen, in real time) causes the activation of the social norms of looking behaviour. There were no differences in eye movements directed to the face between the pre-recorded and one-way conditions which would have supported this. Previous work has suggested that the reduction in social gaze observed when viewing others in real time may be due to the operation of the social norms of looking behaviour causing gaze avoidance of “real” people (Gregory et al., 2015). Certainly at the neural level, it would seem that viewing others in real time is a qualitatively different experience to viewing them as a pre-recording or photograph (Cavallo et al., 2015; Pönkänen et al., 2011; Redcay et al., 2010), and the social psychology literature has presented many studies showing the mere presence of another person can alter performance on a range of tasks (e.g., Markus, 1978; Platania & Moran, 2001; Rajecki et al., 1977; Ukezono et al., 2015; Zajonc, 1965). However, while this may well be the case, our results demonstrate that specifically in terms of social gaze during a real-time social scenario, it is the inability of the participant to interact with the stimulus rather than the lack of social presence which causes participants to increase the amount of attention allocated to the faces of those viewed onscreen.
Previous research using dynamic video stimuli such as this has typically compared the face with body or other important elements of the scene (e.g., Gregory et al., 2015; Kuhn & Land, 2006; Kuhn et al., 2016), as we did in our main analysis. Our exploratory analysis on the facial features (eyes, lower face [encompassing nose and mouth area], and outer facial features and head area) did not suggest that individual facial regions were processed differently under the three viewing conditions. Contrary to work showing the eyes to be the most fixated facial region (Birmingham et al., 2008; Foulsham, Cheng, Tracy, Henrich, & Kingstone, 2010), we found that the lower face attracted greater dwell time than the eyes for all participants, although both total dwell time and the proportion of fixations to the eyes increased during the EC period, although only to the extent that lower face and eyes were fixated equivalently. Taken together with the whole scene analysis, these results show that differences between the groups in terms of attention to the face were not driven by differences at the specific region level. Rather, participants in the two-way, interactive condition were more inclined to avoid the face as a whole compared with the other groups, rather than specifically the eyes. Recent research has argued against a bias towards the eyes in social scene viewing as ubiquitous, having been shown to be dependent on stimuli and task (Peterson & Eckstein, 2012; Vo, Smith, Mital, & Henderson, 2012) and individual differences in participants, such as autistic traits (Freeth, Foulsham, & Kingstone, 2013) or face recognition ability (Bobak, Parris, Gregory, Bennetts, & Bate, 2017) In addition, the content of the scene used in this task may have contributed to the wide distribution of fixations observed. The lone actor sat side on to the webcam and only turned her head to gaze at the camera during the EC period. She gazed downwards at a clipboard for the remainder of the time, and it is possible that her eyes were not fixated to a greater extent because they may not have been perceived as important social cues during this phase of the scene compared with if she had been facing the camera (Vo et al., 2012) or had she been involved in an interaction with another actor (Birmingham et al., 2008), situations where longer dwell time to the eyes has been found. However, even in that period, the lower face attracted an equal number of fixations, possibly in anticipation of speech which would be a possibility unique to dynamic scenes. Although overall dwell time and proportion of fixations were often no greater for the eyes than other elements of the scene, we did find that the average fixation duration was longer for the eyes compared with the other facial IAs, particularly, during the EC period. This finding is likely to reflect the high biological and social relevance of the eyes (Adolphs, 2008) and, therefore, the increased processing of this facial region compared with the others (Henderson, 2003).
Finally, it was notable that participants in the pre-recorded group showed longer fixation durations relative to the other groups, regardless of EC condition or IA. It is possible that the pre-recorded group experienced reduced levels of stress or anxiety compared with the groups who believed the scene was live, as shorter fixation durations have been found in participants with social phobia when viewing social stimuli (Horley et al., 2003). Alternatively or perhaps in addition, participants in the live groups may have experienced increased cognitive load (Matthews, Reinerman-Jones, Abich, & Kustubayeva, 2017) compared with the pre-recorded group, given the additional manipulations employed for these participants, which could have reduced their fixation durations compared with the pre-recorded group. Although further research may be required to determine the underlying mechanism responsible for this, it is clear that the knowledge that the scene was pre-recorded was exerting an influence at a global level for these participants.
Our results can explain why previous researchers have shown increased social gaze when viewing others onscreen compared with real life. This effect may have less to do with online social presence (or a lack of it) but more to do with the potential for reciprocity between participant and confederate which is an entirely different issue. This study is the first to attempt to isolate the independent contributions of these mechanisms.
Limitations and future directions
It is important to note that the bid for eye contact resulted in an increase in dwell time to the face for all participants, relative to where the confederate looked away from the camera. There are several possible reasons for this. First, the movement itself, which involved a turn of the head through 90° towards the camera would have acted as a movement cue in a scene where the confederate otherwise looked in one direction. Movement has the ability to attract attention regardless of the nature of the stimulus (Abrams & Christ, 2003), so this could account for the increased dwell time to the face during this period. Second, the period of direct gaze by the confederate was only 4 s in length. Previous theorists have suggested that a brief acknowledgement by one unacquainted individual to another is considered appropriate behaviour, whereas prolonged gaze at another is not—this is the basis of the theory of civil inattention (Goffman, 1963; Zuckerman et al., 1983). Had the confederate prolonged this period of direct gaze, overall gaze avoidance might have been evident as the participants attempted to reduce the “social risk” which results from making direct gaze with a stranger.
We are the first group to adopt a live viewing paradigm to study onscreen social attention. As such, many questions remain unanswered. Indeed, as we have shown the pattern of social attention deployed to live, interactive scenes is qualitatively different to that found when watching pre-recorded stimuli, our findings may necessitate the re-evaluation of several decades of social attention research. This previous research has not been in vain. Rather, it has provided an understanding of how people view social stimuli without the constraints of social norms under carefully controlled conditions. A critical task of future social attention research, if it is to be ecologically valid, is to revisit the findings of lab-based studies of the past and adapt them to include manipulations of the genuine social pressures experienced in everyday social life to assess whether effects persist or are moderated by social norms of looking.
The “live lab” paradigm offers many benefits over using mobile eye trackers in genuine social situations. Using screen-based eye trackers allows for more stimulus control, increased sensitivity, and accuracy, negates the need for confederates to be present at each testing session, and allows for more fine-grained data analyses than is possible with a mobile eye tracker. The “live lab” paradigm is no more time-consuming in terms of data collection than mobile eye tracking, and the data analysis is significantly swifter as no hand coding is required. One potential drawback is that a significant minority of participants fail to believe the deception involved and researchers need to design their studies to minimise the problem. A possible limitation of this however is that in excluding participants who do not believe the manipulation may represent a form of selection bias, as these participants may possess some particular characteristics which prevent their belief in deception and therefore which are not represented in the final sample. Nevertheless, we believe that the “live lab” paradigm offers a flexible, alternative paradigm to researchers of social psychology in a range of sub-disciplines not limited to those using eye tracking, where social norms may influence behaviour or cognition but where tight experimental control is desired.
Conclusion
Our results show that the ability for participants to interact with the social stimulus they view onscreen results in a reduction in social gaze compared with when the same stimulus is viewed without the means for interaction, particularly, when an interaction is immediately imminent. Mere social presence was not sufficient to cause this reduction effect. Our findings support the dual function of gaze theory in that when participants’ gaze can act as a signal to the individuals they are viewing, they look less at the face of that individual than when social signalling is impossible, either because of technological limitations or because the scene is pre-recorded, as they adhere to “civil inattention.” We suggest that previous research has shown increased social attention to onscreen others due to this inability to interact, not due to a lack of social presence. Given these findings, reassessment of over a decade of social attention research may be warranted. We conclude by suggesting that adopting a “live lab” paradigm may offer researchers an ecologically valid framework for exploring social psychological phenomenon while maintaining high levels of experimental control.
Footnotes
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
This work was funded by an award to N.G. from the Experimental Psychology Society Small Grant Scheme.
