Abstract
In the context of rapidly developing technologies and widespread online access, it is important to understand how our perception of images on a computer screen may vary from traditional in-person encounters. This research compared the perception of subjects in view of eight paintings, presented either on a computer monitor or as a printed reproduction. Both stationary and mobile eye-tracking technologies were used to analyze the viewing patterns of both forms of engagement. Results suggested that subjects engaging with physical works tended to exhibit more varied fixational patterns than those viewing the same works on a computer monitor. Data showed parity in the high degree of correlation between viewing times and personal preference, regardless of viewing medium. The results indicate that the modalities through which we engage with works of art matter, and that a single image can resonate across an array of media.
Introduction
Recent global health events have emphatically underscored the importance of providing alternative means for the public dissemination of art and cultural heritage. The online availability of gallery collections allows for opportunities for access that are simply not possible within the physical confines of a typical museum setting. Beyond recent physical constraints ushered in by the pandemic, perennial considerations include reaching audiences where they already are, round-the-clock access regardless of an institution's logistical constraints, the ability to accommodate visitors around the world, and in many instances, greater access to those who may not be able to afford access costs and fees (King, Smith, Wilson, & Williams, 2021; Marty & Buchanan, 2021; McGrath, 2020). The online exhibition of works, therefore, presents an ongoing opportunity to define visitorship in the digital age (Knight Foundation, 2020; Proctor, 2010). Now a couple of decades since the earliest examples of online viewing, it is important to not only celebrate its emerging possibilities but also to consider how the experience of viewing works may change in contexts outside conventional gallery presentation. Among myriad considerations, this paper asks how the perception of digital images may differ from a work's more traditional gallery setting. More specifically, this study asks if paintings presented on a computer monitor are viewed differently than viewing the same works in person.
As one mode of studying visual perception, the study of eye movements provides an avenue into how aesthetic perception can be understood, whose repercussions can straddle both the digital humanities and psychology. Eye tracking's dual relevance to the arts and the sciences was inherent from its inception over one hundred years ago and throughout the 20th century (Buswell, 1935; Huey, 1908/1979; Rosenberg & Klein, 2015). Prior to the relatively recent proliferation of digital technologies, measuring eye movements was often performed using scaled reproductions, such as on a printed card or a computer monitor. This was largely due to technological constraints, as eye-tracking devices were often optimized for heavily prescribed testing parameters, including limited image sizes and fixed viewing positions. In recent decades, however, the development of modern tools has allowed for greater flexibility in how we can capture a person's visual attention (Al-Rahayfeh & Faezipour, 2013). Taking advantage of such technologies, and given the importance of scale in painting, this research examines how images may or may not be viewed differently within two common presentation scenarios: in view of large canvases versus digital projection on a computer monitor. More specifically, this study compares how ambulant viewing of relatively large-scale works compares to viewing the same pieces on a computer monitor while seated. As relatively large paintings often provoke some degree of physical movement in their examination, this paper implicitly questions if analyses gained from conventional eye-tracking of a seated viewer sufficiently describes our typical engagement with paintings and images within a conventional gallery setting.
In order to compare eye movements of subjects in view of images at large versus monitor scales, this study compares the results of two forms of eye tracking. To capture the perception of stimuli displayed on a computer monitor, a stationary eye tracker was used in which the viewer is seated in front of a desk. In contrast to these recordings, eye movements of subjects viewing images at relatively larger scales were captured using a combination of wearable eye-tracking spectacles and spatial sensors. This paper will introduce basic concepts of eye tracking, outline methods for testing, and discuss its findings.
Context
Although highlighted by recent issues of in-person attendance to museums during global health crises, the rapid rise in the online dissemination of art and cultural heritage has been underway for the past couple of decades, largely enabled by contemporary technologies, digital publishing practices, and changes in how people engage with cultural experiences. Major institutions such as the Museum of Modern Art (MoMA NYC) have invested heavily in the online publication of works, without almost half of their entire collection available online, while the Louvre's entire collection can be found online, with daily updates. In addition to online publication, cultural institutions are now increasingly making available works at extraordinary resolutions and with open access. These trends of digital exhibition during a period of both uncertainty in visitorship and exponential technological advancements have been highlighted by a number of recent studies (Galani & Chalmers, 2010; King et al., 2021; Marty, 2008; McGrath, 2020).
In light of these innovations, it is critical to understand how such developments may affect the reception of art in relation to conventional in-person viewing. Although many new forms of online exhibition involve immersive engagements such as virtual tours and interactive interfaces, conventional still images require minimal to no instruction for their viewing, and obviate many technical hurdles involved in browsing, publishing, and hosting artworks that incorporate more novel modes of experience.
Environmental Influences and Aesthetic Engagement
Although the immediate context of this research responds to the rapidly emerging forms of digital presentation in art and cultural heritage, it is ultimately situated within a much larger inquiry into the influences of physical context in aesthetic appreciation. Similar to the basic question asked by the current study, prior research has compared degrees of viewer engagement with works of art by studying how the perception of art may be influenced by museum versus laboratory settings (Locher, Smith, & Smith, 2001; Specker, Tinio, & van Elk, 2017). Brieber, Nadal, and Leder (2015) found no correlation between physical context and art appreciation in participants viewing photographs in either setting. Responding to this research, however, Grüner, Specker, and Leder (2019) contended more additional care was required in selecting artwork medium, as well as disentangling testing variables including physical setting and concepts of genuineness of art material. In their own study, they found a correlation between the viewing of works in a museum, versus a laboratory, corresponding to increased viewer valuation and interest. They suggested that an overall feeling (i.e., white cube effect) could condition viewers to engage more profoundly. Interestingly, they did not find any evidence of an inverse white cube effect, in which the act of viewing a work of art outside of a museum setting would “elevate” such a context. In both studies, participant engagement was assessed on a numeric grading scale.
In considering the influence of environment on art appreciation beyond a museum-laboratory binary, Pelowski, Forster, Tinio, Scholl, and Leder (2017) identified important psychophysical factors that could support the increased engagement of works within a museum setting. They noted that the act of standing or moving in front and around works of art could trigger cognitive behaviors that are less provoked in seated observation. Similarly, Kapoula, Lang, and Locher (2014) found correspondence between anteroposterior sway, as is typically present in the untethered viewing of works of art, with greater ranges of eye movement.
Although not strictly a matter of environmental setting, the issue of physical materiality is often implicit in viewing conditions in museums and laboratories. For example, empirical testing of art appreciation within museum and gallery settings are often conducted using actual paintings on exhibition. Contrarily, it is common in laboratory testing procedures to use reproductions, scaled facsimiles, and computer monitor projections to present subjects with works of art. Developing their study on Currie's (1985) notion of a “transferability thesis” in which the aesthetic value of a work pertains between original and reproduction, Locher et al. (2001) posited a “facsimile-accommodation” hypothesis in which viewers could appreciate the attributes of artworks despite consciously viewing them as reproductions. Using numeric grading scales from participant questionnaires, their results supported their accommodation thesis, though the greater engagement was shown in original works when viewed within museum settings.
Developing from these broader questions of environment and aesthetic appreciation, the current study responds more specifically to the above precedents by focusing on how patterns of eye movements may or may not reflect the myriad factors inherent in an ostensibly simple distinction between viewing works within a gallery setting versus on a computer monitor.
Eye-Tracking Research and the Museum
After decades of research within the confines of the laboratory, there is now a rapid increase in studies that examine the nature of eye movements within more natural conditions, with the ambition to better understand visitors’ experiences of art at the very sites of typical engagement (Eghbal-Azar, 2016; Eghbal-Azar & Widlok, 2013; Han, 2021; Reitstätter et al., 2020; Santini et al., 2018). Wooding, Mugglestone, Purdy, and Gale (2002) installed a viewing booth at the National Gallery London in order to gather perceptual data of visitors, though their study involved viewing digital facsimile projections. Overcoming the presentational limitation of Wooding's digital images, Bachta, Stein, Filippini-Fantoni, and Leason (2012) conducted eye-tracking analysis using an actual painting situated on a gallery wall, with participants seated directly in front of the viewed stimulus. Incorporating ambulation to their research, Heidenreich and Turano (2011) utilized a mobile eye-tracking (MET) unit to understand the viewing behaviors of paintings, while Filippini Fantoni, Jaebker, Bauer, and Stofer (2013) provided an early comparison between MET and the stationary eye-tracking study of Bachta et al. (2012). Using a comparable MET apparatus, Walker, Bucker, Anderson, Schreij, and Theeuwes (2017) differentiated bottom-up versus top-down perceptual predilections between children and adults. To obviate the inconveniences inherent in setting up an eye-tracker within a gallery setting, Dare, Brinkmann, and Rosenberg (2020) developed a calibration-free tracker that could be placed discreetly underneath a painting, avoiding the need for a cumbersome set up that could potentially interfere with natural viewing behavior, albeit at the expense of limiting the range of a viewer's position and movement.
A number of recent exhibitions and projects have analyzed visitors’ eye movements for immersive forms of engagement including ARoS Art Museum's “The Art of Looking” (2016), 1 Cleveland Museum of Art's Gaze Tracker (2017), 2 and the M Museum Leuven's ongoing public exhibitions that implement eye tracking for various feedback. 3 Beyond the application of eye tracking to measure the movement and location of one's gaze of art, Dondi, Porta, Donvito, and Volpe (2022) concluded that the use of eye tracking was particularly auspicious in light of COVID-19, in its capacity as an assistive input device, whereby visitors could interact with exhibitions without the need for the contact of physical interfaces such as buttons, screens, or a mouse.
Within this landscape of research, this paper compares how the same image may or may not be seen differently according to medium. The specific methods of this research are situated within precedent works that provide important context. Quiroga, Dudley, and Binnie (2011) compared the perceptual behavior between subjects in view of John Everett Millais’ Ophelia (1851–52) at Tate Britain in London against subjects in view of the same work on computer monitor 4 and seated within the settings of a laboratory. Although Quiroga evidenced disparities of viewing behaviors between the two different participant groups, they noted that the discrepancy of viewing environments could have contributed to perceptual differences, as well as the role played by the possible “aura” of the original work, most notably conceived by Benjamin (1935/2008). The empirical consequences of aura, genuineness, and originality are manifold, as Locher et al. (2001) illustrated divergences of aesthetic appreciation in viewing patterns between art-laypersons and art-professionals. Likewise, Newman and Bloom (2012) reference the idea of “contagion” in the context of art valuation, which in turn may affect how we physically engage with such works. As discussed in the above precedents, the issue of originality that was inherent in Quiroga's study can actually involve a multitude of testing variables, and as such, this study circumvents these issues by isolating viewing of both monitor and large-scale images to controlled environments, while also using printed reproductions as stimuli.
Brieber, Nadal, Leder, and Rosenberg (2014) compared viewing behaviors between subjects examining exhibited photographs at Wien Museum MUSA with the same works displayed on a computer monitor. Their research demonstrated that, on average, not only were viewing times longer in a gallery setting but also that there was a more positive art experience expressed by the same subject group. Similarly, Locher, Smith, and Smith (1999) observed that the act of viewing an original painting factored into a viewer's general aesthetic awareness when compared to encountering the same work as a digital projection, though in their study, the definitiveness of aesthetic “pleasingness” was less pronounced. In their research, a “pictorial sameness” was evidenced between the two subject groups, with the theory that individuals were able to look past differences in media in order to consult the core message of the artwork.
Like the above precedents, the current paper compares how people see paintings when viewed as relatively large-scale reproductions as opposed to displayed on a computer monitor. However, this research looks to distinguish viewing behaviors within more comparable testing scenarios. More specifically, both in-person and digital testing was conducted within laboratory settings, and to promote parity of experience, printed reproductions were used as proxy stimuli for original paintings in order to eliminate the question of aura.
Prior to testing, it was assumed that the viewing of physical canvases would require increased physical and attentional effort by the viewer, as they occupied a larger field of view relative to the size of a computer monitor, as already suggested by Garbutt et al. (2020). As a consequence of this increased effort, it was also assumed that viewers examining printed canvases would cover a smaller visual footprint within the viewed painting, though attending more to details, in relation to the viewing of images on a digital screen. It was supposed that the discrepancy in visual coverage would only be exacerbated as the difference in the area increased. Further to this, it was expected that there would be an increased correlation of areas of interest between different viewers examining the same image on a monitor, versus those of a large canvas, as ambulation would introduce an additional variable that would further distinguish one viewer's encounter with that of another. Confirmation of the above discrepancies could highlight inherent differences in how art is viewed in-person versus on a computer monitor, thus articulating the inherent implications of physical interaction with the reception of works of art.
Methods
Eye-tracking recordings were taken from 31 participants, assigned to one of two groups (16 in Group A and 15 in Group B). The slight discrepancy in group populations occurred due to technical errors in hardware technologies. Each group viewed the same set of eight paintings, whose sequence was randomly assigned for each recording session. Those in Group A (“monitor group”) viewed eight paintings displayed on a computer monitor, while participants in Group B (“canvas group”) viewed the same paintings printed on canvas at a relatively large scale as indicated in Table 1. Participants were asked to view each work in silence, and afterward, answer a series of questions regarding the respective work. Viewing time was left to the subject's discretion in the case that there could be meaningful differences in lengths of viewing between physical and digital presentations. In order to end each recording, those in Group A were instructed to verbally notify the examiner to discontinue the current display, while those in Group B were asked to return to a starting point whose location was away from the view of the reproduced painting. For this latter group, viewing time was counted as the span between first and last views of the stimulus, as approach and departure times were excluded.
List of All Stimuli and Printed Dimensions.
Note. List of stimuli presented to all subjects, presented in randomized order. Reproduced dimensions applicable for Group B (canvas group). Group A (monitor group) stimuli presented on a 24″ monitor with screen dimensions at approximately 14 × 21″. A black border was used on monitor displays.
Participants
Participants were comprised of an array of individuals within a university setting, including undergraduate and graduate students, as well non-teaching staff. Their ages encompassed a range of 18–40 years and consisted of 18 females and 13 males. Backgrounds of the study were diverse and were not biased toward one specific discipline, including but not limited to fields of art, architecture, computer science, engineering, chemistry, and those without a declared field of study. Participation was entirely voluntary and involved both remuneration and academic course credit where applicable. All testing procedures were performed in accordance with the protocols and ethical standards of the Institutional Review Board at Lehigh University (Bethlehem, PA). Informed consent was obtained from all participants prior to testing.
Overview of Stimuli
A total of eight images were used for eye-tracking tests and were chosen as a result of four principal criteria. Firstly, it was critical that a sufficiently high-resolution scan was publicly obtainable. In consideration of their physical display on printed canvas, the perception of lower quality reproductions could have negatively affected the legitimacy of tests. Secondly, within the pool of available images, a variety of aesthetic styles were chosen so as to generalize the overall range of viewed works. Thirdly, it was important that the selected images not fall under restrictive copyrights given that this research was intended for publication. Finally, there was a preference for relatively large-scale works, to more strongly differentiate the physical scale of printed material from the viewing size of a computer monitor. However, one control stimulus was chosen (stimulus #5) at a scale comparable to that of the utilized monitor (24″ monitor, approximately 14 × 21″) in the case that discrepant viewing behaviors could result from either disparate viewing sizes or simply presentation medium.
Technical Setup
It was important to conduct tests from both groups within comparable environments, so controlled settings were used by limiting any environmental distractions and foot traffic. Beyond this general control, however, the distinct viewing modes warranted unique technical apparatuses and physical constraints.
For Group A, digital images were shown on a 24″ monitor, placed approximately 24″ away from the participant. Eye-tracking information was captured using a Gazepoint GP3 HD, recording at a maximum sampling frequency of 150 Hz. Analysis and presentation of gaze data were processed using custom scripts written in Python, including regression analysis in Scikit-Learn (sklearn). The setup for eye tracking on a computer monitor was conducted along with standard methods, consisting of a subject seated in front of a stationary monitor with an accompanying stationary bar-type camera mounted immediately below.
Testing apparatuses and setup were largely distinct in Group B and catered for untethered viewing. Printed reproductions were produced on a large-format printer, on a canvas-like substrate, mounted on stretcher bars, and hung on a bare white wall, so as to emulate conventional gallery presentation. Printed works were reproduced near or at the original scale. The principal determining factors for canvas sizes included a maximum printing width of 60″ (no limitations of print length along with the longer side), availability of common stretcher bar lengths at regular 1″ increments, and the transportability of stretched printed canvases. Figure 1 5 shows the basic testing scenario for this group of subjects.

Testing set up for Group B..
Mobile eye tracking was achieved by using a combination of data from various sensors, including camera-mounted spectacles (Pupil Labs Invisible) which recorded a viewer's relative pupillary movement, a Visual Simultaneous Localization and Mapping (VSLAM) device (Intel T265) that recorded head orientation and relative location, while the global registration of head location was guided by three synchronized RGBD cameras (Microsoft Azure Kinect DK). As these products operate within their own applications, their data integration was performed asynchronously using custom-written scripts in C++ and Python.
Procedures and Measures
Prospective participants in both testing groups were first introduced to the basic research questions and procedures in order to receive informed consent. Thereafter, those in Group A were seated in front of a computer monitor with an accompanying stationary eye-tracking camera, while participants in Group B were fitted with eye-tracking spectacles, a hat mounted with a VSLAM camera, and a computer tablet worn over their shoulder or held in their hands. Those in Group B were shown a prescribed walking route as indicated by a taped path along with the testing floor. Unlike the viewing process of those in Group A in which the investigator could turn off the computer display remotely between successive scenes, the process of “advancing” from one physically mounted canvas to the following required the subject to leave the immediate viewing area.
For both groups of participants, a questionnaire was shown prior to commencing the first eye-tracking recording and was then issued after each viewed painting. The form consisted of three questions, whose principal function was to motivate the viewer to visually engage with each scene. The first question asked the participant to rate the recently viewed work on a scale from 1 (lowest) to 10 (highest). The next question asked the viewer to list three memorable aspects of the painting, while the final question asked the participant what they felt was the artist's intended message. There was no time restriction for filling out questionnaires, as it was encouraged for participants to actively process their recent visual encounters. Though untimed, the range of completing a questionnaire ranged from approximately 30 s to 5 min.
Basis of Eye-Tracking Data
The essential information registered by an eye tracker consists of gaze points, which represent the location of a subject's gaze at a single moment in time. The overall quantity of gaze points in a recording depends on the total length of viewing time as well as the frame rate of the eye tracker, typically registered at 60–200 Hz. From this information, other metrics can be extrapolated. Most importantly for this study, these include fixations, which refer to areas populated with a dense cluster of gaze points 6 and generally interpreted to signify a moment of concentrated visual engagement. A fixation map refers to a form of visualization in which the aggregate of fixations, whether within a single viewer or a group of viewers, are illustrated across a specified period of time.
Analysis
Initial Results
Figure 2 illustrates the result of consolidating registered fixations provided by all subjects in Groups A and B onto one of the eight total stimuli (#2), in order to indicate the net distributions of visual attention. 7 Using conventional circle graphics whose radii correspond to the overall dwell times of their respective fixation, cumulative areas of visual interest are represented by areas of more opaque graphic overlays. It is important to recall the participant pool discrepancy, with Group A comprised of 16 subjects and Group B with 15 subjects. Moreover, while the vast majority of eye-tracking sessions were successful, not all produced usable results. Table 2 outlines the total count of successfully registered recordings per group, stimulus, and average recording time. The same table shows that total average viewing times were significantly longer for those viewing large-scale printed reproductions; hence, there was a consequent increase in total fixations for Group B over those of Group A for each stimulus case.

Basic fixation maps for Groups A and B, stimulus #2.
Count of Successful Recordings and Average Total Viewing Time per Subject Group.
Note. Overview of total recordings per stimulus and group. Not all subjects registered a successful eye-tracking recording, either due to technical errors or eye-tracking data with insufficient confidence registration. Average total view times for Group B were longer in all stimulus cases.
Length of Viewing and Preference
For all eight stimuli recordings, Group B viewing times were significantly longer than their Group A counterparts. On average, Group B times were +141.36% longer (Table 2), with the minimum increase at +92.32% (stimulus #8) and the maximum increase at +201.52% (stimulus #1). Resulting from participant responses from questionnaires issued after viewing each painting, Figure 3 shows the correlation between participants’ preference ratings with their associated viewing times. In both groups, there was a clear correlation in which higher-ranked paintings were viewed for longer periods of time, as corroborated by the positive slope of the corresponding linear regression lines. The lowest-ranked work (stimulus #5), which was the same in both groups, was an outlier in terms of its viewing time-to-preference rating. Favorability scores of individual works were similar, though not identical between the two subject groups. As well, results showed some moderate correlation between a stimulus’ physical area and an average length of engagement with an R2 value of 0.4678 (Figure 4).

Average total viewing time versus average personal score per stimulus.

Correlation between viewing time and physical size of pointing stimuli in Group B.
Comparing Fixation Distributions
Upon preliminary inspection of eye-tracking data, it was apparent that the prior posited assumption regarding discrepancies of “visual footprints” between participant groups was too imprecise to prove as a useful metric for analysis. For example, as Figure 2 shows, areas of fixational absences would not be accounted for by simple area comparisons such as those provided by convex or concave hulls. Instead, comparisons of fixational density distributions were pursued.
As there was a significant disparity in viewing times between the two groups, additional analysis was produced using culled time scales so that comparisons could be commensurate. The additional analyses included viewing ranges of 2.5 and 10 s, as well as an additional scale ranging from 16 to 24 s depending on the respective total view times for each stimulus, as specified in Table 3. The variability of the latter time scale was the result of comparing a minimum sampling threshold of at least 11 participants from each group within a consistent recording time (16–23 s). In all cases, this was determined by the recording times of Group A, as their viewing times were considerably shorter than their respective Group B counterparts (Table 2). Figure 5 illustrates the distribution of fixations with regard to stimulus #7, across the aforementioned time scales.

Fixation maps over total and culled time scales in Groups A and B.
Recording Count per Culled Time Scale, per Group.
Note. Overview of recording count after implementing culled time scales. Variable scales were determined on a per stimulus basis, with the limiting factor of Group A recordings given their relatively shorter total recording times. Data for stimuli #5 and #8 are included for reference only, as their resultant gaze data did not yield fixations with sufficient levels of confidence.
Upon examining eye-tracking data from time-culled sets, the general distribution of fixations was largely comparable and centered on narrative focal points as well as areas of compositional distinction. For example, in stimuli with human figures, subjects’ attention in both groups tended toward the faces of central characters, while in nonfigural works, participants gravitated toward areas of contrast in color, light, texture, and form, corresponding to typical top-down and bottom-up viewing behaviors (Fontoura, Schaeffer, & Menu, 2019; Massaro et al., 2012). Beyond this broad correlation, however, there was an apparent and persistent difference in the distribution of fixation clusters. More specifically, the results from Group A yielded fixation center points with moments of relatively denser spatial aggregation versus their distribution in their Group B counterparts. In other words, fixation clusters from Group A recordings appeared to be relatively denser, with emptier areas of separation between other neighboring clusters.
In order to distinguish the relative “tightness” of fixation clusters, Delaunay triangulation was performed on the cumulative set of fixation center points for each group and stimulus, thereby producing a set of proximal distances (Figure 6). A polynomial regression line which was based on the frequency of Delaunay edge lengths was then calculated in order to determine key distribution features. The principal metric of this analysis was to ascertain the “maxima” value of the regression line (degree = 6), and to then compare the associated Delaunay edge length between both groups. A lower maxima edge length value would indicate a greater spatial concentration of fixation center points at areas of aggregation. As these calculations were based on the same scale of analysis (i.e., a computer monitor) rather than the scale of stimuli sizes, the size discrepancies between printed reproductions and the testing monitor were obviated.

Basic fixation maps for Groups A and B, stimulus #2.
As Figure 7a illustrates, the distribution of Delaunay edge lengths followed a common pattern across all culled times scales, with a greater concentration of distances toward a relatively shorter length (x-axis toward origin) and following a positive skew (i.e., lean toward origin), and precipitously falling at lengths either shorter or greater than the regression maxima. The corresponding frequency of edge length values (y-axis) was deemed less consequential in comparison, as their value was directly affected by the total registered recordings (i.e., participants) per analysis, and was thus variable as indicated by Table 3. Figure 7b shows in greater detail the metrics involved in regression analyses.

(a) Regression graphs for Delaunay edge length frequency in Groups A and B. (b) Detail of regression graphs for Delaunay edge length frequency, stimulus #1 at 10 seconds cull time scale.
Figure 7a shows that there was a salient tendency for subjects in Group A to view stimuli with clusters of more densely aggregated fixations (i.e., lower Delaunay edge length at regression maxima), supporting the prior assumption of increased correlation of fixation locations from participants Group A versus those in Group B. This tendency was evident in all three culled time scales for all eight stimuli, with the exception of a single analysis (stimulus #3 at 23 s). This trend was not as apparent when comparing the distribution of fixations at full viewing times, though this time scale was not considered significant given the considerable increase in total average viewing time in Group B relative to those of Group A.
There was also a tendency for differences in Delaunay edge length at maxima to converge as viewing times increased (Table 4 and Figure 8), with relevant cases included in six of the eight stimuli sets (stimuli #2, 3, 5, 6, 7, 8). The two remaining exceptions included stimulus #1, which demonstrated a viewing time trend with negligible change (linear regression of slope −0.0207), and stimulus #4, whose regression analyses involved maxima values measuring below the minimum scale threshold (i.e., maxima off-chart).

Decreasing difference in maxima x-value (most common Delaunay edge length) as viewing time increase.
Change of Rate of Fixation-Based Delaunay Edge Length at Regression Maxima.
Note. Negative values denote a decrease in most frequent Delaunay edge length group from Group A to Group B.
*Strongly aberrant and coincident values due to nature of regression analysis, whereby calculated maxima may occur out of normal range.
**Stimulus #4 results were inconclusive as maxima were out of normal range, as indicated by significant percentage changes.
Contradicting the assumption that was made prior to testing regarding the correlation between canvas size and viewing behavior, there was no such evidence for any relationship when comparing the physical area increase of the printed canvas over the size of the testing computer monitor versus the Delaunay edge length increase expressed as a percentage (Table 5). As indicated by Figure 9 using data at both 2.5 and 10 s of viewing time, no correlative trend was determined such that the size of the viewed canvas could indicate the corresponding increase in the change of Delaunay edge length maxima between participant groups. Corroborating the above finding as shown in Figure 7a regarding the shifting Delaunay edge length with increased viewing time, in all but one case (stimulus #5), there was a regular decrease in the difference of edge length maxima from 2.5 to 10 s time scales. The reverse trend in stimulus #5 occurred in the “control” scene in which the printed canvas of Group B was produced at a comparable physical scale to the size of the computer monitor of Group A. The corresponding R-values remained low when correlating results from Table 5 using Pearson analysis, with results of r = −.14 at 2.5 s and r = −.52 at 10 s.

Relative increase in stimulus surface area versus relative increase in Delaunay edge length maxima at 2.5 and 10 second time scales.
Correlation Between Relative Surface Area and Relative Increase of Delaunay Edge Length at Maxima at 10 s.
Note. *Stimulus #4 results were not included in correlation test, as edge length at maxima values were off-scale. **Stimulus #5 acted as a control scene in which the printed canvas size was as close to the computer monitor dimensions as allowed with canvas materials, resulting in a surface area increase of close to 0%.
Case Study #1: Botticelli. To more clearly illustrate the observed trends between Groups A and B, two cases are presented in greater detail. They have been chosen namely for their stylistic contrasts while the results of eye-tracking analyses remained consistent.
Stimulus #1 involved Sandro Botticelli's Birth of Venus (1484–1486). 8 The scene involves a central figure standing on a floating shell, with secondary figures flanking each side. The rendering style is figurative, though not photorealistic. As it has been well-documented that visual attention of human figures tends toward faces, it is worth noting that in this painting, all four faces occupy roughly the same elevated horizon toward the top of the composition. Although the printed reproduction used for testing was not shown at the original scale, it was nevertheless presented to subjects at a relatively large size of 56.5 × 90.0 in (143.5 × 228.6 cm) (Table 1), which resulted in a total surface area of 5,086.8 in2. This area was +1,630.2% larger than that of the computer monitor (21 × 14 in − 53.3 × 35.6 cm) used in Group A.
The Botticelli painting was ranked relatively high in terms of personal favorability, with a score of 6.69 points in Group A and 7.47 points in Group B, out of 10 points maximum (Figure 3). Consistent with other stimuli, Group B participants viewed the painting reproduction for much longer, with an average time of 118.36 s in contrast to Group A at 39.26 s. As with all other stimuli, culled time frames were analyzed in order to more comparably understand the behaviors of both groups, resulting in additional time scales of 2.5, 10, and 23 s. Figure 10 shows the cumulative distribution of fixations for both groups, including culled time scales at 23 and 2.5 s.

Groups A and B fixation maps stimulus #1 at culled time scales of 23 and 2.5 seconds.
When comparing the results of fixations at 23 s in both groups, it is clear that there was a greater infill of fixation coverage from Group B participants, in other words, there were fewer “pockets” of unattended areas. There was a shared interest in both groups with regard to basic pictorial features, particularly on the face of the central Venus figure. Beyond this, however, the total distribution of fixations in Group B tended more toward the figure on the right, while those in Group A focused more heavily on the two figures to the left. As well, those in Group A consolidated far more attention toward the base of Venus.
The distribution of fixations in Group A appeared relatively more compact in areas of heightened collective interest. As with almost all other stimuli, when examining the comparison of fixations at 2.5 s, this behavior becomes even more explicit. At this time scale, the separation of areas of interest is more pronounced, yet preface the forthcoming fixation patterns at longer viewing times. In addition to areas of heightened attention, this time scale also illustrates the relatively higher compactness of spatial clustering in Group A, while the distribution of fixations in Group B is relatively more diffused throughout the composition.
Figure 11 illustrates a detail of a Delaunay conversion of fixations. Corresponding to observations regarding differences of spatial densities, there is a more varied hierarchy of edge lengths in the tessellation result of Group A. Relatively tight triangulation occurs at the two face groups, while a line of dense triangles meanders vertically along with the Venus figure. In contrast to these local densities, the triangulation occurring in the space between these two groups of figures is far wider. Contrary to this hierarchy, the tessellation of Group B is more consistent in spacing. Although dense triangles occur at the face of the Venus figure, the remaining triangulation is processed at a relatively similar scale, thereby producing a narrower hierarchy of Delaunay edge lengths. Accordingly, this distinction implies that the collective gaze of participants in Group A converged with higher frequency, while those in Group B were relatively more diffused. Regression analysis on the distribution of Delaunay edge lengths was performed to calculate this visual assessment.

Detail of Delaunay triangulation from fixation center points of stimulus #1 from groups A and B at 23 seconds.
The resultant analysis for stimulus #1 corroborated the above observations. As Figure 12 outlines, the regression curve maxima that corresponds to Group B (blue) at both 23 and 2.5 s were a higher value relative to those of Group A (yellow). This is evident as the “peaks” of the regression curves for Group B are shifted to the right in relation to those of Group A along with the x-axis.

Regression graphs for Delaunay edge length frequency of stimulus #1 at 23 seconds and 2.5 seconds.
Case Study #2: Richter. To ensure that the observed distinction of fixation clustering was applicable beyond its specific style or genre, the next case study is provided using Gerhard Richter's Abstract Painting (849-2). 9 Contrary to the previous painting, Richter's work is rendered in the abstract. The composition of the painting is dominated by large swathes of pigment, primarily in magenta tones, with white, blue, and other undertones spread throughout. The reproduction used for eye-tracking recording was produced at a smaller than original scale at 45.9 × 60 inches (116.6 × 152.4 cm), whereas the original work measures 102.4 × 133.9 (260 × 340 cm) (Table 1). The reproduction was produced at +836.33% of the physical size of the computer monitor viewed by Group A subjects.
In line with all other stimuli, the average total viewing time of Richter's work for Group B was considerably longer at 84.39 s versus 41.78 s in Group A (Table 2). Figure 13 shows the fixation maps for both groups at 19 and 2.5 s. Both time scales exhibit a similar distribution of fixations, namely toward the center and center-top of the painting. At 2.5 s in Group A, there are three small fixation clusters toward the bottom of the composition that are not present in Group B. However, at 19 s, both groups develop these clusters at comparable locations. As with the previous case study, the fixations of Group B tend toward greater diffusion, and conversely, those of Group A developed with more heightened cluster densities.

Groups A and B fixation of stimulus #6 at culled time scales of 19 and 2.5 seconds.
When examining the overall coverage of fixation center points as Delaunay edges at 10 s (Figure 14), the increased spatial clustering of Group A becomes apparent, particularly when the entire composition is viewed. In Group A, there are clear aggregations of shorter length edges alongside those that are much longer, resulting in a more varied hierarchy when compared to the tessellation set of Group B. Although both sets contain areas of high and low fixation concentrations, the degree of variation distinguishes one participant pool from the other.

Detail of Delaunay triangulation form fixation center points of stimulus #6 from groups A and B at 10 seconds.
Using the same parameters for polynomial regression as in the previous case study, we see a comparable trend in which the x-value corresponding to each maxima is greater for those in Group B (blue) versus those of Group A (yellow) (Figure 15). As with the general trend for most stimuli comparisons, this distinction became less pronounced as time scales increased.

Regression graphs for Delaunay edge length frequency of stimulus #6 at 19 seconds and 2.5 seconds.
Discussion
The principal aim of the current study was to use eye tracking in order to understand whether paintings were viewed differently when presented to a seated observer viewing a computer monitor, versus when the same work was viewed freestanding as a much larger printed reproduction. As the above data indicate, a salient difference was found in the densities of fixation patterns. Supporting the prior assumption that those viewing a computer monitor would view works with a higher degree of correlative areas of interest, it was found that there was indeed an increase in the range of fixation distances from such participants, indicating more concentrated areas of common visual interest. Comparisons of Delaunay edge lengths between fixation points were used to corroborate these results. Although basic visual patterns of those viewing the same works at a large-scale corresponded at key figurative moments within a painting, there was a higher degree of fixation dispersion, suggesting lower correlative visual interest within this respective group. Contrary to the prior assumption that an increase in canvas size would exacerbate this difference of fixation clustering between participant groups, no such relationship was determined.
Correlation of Viewing Time to Personal Preference
Although not an initial question of the current study, it was significant to observe a direct correlation between length of viewing time and the physical mode of encounter with a work. This result supported the finding of Brieber et al. (2014), whose method of research at the Wien Museum MUSA aligned with the mode of viewing of Group B. Given the data, there was clear evidence that physically moving about an image significantly increased total view times, as opposed to merely sitting in front of a digital projection on a monitor. Although this may be due to the variables involving the relatively much larger stimuli sizes of Group B, it is important to highlight that the longer viewing trend also presented in the view of stimulus #5, whose printed dimensions were nearly identical to the area of the testing monitor. Although there was only one stimulus at such a scale, the significant durational difference of +116.49% (Table 2) was compelling.
As a supplementary finding, the correlation between relative viewing time and personal preference was also noteworthy (Figure 3), particularly as it pertained to similar measures in both testing groups. There was a salient correlation between personal preference of a work with the average total time that it was viewed. Although in agreement with Brieber et al. (2014) regarding longer times for in-person viewing, this study did not register any association with gallery-type viewing and increased likeability of images as their study suggested. As Figure 3 indicates, the overall preference ranking of works was largely comparable whether viewing paintings on a computer monitor or as physical reproductions.
Increase in Clustering
The principal conclusion of the above analyses regarded the distinction in spatial clustering trends of fixation points between both groups. There was a persistent pattern for fixation aggregations to cluster more tightly in subjects viewing works on a computer monitor (Group A) versus those viewing large-scale printed reproductions (Group B). This distinction was more pronounced at shorter viewing times (at or toward 2.5 s), and diminished as viewing times increased, though the rate of change was variable depending on the stimulus (Figure 8). There was no verified factor that correlated the rate of spatial clustering over time, as comparisons with painting size, overall view times, and personal likeability responses yielded no conclusive correspondence. It is certainly plausible that artistic factors inherent in a painting's composition would also affect the rate of change of fixation cluster aggregations, though this set of variables was not explicitly accommodated for in this study. Nevertheless, this paper asserts that the physical mode of engagement with works had a salient and demonstrated effect on spatial hierarchies of fixation aggregations. Monitor-based viewing regularly resulted in fixations with more pronounced spatial clustering versus canvas-based viewing, whose associated fixations were more regularly spatially distributed.
In being seated with a relatively fixed head position, the nature of Group A's physical engagement was far simpler with fewer opportunities for personal variation. Participants in this group were seated at a prescribed distance away from the computer monitor. Given the relatively familiar size of a 24″ computer monitor, little to no physical accommodation was required during the course of viewing, such as major vestibulo-ocular movement or head sways. All Group A subjects were seated in order to encourage stillness due to the stationary placement of the eye-tracking camera.
In sharp contrast, subjects in Group B had few restrictions on how they could visually engage with works. The only directive was that they follow a walking path taped onto the floor of the testing area for the purposes of spatial calibration. This path led to the viewing stimulus, after which they could move about the scene at their personal discretion. This freedom of movement, coupled with significantly longer viewing times, inevitably led to higher variations of visual engagement. The lower degrees of compactness in fixation clusters of Group B participants reflected the diversity of the subjects’ physical encounters. In principle, this conclusion resonates with the correlation suggested by Garbutt et al. (2020) between the physical size of a painting and its average viewing time, given that there is more “visual territory” to process. This finding was evident in moderate correlation in the current study as presented in Figure 4. However, no compelling relationship was found between the physical size of an image and the rate of fixation clustering over time.
In regard to these observations, the role of parafoveal perception should be considered. The computer monitor viewed in Group A encompassed a much smaller area of the subjects’ total field of view when compared to the visual coverage by printed reproductions in Group B. As such, key narrative and compositional elements could be more easily cued in parafovea, resulting in greater consolidation of visual attention, particularly at shorter viewing times as supported by the testing data.
Contrary to the prior assumption regarding an exacerbation of discrepancy trends between participant groups as printed canvas sizes increased in Group B, no such correlation was found. Across a wide range of reproduced stimulus size increases from ∼0% to ∼1,600% in relation to the surface area of the viewed computer monitor, the lack of correspondence suggested that the above findings regarding varying cluster densities between groups were nonscalar, at least within the physical sizes presented in the current study. This finding corresponds to the conclusions of Clarke, Shortess, and Richter (1984), in which stimulus size did not necessarily provide a linear basis upon viewers’ aesthetic preferences and engagement with works of art.
Despite the relatively narrow sampling of eight stimuli, it was significant that the above trends were produced across an array of artistic styles. From relatively realistic figurative scenes to purely abstract compositions, the set of included paintings demonstrated consistent results in cluster fixation aggregations, as well as viewing time and preference scale correlations.
Limitations and Future Study
Although the current study was originally conceived in order to compare the gaze of viewers in a typical gallery setting versus those browsing works online, subject behaviors were recorded in environments that necessarily reduced the number of variables that are otherwise present in more natural viewing scenarios. Perhaps as the most notable example, curation of the physical space in which a work is presented remains a critical part of its reading. Due to the quantity of factors considered by the current analysis, environmental curation and its repercussions were not integral to the subsequent analysis.
Beyond physical curation in terms of architecture, landscape, or accompanying works, the typical “action” of viewing works at a gallery can include a bevy of variables that were unaccounted for in this study. Carbon (2017) identifies a number of important considerations when analyzing the viewing behaviors of visitors with “natural” or “ecologically valid” testing scenarios, such as how viewing time increased when attending to work in a group. Carbon also makes mention of the importance of considering the physical return of visitors to a work of art after an initial view. This critical real-world aspect was not addressed by the current study given the difficulty of ensuring parity with seated subjects viewing a computer monitor.
In line with the innumerable variables at play in typical art viewing conditions, rarely, if ever, do people engage with a single work within a fully controlled environment. Unlike the finding of Smith, Smith, and Tinio (2017) in which viewers attended to paintings for an average of 27.2 s, average viewing times in the current research extended well beyond at 92.0 s, which was most likely due to the laboratory settings, including subjects’ cognizance of follow-up questions regarding artworks.
In a similar vein, the presentation of works on a computer monitor in Group A did not involve many of the nuances that are inherent to modern online gallery viewing. Cultural institutions are developing sophisticated online and digital browsers that extend well beyond simple image slides (King et al., 2021; Windhager, Salisu, & Mayr, 2019). In exploring online collections, a visitor is not suddenly presented with an image, rather, one needs to traverse a series of “spaces” in the form of web pages and user interfaces which may ultimately affect viewing behavior. Complex interfaces were avoided in this study in order to reduce the quantity of possible variables that could ambiguate analysis.
A regrettable repercussion of selecting stimuli based on criteria requiring extremely high-resolution images, ease of accessibility of digital files, and open usage rights for research and publication, was that there was a clear bias toward Western works produced by relatively well-known artists or housed within renowned institutional collections, as such public availability often requires significant resources and distribution infrastructures. The current research could be improved by implementing a range of works from a wider range of artists, periods, cultures, and styles.
The material characteristics of a painting serve a vital aesthetic role. From qualities such as impasto to the subtle variance of sheen with the use of pigments, such aspects of painted works were not present in the printed reproductions of Group B, and therefore, they were not produced with the same fracture inherent in actual paintings. It is certainly plausible that such surface features could influence a viewer's visual and physical engagement.
Finally, an inherent limitation to the methodology conducted by the current research involves the implementation of two different eye-tracking apparatuses, involving a stationary eye-tracker for Group A and a combination of wearable eye-tracking spectacles, a head-mounted VSLAM unit, and RGBD cameras, for Group B. This distinction was necessary in order to track the visual gaze of two subject pools with significantly different parameters for physical and perceptual engagement. As such, the setup for one apparatus could not be calibrated to the other.
Conclusion
Online publication of cultural heritage has been a major concern for institutions since the early days of public internet access. This development not only pertains to questions of exhibition and archiving but also has the potential to accommodate for inequities of access to cultural heritage, expanded research possibilities, and opportunities in which digital forms of the exhibition can supplement conventional in-person visitation during periods where physical access is not possible. In light of these concerns, this paper looked to compare the perceptual behaviors of subjects in view of paintings, either in-person using printed reproductions or as digital projections on a computer monitor.
From the resultant data and analysis, the viewing behaviors between these two groups of subjects demonstrated both convergent and divergent patterns. Subjects in both groups tended to view works that were more personally favored for relatively longer periods of time. In contrast, average viewing times for in-person observation were significantly longer than those viewing on a monitor. Most importantly, the analysis showed that there was increased clustering of eye-tracking fixations when viewing works on a monitor and a corresponding decrease in the variance of how participants visually navigated across the space of images. This research asserts that the added complexities of viewing a painting in-person, including head and ambulatory movements, encourage a higher degree of personalized visual navigation within a scene. Through the required physical movement around a canvas, the presentation of the painting's image is constantly unfolding in front of the viewer, and must therefore be processed across a broader field of visual conditions, and further multiplied by a longer space of viewing time.
Despite the findings of this study, there is no implied argument that viewing a work on a computer monitor is somehow a lesser form of in-person viewing, but rather, it is argued that it serves as simply another opportunity for aesthetic engagement. Although there may be no substitute for viewing a painting “in the flesh,” this contention should not be conflated with the idea that art cannot communicate in altered forms. Instead, modes of interaction such as online or digital viewing should not necessarily be understood as a “reduction” to viewing a work in its original format, but as an additional form of how art can reach wider audiences and resonate beyond past, current, or future technologies.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
