Abstract
A major objective of perception is the reduction of uncertainty about the outside world. Eye-movement research has demonstrated that attention and oculomotor control can subserve the function of decreasing uncertainty in vision. Here, we ask whether a similar effect exists for awareness in binocular rivalry, when two distinct stimuli presented to the two eyes compete for awareness. We tested whether this competition can be biased by uncertainty about the stimuli and their relevance for a perceptual task. Specifically, we have stimuli that are perceptually difficult (i.e., carry high perceptual uncertainty) compete with stimuli that are perceptually easy (low perceptual uncertainty). Using a no-report paradigm and reading the dominant stimulus continuously from the observers’ eye movements, we find that the perceptually difficult stimulus becomes more dominant than the easy stimulus. This difference is enhanced by the stimuli’s relevance for the task. In trials with task, the difference in dominance emerges quickly, peaks before the response, and then persists throughout the trial (further 10 s). However, the difference is already present in blocks before task instruction and still observable when the stimuli have ceased to be task relevant. This shows that perceptual uncertainty persistently increases perceptual dominance, and this is magnified by task relevance.
In binocular rivalry, two distinct stimuli are presented to the two eyes resulting in perception to alternate between the two (Wheatstone, 1838). Many factors have been shown to influence the relative dominance of either stimulus. This includes comparably simple features such as contrast, size, spatial frequency, and retinal position (Blake et al., 1992; Kang, 2009; Levelt, 1965; O’Shea et al., 1997) but also configural aspects such as the adherence of a rivaling figure to principles of perceptual organization such as Gestalt (Alais & Blake, 1999; De Weert et al., 2005). Besides factors pertaining to the stimulus as such, dominance in binocular rivalry is also affected by associations of the stimuli to other domains, such as to other sensory modalities (Einhäuser et al., 2017a; Kim & Kim, 2019) or to reward and punishment (Marx & Einhäuser, 2015; Wilbertz et al., 2014). Moreover, attention influences binocular-rivalry dynamics in at least two distinct ways. First, attention to the binocular-rivalry stimuli in general increases the frequency of rivalry alternations (i.e., shortens dominance durations; Paffen et al., 2006), while the withdrawal of attention from the rivalling stimuli can abolish rivalry in full (Brascamp & Blake, 2012). Second, attending to a specific stimulus increases its perceptual dominance relative to the other, unattended, stimulus (Marx & Einhäuser, 2015; Ooi & He, 1999). Related to the deployment of attention, there is at least some degree of volitional control on temporal dynamics of rivalry (van Ee et al., 2005), though it has been argued that the ability to control which specific stimulus or perceptual interpretation is perceived is weaker for binocular rivalry than for other forms of multistability, such as ambiguous figures (Meng & Tong, 2004).
Even without explicit or even volitional attentional deployment to a stimulus, its relevance for a task can enhance its dominance relative to the other stimulus: Chopin and Mamassian (2010) combined binocular rivalry with a search task. They presented two gratings of different orientation, and one of the two gratings contained a contrast-defined target. In some trials, observers had to search for the target, while in the others, they reported their percept of the two gratings. As critical manipulation, during the middle blocks of their experiment, the target occurred always in the same grating. The authors found that the dominance of this grating increased relative to the other grating, and this bias in dominance persisted even after the target-grating association was removed. Interestingly, this increase in dominance was restricted to the first period of the trials in which perception was reported. This shows that implicit learning of a stimulus’ task relevance enhances its dominance in rivalry persistently.
When studying binocular rivalry in combination with other factors or tasks, potential task interference and the reliance on subjective report become critical factors. Withdrawing attention from one or both rivalling stimuli, for example, often involves a secondary task, which interferes not only with the perception but also potentially with reporting the current perceptual state. Even though most of the aforementioned studies use one or more measures to exclude the possibility that observers merely report what they believe they are expected to (e.g., by using catch trials, or tasks that can only be executed when a specific stimulus is visible), it is hard to rule out such biases in full. This is especially so, when studying volitional control or coupling the preferential perception of one stimulus to explicit reward or higher performance. Moreover, when the task of reporting the perception is separated in time from the manipulation of interest, the temporal resolution for studying effects of the manipulation is limited. As a complementary approach, no-report paradigms have been suggested and used in studying binocular rivalry and related forms of multistability (e.g., Fahle et al., 2011; Frässle et al., 2014; Naber et al., 2011; Wilbertz et al., 2018; see Koch et al., 2016; Overgaard & Fazekas, 2016; Tsuchiya et al., 2015 for reviews). By avoiding the necessity to query observers about their perceptual state, but instead continuously inferring their perception from physiological measures such as eye movements, task interference or biases are circumvented. In addition, no-report paradigms can provide a more fine-grained temporal resolution, as they do not rely on discrete button presses, and may more faithfully capture nondiscrete aspects of perception (Fahle et al., 2011; Naber et al., 2011). Finally, they prevent any influence the action of percept reporting may have on multistable perception (Beets et al., 2010; Veto et al., 2018; Wohlschläger, 2000).
Presumably due to the challenges associated with bias-free querying the percept in such contexts, only few studies in binocular rivalry—and in multistability in general—have addressed how the possible consequences of a specific percept influence its perceptual dominance. This issue is, however, of high relevance, as it points to the more general question of how (expected) outcome or value (here the value associated with a certain perceptual interpretation) affects perception. In binocular rivalry, this influence can be studied in a most direct way, as no stimulus changes are required to see impact on perceptual representations. Studies that did address effects of outcome on multistable perception so far have mostly considered motivational value (i.e., reward or punishment) or task relevance as detailed earlier. Perceptual consequences such as the reduction in uncertainty or the gain in information have been addressed in greater detail in studies on the control of voluntary eye movements (for reviews, see Eckstein, 2011; Gottlieb, 2018; Schütz et al., 2011; Tatler et al., 2011). While studying eye movement control has its merits in its own regard, perceptual and motor components are necessarily intertwined, which makes it difficult to isolate purely perceptual effects in those paradigms.
Here, we use a no-report paradigm to address whether the information contained in the stimulus itself has an impact on perceptual dominance and whether this is modulated by relevance for the task. As a proxy for perceptual uncertainty, we use the visibility of a subtle modulation on top of a high-contrast visual stimulus. This notion to reduce certainty about a visual stimulus by reducing its visibility is conceptually similar to studies testing Bayesian cue integration, where spatial reliability of a stimulus is reduced by blurring (e.g., Alais & Burr, 2004), which is often interpreted as an increase of uncertainty about this stimulus (e.g., Knill & Pouget, 2004). Specifically, we present observers with high-contrast gratings that have a pattern embedded on them that is either of high visibility (low noise) or low visibility (high noise). If perception aims at gathering as much information as possible about the stimuli, we expect the low-visibility pattern to be dominant longer when brought in rivalry with the high-visibility pattern. Conversely, if perceptual dominance were driven by the salience or vividness of patterns on a given stimulus, we would expect the grating with the high-visibility pattern to be more dominant when brought in rivalry with a low-visibility pattern. Given the results on task relevance (Chopin & Mamassian, 2010), we expect in any case that the dominance of the low-visibility pattern increases, once the information contained in the pattern is required for the task. As visibility in the present context refers to the visibility of the pattern, not of the overall stimulus, we for notational simplicity refer to the low-visibility versions as “(perceptually) difficult” and to the high-visibility versions as “(perceptually) easy” to avoid any confusion with the overall stimulus reaching awareness in binocular rivalry (i.e., with the stimulus becoming visible).
Methods
Participants
Twenty-four volunteers (20 women, 4 men, age range 18–29 years), who were naïve to the hypotheses, participated and were included in the analysis; one further participant was excluded prior to any analysis as they indicated after the first block that required a response that they had misunderstood the task. All gave written informed consent prior to participation and were reimbursed by either course credit or 8€/h. The number of participants was determined prior to the experiment as the minimum required to counterbalance conditions as described in the Procedure section.
Setup
For dichoptic presentation, we used a Wheatstone-type stereoscope with cold (infrared-transparent) mirrors. In this setup, one mirror is placed in front of each eye at a 45° angle to the frontoparallel plane such that the mirrors meet directly at the observer’s head. Monitors are placed at the sides for stimulus presentation with their screens perpendicular to the frontoparallel plane. The path from each of the participant’s eyes via the mirror to the screen is covered by a metal cone extending toward the screen and covered from the inside with black felt. A slightly modified version of this setup including a construction instruction has been described in detail recently (Brascamp & Naber, 2017).
The stimulus for each eye was presented on a 27-inch (LG 27GK750F) monitor, operating at 1,920 × 1,080 pixels resolution and 120 Hz refresh rate. The distance along the line of sight from eye to monitor was 30 cm.
Eye position of the left eye was tracked at 1000 Hz using an Eyelink-2000 eye tracker (SR Research Ltd, Ottawa, ON, Canada) that was mounted behind the mirrors (and thus invisible to the observer) exploiting the mirrors’ transparency for infrared light.
The experiment was implemented in MATLAB R2017B (The MathWorks Inc., Natick, MA, USA), including the psychophysics and Eyelink toolboxes (Brainard, 1997; Cornelissen et al., 2002; Pelli, 1997).
Stimuli
We created oriented noise stimuli at two different noise levels (high, low), two different orientations (leftward tilted, rightward tilted), and two different colors (blue, red) by the following procedure. First, we generated a white noise pattern in which each pixel was drawn independently from a uniform distribution between 0 and 1. This pattern was then convolved with a Gabor of +45° or –45° orientation (spatial frequency: 0.79 cyc/deg, envelope standard deviation: 1.5 deg). The result was normalized to the dynamic range from 0 to 1. This orientation-filtered noise served as basis for high-noise and low-noise stimuli. For the low-noise stimuli (Figure 1A), the result was multiplied with a vertical square-wave grating of wavelength 4.8 deg, which had an on-phase only for the respective color channel (red or blue). For the high-noise stimulus (Figure 1B), the orientation-filtered noise was first pointwise added to the original pixel-noise stimulus, where the former had a weight of 30% and the latter of 70% and then also multiplied with the square-wave grating. The maximum luminance in the on phase for both colors was 12 cd/m2 and in the off phase was below 0.5 cd/m2.

Stimuli. (A and B) Stimuli used in the present experiment. Top: blue grating, bottom: red grating; odd columns: leftward tilt; even columns: rightward tilt. (A) low-noise (“easy”) gratings, (B) high-noise (“difficult”) gratings. Only a part of each grating is shown, corresponding to the part visible at a given frame (see also Movie 1). (C) Luminance along a vertical line for the top-left grating of Panel A. (D) Luminance along a vertical line for the top-left grating of Panel B; scale matches Panel C. (E) Average amplitude spectrum of all vertical lines in all colored phases of the low-noise gratings. Vertical line denotes spatial frequency of Gabor used to generate stimuli (divided by sqrt(2) because of the projection to the vertical). (F) Average amplitude spectrum of all vertical lines in all colored phases of the high-noise gratings; scale matches Panel E. (G–J) As Panel D through F for horizontal scan lines. The value of the DC component (f = 0), that is, the stimulus mean, is omitted from Fourier spectra.
During the experiment, gratings drifted outward (i.e., in nasotemporal direction; to the left in the left eye and to the right in the right eye) at a speed of 14.3 deg/s. The part leaving the screen on one side reappeared on the other, as if the grating were mounted on a drum. To provide a plane of constant perceived depth and to avoid extremely peripheral stimulation, stimuli were surrounded by a high-contrast random noise pattern that was identical in both eyes. The pattern consisted of squares of 1.3 degree wide squares that were randomly set to either black (<0.5 cd/m2) or white (240 cd/m2). This left the central 28.3 × 27.3 degrees of the grating visible (Movie 1 in the Supplemental Material). We chose this particular pattern, as a combination of color and motion and using a square-wave (or in some cases near-square-wave) grating in our hands most robustly induces rivalry (cf. e.g., Einhäuser et al., 2017a, 2017b), when using the large field drifting stimuli that are needed for inducing robust optokinetic nystagmus (OKN).
Stimulus Characterization
When considered on a pixel level, high-noise stimuli had a root-mean-square (RMS) contrast (i.e., standard deviation of luminance divided by mean luminance) of 110%, and low-noise stimuli had an RMS contrast of 91%. However, this difference almost exclusively arose from the pixel-wise white noise; already taking the average over neighboring pixels (i.e., filtering the stimuli with a 2-by-2 pixel wide filter) and computing RMS contrast thereafter resulted in about equal RMS contrast levels (88% for high noise, 90% for low-noise). Because a pixel spanned about 0.06 degrees of visual angle (3.6 arc min), it is unlikely that the former difference contributed substantially to the relevant stimulus strength.
Although we did not formally test visibility, the difference in visibility of the tilted grating is quite striking between high-noise and low-noise stimuli. This is illustrated by considering cross sections through the stimuli: When taking a vertical cross section (through the colored part), in the low-noise grating, the oscillation of the grating (in this case the projection on the vertical axis) is clearly noticeable (Figure 1C), while this is not the case for the high-noise stimulus (Figure 1D). We quantified this by taking the Fourier spectrum of these lines and averaging them over lines and stimuli. As expected, we found a peak in the amplitude spectrum at 0.56 cyc/deg, which corresponds to the projection of the Gabor’s spatial frequency on the coordinate axis (0.79 cyc/deg/sqrt(2)). However, this was substantially more pronounced in the low-noise case (Figure 1E) than in the high-noise case, where the constant (“white”) offset dominated (Figure 1F). A similar pattern was obtained for the horizontal section (Figure 1G and H), with more power at 0.56 cyc/deg for low-noise (Figure 1I) than for high-noise (Figure 1J). However, in either case, the horizontal sections were dominated by the overall vertical (“carrier”) grating, which added additional peaks to the power spectrum at the odd multiples of the spatial frequency of the carrier grating (i.e., at 0.21 cyc/deg, 0.63 cyc/deg, 1.05 cyc/deg, etc.) as expected from the Fourier transform of a square wave (∼sin(x)+sin(3x)/3+sin(5x)/5+⋯). These peaks dominated the power spectrum, which is consistent with the strong subjective visual impression of a vertical square wave grating on which the tilted colored gratings are mere add-ons. Because the gratings were drifting at constant velocity in horizontal direction, these carrier frequencies (multiplied by the speed) will also dominate the temporal spectrum at any given point.
Procedure
The experiment consisted of three phases split in a total of 16 blocks. In the first phase (Blocks 1 through 4) and in the last phase (Blocks 13 through 16), observers merely watched the stimuli while their eye position was tracked. In the second phase (Blocks 5 through 12), observers in addition had to indicate by a button press, whether the pattern on the two drifting gratings presented to the two eyes were tilted in the same direction (both eyes +45° or both eyes –45°) or in different directions (+45° in one eye and –45° in the other). By design, this judgment was only possible once both gratings had been perceived. Instruction about the task was only given prior to block five, such that in the first phase, observers were not made aware of the different patterns or the subsequent judgment task. Within a block, the noise levels of the red and the blue grating were held constant, while the assignment of color to eye and of color to the tilts was balanced such that each of the eight possible combinations (2 tilt directions for red × 2 tilt directions for blue × 2 eyes) occurred equally often in a block. There were 16 trials in each block with report (5–12) and 8 trials in each of the other blocks (1–4, 13–16). In blocks with report, trials lasted 10 s after the response; in blocks without report, the whole trial lasted 10 s. Each trial was preceded by a 0.5 s blank period in which only the high-contrast frame, but no grating was presented. The order of the noise level to color assignment was counterbalanced across observers in each set of four blocks, with the order reversed in Blocks 13 through 16 relative to Blocks 1 through 4 and in Blocks 9 through 12 relative to 5 through 8 (i.e., an ABCDDCBA design). The back two buttons of a gamepad were used for the responses, with half of the observes using the left button to report “same tilt” and the right button to report “different tilt” with the assignment reversed in the other half. Before each block, the eye tracker was calibrated, and the calibration validated with a nine-point calibration for the area of the screen in which the gratings were to be presented.
All procedures were evaluated by the local ethics board (Ethikkommission Fakultät HSW) who decided that no in-depth evaluation was necessary (case no. V-228-PHKP-WET-Aufmerksamkeit-16102017).
Analysis
All data are available at https://doi.org/10.5281/zenodo.4575552. For all analyses, we distinguished four phases of the experiment (factor “experiment phase”):
“ “ “
Within each experiment phase, we compared four levels of perceptual difficulty (factor “difficulty”)
“ “ “ “
Behavior
To check whether the noise levels indeed manipulated difficulty as expected, we measured the percentage of correct responses and the median time an observer needed for a correct response. The dependence of these measures on the factor difficulty was assessed by a 1-factor repeated-measures analysis of variance (rmANOVA).
Eye-Movement Data, OKN Slow-Phase Gain
We computed the slow-phase gain of the OKN. As gratings drifted horizontally, this analysis was restricted to the horizontal axis. As a first step, we detected OKN fast phases in eye-movement signal by using the eye tracker’s saccade detection algorithm at thresholds of 35 deg/s for velocity and 9500 deg/s2 for acceleration. This exploits the fact that the dynamics of OKN fast phases is highly similar to that of saccades (Garbutt et al., 2001). These fast phases were treated as missing data. For each slow phase (the eye position between two subsequent fast phases), we fitted a linear function in the horizontal eye-position data to obtain the horizontal velocity. By dividing this horizontal velocity by the horizontal stimulus velocity, we obtained the slow-phase gain. The sign of the gain for each trial was chosen such that a positive gain corresponded to a slow phase in the drift direction of the red grating, and a negative gain corresponded to a slow phase in the drift direction of the blue grating. Perfect tracking of an exclusive “red-grating” percept would therefore result in a gain of +1, of an exclusive “blue-grating” percept in a gain of –1.
Perceptual Dominance
From the gain, we defined a measure of relative dominance. For each observer and condition (difficulty × experiment phase), we determined the duration during which the gain was positive (red grating dominant) and the duration during which the gain was negative (blue grating dominant). We subtracted the two values from each other and divided them by their sum (corresponding to the aggregate duration of all slow phases). This results in a measure of relative perceptual dominance, which is +1 if the red grating is dominant throughout, –1 if the blue grating is dominant throughout, and 0 if both gratings are dominant for the same amount of time. We will refer to this measure as relative dominance.
Statistical Analysis
For each observer, difficulty level, and experiment phase, we determined the mean gain and the relative dominance. We subject them to a two-factor rmANOVA with factors difficulty and phase. For the follow-up tests, we restricted our analysis to the main contrast of interest, the difference between “blue-difficult” and “red-difficult.” For each experiment phase, we tested whether this difference was different from 0 (following up on a main effect of difficulty and a difficulty × experiment phase interaction, as this is equivalent to testing the difference between these conditions), and whether this difference itself depended on the experiment phase (following up on a main effect of experiment phase and a difficulty × experiment phase interaction).
Dominance Durations
We defined a dominance phase as the period between subsequent sign changes of the OKN slow phase (i.e., from the start of the first slow phase after onset or a sign change to the end of the last slow phase before a sign change or the trial end). For each observer, stimulus, and condition, we computed the median duration of these phases. Separately for each stimulus and phase, we computed the dependence of this median dominance duration on condition by means of an rmANOVA. Where appropriate, we followed up on the ANOVAs by comparisons of each condition to the both-easy condition.
Temporal Dynamics of the Gain
In addition to considering the aggregate data, we analyzed how the gain developed over time by considering event-triggered averages in gain for the two events of relevance: trial onset and response. For statistical analysis, we considered only the differences between the time course in “red-difficult” and “blue-difficult,” which we compared by a paired t test at each time point and correcting for multiple comparisons at an expected false discovery rate (FDR) of 5% using the procedure by Benjamini and Hochberg (1995).
Results
Behavior
As a manipulation check, we analyzed whether the orientation judgment was indeed more difficult when one or two of the gratings was of high noise. Observers performed the same/different judgment near ceiling in all conditions (Figure 2A) with no difference between conditions, F(3, 69) = 0.53, p = .660. However, there was a main effect of condition on response time, F(3, 69) = 4.83, p = .004 (Figure 2B) with the response in the “both-easy” condition faster than in all other conditions, all |t(23)| > 2.52, all p < .019, with no differences among those, all |t(23)| < 0.41, all p > .689. This confirms that the experimental manipulation worked as intended: When the signal strength in at least one of the stimuli was low, the task became more difficult (i.e., required more time) than when the signal strength was high for both stimuli.

Behavior. (A) Percentage of correct judgments by condition. (B) Median response time for correct trials by condition. Error bars denote standard error of the mean across observers.
Relative Dominance
We found a main effect of difficulty, F(3, 69) = 31.3, p < .001, and of experiment phase, F(3, 69) = 3.33, p = .024, on relative dominance as well as an interaction between these factors, F(9, 207) = 3.79, p < .001 (Figure 3A). We focused the follow-up tests on the contrasts of interest, namely the difference between red-difficult and blue-difficulty. If difficulty increases perceptual dominance, we expect this difference to be larger than 0. This was indeed the case numerically for all phases (Figure 3B). The difference was significant in the phase before the comparison task had been instructed, “no-task I”, t(23) = 2.85, p = .009, as well as before the response, t(23) = 7.56, p < .001, and after the response, t(23) = 6.14, p < .001, in trials with the task. For the trials without the task after the blocks with task, a similar trend was observed, “no-task II”, t(23) = 2.00, p = .057. Importantly, there was an effect of experiment phase on this difference, F(3, 69) = 5.35, p = .002; note that this could be different from the main effect in the two-factor ANOVA as it does not involve the “both-easy” and “both-difficult” condition. In the follow-up tests, we observed no significant difference between before the response and after the response in trials with task, t(23) = 1.80, p = .08. Comparing the response trials to the no-task trials, we found differences between before the response and each of the no-task phases, compared to “no-task I”: t(23) = 4.88, p < .001; to “no-task II”: t(23) = 2.39, p = .026, while for the “after-response,” this held only for the comparison to “no-task I,” t(23) = 2.78, p = .011, and not to “no-task II,” t(23) = 1.86, p = .076. There was no difference between the two “no-task” phases, t(23) = 0.01, p = .992. Taken together, conducting the comparison task clearly augmented the relative dominance of the more difficult-to-perceive grating, but the effect was already apparent before any task has ever been required and persisted within a trial after the response had been given.

Perceptual dominance as function of perceptual difficulty. (A) Relative dominance of the red grating as a function of condition and experiment phase (0 implies equal dominance of both gratings). (B) Difference in relative dominance between “red-difficult” and “blue-difficult” condition. (C) Gain in direction of the red grating (+1 implies red grating dominant, –1 blue grating dominant). (D) Difference in gain between “red-difficult” and “blue-difficult” condition. In all panels, error bars denote standard error of the mean across observers.
Gain
As an alternative measure of perceptual dominance, we computed the average OKN slow-phase gain. Rather than considering only the slow-phase direction, and thus assuming a binary decision between red dominance and blue dominance, this measure weighs phases with lower absolute gain, which presumably correspond to less exclusive or vivid percepts, less than phases of high gain, which presumably correspond to (near) exclusive dominance. The result pattern was remarkable similar to the relative-dominance measure (Figure 3C): there was a main effect of difficulty, F(3, 69) = 44.9, p < .001, a trend to a main effect of experiment phase, F(3, 69) = 2.70, p = .052, and an interaction of the factors, F(9, 207) = 4.16, p < .001. Following up on the interaction, we followed the same procedure as for the relative-dominance measure and analyzed the differences between “red-difficult” and “blue-difficult.” We found them to be different from 0 for all experiment phases except “no-task II,” t(23) = 1.86, p = .076; all other t(23) > 4.48, p < .001 (Figure 3D). The difference depended on experimental phase, F(3, 69) = 5.86, p = .001. In blocks with a task, we found no difference between before and after the response, t(23) = 1.76, p = .092. The “before-response” phase was different from both no-task phases, I: t(23) = 3.26, p = .003; II: t(23) = 3.06, p = .006, while the difference between “after-response” and “no-task” was significant only for “no-task II,” t(23) = 2.51, p = .020, but not for “no-task I,” t(23) = 1.94, p = .065. There was no difference between the two no-task phases, t(23) = 1.04, p = .311. In sum, this corroborates the relative-dominance analysis: The “more difficult” grating dominateed more and this depended on the task, but the effect was present prior to any instruction on this task.
Time Course
Numerically, we found in both measures reported earlier that the dominance of the “more difficult” grating is more prominent prior to the report than after the report. While these differences were not significant at a 5% level, the trend inherent in them merits a more detailed, temporally resolved analysis over the whole duration of a trial. To this end, we aligned all trials with report to the time of report and averaged the gain separately for the four difficulty conditions in a time window starting 2 s before the response until the end of trial. We found that the gain was consistently higher for the “red-difficult” condition than for the “blue-difficult” condition with the other two conditions having gains consistently in between (Figure 4A). The difference between “red-difficult” and “blue-difficult” was significant at all time points in the time window (all p < .026; by definition all points remain significant after an FDR correction, if all values are below the set expected FDR) and peaked 684 ms prior to the report (Figure 4B). The fact that we observed differences already close to trial onset, which were indeed similar to those at a trial’s end, was a consequence of the blocked conditions; there was a preference for the red grating to be dominant in “red-difficult” blocks and a preference for the blue grating in “blue-difficult” blocks. The peak around 600–700 ms prior to the response likely reflects the necessity to extract information from the difficult grating before a response can be made.

Time course of gain. (A) Mean gain for the four different difficulty levels in the task phase (Blocks 5–12) aligned to the time of report. Positive numbers correspond to dominance of the red; negative numbers correspond to dominance of the blue grating. (B) Difference between gain in “red-difficult” and “blue-difficult” condition of Panel A, mean, and SEM across observers. (C) Difference in gain as in Panel B but aligned at trial onset and depicted for the different experiment phases. Markers at top denote time points with differences significantly different from 0 at an uncorrected alpha level of 5% and at an expected false discovery rate (FDR) of 5%, respectively. FDR-adjusted values are not shown for the no-task II phase, because no individual time point reached significance at this corrected level.
Aligning the curves to trial onset rather than to the response yielded a similar picture for the difference between “red-difficult” and “blue-difficult” in trials with task (Figure 4C, black): This difference was numerically above 0 throughout the trial and significantly so at trial onset. After about 300 ms, the difference quickly rose and had a peak about 1.5 s wide. With the exception of a short phase between 1,776 ms and 1,816 ms after trial onset, the difference was significantly larger than 0 at an expected FDR of 5% from 289 ms after trial onset until the end of the analysis window (10 s after trial onset to match the no-task phases). This supports the notion that in addition to a response-related early peak, the relative dominance of the perceptually difficult grating persisted at a rather constant level throughout a trial. In contrast, the no-task conditions showed substantial fluctuations throughout the trial (Figure 4C, gray). These remained clearly above 0 on average and the no-task I phase occasionally numerically even exceeded the phase with task. However, there was a substantial fraction of individual time points at which the difference was indistinguishable from 0 at an uncorrected 5% level—to an extent that no individual timepoint reached significance at an FDR-adjusted 5% level in the no-task II phase (Figure 4C, top). In sum, this shows that perceptual difficulty induced a persistent increase in dominance, which was, however, more robust and stable when the stimuli were task relevant.
Onset
The first percept after stimulus onset in multistability frequently differs in its properties from later percepts (e.g., Carter & Cavanagh, 2007; Chong & Blake, 2006; Hupé & Pressnitzer, 2012). In binocular rivalry, there is typically a bias toward the stimulus of larger strength at onset (Mamassian & Goutcher, 2005). Hence, we asked for blocks in which stimuli of distinct difficulty were presented (i.e., red easy and blue-difficult or vice versa) how frequently the first percept corresponded to the difficult stimulus (i.e., how often the first OKN slow phase went in the direction of the difficult stimulus). For the no-task I phase, we found the difficult stimulus to dominate first in 52.1% (SD: 9.0%) of trials, a fraction indistinguishable from chance, t(23) = 1.14, p = .267. Similarly, there was no significant effect in the no-task II phase with 50.3% (10.1%) of trials with the difficult stimulus dominating at onset, t(23) = 0.118, p = .907. In the trials with task, there was a trend toward the more difficult stimulus being perceived more frequently at onset (52.5%, SD: 6.4%), which however did not reach significance, either—t(23) = 1.89, p = .071. We also found no difference between the phases, F(2, 46) = 0.41, p = .668. Noting that this analysis was based on only a small set of data (each trial contributes only one data point) and that the continuous analysis (Figure 4C) has shown a difference early after onset, we cannot fully exclude that there was some bias toward the more difficult stimulus for the very first percept after onset, but—if anything—it was subtle compared with the effects later in the trial.
Learning Over Trials
The data aggregated over each experiment phase had shown that there was some bias toward the more difficult stimulus in blocks prior to the task, which was then enhanced in blocks with task. This raises the question as to whether there is a gradual buildup of this bias and/or its enhancement (i.e., some form of learning) over the course over the blocks or whether instead usefulness strengthens the dominance of the difficult stimulus immediately. Similarly, we have seen persistence of the bias after the task, but the analysis so far has left open whether there is an immediate drop-off to a pretask level or a gradual decline. Hence, we asked how the differences in relative dominance and gain developed across trials. Because we can only consider trials with distinct difficulties for red and blue grating, we have 16 relevant trials in the no-task I phase, 16 in the no-task II phase and 64 in the phase with task. We correlated the serial position (1…16 or 1…64) of a trial among these trials against the mean relative dominance or mean the gain, where we chose the sign such that positive values imply greater dominance of the more difficult grating. For these measures, we found no significant correlation in the “no-task I” phase, relative dominance: r(14) = .119, p = . 661, Figure 5A; gain: r(14) = .076, p = .779, Figure 5B. For the phase with task, we found a small but significant correlation for the time before the response, relative dominance: r(62) = .395, p = .001, Figure 5C; gain: r(62) = .258, p = .039, Figure 5D, but not after the response, relative dominance: r(62) = −.018, p = .888, Figure 5E; gain: r(62) = -.088, p = .489, Figure 5F, indicating that the effect before the response increased with time on the task. For the “no-task II” phase, we observed a significant negative correlation, relative dominance: r(14) = −.672, p = .004, Figure 5G; gain: r(14) = −.595, p = .015, Figure 5H, indicating that the effect declined after the stimuli had become task irrelevant.

Learning over trials. Relative dominance (top row) and gain (bottom row) in the direction of the difficult grating plotted against the serial order of the trial, including only blocks where one difficult and one easy grating was shown. Mean and SEM across observers; optimal linear regression. (A, B) No-task I phase; (C, D) blocks with task, data before response; (E, F) blocks with task, data after response; (G, H) no-task II phase.
Dominance Durations
So far, we have considered only the relative dominance—either measured directly or in terms of average gain—of the two stimuli. However, an increase in relative dominance can be caused by two distinct effects, a decrease of dominance durations for one stimulus or an increase of dominance durations for the other. Moreover, even if the relative dominance remains unchanged, both stimuli can concurrently increase or decrease their dominance durations. The difference is of theoretical importance, as a change in stimulus strength (e.g., contrast) of one stimulus primarily leads to a decrease in the other stimulus’ dominance durations (second proposition of Levelt, 1965) and concurrent increases in stimulus strength lead to a shortening of dominance durations (Levelt’s fourth proposition). If an increase in perceptual difficulty is equivalent to a change in stimulus strength, we should observe these effects for perceptual difficulty as well; in turn, if we do not find these patterns, we can exclude that the observed effects of perceptual difficulty are primarily driven by an unintended difference in stimulus strength.
In the no-task I phase, we found a main effect of condition on dominance duration for the red stimulus, F(3, 69) = 4.10, p = .001, and a corresponding trend for the blue stimulus, F(3, 69) = 2.32, p = .083. For the blue-difficult condition, there was a trend to an increase in dominance duration for the blue stimulus, t(23) = 1.97, p = .061 (Figure 6A) and a decrease for the red stimulus, t(23) = 2.81, p = .010 (Figure 6B) relative to the both-easy condition. For the red-difficult condition, neither the dominance duration of the blue stimulus, t(23) = 0.18, p = .858, nor of the red stimulus, t(23) = 1.22, p = .236, was different from the both-easy condition. The pattern suggests that the observed changes in relative dominance occurred as a consequence of one stimulus becoming dominant for longer and the other becoming dominant for shorter periods of time. An increase in stimulus strength according to Levelt’s second proposition, in contrast, should shorten only dominance durations in the other stimulus. For neither stimulus did we observe a difference in dominance duration between the both-easy and the both-difficult condition—both t(23) < 0.63, p > .538—which Levelt’s fourth proposition would predict for an increase in stimulus strength. Hence, equating the increased stimulus’ difficulty with larger stimulus strength is not consistent with Levelt’s second proposition, and we find no evidence that Levelt’s fourth proposition is fulfilled, either.

Dominance durations. Median dominance durations per stimulus and condition, top row: blue stimulus, bottom row: red stimulus. (A, B) no-task I; (C, D) blocks with response, before response; (E, F) blocks with response, after response; (G, H) no-task II. All data mean and SEM across observers, dotted lines at reference condition (both-easy) to ease comparison.
In trials with task, we considered dominance periods that started before the response separately from those that started after the response. In all cases (both stimuli; before and after response), we found significant main effects of condition, all F(3, 69) > 12.4, all p < .001. Before the response, there was a decrease in dominance duration of the easy stimulus and an increase for the difficult stimulus compared with the “both-easy” case, all t(23) > 2.20, all p < .038, for both stimuli (Figure 6C and D). After the response, there was an increase in dominance duration for the blue stimulus (Figure 6E) in the blue-difficult condition, t(23) = 4.31, p < .001, but no decrease for the red-difficult condition, t(23) = 0.712, p = .480. For the red stimulus (Figure 6F), there was both an increase in dominance duration for the red-difficult condition, t(23) = 3.78, p = .001, and a decrease for the blue-difficult condition, t(23) = 2.99, p = .007. Significant differences between the both-difficult and the both-easy condition were observed only for the blue stimulus after the response, t(23) = 3.40, p = .002; all other t(23) < 1.94, p > .065.
For the no-task II phase, we did not observe a significant main effect on dominance durations for either stimulus—blue stimulus: F(3, 69) = 0.465, p = .708, Figure 6G; red stimulus: F(3, 69) = 2.44. p = .072, Figure 6H—although the pattern for the red stimulus was similar to the no-task I phase.
The lack of robust evidence for consistency with Levelt’s second proposition still holds for a modified version that posits that “Increasing the difference in stimulus strength between the two eyes will primarily act to increase the average perceptual dominance duration of the stronger stimulus” (Brascamp et al., 2015, p. 25). For the no-task I phase—if anything—it was the somewhat weaker (blue) stimulus that increased in dominance, while the somewhat stronger (red) stimulus decreased in dominance. For the condition with task, the blue stimulus tended to be stronger on average, but a combination of dominance duration increases and decreases contributed to the effects on relative dominance. Similarly, in the no-task II phase, we found no significant effects on dominance durations, indicating that increases in relative dominance and gain resulted from a combination of the dominance durations for the difficult stimulus becoming longer and shorter for the easy-stimulus relative to the both-easy and the both-difficult conditions.
In sum, under the assumption that perceptual difficulty corresponds to stimulus strength, we found little support for Levelt’s propositions, in their original and in their modified form, and violations of them in some cases. In turn, provided the validity of Levelt’s propositions, the pattern as to how dominance durations changed to yield differences in relative dominance was largely inconsistent with a mere effect of stimulus strength.
Aggregated Dominance Time Until Response
Having extracted the dominance durations allows us a further assessment of stimulus difficulty, by asking for how long a stimulus needed to be dominant prior to a response. A successful manipulation of difficulty predicts that a difficult stimulus has to be seen longer prior to the response than an easy stimulus. That is, the blue stimulus should have longer aggregated dominance times until the response in blue-difficult and both-difficult conditions than in the red-difficult and the both-easy conditions. Conversely, the aggregated dominance times until the response for the red stimulus should be longer in the red-difficult and the both-difficult condition than in the blue-difficult and the both-easy condition. We therefore measured the aggregate duration a particular stimulus was dominant before the response and compared it between the conditions. To this end, in each trial, we used the dominance durations (computed as aforementioned) that started prior to the response and summed them separately for each stimulus (color). For this calculation, the dominance phase during which the response was given was cut at the response such that only durations up to the exact time of response were summed. The resulting data were averaged across all trials of each condition, resulting in 8 values per participant (2 colors × 4 conditions).
The easy blue stimulus had to be seen for 1.04 s (SD: 0.39 s) in the both-easy condition and for 1.00 s (SD: 0.49 s) in the red-difficult condition before a response was executed, a numerical difference that did not reach significance, t(23) = 0.54, p = .592 (Figure 7A). The difficult blue stimulus had to be seen for 1.38 s (0.55 s) in the blue-difficult and for 1.20 s (0.43 s) in the both-difficult condition prior to the response. While all four pairwise comparisons between conditions with the difficult blue stimulus (i.e., blue-difficult, both-difficult) and conditions with the easy blue stimulus (i.e., both-easy, red-difficult) showed significantly larger summed durations for the difficult blue stimulus, all t(23) > 2.51, all p < .020, the difference between the blue-difficult condition and the both-difficult condition was also significant, t(23) = 2.38, p = .026. Performing the same analysis for the red stimuli (Figure 7B) resulted in a qualitatively similar pattern. The time the red stimulus had to be seen up to the response was longer for both-easy, 1.00 s (0.43 s), than for blue-difficult, 0.81 s (0.37 s), a difference that was significant, t(23) = 3.10, p = .005. Similarly, there was a difference between red-difficult, 1.24 s (0.38 s), and both-difficult, 1.02 s (0.42 s), t(23) = 6.05, p < .001) and pairwise differences between each condition with easy red stimulus (both-easy, blue-difficult) and difficult red stimulus (red-difficult, both-difficult), except the difference between both-easy and both-difficult, t(23) = 0.43, p = .674; all other t(23) > 4.21, p < .001. Consistently across both colors, the aggregate duration a stimulus was dominant before the response was smallest if the stimulus was easy and the other stimulus difficult, larger, if both stimuli were easy, at least numerically even larger, if both stimuli were difficult and largest if the stimulus itself was difficult and the other easy. This indicates that indeed the difficult stimulus had to be seen longer than the easy stimulus prior to a response, verifying further that difficulty was indeed successfully manipulated. However, the difference in dominance is not fully explained by the need to aggregate information over time, in which case we would expect no difference between the condition in which only the stimulus under consideration is difficult and the both-difficult condition, contrary to our observation.

Aggregated dominance time until response. Aggregated time a stimulus was dominant prior to the response. (A) Blue stimulus; (B) red stimulus. All data mean and SEM across observers.
Discussion
Using a no-report paradigm, we found that the dominance of a stimulus in binocular rivalry increases, when the stimulus properties are perceptually more difficult to discern. This effect increases when the manipulated property of the stimulus is task relevant. However, it exists prior to any task to be conducted with the stimuli and persists after the task is completed. Hence, perceptual difficulty increases perceptual dominance, and this effect is enhanced if a stimulus is task relevant.
Our observation that the difficulty of the perceptual pattern increases its dominance fits well with findings that consider the outcome associated with perceiving a particular stimulus. This is most evident in the case of reward, where stimuli that are associated with positive outcome (reward) are enhanced in dominance and those associated with negative outcome (punishment) are reduced in dominance (Marx & Einhäuser, 2015; Wilbertz et al., 2014). Interestingly, such reward effects have not been found when observers were unaware of the stimulus-reward contingency (Wilbertz et al., 2017), which suggests that a (meta-cognitive) awareness of a stimulus-outcome association is critical for reward-based effects on multistable perception. In this respect, the effects of perceptual difficulty, which already occur—albeit weaker—when no task is to be conducted with the stimuli, may be conceptually distinct from effects of motivational value.
When considered on the scale of a single pixel, the RMS contrast of the difficult (less visible, more uncertain) stimuli were larger than of the easy (more visible, less uncertain) stimuli. While we cannot fully rule out that this contributed to the observed differences between difficult und easy stimuli, we consider it highly unlikely to be a major factor for a number of reasons: (a) The difference in contrast only holds at the pixel-scale, corresponding to about 0.06 degrees (3.6 arcmin) of visual angle, which is unlikely to be the relevant scale for dominance of a large-scale visual stimulus, (b) this would neither explain the modulation by the task nor the steady decline after the stimuli have become task irrelevant, and (c) if the difference were due to a difference in stimulus strength, the effects should follow Levelt’s propositions in terms of dominance durations, which they did not. However, we acknowledge that differences in perceptual difficulty will always require physically distinct stimuli, at least if there is no task that picks a specific stimulus dimension on which difficulty is considered. Hence, in the absence of a task on the stimulus, it will in principle not be possible to fully dissociate between physical properties of the stimulus (stimulus strength in a broad sense) and perceptual qualities (such as the perceptual difficulty). The fact that Levelt’s propositions are not fulfilled can then also be interpreted as a limitation of the propositions to certain definitions of stimulus strength. Hence, the demonstration of a modulation by the task is critical to demonstrate that it is not the physical properties alone that determine the difference between the dominance of one stimulus over the other.
To make use of an OKN-based no-report paradigm, stimuli have to robustly induce OKN and therefore span a substantial part of the visual field. With larger stimuli, mixed (piecemeal) percepts become more prominent (Blake et al., 1992), that is, dominance becomes less exclusive. It is likely that piecemealing results in a lower OKN gain. This is the main rationale to use both measures—relative dominance and average gain—as they capture lack of exclusivity differently. For example, if one stimulus were almost exclusively dominant half of the time, and the other would only be slightly dominant the other half, the relative-dominance measure would provide a 50/50 result (relative dominance near 0), while the average gain would still be biased toward the “more exclusive” stimulus. Finding qualitatively very similar result patterns for both measures makes it unlikely that systematic differences in piecemealing are responsible for the result.
The finding that task relevance further enhances the dominance of the stimulus that is harder to perceive is in line with the effect of task relevance as such on perceptual dominance (Chopin & Mamassian, 2010). In Chopin and Mamassian’s study, observers were unaware of the task relevance of the stimulus and nonetheless a profound effect on its dominance was found. Taken together with the results of Wilbertz et al. (2017), this indicates that task-relevance effects are unlikely to be mediated by the implicit reward of better task performance. However, it should be noted that in Chopin and Mamassian’s case, the observers had to deploy attention to parts of the stimulus that was to become the dominant one. So, it is well-conceivable that attentional mechanisms contribute to the effects of task relevance in their case. In our study, we do not measure or manipulate attention explicitly, but given that a lower-visibility stimulus will require more attentional resources, and the enhancement of dominance by attention (Marx & Einhäuser, 2015; Ooi & He, 1999) render an attentional contribution to the task-relevance effect likely. Our experimental design is agnostic about the mechanisms through which task relevance is mediated. Hence, our results would also be consistent with an interpretation in which the effect of perceptual difficulty prior to the task acts akin to stimulus strength (in a broad sense—i.e., without fully respecting Levelt’s propositions, see earlier) and attention due to task relevance enhances this effect, although the readaptation after the task has ended (Figure 5G and H) would imply that attention is reduced gradually. Provided observer’s ability to volitionally control perception in binocular rivalry at least with respect to percept durations (van Ee et al., 2005), it is also possible that in blocks with task, there is some active or strategic component that makes observers try to hold the difficult stimulus until a perceptual decision is possible. It should be noted, though, that by design observers need to have experienced both stimuli before they can make the same/different judgment and the control of a specific stimulus is rather limited in binocular rivalry (Meng & Tong, 2004). Moreover, the presence of an effect prior to any task rules out that the observed effects are purely due to volitional control and the no-report paradigm excludes any volitional bias on the report of the perceptual state. Taken together, attentional factors, and possibly volitional control and the implicit reward of responding more quickly in a trial, may contribute to the increased dominance of the perceptually difficult stimulus. Teasing apart these factors and probe their interaction with explicit reward and other associations, however, remains an interesting issue for further research.
A further interesting difference between the study of Chopin and Mamassian (2010) and the present data is the persistence of the increased dominance throughout the trial. While—in response blocks—the effect of difficulty is maximal immediately prior to the response, it persists at a lower level throughout the trial. In contrast, Chopin and Mamassian detect an enhancement of dominance by task relevance only at the trial onset, which vanishes after about 2.2 s. It is known that the first percept can have different properties from later phases of prolonged viewing in binocular rivalry (e.g., Attarha & Moore, 2015; Carter & Cavanagh, 2007) and other forms of multistability (Hupé & Pressnitzer, 2012) and is particularly susceptible to attentional effects (Chong & Blake, 2006). There is frequently a bias toward perceiving the stimulus with higher stimulus strength (Goutcher & Mamassian, 2006; Mamassian & Goutcher, 2005). In our case, the very first percept is comparably unaffected by perceptual difficulty, again indicating that effects of perceptual difficulty are dissimilar from typical effects of stimulus strength. It is, however, possible that the first dominance phase assessed from the OKN would not be equivalent to the first percept reported (especially if the phase is very short) and indeed we observed a preference for the perceptually more difficult stimulus already soon after onset (Figure 4C). Importantly, when stimuli are task relevant, the increased dominance for the perceptually more difficult stimulus persists and remains at a rather constant level after the response throughout the trial. The difference on average between the “after-response” phase and the “no-task” blocks suggests that this persistent effect is also modulated by task relevance rather than only the response-locked peak. Whether this difference to Chopin and Mamassian arises from the finer-grained temporal resolution in our study by avoiding the need to average over discrete button press events or from the temporal separation between rivalry and search trials in their experiment has to remain open and further highlights the benefit of a continuous response-free measure of dominance. More important, however, the observation that the increased dominance persists across trials (as evidenced by the difference from 0 at trial onset in our blocked design) is well in line with the persistence across trials and blocks also observed by Chopin and Mamassian.
In the vertical direction, the difficult gratings contain more high spatial frequencies than the easy gratings. Experiments with low-pass filtered gratings show that in binocular rivalry removing high spatial frequencies from one stimulus (i.e., blurring) can increase the relative dominance of the other stimulus (Arnold et al., 2007; Fahle, 1982, 1983). However, the same applies to band-pass filtering (Fahle, 1982): Providing more spatial frequency “channels” in general increases the dominance of a stimulus. It is unclear, whether adding more power in some channels compensates for less power in other channels. An investigation of noise stimuli with different power-function amplitude spectra (1/fα) found that stimuli with α=1 dominate over other exponents (Baker & Graf, 2009), including white noise (α=0). We acknowledge that our data cannot fully exclude that the distinct spatial frequency content (or any other stimulus property) contribute to the observed effects. Indeed, when using physically distinct stimuli, effects of stimulus properties can never be fully excluded, unless one shows different modulations by different tasks. For our stimuli, we note that power and subjective perception are dominated by the frequency content along the horizontal direction (i.e., the “carrier” grating). Because on a global scale the easy and the difficult stimulus both are high-contrast gratings, and—if anything—the easy stimulus appears subjectively slightly more salient, it seems exceedingly unlikely that the observed increase in relative dominance for the difficult stimulus results from its “stimulus strength,” that is, some physical stimulus attribute, alone. Even if one would assume this to be the case, this would not explain the modulation by task relevance. So, it remains to speculate what makes the difficult stimulus more dominant in first place. One possible interpretation states that the difficult stimulus carries more (perceptual) uncertainty. The role of perceptual uncertainty has been addressed in studies about the voluntary control of eye movements, which overall produced mixed results. While the latencies of saccadic eye movements are facilitated by the introduction of a perceptual task (Bieg et al., 2012; Montagnini & Chelazzi, 2005; Trottier & Pratt, 2005), they are not modulated by uncertainty in the stimulus (Wolf & Schütz, 2017). The sensorimotor adaptation of saccades is sensitive to a perceptual task (Schütz et al., 2014) and to some extent also to perceptual uncertainty (Gerardin et al., 2015; Wolf et al., 2019). In visual search, eye movements minimize uncertainty and maximize information gain under some conditions (e.g., Najemnik & Geisler, 2005; Peterson & Eckstein, 2012) but not under others (e.g., Nowakowska et al., 2017; Verghese, 2012). Overall, it seems that perceptual uncertainty has the potential to influence eye movement control, but not under all circumstances. For associative uncertainty, it has been shown that attention—as measured by eye movements in visual search—is held longer at an item, when it is related to high uncertainty (i.e., it is a less predictive cue about a subsequent reward than other cues; Koenig et al., 2017), so it is conceivable that the higher (perceptual) uncertainty yields more perceptual dominance through similar mechanisms. If true, this would predict that perceptual dominance in Binocular Rivalry is also increased by associative uncertainty, for which we have seen initial evidence (Einhäuser et al., 2016), whose further exploration remains an avenue for further research.
Footnotes
Acknowledgements
The authors are grateful to Chris Paffen and a further anonymous reviewer for helpful comments on previous versions of the article.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research was supported by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation)—project number 222641018—SFB/TRR 135, B2 and B4.
Supplemental Material
Supplemental material for this article is available online.
