Abstract
When presented with two different target–penalty configurations of similar maximum expected gain (MEG), participants prefer aiming to configurations with more advantageous spatial, rather than more advantageous gain parameters—perhaps due to the motor system’s inherent prioritisation of spatial information during movements with high accuracy demands such as aiming. To test this hypothesis, participants in the present studies chose between target–penalty configurations via key presses to reduce the importance of spatial parameters of the response and performance-related feedback. Configurations varied in spatial (target–penalty region overlap) and gain parameters (negative penalty values) and could have similar or different MEG. Choices were made without prior aiming experience (Experiment 1), after aiming experience provided information of movement variability (Experiment 2), or after aiming experience provided information of movement variability and outcome feedback (Experiment 3). Overall, configurations with advantageous spatial or gain parameters were chosen equally (Both-Similar condition) in all experiments. However, average behaviour at the group level was not reflective of the behaviour of most individual participants with three subgroups emerging: those with a value preference, distance preference, or no preference. In Experiments 1 and 2, these individual differences cannot be explained by MEG differences between configurations or participants’ movement variability, but these variables predicted choice behaviour in Experiment 3. Further in the Both-Different condition, participants only selected the larger MEG configuration at a level above chance when both variability and outcome information were given prior to the key press task (Experiment 3). In sum, the data indicate that prioritisation of spatial information did not emerge at the group level when performing key presses and more optimal behaviour emerged when information regarding movement variability and outcome feedback were given.
Keywords
Introduction
Fundamental to the human experience is deciding how to interact with different objects in the environment. When deciding how to act, it is important to understand both the positive and the negative consequences associated with each action and how likely each of these consequences are for a given action. For example, consider this daily scenario of parking a car when only two parking spots remain—one option is to park in a small space beside an old van and the other option is to park in a larger space beside an expensive sports car. The probability of the negative outcome (i.e., hitting the parked car) is high if one parks in the small space and low if one parks in the large space. However, if this negative outcome does occur, it will be more expensive to repair the other car if the driver hit the sports car than the old van. The question is, in which spot will the driver choose to park?
The above scenario in which the probability and potential gain/loss of each option is considered during a movement decision is captured by the maximum expected gain (MEG) model (Trommershäuser et al., 2003a, 2003b). The MEG model is a prescriptive model that specifies the visuomotor strategy that a rational actor should use to maximise gain when acting in the presence of stimuli with extrinsic value. That is, this visuomotor strategy prescribes which of two or more potential stimuli a rational actor should select to aim to as well as the location at which the actor should select to aim within the space of the stimuli. The model considers both the values of a given outcome and the probability of each outcome based on the spatial parameters of the action context. For example, participants may be presented with a target that overlaps with a penalty region (i.e., a target–penalty configuration) (see Figure 1 for examples). The target region has a positive gain, the penalty region has a negative gain, the overlapping region has a gain which is the sum of the positive and negative gains, and the region outside the target–penalty configurations has no gain. The key spatial parameters in determining the probability of each outcome are as follows: (a) the size of the exposed and overlapping areas of the target and penalty regions as determined by the distance between the centres of each region; and (b) the actor’s own movement endpoint variability. The actor’s own movement endpoint variability is an important factor because, across a series of movements, this variability generates a distribution of actual endpoints around the aimed-to location. Thus, as a result of this variability, there may be times when the actor will contact an overlapping or penalty region even when the actor is aiming to the target location.

Target–penalty configurations, trial procedure, and stimuli used for each task. (a) Target stimuli were unfilled circles with a green outline, while penalty stimuli were either filled magenta or red depending on their penalty value (target location is shown as an open black circle, and penalty region is shown as a solid grey circle in the panels). The target could appear either left or right of the penalty region. Each target–penalty configuration had an associated distance/overlap value and penalty value. (b) In the key press task, two target–penalty configurations were presented side-by-side and participants responded by pressing one of two buttons on a keyboard. (c) In the single-target aiming task, only the target circle was presented and participants aimed to and touched the green circle as fast and accurately as possible. (d) In the target–penalty aiming task, only a single target–penalty configuration was present and participants aimed to and touched the configuration as fast and accurately as possible. The reach endpoint determined the characteristics of the feedback received. If the reach endpoint was inside the target region, the colour associated with the immediate feedback was green. If the reach endpoint was outside the target region, the colour associated with the immediate feedback was different depending on the region that was contacted (see task description for details). Note that each experiment included each task; however, the order in which the tasks were performed varied across the experiments. See text for details.
In this experimental task, the acting participants must select an endpoint to aim to within the configuration that they believe will maximise the reward across a series of aiming movements. The potential gain associated with each possible chosen endpoint in the environment can be calculated using an estimate of the participant’s endpoint variability along with knowledge of the values and spatial layout of the stimuli. According to the prescriptive model, an optimal actor should select the endpoint with the MEG. Furthermore, when presented with two or more potential target–penalty configurations, the actor should select the target–penalty configuration that has the largest MEG. For instance, if the spatial parameters of two configurations are the same, then the actor should choose the configuration with the smaller penalty value because it has the larger MEG. Alternatively, if the penalty parameters across two configurations are the same, the actor should choose the configuration with the larger exposed target area/less overlap between the target and the penalty regions (i.e., larger distance between target centre and penalty centre), given that it will have a larger MEG due to the increased probability of contacting the target and decreased probability of contacting the penalty region.
Observed choice behaviour in these studies can be compared with the prescriptions of the MEG model, 1 and adherence to or deviation from what would be considered “optimal” behaviour according to the prescriptive MEG model can give some indication on how probability and value information are processed and weighted by the individual for the decision. In one such study, it was found that, when presented with two configurations that had different MEGs due to variations in the spatial parameters or penalty value, participants could select the configuration with the larger MEG (Trommershäuser et al., 2006). This optimal choice behaviour has been taken as general evidence that actors can account for changes in both probability and value information when performing visuomotor tasks to target–penalty configurations (see also Gepshtein et al., 2007; Neyedli & Welsh, 2013, 2014)
Although there was general support for the conclusion that actors often make optimal choices, more recent evidence suggests that probability and value information are not always integrated in an optimal manner during action selection (Jarvstad et al., 2014; Neyedli & Welsh, 2015b; see also Neyedli & LeBlanc, 2017; Neyedli & Welsh, 2014). For instance, similar to Trommershäuser et al. (2006), Neyedli and Welsh (2015b) presented participants with a pair of target–penalty configurations and asked the participants to aim to and touch the configuration which would allow them to accumulate the most points over the experiment. In the critical conditions, the experimenters varied the distance between the target and the penalty regions and/or penalty value such that the configurations could either have a relatively large difference in MEG (>5 points) or the configurations could have a similar MEG (<5 points). For the configurations with a large difference in MEG, the authors analysed the configurations that differed based on distance parameters and penalty value parameters separately which differed from Trommershäuser et al. (2006) who grouped the different types of configurations for the analysis. When participants had a short preview time of 400 ms (similar to Trommershäuser et al., 2006), the participants selected the configuration with the higher MEG at a level greater than chance when configurations differed in the distance between the target and the penalty regions (i.e., the spatial feature related to probability). Selection was not different from chance when configurations only differed in penalty value. When a longer preview of the pair was given (i.e., 2000 ms), participants’ selection improved in the configurations where the penalty value differed between the configurations to a level above chance, but selection of the configuration with the higher MEG in this condition was still lower than for the distance configurations. Overall, these data indicated that participants more effectively weighted the distance parameter than the penalty parameter when choosing between configurations with the goal of maximising gain.
Returning to the Similar MEG condition, to create pairs of configurations with similar MEGs, one configuration would have to have a small distance between regions but low penalty value, whereas the other configuration would have to have a large distance between regions but high penalty value. Because the MEGs are equal in these configurations, selection of one configuration over the other configuration on a given trial and across a series of trials should be at chance. Therefore, any bias in their selection (e.g., more often selecting the configuration with the larger distance and higher penalty than the configuration with the smaller distance and lower penalty) should reflect which parameter the participant prioritised (see Note 1 for a working definition of bias). The results of Neyedli and Welsh (2015b) revealed that participants preferred to aim to the configuration with the larger distance between the target and the penalty regions rather than one with the lower penalty value. That is, participants seemed to put a greater weight on the spatial (probability) parameter compared with the value parameter when the gains of the configurations were similar. The authors suggested that this spatial preference emerged because it was less efficient to make visuomotor decisions based on value information relative to probability information, a result that may have led to the identified choice bias (see also Neyedli & Welsh, 2014).
The relative efficiency with which probability information was used, compared with value information, may be due to spatial properties having a larger or more direct influence on motor planning than value information. In these action decision-making tasks, spatial information related to the task goal likely has a direct influence on movement planning through the dorsal visual stream (Cisek & Kalaska, 2010; Goodale & Milner, 1992; Pisella et al., 2000). However, to evaluate the value information, participants first had to identify the colour of the penalty region and associate that colour with a value because the colour determined the penalty region’s identity and penalty value (which changed trial-by-trial). Importantly, both colour information and object identity are thought to be processed via the ventral visual stream (Cisek & Kalaska, 2010; Goodale & Milner, 1992). Furthermore, the behavioural relevance associated with the identity of the penalty region may be represented in the ventrolateral prefrontal cortex before influencing action selection via projections to the dorsolateral prefrontal and premotor cortices (Sakagami & Pan, 2007). This indirect route for value processing, compared with the more direct route representing spatial information, may take longer to influence visuomotor strategies. Thus, when performing the action selection and decision-making task with aiming movements outlined above, there may be an inherent prioritisation of spatial information by the motor system due to the efficiency with which the spatial information is utilised for programming and controlling action (a view perhaps consistent with a context-dependent or bounded rationality approach to decision making proposed by Simon, 1955).
A core prediction that follows from the above hypothesis is that the spatial accuracy demands of the motor response led to the prioritisation of spatial information in the decision-making process. That is, because the spatial locations of the reaching limb and the target need to be assessed prior to and continually during the movement to perform an accurate reaching response (Heath, 2005; Welsh & Pratt, 2008; Woodworth, 1899), spatial information should be prioritised when aiming movements are required. As such, to increase the potential for task success, the actor should increase the weighting or prioritisation of the perceptual (spatial) information that can lead to efficient specification of action parameters during movement planning and facilitate efficient visual-based online corrections during control.
The accuracy of specific types of actions (e.g., aiming, grasping, and key pressing), however, is dependent on different features of the object and environment. As a result, certain spatial features will or will not be prioritised in different action contexts (see Hommel, 2010). Hence, an additional prediction that comes from the above hypothesis is that the bias towards the spatial/probability information may disappear if there is no direct interaction with the object and the success of the motor response used as the output of the decision processes is not entirely determined by the endpoint accuracy (e.g., a vocal or key press response). Stated another way, because the initial assessment of spatial information for planning and continuous assessment of spatial information for online control is not required for achieving success of ballistic movements such as key presses, the bias towards spatial information should decrease or disappear when the response is a key press.
Although this prediction has not explicitly been explored in the action decision-making literature, there are relevant studies from the action selection literature that have explored the relative salience or prioritisation of different stimulus features during different response modalities such as key presses, aiming, and grasping movements. The results of these studies indicate that when the spatial accuracy demands on the motor system are decreased, the salience of specific spatial features of the stimuli decrease as well (Welsh & Pratt, 2008; Welsh & Zbinden, 2009), contributing to the evidence that different features are prioritised in different action contexts (e.g., Bekkering & Neggers, 2002; Fagioli et al., 2007; Wykowska et al., 2009; see also Hommel, 2010 and Welsh & Weeks, 2010 for an expanded theoretical account). Thus, it is possible that the relative prioritisation of the spatial over value (colour) information associated with the target–penalty configurations should not emerge in the present action decision-making tasks in which an actor must choose between alternative prospects using key press responses because task success is not as reliant on spatial information when performing key press responses.
Although this response mode specific prediction has not been directly tested, some studies of action decision-making have used key press responses. Interestingly, Trommershäuser et al. (2006) found that participants selected the configuration with the larger MEG on >70% of trials when indicating their responses via a key press. Notably, while this selection indicates optimal selection behaviour in the key press task, the authors did not report if differences arose when configurations independently varied in probability or penalty value. In another example, Wu et al. (2009) found that when choosing via a key press between configurations of varying probability and positive gain, a preference emerged for the configuration with the smaller spatial area, but higher value. Thus, a bias towards higher value emerged when key press responses indicated the choice, rather than towards higher probability like that seen when the actors aimed towards the targets (Neyedli & Welsh, 2015b).
Although the above two studies provide some insight into key press decision-making in the presence of varying probabilities and gains, a few caveats should be presented. First, the range of differences in MEG between configurations were much greater in the study by Trommershäuser et al. (2006; i.e., 2–69 points) compared with the study by Neyedli and Welsh (2015b; i.e., 1–25 points). Thus, it is possible that the methods of Trommershäuser et al. were not sensitive enough to capture a bias because the choice was relatively trivial in nature. Second, in Trommershäuser et al. and Wu et al. (2009), key press decisions were made after extensive experience performing reaching decisions to the same set of configurations. Thus, experience may have affected decision-making during the key press task. Furthermore, in Trommershäuser et al., after each key press choice, performance-related feedback was given to participants by the computer programme which simulated that participant’s reach trajectories from their performance on a previous aiming task. The prior motor experience and trial-by-trial feedback on the potential outcome of the choice are important considerations. Indeed, Neyedli and Welsh (2013) have identified that experience and trial-to-trial feedback facilitate participants in selecting more optimal visuomotor strategies when aiming to single target–penalty configurations (see also, Barron & Erev, 2003; Neyedli and Welsh, 2015a, 2015b). Thus, the choice behaviour in Trommershäuser et al. may have been more optimal due to the augmented feedback provided via the computer-simulated performance, something that is not present in other contexts in which the responses do not generate feedback and/or spatial accuracy demands are low. Last, the studies were not specifically designed to assess if a preference existed between configurations with high probability or low negative value while maintaining the similar MEGs. Therefore, it remains unclear if key presses in isolation, without influence from prior motor experience and concurrent feedback, lead to an equal weighting of probability and value characteristics, or if a choice bias would emerge.
The purpose of the present investigation was to determine if there were differences in how efficiently participants processed and used probability and negative value information during a key press decision-making task. This design approach will test the hypothesis that the efficiency differences and spatial preference present in Neyedli and Welsh (2014, 2015b) arose via mechanisms related to the spatial constraints of the (aiming) response mode. To that end, participants were presented with pairs of target–penalty configurations and asked to choose, via a key press, which configuration they believed would allow them to accrue the most amount of points over the course of the experiment if they were to aim to it. The specific conditions were chosen to mirror that of Neyedli and Welsh (2015b). In the Different MEG conditions, each configuration could vary in probability alone, negative value alone, or both. In the Similar MEG condition, each configuration had a different probability (spatial overlap between the target and the penalty regions) and negative penalty value such that one configuration had a large distance between the target and the penalty regions but more negative penalty value and the other configuration had a small distance but a smaller negative penalty value. If the probability (spatial) parameters are only prioritised when aiming to the configurations, then no bias towards selecting the configuration with the advantageous spatial parameter should emerge in this study where the selection is made via key presses.
As a secondary research objective, the role of experience and feedback in a prior aiming task has on later key press choice behaviour was explored. To that end, separate experiments manipulated the order of the key press choice task relative to aiming tasks. Experiment 1 was designed to determine if participants would not show any prioritisation or would more heavily weight spatial information without having any direct experience of and feedback from the aiming task. Hence, in Experiment 1, participants completed the key press selection task prior to gaining any experience with the aiming tasks. Experiment 2 was designed to investigate whether or not gaining experience aiming and an understanding of motor variability prior to the key press selection task would lead to a higher weighting of the spatial information. Recall that movement endpoint variability is an important factor in determining the probability of contacting any of the regions on a given aiming movement. To separate the potential role of experience and knowledge of movement variability from gain/loss feedback from the target–penalty aiming task, participants in Experiment 2 only completed a series of aiming movements to a single target with no overlapping penalty region prior to key press task. Finally, in Experiment 3, participants gained both aiming experience and response-produced feedback about their selection prior to completing the key press selection task by aiming to actual target–penalty configurations and receiving feedback about the points they earned on that aiming trial. Thus, the between-experiment manipulations were specifically designed to separate the influence of aiming experience and an understanding of motor variability (Experiment 2) from the influence of feedback acquired from the gain landscape (Experiment 3) on later key press choice behaviour. Differences in the preference for a given parameter across the different studies will provide insight into the potential role task experience and performance feedback will play in shaping the selection process.
Experiment 1
The purpose of Experiment 1 was to determine if participants display optimal choice behaviour (as predicted by the prescriptive MEG model) and if choice biases emerge when expressing their decisions via a key press response. Importantly, the key press choices in Experiment 1 were performed before gaining aiming experience and, therefore, before receiving feedback regarding their motor variability or the gain associated with each region in different target–penalty configurations. In terms of research predictions, if probability information is simply prioritised during this decision-making task (Neyedli & Welsh, 2014) regardless of how the decision is made (i.e., responses with high or low spatial constraints), then efficiency and choice biases will arise. Specifically, in the Different MEG condition, participants will choose the higher MEG configuration at a level greater than chance when configurations vary in probability, but not negative value. Furthermore, in the Similar MEG condition, a preference for the high probability configuration will emerge. However, if the weighting of the spatial/probability information is only increased when the spatial constraints of the response are high (i.e., as in previous studies using aiming responses, Neyedli & Welsh, 2015b), then no efficiency or choice bias should arise between spatial and value components in this key press task because there are little-to-no spatial constraints. That is, in the Different MEG conditions, the configuration with the larger MEG should be chosen at a proportion greater than chance (i.e., 50%). Furthermore, in the Similar MEG condition, participants will choose the high probability and low negative value configurations with equal likelihood.
Methods
Participants
A total of 17 right-hand-dominant participants with normal or corrected-to-normal vision (11 females, 6 males, mean age = 20 years) were recruited from the University of Toronto community. Participants were compensated a minimum of 15 CAD and were given a bonus based on their performance in the target–penalty aiming task (M = 5.10 CAD; see below for details). All experimental conditions were completed in a single 1.5-hr session. Each participant provided written informed consent prior to participation in the study. The procedures were approved by the University of Toronto Ethics Review Board, and this work was completed in accordance with the Declaration of Helsinki.
Apparatus
Participants sat approximately 40 cm in front of a 22″ ViewSonic LCD touch screen monitor (resolution: 1680 × 1050 pixels, 3.5 pixels/mm). A standard keyboard was placed in front of the screen and approximately 5 cm below the bottom of the screen. Participants used the “<” and “>” to indicate their choice in the key press task and the “<” acted as a home button in the aiming tasks. These buttons were aligned with the centre of the screen. MATLAB (version 2010b; MathWorks, Natick, MA, USA) controlled all experimental events.
General procedure
Each testing session consisted of three tasks. These tasks were completed in the same order for each participant. In the key press task, participants chose between target–penalty configurations of varying MEG via a key press. In the single-target aiming task, participants aimed to the centre of a single target. Last, in the target–penalty aiming task, participants aimed to a target–penalty configuration and acquired points that were converted to money. The details of each task are presented below.
Key press task
In Experiment 1, participants first completed the key press task. Each trial began when the spacebar was pressed. Following the button-press, a 9 mm × 9 mm white fixation cross was displayed in the centre of the screen for 1,000 ms. Coincident with the disappearance of the fixation cross, the outline of a 115 mm × 80 mm blue rectangle appeared. Two target–penalty configurations were presented within the blue rectangle 500 ms later, one in each half of the outline (see Figure 1). Each target–penalty configuration consisted of an overlapping target and penalty circle. Both the target and the penalty circles had a radius of 9 mm. The circles were distinguishable by colour—the target circle had a green outline and open centre and the penalty circle was a solid circle that was either magenta or red, depending on its value.
Participants viewed the pair of target–penalty configurations for 2,000 ms before an auditory tone was presented. Upon presentation of the tone, participants were required to choose the left or right configuration, within a time constraint (described below), by pressing the “<” or “>,” respectively. Participants were instructed to choose the configuration that would earn them the most points (and subsequently monetary reward) across a series of aiming movements. Notably, although participants did not aim to each configuration, they were told to base their choices on the knowledge that, if aiming to the configuration, landing inside the target region would earn 100 points and landing inside a penalty region would reduce points and that the amount of the reduction (−100 or −500 points) would be indicated by the colour of the penalty circle. After the choice was made, no reward feedback was given because participants did not aim to the configurations and therefore there was no actual movement endpoint to use to determine points for that trial.
To understand how participants evaluated and chose between target–penalty configurations, both the value of the penalty regions and the distance between the target and the penalty regions were experimentally manipulated. The manipulation of penalty value and distance allows for the presentation of configurations with differing MEG. Here, varying the colour of the penalty circle represented a penalty value manipulation, while varying the distance between the target and the penalty centre represented a probability manipulation. Probability is operationally defined with respect to the probability of contacting the positive target region—a large distance between the target and the penalty regions (a small overlap) indicates a relatively high probability of landing in the target region and gaining points and vice versa. In the present experiment, two penalty and three distance values were used to generate a total of six configurations (see Table 1 and Figure 1a) with different MEG values. For example, configuration “c” (Table 1) with low penalty value (e.g., −100 points) and high probability (e.g., large distance between the target and the penalty centre: 1.25 radii apart) has a higher MEG than configuration “d” (Table 1) with high penalty value (i.e., −500 points) and low probability (e.g., small distance: 0.75 radii apart). Notably, while distance is consistent across participants, the actual probability is different for each individual and is dependent on their endpoint variability.
List of the target–penalty configurations used in all experiments.
The letters are the configuration labels referred to in the text and other tables (see also Figure 1a).
These six target–penalty configurations were then used to create 10 configuration pairs with each pair belonging to one of four conditions (see Table 2). The specific conditions mirrored the design of Neyedli and Welsh (2015b). In the Penalty-Different condition (Pairs 1–3), configuration pairs differed in MEG by varying the penalty value between configurations, while the distance between regions (i.e., probability) was the same across the configurations. In the Distance-Different condition (Pairs 4–7), configurations differed in MEG by varying the distance between the regions (i.e., probability), but not penalty value, between configurations. In the Both-Similar condition (Pairs 8 and 9), configuration pairs had different penalty values and distances, but the penalty values and distances were varied such that the MEG of each configuration was similar. Specifically, one configuration had a larger distance between regions but large negative penalty value and the other had a small distance but a smaller negative penalty value. Here, similar MEG is defined as a MEG difference of less than 5 points between configurations (Neyedli & Welsh, 2015b). Because the MEGs were similar between configurations in this condition, the participant would be able to choose randomly between these two configurations and still achieve the goal of gaining as many points as possible across the study. Therefore, any bias in the chosen prospect in this condition may provide an indication of the preference (i.e., low penalty value or high probability) that each participant had. Finally, in the Both-Different condition (Pair 10), configuration pairs had different penalty values and distances leading to a different MEG value for each configuration. In this condition, the configuration with the larger distance between regions (and larger penalty value) had a larger MEG than the configuration with the smaller distance between the regions (and smaller penalty value).
List of configuration pairs used in all experiments and the average of the absolute MEG difference between configurations in each pair.
Standard deviation is presented in brackets. Note that a “Different” MEG relationship refers to a difference equal to or greater than 5 points, whereas a “Similar” relationship refers to a difference of less than 5 points.
A practice block consisting of 20 trials was provided to allow participants to understand the task and the timing constraint associated with their response (i.e., how quickly they had to respond after the auditory tone). The time constraint was tapered such that the first 5 trials were to be completed in less than 1,000 ms, the next 5 trials in less than 750 ms, and the final 10 trials in less than 400 ms. Following the practice block, 400 experimental trials were completed. Each configuration pair was presented 40 times in a random order. If participants did not respond within 400 ms of the auditory imperative, a feedback message appeared (i.e., “Please respond faster”) and the trial was repeated.
Single-target aiming task
Following the key press task, participants performed the single-target aiming task. The first purpose of this task was to record actual response endpoints for use in calculating the participant-specific motor (i.e., movement endpoint) variability. These motor variability values were then used in later simulations that verified the MEG calculations of the configurations used in the key press task (i.e., to verify that “Different” and “Similar” conditions indeed had different or similar MEGs for each individual). The second purpose was to provide each participant with aiming practice before they performed the final target–penalty aiming task in which they actually gained a reward.
Each trial started when participants had the “<” key continuously depressed. The temporal structure and order of stimuli presentation in each trial were identical to the key press task. The key difference was that, in this task, only a single open green target circle (radius 9 mm) was presented in the centre of the blue rectangle instead of two target–penalty configurations (see Figure 1). After the auditory imperative tone sounded, participants were required to reach to the centre of the target circle as fast and accurately as possible.
A reaction time (RT) and movement time (MT) constraint was imposed like that of the key press task. Specifically, if the sum of the RT and MT was greater than 750 ms, then a feedback message was displayed (i.e., “Please respond faster”) and the trial was repeated. RT was defined as the time between the auditory imperative and the time of “<” key lift off. MT was defined as the time between “<” lift off and screen touch. Participants performed 50 trials in this task.
Target–penalty aiming task
Following the single-target aiming task, participants performed the target–penalty aiming task. The temporal structure and order of stimuli presentation were identical to the previous two tasks. However, one target–penalty configuration was presented in the centre of the blue rectangle each trial (see Figure 1). In this task, participants were instructed to reach to a location in the configuration that they believed would maximise the amount of points they would receive. Furthermore, to encourage rapid responses, a timing constraint was imposed: if the sum of their RT and MT was above 750 ms, they lost 700 points. If participants were within the time constraint, the blue rectangle filled with the colour of the region participants landed in and the associated point value was displayed. If participants were outside the time constraint a feedback message and associated point value was displayed (i.e., “Please respond faster. −700 points”). After the score for the trial was displayed, the cumulative score was displayed (see Figure 1). Participants performed 12 trials in this task corresponding to 2 trials per target–penalty configuration used (see Table 1) and their points were converted to a performance-related monetary bonus (100 points = 0.50 CAD).
Configuration pair creation and validation
It is important to note that the actual MEG of the specific target–penalty configurations used in the initial choice key press task could not be estimated until a measure of motor variability was obtained through the single-target aiming task. That is, because participants performed the key press task before the single-target aiming task, the MEGs of the configurations used in the key press task were estimates based on the participant specific data of Neyedli and Welsh (2015b) and not the motor variability obtained in the present experiment. This design feature was necessary because the main purpose of Experiment 1 was to assess the choice of actors who did not have an aiming-based representation of the task and without having reward feedback after each trial.
To confirm that the global configuration pairs used for each participant in the key press task lead to the conditions outlined in Table 2 (i.e., Penalty-Different, Distance-Different, Both-Similar, Both-Different), the MEG for each target–penalty configuration was calculated using the MEG model (Neyedli & Welsh, 2015b; Trommershäuser et al., 2003a, 2003b). Data from the single-target aiming task were used as an estimate of each participant’s endpoint distribution. Before the endpoints on single-target aiming data were used in the calculation of the MEG for each target–penalty configuration, <4% of outlier trials were removed. Endpoint outliers were considered to be any trial with an endpoint that was 2.5 standard deviations above or below the participant-specific mean in any dimension (i.e., x or y). The difference in MEG between configurations in each pair was then calculated. The outcome of these calculations is presented in Table 2 and support the use of these configuration pairs in each specified condition. Specifically, the MEGs of the two configurations in the Both-Similar condition were similar (<5 points different), whereas the MEGs of the configurations in each of the other conditions were different (equal to or >5 points different).
Data reduction and analysis
The data analysis focused on the key press task as there were no experimental manipulations in the single-target aiming and target–penalty aiming tasks and the data from the key press task directly addresses the research predictions. The primary dependent measure was the proportion of trials participants chose the configuration with the larger MEG in the Penalty-Different, Distance-Different, and Both-Different conditions. Notably, this measure could not be used for the configuration pairs with similar MEGs in the Both-Similar condition (Pairs 8 and 9). As such, the proportion of trials participants chose the configuration with larger distance was calculated (i.e., configuration “e” of Pair 8 and configuration “f” in Pair 9), thus a proportion greater than .5 indicates a preference for a larger distance between the target and the penalty configuration, while less than .5 indicates preference for a lower penalty value. Note that these same relationships also held, to a certain degree, for the Both-Different configurations because the configuration with the larger MEG had the larger distance/higher penalty than the configuration with the lower MEG which had the smaller distance/lower penalty. The proportions for each configuration pair were then averaged based on the condition (i.e., Penalty-Different, Distance-Different, Both-Similar, Both-Different) and converted to a percentage. Before the proportions were calculated, trials were removed if participants pressed a response key other than “<” or “>” (average exclusion <1%).
One-sample t-tests were used to compare the average proportions of choice in each condition to the probability of choosing the configuration with the larger MEG or distance based on chance alone (i.e., 50%). For this analysis, the alpha level was Bonferonni corrected to .0125 to adjust for four comparisons and a bootstrapping approach with 5,000 runs was used to obtain 99% confidence intervals. Furthermore, a Pearson correlation was used to determine if there was a relationship between the choice behaviour in the Both-Similar and Both-Different conditions. This correlation indicates if choice behaviour was influenced by a consistent preference for low penalty or high distance configurations regardless of the MEG difference between configurations. The alpha level for this analysis was set at .05 and a bootstrapping approach with 5,000 runs was used to obtain 95% confidence intervals.
Results and discussion
Figure 2a displays the proportion of trials that the configuration with the larger MEG (i.e., in the Penalty-Different, Distance-Different, and Both-Different conditions) or larger distance (i.e., in the Both-Similar condition) was chosen for each participant. This figure demonstrates that participants consistently chose the configuration with the larger MEG in the Penalty-Different and Distance-Different conditions. In contrast, there was much less consistency in the chosen configuration across participants in the Both-Similar and Both-Different conditions where both parameters differed across configurations. These observations were confirmed via the statistical analyses which showed that participants chose the target–penalty configurations with larger MEG at a level that was well above chance in both the Penalty-Different, M = 90%, SD = 13%, t(16) = 12.53, p < .0125, dz = 3.04, CIBS = [82, 97], and Distance-Different, M = 93%, SD = 23%, t(16) = 7.85, p < .0125, dz = 1.9, CIBS = [76, 100], conditions. Therefore, when either only the penalty or distance varied between configurations, participants were able to evaluate and choose the configuration with larger MEG. However, participants’ choice behaviour did not differ from chance in the Both-Similar, M = 47%, SD = 39%, t(16) = 0.28, p > .0125, dz = 0.07, CIBS = [25, 70], and Both-Different, M = 69%, SD = 40%, t(16) = 1.99, p > .0125, dz = 0.48, CIBS = [43, 91], conditions. While the finding that choice behaviour did not differ from chance in the Both-Similar condition was not surprising (because the MEGs of the configurations were essentially the same), the finding that choice did not differ from chance in the Both-Different condition was not expected because there was a clear difference in MEG between the two configurations.

Individual and group average choice proportions in (a) Experiment 1, (b) Experiment 2, and (c) Experiment 3. The proportion chosen indicates the proportion of times the participant chose the configuration with the larger MEG (i.e., Penalty-Different, Distance-Different, Both-Different conditions) or the proportion of times the participant chose the configuration with the larger distance (i.e., Both-Similar condition). Each dot is representative of one participant and the middle horizontal bars represent the means of each condition. Error bars represent 99% between-participant confidence intervals and the absence of an overlap between an error bar and the 50% line represents a reliable effect that can be interpreted inclusive to a test of the null hypothesis (Cumming, 2013) (see text for details).
An examination of each individual participant’s choice behaviour provides insight into why there is not a consistent choice preference overall when both penalty value and distance differ between configurations. Figure 2a reveals that some participants weighted (showed a preference for) value in their decision, whereas others showed no preference or weighted space more strongly. This difference in weighting is revealed through the determination that some participants display a penalty (closer to 0%), distance (closer to 100%), or no preference/mixed choice (closer to 50%) in the Both-Similar and Both-Different conditions.
To determine if these preferences were consistent for individuals in the Both-Similar and Both-Different conditions, a Pearson correlation was used to understand if a relationship existed between choices made in each condition (Figure 3a). A significant, linear relationship (r = .77, p < .05, CIBS = [0.58, 0.93]) revealed that participants tended to make similar choices across the two conditions—those who tended to choose configurations with the lower penalty value (or higher distance) during the Both-Similar condition were more likely to choose the lower penalty value (or higher distance) configuration during the Both-Different condition and vice versa.

Individual participant’s proportion of chosen configurations with larger MEG in the Both-Different condition as a function of their proportion of chosen configurations with larger distance in the Both-Similar condition for (a) Experiment 1, (b) Experiment 2, and (c) Experiment 3. Each dot represents a single participant. The Pearson correlation (r) for each experiment is presented within each subplot.
It is important to note here that the preferences of the individuals would lead to quite different outcomes in the separate conditions. To consider the Both-Similar condition first, different preferences (or lack thereof) in the Both-Similar condition may have occurred because the MEGs of the configurations in this condition were similar and thus would not have a long-term implication for point gain across the experiment. However, different preferences in the Both-Different condition would lead to different long-term point gain across the experiment. For the Both-Different condition, a preference for configuration with the lower value in this condition would be opposite to what would be predicted for a rational decision-maker according to the MEG model because the individual is selecting the configuration with the lower MEG. Participants without a strong preference (i.e., near 50%) in the Both-Different condition are likewise not demonstrating completely rational choices.
Overall, four main results emerged from Experiment 1. First, participants chose the configuration with the higher MEG at a level above chance when configurations varied in distance or penalty value in isolation (i.e., Distance-Different and Penalty-Different conditions). Second, an overall group-level choice bias towards larger distance or lower penalty value did not emerge when participants were presented with configurations of similar MEG (i.e., Both-Similar condition). Thus, when performing a task in which a reliance on spatial information is unnecessary to accomplish the goal (i.e., a key press response), consistent choice biases did not arise across participants between spatial and value parameters—something that may have led to a higher proportion of penalty preferences in the present compared with previous experiment in which aiming responses were executed in the selection task (Neyedli & Welsh, 2015b). Third, when both distance between regions and negative gain varied between configurations and the MEG was different (i.e., Both-Different condition), participants did not choose the higher MEG option at a level greater than chance. Last, an analysis of individual differences in the Both-Similar and Both-Different conditions revealed three groups of participants: those with a value preference, those with a distance preference, and those with no preference.
The sub-optimal choice behaviour (as characterised by the prescriptive MEG model) in the Both-Different condition may have emerged due to making decisions prior to gaining performance-related aiming feedback that has been shown to facilitate optimal visuomotor strategies (Neyedli & Welsh, 2013, 2014, 2015a; see also Barron & Erev, 2003). That is, when performing aiming movements towards target–penalty configurations as in previous work, participants receive both an understanding of their motor variability and the gain associated with landing in each region. This information may facilitate the appropriate weighing of the probability and value associated with each configuration. The possibility that the lack of performance-related feedback influenced choice behaviour is explored in the next two experiments.
Experiment 2
Experiment 2 was designed to determine if providing aiming experience and associated knowledge of movement variability (without performance-related feedback) prior to the key press choice task would lead to the more optimal choice behaviour in the Both-Different condition (i.e., choosing the configuration with the higher MEG in a greater proportion of trials). Thus, prior to performing the choice task, participants performed the single-target aiming task—a task in which participants aim to a target presented in isolation. The prior motor experience provided participants with a better understanding of their motor variability. Because participants did not perform the target–penalty aiming task prior to the key press task, they did not gain experience with the gain landscape and implications of contacting the different regions (cf. Trommershäuser et al., 2006; Wu et al., 2009). Thus, the present experiment will provide an indication of how an understanding of motor variability in isolation influences the expression of choice behaviour when the subsequent choice response has little to no spatial constraints. In terms of research predictions, if participants require an understanding of their movement variability to display more optimal choice behaviour and consistently choose the option with the higher MEG, then providing participants with aiming experience before the key press choice task will lead to the configuration with the larger MEG being chosen at a proportion greater than chance in the Both-Different condition.
Methods
Participants
A total of 23 right-hand-dominant participants with normal or corrected-to-normal vision (9 females, 14 males, mean age = 20 years) were recruited from the University of Toronto community. The data from two participants were removed from the final sample for the following reasons: (a) the single-target aiming data was not recorded for one participant and (b) 15.25% of the key press choice data were not recorded properly for the other participant because they pressed the wrong keyboard keys. All participants were compensated a minimum of 15 CAD and were given a bonus based on their performance in the target–penalty aiming task (M = 4.20 CAD). All experimental conditions were completed in a single 1.5-hr session. Each participant provided written informed consent approved by the University of Toronto Ethics Review Board, and this work was completed in accordance with the Declaration of Helsinki.
Apparatus, procedure, and analysis
The apparatus, procedures of each task, and data analysis were identical to that of Experiment 1. However, the order in which each task was performed differed. Specifically, in Experiment 2, participants performed the tasks in the following order: single-target aiming task, key press task, and, finally, target–penalty aiming task.
Before the single-target aiming data were used for the MEG calculations, <3% of outlier trials were removed. Table 2 reveals that the MEGs of the two configurations in the Both-Similar condition were similar (<5 points different), whereas the MEGs of the configurations in each of the other conditions were different (equal to or >5 points different). Less than 2% of trials were removed from the key press task due to participants pressing a response key other than “<” or “>.”
Results and discussion
Examination of the individual participant data from Figure 2b indicated that, consistent with Experiment 1, every participant predominantly chose the configuration with the larger MEG in the Penalty- and Distance-Different conditions, whereas the preferred configurations were less consistent across participants in the Both-Similar and Both-Different conditions. These observations were confirmed in a statistical analysis. In both the Penalty-Different, M = 97%, SD = 6%, t(20) = 45.67, p < .0125, dz = 9.96, CIBS = [94, 99], and Distance-Different, M = 98%, SD = 3%, t(20) = 91.23, p < .0125, dz = 19.91, CIBS = [97, 99], conditions, participants chose the target–penalty configurations with larger MEG at a level above chance. These analyses indicate that participants were able to evaluate and choose the configuration with the larger MEG when the penalty and distance values varied in isolation. However, in the Both-Similar, M = 32%, SD = 35%, t(20) = 2.37, p > .0125, dz = 0.52, CIBS = [13, 52], and Both-Different, M = 57%, SD = 44%, t(20) = 0.68, p > .0125, dz = 0.15, CIBS = [31, 80], conditions, participants’ choice behaviour did not differ from chance.
When looking at the proportion chosen at the level of each participant (Figure 2b), it is revealed that certain participants displayed a penalty, distance or no preference in the Both-Similar and Both-Different conditions. Thus, consistent with Experiment 1, when both the penalty and distance vary simultaneously between configurations, there is not a consistent choice preference across participants. A significant linear relationship between the proportion chosen in the Both-Similar and Both-Different conditions (r = .81, p < .05, CIBS = [0.66, 0.94]) revealed that participants who chose configurations with the lower penalty value (or higher distance) during the Both-Similar condition were more likely to choose the lower penalty value (or higher distance) configuration during the Both-Different condition (Figure 3b). For the group of participants who chose the configurations with lower penalty value in the Both-Different condition, this strategy would lead to a relative reduction in gain. Thus, this subset of participants displayed a preference for lower penalty configurations even when it was non-optimal according to the MEG model.
Overall, the results of Experiment 2 revealed that providing participants with motor experience and self-generated information of their motor variability prior to performing the decision-making task had little, if any, influence on choice behaviour. That is, consistent with Experiment 1, participants chose the configuration with the larger MEG in the Penalty-Different and Distance-Different conditions and there was no consistent choice bias across participants towards spatial or value parameters in the Both-Similar condition. Furthermore, participants displayed sub-optimal behaviour (as characterised by the MEG model) in the Both-Different condition which suggests that the feedback gained from the single-target aiming task was not sufficient to facilitate optimal choice behaviour. Although aiming in the single-target condition provides the participant with knowledge of their specific endpoint variability, it does not provide feedback/information of the gain landscape because participants did not interact with the different regions of the target–penalty configurations. Importantly, both types of feedback may be necessary for participants to fully understand the gain associated with each configuration and therefore weigh probability and negative value in an optimal manner (Neyedli & Welsh, 2013, 2014, 2015a; Barron & Erev, 2003; c.f. Trommershäuser et al., 2003a). This possibility was explored in Experiment 3.
Experiment 3
The goal of the prior aiming experience given in Experiment 2 was to provide participants a better understanding of their endpoint variability in relation to the size of the target circle, a main factor that determines the probability of landing inside each region of the target–penalty configuration. Participants, however, did not receive feedback about all the potential value-based outcomes in the aiming environment and so may not have appropriately understood or weighed value information. That is, although participants had explicit knowledge of the values associated with each penalty colour, they may not have fully appreciated the consequences associated with landing in the penalty region. Thus, the third experiment was designed and conducted to determine if participants would make more optimal choices in the Both-Different condition if they gained performance feedback related to their movement variability and the gain landscape of the presented target–penalty configurations. To that end, prior to performing key press choices, participants performed 50 trials of the single-target aiming task and 60 trials of the target–penalty aiming task. In terms of research predictions, if providing aiming experience to target–penalty configurations leads to an appropriate weighting of value information, then participants should display optimal choice behaviour according to the prescriptive MEG model. That is, participants should choose the configuration with the larger MEG a proportion greater than chance in the Both-Different condition.
Methods
Participants
A total of 23 right-hand-dominant participants with normal or corrected-to-normal vision (13 females, 10 males, mean age = 22 years) were recruited from the University of Toronto community. The data from three participants were removed from the final sample for the following reasons: (a) one participant did not complete the full experimental session; (b) one participant did not follow task instructions; and (c) the MEG relationships between configuration pairs for one participant were not consistent with the initially designed conditions. Regarding this last excluded participant, the MEG difference (derived from the simulations of their movement variability in the single-target task) in the Both-Similar condition was greater than 25 and the MEG relationship in the Both-Different condition was in the opposite direction to what the rest of the participants experienced. Thus, only 20 participants were considered in the final analysis. Participants were compensated a minimum of 10 CAD and were given a bonus based on their performance in the target–penalty aiming task (M = 3.30 CAD). All experimental conditions were completed in a single 1.5-hr session. Each participant provided written informed consent approved by the University of Toronto Ethics Review Board, and this work was completed in accordance with the Declaration of Helsinki.
Apparatus, procedure, and analysis
The apparatus, procedures of each task, and the analysis were identical to that of Experiments 1 and 2, but the order in which each task was performed differed. Specifically, in Experiment 3, participants performed the tasks in the following order: single-target aiming task, target–penalty aiming task, and finally, key press task. This order was set to determine how interacting with target–penalty configurations and receiving feedback on performance prior to the key press task influenced participants’ key press choices. Furthermore, to ensure participants had a good estimate of their performance in the target–penalty aiming task, they performed 10 instead of 2 aiming movements to each configuration from Table 1. Thus, a total of 60 target–penalty aiming movements were performed.
Before the single-target aiming data was used to calculate MEG, <3% of outlier trials were removed. Table 2 reveals that the MEGs of the two configurations in the Both-Similar condition were similar (<5 points different), whereas the MEGs of the configurations in each of the other conditions were different (equal to or >5 points different). Less than 1% of trials were removed from the key press task due to participants pressing a response key other than “<” or “>.”
Results and discussion
An examination of the individual participant data from Figure 2c indicated that almost every participant predominantly chose the configuration with the larger MEG in the Penalty-Different and Distance-Different conditions. In the Both-Similar condition, the chosen configurations were less consistent across participants. In contrast to Experiments 1 and 2, however, almost every participant predominantly chose the configuration with the larger MEG in the Both-Different condition. That is, only two participants chose the configuration with the lower MEG on less than 50% of the trials. Statistical analyses revealed that participants chose the target–penalty configurations with larger MEG at a level above chance in the Penalty-Different, M = 89%, SD = 14%, t(19) = 12.8, p < .0125, dz = 2.86, CIBS = [81, 95], Distance-Different, M = 98%, SD = 3%, t(19) = 79.82, p < .0125, dz = 17.85, CIBS = [96, 99], and Both-Different, M = 85%, SD = 22%, t(19) = 7.12, p < .0125, dz = 1.59, CIBS = [71, 96], conditions. These analyses provide evidence that participants were able to evaluate and choose the configurations with the larger MEG when penalty and distance differed in isolation and when both penalty and distance varied simultaneously, something that did not occur in Experiment 1 or 2. In the Both-Similar condition, M = 59%, SD = 33%, t(19) = 1.21, p > .0125, dz = 0.27, CIBS = [40, 77], participants’ choice behaviour did not differ from chance.
Consistent with the previous experiments, however, there still appeared to be a range of individual choices in the Both-Different condition. Indeed, a significant linear relationship between the proportion chosen in the Both-Similar and Both-Different condition (r = .70, p < .05, CIBS = [0.52, 0.86]) revealed that participants who were more likely to choose the configurations with the lower penalty value (or higher distance) during the Both-Similar condition were more likely to choose the lower penalty value (or higher distance) configuration during the Both-Different condition (Figure 3c). Thus, although participants chose more optimally in the Both-Different condition overall, there still appeared to be differences in preference at the individual level. Unlike the previous experiments, there were far fewer participants (only two participants) displaying a consistent sub-optimal penalty preference in the Both-Different condition. Most of the participants with a penalty preference in the Both-Similar condition actually displayed optimal distance preference behaviour in the Both-Different condition by choosing the configuration with the larger distance (and MEG). This reduction in participants who displayed a consistent penalty preference also explains the reduction in the proportion of explained variance in this regression model relative to the regression models from other experiments (Figure 3). Thus, the data suggest that gaining actual experience with the penalty configurations and the consequences of the performance facilitates better choice behaviour.
Overall, the results of Experiment 3 revealed that participants generally displayed more optimal choice behaviour, in accordance with the MEG model, in all conditions in which there was a difference in the MEGs—something that did not occur in Experiments 1 and 2. That is, Experiment 3 was the first experiment in this series in which participants consistently chose the higher MEG configuration in the Both-Different condition. These results provide evidence that both an understanding of motor variability and the consequences of landing within each region facilitate the optimal weighing of probability and value information. One consistent finding across all experiments, however, was the individual differences in choice preference. The set of supplementary analyses presented in the following section were conducted to explore these individual differences.
Additional analyses exploring individual differences
Although the current experiments were not specifically designed to understand why the individual differences in preference arose in the Both-Similar and Both-Different conditions (i.e., this finding was not expected while designing the experiment), there are a number of factors that may have influenced participants’ preferences which we explore here, post hoc. One factor that may have promoted inter-subject variability in choice behaviour is the between-subject differences in motor variability. Recall that the probability of landing inside each region of a target–penalty configuration is a product of the spatial relationship between regions (i.e., distance) as well as the participant’s motor variability. Thus, differences in motor variability, in general, may influence choice preference such that participants with greater variability may show a probability (spatial) preference and choose the target-penalty configuration with the most exposed target region to increase their chances of avoiding the negative consequences of penalty region contact.
Furthermore, given that MEG varies with motor variability, MEG differences between configuration pairs in the Both-Similar and Both-Different conditions varied across participants. This variation is particularly relevant in the Both-Similar condition as either the configuration with the low penalty parameter or the configuration with the high distance parameter had the slightly higher MEG for each participant. For example, for some participants, configuration “a” had a larger MEG than configuration “e,” whereas for some participants, “e” may have had a slightly larger MEG than “a” (see Table 2). Although it is unlikely that participants picked up on these small differences in MEG between configurations (i.e., <5 points), it remains a possibility that participants chose the configuration with the larger MEG on an individual basis as determined by their personal motor variability. Furthermore, the MEG differences in the Both-Different condition were, on average, greater than 10 points. It is possible that in this condition, the proportion in which a participant chose the higher MEG configuration may have been driven by the magnitude of the MEG difference between configurations.
To explore whether any of these factors were associated with the individual’s preference for probability or value, Pearson’s correlations were used to explore the possibilities that inter-subject motor variability and inter-subject MEG differences between configurations were related to choice behaviour. We ran two sets of correlations. First, we correlated motor variability in the x (horizontal) dimension (VEx: standard deviation of endpoint position in the x dimension) with choice behaviour in the Both-Similar and Both-Different condition. Second, we correlated the relative MEG difference between configurations in the Both-Similar condition (i.e., Pairs 8 and 9) and Both-Different condition (i.e., Pair 10) with the choice behaviour when participants were presented with each respective pair. The results of these analyses are presented in Table 3.
Summary of Pearson’s correlations performed in all experiments.
The upper and lower bounds of the bootstrapped 95% confidence interval are presented for each correlation. The bootstrap was run 5,000 times.
Significant relationship (p < .05).
The analyses revealed that in Experiments 1 and 2, there was no relationship between a participant’s motor variability and their choice behaviour in the Both-Similar or Both-Different conditions. Furthermore, there was no relationship between the relative MEG difference between configurations in Pair 8, 9, or 10 and their choice behaviour when presented with the respective pairs. Interestingly, in Experiment 3, there was a significant correlation between a participant’s motor variability and their choice behaviour in the Both-Similar and Both-Different conditions indicating that participants with higher motor variability were more likely to choose the configurations with the lower penalty values in each condition. Furthermore, the relative MEG difference between configurations and choice behaviour were significantly correlated in Pair 8, but not in Pair 9 or 10. The relationship for Pair 8 revealed that participants were more likely to choose the configuration with the lower penalty value when this configuration had the larger MEG and vice versa. Thus, overall, choice preference cannot be explained by individual differences in motor variability or MEG differences between configurations in Experiment 1 or 2, but those variables may explain choice behaviour in Experiment 3. This result suggests that participants did not strategically use knowledge of their motor variability to choose between options until they were given feedback regarding their motor variability and the gain landscape. In addition, the availability of both types of feedback enabled the best representation of the MEG differences between configurations even when their difference was small (i.e., < 5 points).
General discussion
The goals of the present experiments were to determine whether choice biases and optimal behaviour emerged when choosing between target–penalty configurations via a key press response and if this behaviour was modified after providing varying levels of aiming experience. Three main conclusions can be drawn from the three experiments. First, when performing a task in which the response has a low spatial accuracy constraint (i.e., the key press response), no consistent choice preference towards probability or gain information emerged across all participants. Second, there was a wide variation in individual differences in choice preference in the Both-Similar and Both-Different conditions. And third, when both probability and gain information varied between configurations, optimal selection according to the prescriptive MEG model only occurred at a level greater than chance when participants experienced aiming to the target–penalty configurations. Each conclusion will be discussed across the following sections.
Choice biases when choosing via a key press response
A key press response was chosen as the response mode in the present experiments for two reasons. First, the spatial accuracy constraints associated with completing the response are low. Second, when performing key press responses, performance-related feedback, such as endpoint accuracy, endpoint precision, and outcome, is not obtained by the actor. Therefore, we were able to determine if a bias towards probability information reported in the previous experiments (Neyedli & Welsh, 2014, 2015b) was due to the constraints of the action itself or if the bias is inherent to tasks in which participants make visuospatial decisions such as when they choose between target–penalty configurations in the current study. Participants in the present experiments chose the configuration with the larger MEG at a level above chance when configurations varied in distance or penalty value in isolation (i.e., Distance-Different and Penalty-Different conditions). Furthermore, when presented with configurations that had equivalent MEG (i.e., Both-Similar condition), participants, as a whole, did not consistently choose the configuration with the more favourable probability or negative gain parameter. In other words, there was not a strong bias to select the configuration with the more advantageous spatial parameters as observed when participants were aiming to the configurations. Taken together, these findings show that when the response mode does not have high levels of terminal accuracy demands, participants as a group used both probability and gain information equivalently when making non-trivial and trivial decisions.
The results of the present experiments, in combination with Neyedli and Welsh (2014, 2015b), indicate that the characteristics of the response used to report decisions (i.e., the to-be-performed action) influences the decision-making process itself (see also, Burk et al., 2014; Hagura et al., 2017; Marcos et al., 2015; Moher & Song, 2014). Indeed, to make accurate aiming movements, initial and continuous assessments of the spatial information associated with the target and limb are necessary (Cisek & Kalaska, 2010; Goodale & Milner, 1992; Heath, 2005; Welsh & Pratt, 2008). Thus, a prioritisation of spatial information may be an important feature that characterises all movements with a high spatial accuracy constraint. When performing a key press response, however, there are minimal accuracy constraints and it would therefore be unnecessary for the actor to prioritise spatial over gain information. Indeed, because there is a lack of interaction between the action and the target during a key press response, the planning of this action is not sensitive to the size or location of the to-be-responded-to target—parameters that influence the planning of aiming movements (Fitts & Peterson, 1964; Welsh & Zbinden, 2009). Interestingly, a similar conceptual argument can be drawn from the selective attention and action literature such that the constraints of the motor task itself lead to increased weighting of perceptual features that are relevant to that action during selection tasks (see Bekkering & Neggers, 2002; Fagioli et al., 2007; Wykowska et al., 2009; see Hommel, 2010 and Hommel et al., 2019 for reviews and discussion). Thus, the type of action being performed (e.g., aiming, grasping, and key presses) will lead to the prioritisation of distinct features, and if properties of the stimuli that participants are deciding on overlap with the prioritised features, then decision making may be influenced.
It is important to understand the present results in the context of other relevant decision-making experiments that use key press responses. Specifically, Wu et al. (2009) had participants choose between prospects of varying probability and positive, rather than negative, gain via key press responses. In contrast to the present results, participants displayed a preference for the option with a low probability of high positive gain rather than a high probability of low positive gain. Thus, a preference for more advantageous gain, rather than probability information, emerged when choosing via a key press response. While in the present experiments no distinct preferences emerged overall, our results in combination with Wu et al. provide evidence that, when expressing the results of a decision via a key press response, a preference towards probability information is less likely to emerge. Interestingly, that a distinct overall preference did not emerge in the present experiments, but did emerge in Wu et al., may be explained by differences in the configuration structure across experiments. Specifically, in Wu et al., no region of the presented prospects was associated with negative gain while in the present experiments a region of negative gain was an important feature of the target–penalty configuration. Importantly, negative and positive gain may be processed and evaluated differently for use in decision-making tasks (Chapman et al., 2015; Tversky & Kahneman, 1981)—a factor which may be responsible for the differences in choice behaviour across experiments.
Individual differences in choice behaviour
While no overall consistent choice bias emerged in the Both-Similar and Both-Different conditions, the between-participant variability was very high in these conditions such that individual participants either displayed a penalty, distance, or no preference (see Figure 2). Given the similar MEGs between configurations in the Both-Similar condition, a preference in either direction would not lead to differences in obtained gain. Interestingly, the individual differences that arose in the Both-Similar condition were associated with the choice behaviour in the Both-Different condition. Thus, participants with a penalty preference in the Both-Similar condition were more likely to choose the lower negative gain configuration in the Both-Different condition. This finding is important because this configuration (i.e., configuration “a,” see Table 1) has a lower MEG in the Both-Different condition and therefore represents sub-optimal choice behaviour according to the prescriptive MEG model because the participant would not gain as many points across the entire experiment (an issue that will be explored further in the next section).
To explore whether or not an individual’s motor variability or differences in MEG shaped these individual preferences on choice behaviour, a series of correlations were conducted between participants’ choice behaviour and participants’ VEx and between-configuration MEG differences (separately). The results of this analysis revealed that motor variability and MEG difference between configurations could not explain choice behaviour in Experiment 1 or 2 (see Table 3). This result suggests that factors not captured by our experimental design may have influenced choice behaviour in the current experiments. That is, the spatial accuracy constraints of a key press response do not warrant a prioritisation of spatial or gain information. Thus, the penalty preference and distance preferences that did emerge among groups of participants suggests that something other than task constraints led to a bias at the level of each individual. For example, Della Libera et al. (2017) provide evidence that personality traits related to reward processing and reward sensitivity were predictive of value-based attentional shifts. We speculate that processing and evaluating the gain associated with target–penalty configurations may relate to reward processing (like that found in Della Libera et al.) and factors related to risk taking (i.e., the level of risk seeking/aversiveness) may influence how participants evaluate probability information. Thus, an important next step for future research is to determine what leads to individual differences in choice behaviour in this task and what variants of the task lead to the greatest amount of inter-subject variability (see Goodhew & Edwards, 2019). However, it should be noted that in Experiment 3, individual differences in choice behaviour could partially be explained by differences in motor variability and the MEG difference between configurations. This result in combination with the correlational analyses of Experiments 1 and 2, provide evidence that participants may have strategically used the experience-based knowledge of their movement characteristics to make their choices, but only when both feedback of the motor variability and of the gain landscape were given. How this inter-subject variability led to optimal and sub-optimal behaviour at the group level will be explored in the next section.
The role of feedback regarding motor variability and gain landscape on choice behaviour
A main finding across the present experiments was that choice behaviour in the Both-Different condition was sub-optimal according to the MEG model until feedback regarding both motor variability and the consequences of landing within each region of the target–penalty configuration were given. That is, in both Experiments 1 and 2, participants did not choose the configuration with the larger MEG at a level greater than chance in the Both-Different condition (a condition in which there was a clear MEG difference between the configurations). Only in Experiment 3 did this more optimal choice pattern emerge. This finding is important because these results suggest that the feedback gained when performing aiming movements to a target in isolation provides different (or at least less than sufficient) information than the feedback gained when performing aiming movements to target–penalty configurations. Indeed, aiming to a target in isolation only provides an indication of one’s own movement variability. Although the knowledge of movement variability is a primary factor in determining the probability metrics associated with a configuration, aiming to a target in isolation does not provide knowledge of the penalty values or the consequences of penalty contact. Aiming to target–penalty configurations, however, not only provides an indication of movement variability but also of the consequences associated with landing in each region of the configurations through reinforcement learning (see also LeBlanc et al., 2020). This feedback of the gain landscape allows participants to better appreciate and more appropriately evaluate the value of the penalty regions (for discussion see Neyedli & Welsh, 2013). In the present experiments, the latter seemed to have the greatest influence on choice behaviour.
In contrast to our current findings, Trommershäuser et al. (2003a, 2008) suggest that knowledge of intrinsic motor variability, rather than information gained through feedback and reinforcement learning, is most important for obtaining optimal behaviour when interacting with target–penalty configurations. In their experiments, participants performed aiming movements in a practice phase and test phase and “optimal” selection in this case was determined by analysing the mean endpoints of the movement that the participants chose to execute (as opposed to choosing between configurations as in this study). In the practice phase of Trommershäuser et al. (2003a), for example, only one target–penalty configuration was presented (target = 100 points, penalty = 0 points), and this configuration did not appear in the test phase. The goal of this manipulation was to provide participants an indication of their intrinsic motor variability only while also experiencing the visual conditions similar to those in the test phase. Critically, participants’ aiming performance in the test phase was immediately close to optimal according to the MEG model. That is, their endpoint selection did not get closer to optimal with increasing experience (as the trials within a block increased) and therefore as they accumulated more feedback of the consequences of landing within each region of the configuration. However, it is important to note that subtle differences in endpoint selection within a block of trials may have been masked by the inter-trial endpoint variability. Using an analysis in which the average endpoint selection within each block of trials was compared across blocks may have been more sensitive in revealing if endpoint selection changes with experience (Neyedli & Welsh, 2013).
It is important to speculate why the current experiments and those of Trommershäuser et al. (2003a) provide conflicting results regarding the feedback that facilitates optimal behaviour. First, it is possible that the practice session in Trommershäuser et al. and the single-target aiming task in the current experiments provide different information. Although both provide an indication of motor variability, aiming to a target–penalty configuration may provide more contextual information to the participant. That is, participants may start to understand how their motor variability may lead to endpoints landing close to or inside the penalty region, even though that specific penalty value may not be encountered in the test phase. In the single-target aiming task used in this study, the absence of a penalty region does not allow for this understanding of contextual motor variability. Second, the relative time in which performance-related feedback was given differed between the present experiments and those of Trommershäuser et al. In the present experiments, varying amounts of feedback were given prior to, but not during the key press task. However, in Trommershäuser et al. performance-related feedback was given prior to and during the test phase. It may be possible that when trial-to-trial feedback of both motor variability and the gain landscape are not given, both sets of information are needed prior to performing the decision-making task. Third, there are inherent methodological differences in the decision-making tasks being performed and the approaches used to determine if optimal behaviour was achieved. Specifically, in the present experiments participants were asked to choose between target–penalty configurations and the relative frequency with which participants chose the higher MEG option was calculated. However, in Trommershäuser et al., participants aimed to target–penalty configurations and the endpoint position relative to an optimal endpoint position was calculated. It is possible that one method of determining optimality is more sensitive to detecting feedback-based changes and thus, task differences may account for the discrepant results. The current experiments and the previous literature cannot distinguish between these possibilities and therefore, future research should be designed to parse these ideas apart.
It is important to note that while optimality is determined as performance relative to a prescriptive model (i.e., the MEG model) in this article and others (Gepshtein et al., 2007; Neyedli & LeBlanc, 2017; Neyedli & Welsh, 2013, 2014, 2015a, 2015b; Trommershäuser et al., 2003a, 2003b, 2006), there are other perspectives that incorporate the processing limitations of the system, in addition to potential gains, when determining what is optimal. 2 For example, a bounded rationality approach (Simon, 1955) suggests that the decision-maker may choose a certain prospect over another if it requires the use of less processing resources and/or requires less allocation of processing resources. In the context of the present experiment and as an alternative way to interpret the data generated from the present experiments, it may be possible that participants who were deemed to display sub-optimal behaviour in the Both-Different condition may have processed penalty information more efficiently than spatial information. Recall, in the Both-Different condition, the configuration with the larger penalty value (and larger distance between regions) had a larger MEG than the configuration with the smaller penalty value (and smaller distance between regions). Thus, participants who processed penalty information more efficiently than spatial information may choose the configuration with the smaller penalty value, even if it has a lower MEG, to reduce the processing demands put on the decision-making system. Thus, in this interpretation based on a bounded rationality approach, what is deemed an optimal choice may not maximise gain as set out in the prescriptive approach.
Although re-interpreting the results of the Both-Different condition in a bounded rationality perspective is important, the likelihood that differences in processing efficiency of penalty and spatial information led to choosing the configuration with lower MEG in Experiments 1 and 2 is low for two reasons. First, and as eluded to throughout the manuscript, when performing key press responses, it is unlikely for penalty and spatial information to have differences in processing efficiency. Second, participants viewed the configurations for a minimum of 2,000 ms prior to making a decision. This amount of time should be sufficient to overcome any initial processing differences, which are unlikely to occur, and to evaluate the spatial and value characteristics of each configuration. However, the possibility that some participants more efficiently evaluated penalty value information cannot be determined from the design or results of the current experiments and remains open to further investigation. In future work, incorporating prescriptive models (e.g., the MEG model) with models that emphasise the context of the decision-maker (e.g., bounded rationality) may help to determine how choice behaviour may be adaptable to different scenarios. Indeed, the results of the present experiments suggest that the context of the response mode and the available feedback are important characteristics that impact gain maximisation, and therefore, when behaviour is impacted by the processing limitations of the decision-making system.
Conclusion
There were three main findings in the current study. First, using a selection response with a low spatial accuracy constraint did not lead to consistent choice biases towards probability or gain information across all individuals. Given the absence of consistent choice biases in the present experiment (in combination with the results of Neyedli & Welsh, 2014, 2015b), we propose that the response mode (i.e., output of the decision process) influences choice behaviour in a decision-making task (see Hommel et al., 2019). Importantly, the present results add to the literature suggesting that there is a high degree of interaction between the action and decision-making systems (see Cisek & Pastor-Bernier, 2014; Wispinski et al., 2018) and, we recommend that when designing decision-making experiments, the response mode should be carefully considered. The second main finding was that individual differences in choice behaviour were prominent in all experiments and, for the most part, our experimental design and recorded data (e.g., endpoint variability) could not capture why these differences arose. Thus, future work should explore the conditions that lead to varying inter-subject variability and should not focus all analyses on group-level performance but examine if different “sub-groups” of participants appear in the data (Goodhew & Edwards, 2019). Last, by varying the amounts of performance-related feedback participants received prior to the decision-making task, we were able to provide evidence that both feedback regarding motor variability and gain landscape were necessary to obtain more optimal choice behaviour (as characterised by the MEG model) and that if sufficient performance-related feedback is given prior to the task, providing feedback during the task may be unnecessary to obtain optimal behaviour. Furthermore, this result evinces that each type of feedback (i.e., regarding motor variability and gain) may provide the actor with distinct, but important, information in determining the MEG of each configuration (see Neyedli & Welsh, 2013) and is consistent with both cognitive (e.g., Barron & Erev, 2003) and motor decision-making experiments (e.g., Neyedli & Welsh, 2013), which suggest that performance-related feedback is necessary to obtain optimal behaviour.
Relating these findings back to the example that opened the manuscript, it seems that, when deciding whether to park in a small space beside an inexpensive van and a large space beside an expensive sports car, not everyone will choose the same parking spot. Different individuals will have different strategies for choosing their parking spot—some will weigh the size of the spot more heavily, some will weigh the type of neighbouring car more heavily, and some will not show a preference. Finally, it seems that only after experience and receiving performance-based feedback will people be more likely to make optimal choices—which might unfortunately mean that some novice drivers might have to dent the doors and bumpers of a number of old vans before learning that parking in the wider spot beside an expensive car may be more optimal in the long run.
Footnotes
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research was supported by Discovery Grants to H.F.N. and T.N.W. and a Post Graduate Scholarship to J.X.M. all from the Natural Sciences and Engineering Council of Canada.
