Abstract
In this study we investigated whether the accuracy of intraset repetitions in reserve (RIR) predictions changes over time. Nine trained men completed three bench press training sessions per week for 6 weeks (following a 1-week familiarization). The final set of each session was performed until momentary muscular failure, with participants verbally indicating their perceived 4RIR and 1RIR. RIR prediction errors were calculated as raw differences (RIRDIFF), with positive and negative values indicating directionality, and absolute RIRDIFF (absolute value of raw RIRDIFF) indicating error scores. We constructed mixed effect models with time (i.e., session) and proximity to failure as fixed effects, repetitions as a covariate, and random intercepts per participant to account for repeated measures, with statistical significance set at p ≤ .05. We observed a significant main effect for time on raw RIRDIFF (p < .001), with an estimated marginal slope of −.077 repetitions, indicating a slight decrease in raw RIRDIFF over time. Further, the estimated marginal slope of repetitions was −.404 repetitions, indicating a decrease in raw RIRDIFF as more repetitions were performed. There were no significant effects on absolute RIRDIFF. Thus, RIR rating accuracy did not significantly improve over time, though there was a greater tendency to underestimate RIR in later sessions and during higher repetition sets.
Keywords
Introduction
Achieving the appropriate stimulus on a resistance training set is paramount for maximizing adaptations while mitigating excess fatigue. Traditionally, resistance training programs have assigned load using a percentage of one-repetition maximum (1RM) or RM zones (Fleck & Kraemer, 2004); however, both approaches may lead to an inappropriate stimulus. Specifically, past data have shown that the number of repetitions that can be performed at a specific percentage of 1RM is highly individual. For example, Cooke et al. (2019) reported that trained men performed a range of 6–28 repetitions during a set to failure at 70% of 1RM on the back squat. Thus, with a prescription such as 4 (sets) × 8 (repetitions) at 70% of 1RM, some individuals would fail on sets while others would receive too low a stimulus. Further, RM zones require individuals to perform sets to muscular failure, which has been shown to elongate temporal recovery compared to non-failure training (Morán-Navarro et al., 2017; Vieira et al., 2021), potentially impairing weekly training volume and frequency. Autoregulation methods (i.e., methods by which a trainee can adjust training variables in real time to accommodate for day-to-day differences in recovery and performance abilities), such as the repetitions in reserve (RIR) scale, can be used to account for problems with percentage-and RM zone-based loading.
The RIR scale, which has a strong inverse relationship with barbell average concentric velocity (ACV) (Zourdos et al., 2016), was developed for individuals to gauge the number of RIR during or after a resistance training set. As one example, training prescription could be 4 (sets) × 8 (repetitions) at 2 RIR, stipulating that the individual should choose a load with which they believe that eight repetitions will be completed with 2 RIR. Theoretically, this autoregulation strategy will precisely control for proximity to failure, mitigating the issues of percentage-and RM zone-based loading. However, the utility of RIR is predicated upon rating accuracy. A recent meta-analysis from Halperin et al. (2022) reported that individuals tended to underestimate RIR, on average, by .95 repetitions, although these authors noted considerable heterogeneity that was likely due to various factors impacting RIR accuracy. Specifically, Halperin et al. (2022) reported that RIR prediction accuracy improved when ratings were made closer to failure, in later sets, and when fewer repetitions were performed in a set.
Another potential factor moderating RIR prediction accuracy is training experience. While Halperin et al. (2022) did not show RIR prediction accuracy to improve with training experience, cross-sectional data from Zourdos et al. (2021) and Ormsbee et al. (2019) showed that experienced lifters reported significantly (lower RIR ratings following squat and bench press 1RMs, respectively. Further, Zourdos et al. (2016) observed a stronger correlation between RIR and ACV in experienced (r = .88) than non-experienced (r = .77) participants. In addition, Hackett et al. (2018) reported that in two sessions performed 48 hours apart, trained participants who predicted RIR after 10 repetitions during a leg press set at 80% of 1RM improved their RIR predictions from an error of 3.1 to 1.9 repetitions over the two sessions. Moreover, Steele et al. (2017) found that when RIR predictions were made before a chest press set, participants with >36 months of training experience predicted RIR to an accuracy of .55 repetitions. However, Steele et al. (2017) also observed that participants with only .5–6 months of training experience had a difference of 3.90 repetitions between their predicted and actual number of repetitions on the chest press. Despite data from Zourdos et al. (2021), Ormsbee et al. (2019), and Steele et al. (2017) suggesting that RIR prediction accuracy improved with training experience, to our knowledge there is no longitudinal study to date that has examined whether RIR prediction accuracy improves over time. The most closely related study is that of Davies et al. (2022) assessing accuracy of RIR ratings before and after a 6-week bench press training study, during which RIR was estimated at the conclusion of each set performed. Importantly, however, intraset RIR accuracy was not assessed throughout the study, only at pre- and post-testing; thus, session-to-session and week-to-week changes remain to be investigated.
Therefore, the purpose of our study was to assess whether intraset RIR prediction accuracy in the bench press improves over a three day per week 6-week training program (following 1 week of familiarization) in trained men. We hypothesized that intraset RIR prediction accuracy would improve over time.
Methods
Experimental Design
All participants in this study trained the bench press three times per week on nonconsecutive days (i.e., Monday, Wednesday, and Friday). Repetitions performed each day during a given week decreased in a daily undulating fashion (Klemp et al., 2016). In the final set of each session in weeks 1–6, participants verbally provided intraset RIR predictions of 4 RIR and 1 RIR before continuing sets to momentary muscular failure. Failure was defined as either the inability to complete a repetition despite maximum effort to do so, or the participant completing a repetition but not being comfortable attempting another. Additionally, participants also performed the squat in a group-specific fashion prior to the bench press in each session. However, due to multiple participants reporting minor muscle pain following squatting to failure, the squats to failure were discontinued after only a few participants completed squatting to failure in week 1. Therefore, squat RIR accuracy data could not be obtained for this study.
Participants reported to the laboratory a total of 22 times over seven consecutive weeks. First, participants participated in pre-testing including anthropometrics and bench press 1RM testing 48–72 hours prior to beginning the program. The initial week (week 0) constituted an introductory microcycle and the following 6 weeks formed the main training program.
To minimize nutrition-related confounders, participants ingested seven grams of a branched chain amino acid supplement (BCAA, Core Nutritionals, LLC, Arlington, Virginia, United States of America, 22,203) containing 3.5 g of leucine, 1.75 g of isoleucine, 1.75 g of valine, and 2.5 g of glutamine 30 minutes prior to each training or testing session. In addition, participants ingested 42 grams of whey protein (Core Pro, Core Nutritionals, LLC, Arlington, Virginia, United States of America, 22,203) containing 3.5 g of leucine immediately following each training session. These supplements and quantities thereof were chosen due to the benefits observed of stimulating muscle protein synthesis prior to and following resistance training (Tang et al., 2009), and 3.5 g of leucine has been indicated to trigger a maximal protein synthetic response (Stark et al., 2012). Both BCAA and whey protein were consumed in powdered form mixed with 10oz of water. Participants were instructed to cease all other supplementation during the study.
Participants
Participant Characteristics.
Note. Data are presented as means (and standard errors of measurement or SEM).
Procedures
Training Program
Training Program.
Note. RIR: repetitions in reserve.
a4 RIR and 1 RIR were predicted during this set.

Repetitions in Reserve-Based Rating of Perceived Exertion Scale. Reprinted with permission from Zourdos et al., 2016.
In the initial pre-training week, participants performed an introductory microcycle consisting of reduced volume. The final set of each session during this initial week was not performed to momentary muscular failure; however, intraset RIR predictions were still provided during these sets for the purpose of familiarization. The full training program for weeks 1–6 can be seen in Table 2. In brief, participants performed sets of 10, 8, and 6 repetitions on Session 1, Session 2, and Session 3 during weeks 1 and 2. In weeks 3 and 4, participants performed sets of 9 (session 1), 7 (session 2), and 5 (session 3) repetitions. In weeks 5 and 6, repetition targets were reduced again to 8, 6, and 4. Participants were asked to select a load, for all sets, that they believed would produce a 1–3 RIR at set termination. Further, the final set of each session was performed until momentary muscular failure. During every set to failure participants verbally indicated when they believed they were at 4 and 1 RIR to predict intraset RIR. Of note, participants were not directed to perform repetitions with maximal concentric intent; thus, habitual concentric intent was utilized.
Training Load Prescription and Adjustments
As noted earlier, to determine load for the initial set in each session, participants were asked to choose a load which they believed would cause the set to end at the predetermined RIR. Then, to determine loading for subsequent sets, participants provided an RIR-based RPE score following each set which determined loading adjustments. Before each set, participants were reminded of the following: (a) the target RIR and repetition range of that session; (b) to use their knowledge of prior performances and warm-up sets to select an appropriate load for the first set of that session; (c) to provide RIR estimations at the conclusion of each set, and; (d) when load selection was the participant’s choice, to expect perceived effort to rise, and therefore RIR to decrease, set-to-set as fatigue accumulates.
Load Adjustment Protocol.
Note. Protocol from Helms et al., 2018.
Measurements
Anthropometrics
Total body mass was assessed via a calibrated digital scale (Mettler-Toledo, Columbus, Ohio, USA). Height was measured via a wall-mounted stadiometer (SECA, Hamburg, Germany). Body-fat percentage was estimated using the average sum of two skinfold thickness measurements acquired from three sites (chest, abdomen, anterior thigh). If any measurement was more than 2 mm different than the previous measure, a third measurement was taken. The Jackson and Pollock equation (Jackson & Pollock, 1978) was used to estimate body-fat percentage, and the same investigator took all measurements.
Bench Press Technique
The bench press was performed in accordance with International Powerlifting Federation standards (International Powerlifting Federation, n.d.). Specifically, participants laid supine on a weight bench, maintaining five points of contact (head, butt, and shoulders in contact with the bench, and both feet flat on the floor throughout the movement). The barbell was removed from the rack and participants held it with arms extended in a stable position. At pre-testing, participants were given the choice of unracking the barbell themselves or having investigator assistance, and this was kept consistent in all testing and training sessions. Investigators issued a start command upon which participants lowered the barbell until it contacted the chest and then pressed upwards until the arms were once again fully extended. The barbell was not required to pause on the chest. Participants waited until a rack command was issued to re-rack the barbell.
One-Repetition Maximum (1RM) Testing
All 1RM testing was performed in accordance with previously validated procedures (Zourdos et al., 2016). Specifically, all participants completed a 5-minute dynamic warm-up followed by a bench press-specific warm-up consisting of the following: as many repetitions as desired with an empty barbell, then 5 repetitions with 20% of their estimated 1RM, followed by 50% of estimated 1RM for 3 repetitions, 70% of estimated 1RM for 2 repetitions, and 80% of estimated 1RM for 1 repetition. Participants were then given 3–5 minutes rest before a final warm-up repetition at a load determined by the investigators (between 85 and 90% of estimated 1RM). Following the final warm-up, participants took 5–7 minutes rest while the investigators determined the load for the first 1RM attempt. Load was increased on each subsequent attempt until a 1RM was reached and 5–7 minutes rest was given between each attempt. On every warm-up and 1RM attempt, RIR-based RPE and average concentric velocity were collected to aid in attempt selection. A 1RM was accepted as valid if one of three conditions were met: (a) Participant reported 10 RPE (0 RIR) and the investigators determined an additional attempt with increased load would be unsuccessful, (b) participant reported a 9.5 RPE (.5 RIR) and then preceded to fail the subsequent attempt with a load increase of 2.5 kg or less, and (c) participant reported a 9 RPE (1 RIR) and failed the subsequent attempt with a load increase of 5 kg or less. Finally, Eleiko barbells and lifting discs (Chicago, Ill., USA) calibrated to the nearest .25 kg were used for all 1RM testing.
Repetitions to Failure and Intraset RIR Prediction
On the final set of each session in weeks 1–6 (18 total sets), participants performed a set to momentary muscular failure during which they provided intraset RIR predictions of 4 RIR and 1 RIR (Table 2). Muscular failure was defined as either (i) the participant was unable to complete a repetition and needed assistance in racking the barbell, or (ii) the participant completed a repetition and was not comfortable attempting another. Thus, while participants were encouraged to reach “momentary muscular failure” as in definition 1, in 15 of 157 analyzed data points participants achieved what is more accurately termed “volitional muscular failure” as in definition 2. In such cases, participants reported RIR-based RPE of 9.5–10. The procedures for the intraset RIR predictions are explained in the following script which was read to each participant before all 18 sets to failure during the study: “The next set is the last set for bench press today. We’ll be taking this set until failure—if you complete a rep, we would prefer that you attempt another repetition if you are comfortable doing so. So, just focus on performing as many repetitions as possible, and we will spot you very closely. During the set, say ‘NOW’ when you feel like you have 4 reps remaining and when you feel like you have 1 rep remaining.”
Based upon the RIR predictions, both the predicted repetitions to failure at each called RIR and actual repetitions performed were recorded. Next, the difference between the predicted and actual repetitions performed (predicted repetitions—actual repetitions) was calculated as the RIR difference (RIRDIFF) for both RIR predictions. For example, if a participant said “now” for the first time during the failure set following the sixth repetition, then said “now” for the second time following the 10th repetition, that means that participant predicted they would go on to complete 10 total repetitions at the time of providing the 4 RIR prediction, and 11 total repetitions when providing the 1 RIR prediction. If the participant actually performed 12 total repetitions, then the RIRDIFF would be −2 repetitions for the 4 RIR prediction (10 predicted repetitions − 12 actual repetitions) and −1 repetition for the 1 RIR prediction (11 predicted repetitions − 10 actual repetitions).
Further, from all RIR predictions both the raw RIRDIFF and absolute value RIRDIFF were calculated. The raw RIRDIFF accounted for directionality (negative values indicate underestimation and positive values indicate overestimation). For instance, an RIRDIFF of −2 would indicate that a participant underpredicted RIR by two repetitions. The absolute value RIRDIFF treated an underprediction or overprediction of two repetitions simply as an RIRDIFF of 2. In this way, the raw RIRDIFF provided how much on average participants under-or overpredicted RIR, while the absolute value RIRDIFF provided a true measure of accuracy (i.e., the overall average prediction error).
Statistical Analyses
Sample Size
Given the high degree of work involved in a resistance training intervention, we chose this sample size based on feasibility (Lakens, 2022), and therefore no formal power analysis was performed. As many participants as possible were recruited given the constraints of the investigators’ time of graduation. Because the sample size was limited, efforts were undertaken to ensure that data were as easy as possible to aggregate for future meta-analysis. Moreover, despite the relatively low number of participants, a large number of observations occurred throughout the intervention, resulting in 18 time points for each participant in this repeated measures design, strengthening the power of the statistical model we utilized.
Primary Analysis
All statistical analyses were performed using ‘R’ software (v 4.0.2; R Core Team, https://www.r-project.4org/). To investigate the accuracy of RIR predictions, separate mixed effects models were generated for raw and absolute RIR accuracy, respectively. Raw RIR accuracy was analyzed with a linear mixed effects model using restricted maximum likelihood estimation (REML). Absolute RIR accuracy, however, was analyzed with a generalized linear mixed effects model with a Poisson error distribution (specified with a log-link function and REML due to positively skewed count data of the dependent variable). Both models included fixed effects for (a) the RIR target at which accuracy was evaluated, (b) the training session in which the prediction was made, and (c) the interaction thereof. Random intercepts were included per participant to account for repeated measures and each model was adjusted for the number of repetitions performed per set, which was included as a covariate.
To address the primary research questions, null-hypothesis significance testing was performed to evaluate the presence of main or interaction effects using the “car” package. A Wald F-test (with Kenward-Roger degrees of freedom) was used for the model investigating raw RIR accuracy, while a Wald Chi-squared test was used for the model investigating absolute RIR accuracy. For both models, type III sum of squares were utilized, and statistical significance was set at p ≤ .05. To supplement the null-hypothesis significance testing, estimated marginal means (and slopes thereof) were calculated for all parameters of interest and presented with 95% confidence intervals. All estimates presented from the Poisson model were back transformed to the response scale prior to marginal averaging using the “emmeans” package to allow for additive comparisons. Finally, prior to performing any tests or extracting model estimates, the quality of model fit was assessed visually using the “DHARMa” package. The data used for this study are available on the Open Science Framework (https://osf.io/uztg2).
Results
Raw RIRDIFF
Mixed Effects Models.
Note. Mixed effects models. RIR: Repetitions in Reserve.
After adjusting for the perceived RIR and the total number of repetitions performed within a set, the estimated marginal slope for time (i.e., per session) was −.077 repetitions (SE = .011, 95% CI: −.099, −.055), indicating a slight decrease in raw RIRDIFF over time (Figure 2A). After adjusting for the total number of repetitions performed within a set and the training session in which the prediction occurred, RIRDIFF at the perceived 4 RIR was .192 repetitions (SE = .146, 95% CI: −.134, .519), while the RIRDIFF at the perceived 1 RIR was .131 repetitions (SE = .150, 95% CI: −.200, .461) (Figure 3A). After adjusting for the perceived RIR and the training session in which the prediction took place, the estimated marginal slope for repetitions per set was −.404 repetitions (SE = .027, 95% CI: −.456, −.351), indicating a slight decrease in raw RIRDIFF as more repetitions are performed within a set (Figure 4A). AB. Main Effect of Time (Sessions) on RIRDIFF. AB. Main Effect on RIR and RIRDIFF. Effect of Repetitions on RIRDIFF.


Absolute RIRDIFF
No significant effects were observed for Time, χ2 (1, n = 258) = .719, p = .397, RIR, χ2 (1, n = 258) = 3.648, p = .056, or the interaction of Time × RIR, χ2 (1, n = 258) = .296, p = .586. Moreover, the total number of repetitions performed per set was not significantly associated with the accuracy of RIR predictions, χ2 (1, n = 258) = .213, p = .644. (Table 4).
After adjusting for the perceived RIR and the total number of repetitions performed within a set, the estimated marginal slope for time (i.e., per session) was .001 repetitions (SE = .012, 95% CI: −.014, .034) (Figure 2B). After adjusting for the total number of repetitions performed within a set and the training session in which the prediction occurred, RIRDIFF at the perceived 4 RIR was 1.107 repetitions (SE = .151, 95% CI: .809, 1.406), while the RIRDIFF at the perceived 1 RIR was .683 repetitions (SE = .109, 95% CI: .469, .897) (Figure 3B). After adjusting for the perceived RIR and the training session in which the prediction took place, the estimated marginal slope for repetitions per set was .013 repetitions (SE = .028, 95% CI: −.041, .067) (Figure 4B).
Discussion
To our knowledge, this is the only longitudinal study to date to examine if RIR prediction accuracy changes over time on a session-to-session basis. We observed significant trends whereby raw RIR accuracy, which indicates directionality of prediction error, tended to decrease, suggesting a greater likelihood of underpredicting RIR, over time and as more repetitions were performed during a set. However, absolute RIR, which indicates magnitude of prediction error, was predicted within an accuracy of ∼1 repetition with no evidence of a change over time or with different repetitions performed, displaying a high degree of accuracy throughout the study. Further, there were no statistically significant differences between RIRDIFF at the perceived 4 RIR and 1 RIR for neither raw nor absolute RIRDIFF, indicating a lack of difference between RIR prediction accuracy at the two examined proximities to failure. Therefore, our hypothesis that RIR prediction accuracy would improve over time was not supported. Overall, these findings suggest that trained lifters can predict RIR on the bench press to within ∼1 repetition at various proximities to failure and that intraset RIR prediction accuracy does not improve on the bench press over 6 weeks in trained men.
Various studies (Hackett et al., 2017; Mansfield et al., 2020; Odgers et al., 2021; Zourdos et al., 2021) have reported that intraset RIR predictions were more accurate when made closer to failure. In contrast, we observed no such difference, neither in terms of raw RIRDIFF (M = .192, SD = .146 vs. M = .131, SD = .150 at 4 RIR and 1 RIR, respectively, main effect of RIR: p = .465), nor absolute RIRDIFF (M = 1.107, SD = .151 vs. M = .683, SD = .109 at 4 RIR and 1 RIR, respectively, main effect of RIR: p = .056). In either case, participants in the present study were quite accurate (i.e., within ∼1 repetition of error) with their predictions. Other studies such as Odgers et al. (2021) reported RIRDIFF scores of <1 during 4 and 1 RIR predictions on the front squat and hexagonal barbell deadlift. However, Zourdos et al. (2021) reported an RIRDIFF mean of 2.05 (SD = 1.73 when trained men predicted 1 RIR during a set of barbell back squats. The notable difference in RIR prediction accuracy in the present study compared to Zourdos et al. (2021) is likely because Zourdos et al. had participants predict intraset RIR during one set to failure at 70% of 1RM in which more repetitions per set were performed (M = 14, SD = 4 vs. M = 7.51, SD = .24 in the present study). Indeed, Zourdos et al. (2021) noted that more repetitions in a set was significantly related (p = .003) to more inaccurate RIR predictions. Further, in the present study, participants made RIR predictions during their third or fourth set with a similar load compared to the first set in Zourdos et al. (2021). In addition, Mansfield et al. (2020) reported more accurate intraset RIR predictions on sets two, three, and four of the bench row and bench press than on the first set. Therefore, it seems the high degree of RIR prediction accuracy in our study can be, in part, explained by participants predicting intraset RIR after the first set.
A recent meta-analysis from Halperin et al. (2022) found that total repetitions performed in a set was related to raw RIRDIFF (β = .47 repetitions, 95% CI = .35–.58). The findings of the present study are consistent with this, as we observed raw RIRDIFF to decrease as total repetitions performed increased (p < .001; β = −.404 repetitions). However, for absolute RIRDIFF, we identified no such relationship. The lack of relationship between total repetitions per set and absolute RIR prediction accuracy could be because repetitions performed in the present study were relatively low compared to other studies (M = 7.51, range = 2–14). Other studies, where a relationship between total repetitions per set and RIR prediction accuracy was observed, such as Zourdos et al. (2021), had a much wider range of repetitions performed per set (i.e., 6–28 repetitions). Although Halperin et al. (2022) found that repetitions per set was related to RIR accuracy, they also reported that this relationship was significantly stronger when >12 repetitions were performed in a set. Therefore, our study revealed that when the difference in repetitions between sessions was small and the target number of repetitions per set was <12, total repetitions per set did not influence RIR prediction accuracy. However, we did observe that repetitions per set were significantly related to raw RIRDIFF, indicating that while overall accuracy was indeed very high, participants were more prone to underpredicting RIR in higher repetition sets.
The novelty of this study was our examination of RIR prediction accuracy over time. However, despite our hypothesis that RIR prediction accuracy would improve, we found no significant decrease of absolute RIRDIFF over time (p = .397). We hypothesized that RIR prediction accuracy would improve over time as multiple cross-sectional studies suggested greater RIR prediction accuracy in trained versus untrained participants (Ormsbee et al., 2019; Steele et al., 2017; Zourdos et al., 2016). Specifically, Zourdos et al. (2016) observed that novice trainees (M = .4, SD = .6 years training experience) reported an average RIR of 1.04 (SD = .43) after a squat 1RM while individuals with at least 5 years of training experience reported an average RIR of .20 (SD = .18). Further, in Ormsbee et al. (2019) individuals with ≥5 years of training experience reported a lower RIR during a bench press 1RM test than novice trainees (M = 1.1, SD = .6 years training experience). Additionally, Steele et al. (2017) found that men and women with >3 years of training experience predicted RIR before a set to an RIRDIFF of <1, while those with <6 months of training experience had an RIRDIFF of ∼5 repetitions across a variety of exercises (e.g., seated row, chest press, leg press, biceps curls, and pulldowns). Importantly, all these studies were cross-sectional and had large gaps of training experience between trained and novice individuals. The meta-analysis from Halperin et al. (2022) did not detect a significant relationship between training status and RIR prediction accuracy, which may be due to the use of trained individuals in studies examining RIR prediction accuracy. For example, Zourdos et al. (2021) failed to observe a significant relationship between training status and RIR accuracy when enrolling a population that was required to back squat for two consecutive years prior to the study with a training age of 2–12 years. In other words, even though Zourdos et al. (2021) had a 10-year range of training experience, the study did not compare trained and untrained individuals. Finally, in a recent study by Davies et al. (2022), no significant improvement in RIRDIFF was observed between pre-and post-testing of their 6-week training study which included many opportunities to practice rating RIR. Of note, the participants in this study had mean years of training experience of 3.5 (SD = 1.9) and 3.1 (SD = 2.0) in the two groups being compared. Therefore, while intraset RIR prediction accuracy may improve with training status, this improvement may occur within the first 2-3 years of a training career and then plateau. Participants in the current study had an average training experience of 2.664 years (SD = .422) and already had absolute RIRDIFF means of 1.058 (SD = .209) and .566 (SD = .163) at the perceived 4 and 1 RIR, respectively, in week 1; thus, the participants may simply have not had much room to improve.
To our knowledge, the only other study that examined RIR accuracy during more than one session was Hackett et al. (2018), which had participants perform three sets each of the chest press and leg press exercises in session 1, then repeated the protocol 48 hours later. The only difference observed between the two sessions was a higher RIRDIFF during set 1 of session 1 on the leg press compared to session 2 (means of 3.1 vs. 1.9 repetitions, respectively; p = .013). Hackett et al. (2018), stated that “most participants (42 of 48) reported having ≥1 year of resistance training experience at the recreational level,” which is quite possibly lower than the training experience of the participants in the present study (M = 2.664, SD = .422 years). As speculated earlier, it appears improvements in RIR prediction accuracy mainly occur within the first 2-3 years of a training career. Another explanation is that participants in Hackett’s study performed the same protocol (3 sets to failure with RIR prediction after 10 repetitions) between sessions, whereas participants in the present study performed a different number of repetitions at a different load in each session. Therefore, participants in Hackett et al. may have improved simply due to learning how many repetitions they could do during the first session.
Unadjusted and Adjusted Mean RIRDIFF Over Time.
Note. Unadjusted data are presented as means (and standard deviations) and represent the RIRDIFF values without adjustment for other variables from statistical modeling. Adjusted data are presented as estimated marginal means [95% confidence intervals] and represent estimated RIRDIFF after statistical modelling and adjusting for the number of repetitions performed.
Limitations and Directions for Further Research
The primary limitation of the current study was our relatively small sample size. However, the large number of observations per participant (18), strengthened the power of the statistical model utilized. Other limitations include only using the bench press, having participants predict RIR only on the last set, and including only trained men. Importantly, it cannot be ascertained from the current study if RIR accuracy would change over time with other exercises or in women, adolescents, elderly, or untrained populations, and future investigators might explore these populations further. The RIRDIFF values in this study may also not apply during single set training as lifters’ RIR prediction accuracy could have benefited from previous sets.
As we observed that raw RIRDIFF changed over time, potentially indicating a difference in over-or under-prediction associated with differences in loading (as heavier loads were used later in the study), future investigators should further examine the role that loading plays in RIR prediction accuracy, as well as whether novice trainees do in fact improve RIR prediction accuracy over time, which we were unable to investigate in our study due to the participant inclusion criteria. Moreover, if such improvements occur, reasonable time scales for RIR prediction accuracy to improve to <1 repetition, on average, should be determined. Additionally, we encourage researchers to include data for both the raw and absolute value RIRDIFF scores in the future. Reporting both scores provides insight into the directionality of prediction error as well as the actual degree of error.
Conclusion
Overall, our data agree with prior research that RIR predictions are accurate to <1 repetition when provided close to muscular failure in low-to-moderate repetition sets. We did not observe a significant improvement in RIR prediction accuracy over time, although that may be because predictions were already accurate in week one of the study, due to the participants’ training status. Interestingly, participants tended to overpredict RIR in the early weeks, but underpredict RIR in later weeks when relative intensity (percentage of 1RM) increased. Practically, these results suggest that experienced trainees (>2-3 years of experience) can likely predict RIR accurately to <1 repetition at a variety of repetition ranges and loads. However, it is still possible that novice trainees may increase RIR prediction accuracy during the first 2-3 years of their training career. Also, trainees should be aware of the potential for loading to influence their RIR predictions. Specifically, trainees may be more prone to overpredict RIR with lighter loads and underpredict RIR with heavier loads.
Footnotes
Acknowledgments
The authors would like to thank the participants for their time and effort, as well as Core Nutritionals, LLC for providing supplements for this study. No funding was received for any part of this work.
Declaration of Conflicting Interests
The author(s) declared the following potential conflicts of interest with respect to the research, authorship, and/or publication of this article: Authors JFR, ZPR, JCP, SRH, EE, ERH, and MCZ are all coaches and writers in the fitness industry.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
