Abstract
We empirically estimate positional “wins above replacement” (WAR) in the National Football League (NFL). Positional WAR measures the value of players in the NFL, by position, in terms of generating wins. WAR is a commonly used metric to evaluate individual players in professional baseball and basketball in the United States, but to the best of our knowledge, this is the first study to construct WAR measures for American football. A key challenge in constructing these measures is that individual statistics for many football players are not as well developed as in baseball and basketball. Related to this point, the productivity of individual football players, perhaps more than players in any other major sport, is highly dependent on context. We circumvent issues related to measuring productivity for individual players by constructing WAR measures at the position rather than individual level. The identifying variation that we leverage in our study is generated by arguably exogenous player injuries and suspensions. Using data from three seasons and all 32 NFL teams, we show that the most valuable positions in the NFL are quarterback, wide receiver, tight end/fullback, and offensive tackle. Perhaps our most surprising finding is that positional WAR for all positions on the defensive side of the football is zero.
Keywords
Introduction
“Wins above replacement” (WAR) and similar data-driven metrics of player productivity are increasingly used by front-office personnel for professional sports franchises to evaluate players (Barnwell, 2014; Fleming, 2013; London, 2014). These metrics are statistical constructs based on individual and team performance measures, and their incorporation into player assessment represents a fundamental shift in how professional franchises evaluate athletes (DuPaul, 2012). Many attribute the move toward statistics-based analyses of athletes to the publication of Moneyball: The Art of Winning an Unfair Game by Michael Lewis in 2004. The book details how the Oakland Athletics used statistical models to evaluate players. The intuition can be explained as follows: Team wins are a function of runs scored and runs allowed, and runs scored and runs allowed are the functions of individual player statistics. Thus, each individual statistic can be expressed as a contribution to team wins or marginal productivity of labor. Although popularized by the book, this type of model was formally presented by Blass (1992) over a decade earlier, with an application to player wages.
Building on the approach popularized in Moneyball, Gerard (2007) discusses the challenges of applying statistical measures of player productivity to team sports, where individual statistics are dependent on teammate performance. He presents a conceptual model and applies it to soccer players in the English Premier League. A similar approach has also been applied to basketball players by Berri, Schmidt, and Brook (2007) and Hollinger (2014). Berri et al. use a regression framework to create a metric called “Wins Produced.” Hollinger’s “Estimated Wins Added” converts individual player statistics into measures of team wins. 1 A conceptual similarity between both of these basketball metrics is that they compare each player’s contribution to a hypothetical replacement player, thus the “Wins Above Replacement” terminology.
All currently estimated WAR measures of which we are aware are based on the idea that individual statistics can be translated into measures of contributions to team wins. However, linking individual player statistics to wins is particularly challenging in American football. As discussed by Gerard (2007), an individual player’s statistics in football depend on the performance of other teammates more so than in other sports. For example, consider a 1-yard touchdown run by a running back. The touchdown is officially attributed to the running back, but how much of that touchdown is due to his performance and how much is due to the blocking performance of the offensive line (the latter being particularly difficult to measure with individual player statistics)? And what if the play was set up by an 80 yard punt return? The interdependence of individual player performance in American football makes it difficult to use individual statistics to measure the marginal productivity of labor.
These challenges notwithstanding, some productivity metrics have been developed for players in the National Football League (NFL). These metrics are summarized by Berri and Burke (2012), but they all assume that individual player statistics are independent of the contributions of teammates. Realizing this issue, Oliver (2014) created a player-productivity metric for NFL quarterbacks called total quarterback rating (total QBR), which is used by ESPN. Oliver (2014) describes total QBR as a statistical measure that incorporates the context and details of available statistics and what they mean for wins. However, total QBR is a proprietary measure and the extent to which it truly accounts for context and player interdependence in the NFL is unknown. A different approach to estimate player productivity is taken by Atkinson, Stanley, and Tschirhart (1988). Instead of using individual player statistics, they use team statistics to estimate the probability of winning. However, due to their aggregating the statistics to the team level, it is not possible to determine the marginal product of labor for individual players or positions.
The contribution of the present study is to add to the quickly growing, empirically driven literature on player productivity by developing a new, fully contextualized measure of player productivity in the NFL at the position level. We call this new measure “positional WAR.” By constructing our WAR measure at the position rather than individual-player level, we can circumvent many of the above-described challenges associated with attributing wins to individual players, most notably because we estimate positional WAR without relying on individual game statistics for players (see subsequently). Positional WAR can be a useful tool for teams looking for quantitative, objective measures of positional value. It can help to inform decisions about how to invest scarce salary-cap dollars, draft choices, and roster spots. 2 Indirectly, positional WAR represents a first step toward the construction of rigorous, summative performance measures for individual NFL players.
The identification strategy that we employ to estimate positional WAR relies on player injuries and, to a lesser extent, suspensions. Games lost by starting players due to injury and suspension are arguably exogenous and inherently unpredictable. By definition, an injured or suspended player must be replaced. We estimate positional WAR by determining how the use of replacement players for injured and suspended starters affects season win totals for all 32 NFL teams across three seasons.
Our estimates of positional WAR measure the marginal contributions of starters compared to replacement players by position. The measures depend on the relative quality of starters and replacement players within positions and thus reflect the scarcity of talent at the position level (in contrast, positional WAR does not directly compare the absolute value of player performance across positions, which exists conceptually but would be difficult to measure because players at different positions do different things). If, for example, a team has three defensive tackles of equal quality yet only two can start, the difference in productivity between the starters and the backup would be small, hence a case with low positional WAR. Alternatively, if a starting wide receiver is significantly better than his second-string backup, positional WAR would be high.
Our analysis reveals some predictable patterns. For example, we find that quarterbacks are the most valuable position by a wide margin. Other positions for which we find significantly positive WAR values are wide receiver, tight end/fullback, and offensive tackle. Perhaps our most interesting finding is that no position group on the defensive side of the football has a significant WAR value. Put differently, on average, teams do not suffer in terms of wins when defensive starters miss games due to injury and/or suspension. This suggests that relative to their replacements, defensive starters on average are not as valuable as offensive starters at several positions.
Data
We construct a data panel for all 32 NFL teams containing performance (wins) and injury/suspension information for 3 years: 2008, 2010, and 2012. 3 We use even-numbered years instead of all years to reduce the data collection burden; collecting the injury and suspension data is labor intensive (see Appendix A). We spaced out the 3 years of data collection to reduce the intra-team correlation in outcomes in the data panel, which improves statistical power.
The independent variables of interest in our analysis measure games lost to injury and suspension, by position, for each NFL team in each year. To build these variables, we first identified the projected starting lineup for each team in each year prior to the first game of the season. 4 We then determined the number of games that each projected starter lost to injury and/or suspension using a variety of sources (e.g., news articles, fantasy football updates, and injury reports). 5 Of course, some starters missed games for other reasons, such as a demotion; however, our identification strategy relies entirely on using variation from games missed due to injury and suspension, so we code up only games missed for these reasons.
To construct the data set, we began by determining the number of projected starters at each position for each team. All defenses use two starting cornerbacks and two starting safeties, and either a Base 4-3 or 3-4 system. 6 On offense, every team has one starting quarterback, one starting running back, and the traditional five-man offensive line. Starting wide receivers, tight ends, and fullbacks vary. For example, some teams do not use a fullback in the starting lineup; some use two-tight ends, some use three-wide receivers, and so on. No team includes more than one fullback, more than two-tight ends, or more than three-wide receivers in the starting lineup. 7
It can be argued that being a starter in football is less meaningful because of the amount of substitutions that occur during a game. For example, in a soccer match, each team is allowed only three substitutions per match; and as a result, almost all minutes are played by starters. Alternatively, in the NFL, a team can substitute on each play so a starter may get significantly less playing time. To determine the importance of being a starter in terms of predicting playing time, we collected data from 2012 on the number of snaps played by starters on a game-by-game basis. On the offensive side of the ball, starting quarterbacks and offensive linemen play 96% of the snaps. Starting wide receivers and running backs play 78% and 61% of snaps, respectively, while tight ends and fullbacks play 68% and 37% of snaps. On defense, starting defensive linemen play 69% of snaps, starting linebackers play 80%, and defensive backs (safeties and cornerbacks) play 91% of all snaps. 8 Thus, despite the prevalence of specialized player packages, starters on offense and defense in the NFL play the overwhelming majority of snaps within a game.
With the list of projected starters at each position for each team in hand, along with the injury and suspension data, we constructed measures of games lost due to injury/suspension for each team at the following positions: Offense: Quarterback, running back, tight end/fullback, wide receiver, interior offensive lineman (center, guard), and exterior offensive lineman (tackle). Defense: Cornerback, safety, linebacker, and defensive line.
We combine some positions to improve statistical power. An example is the tight-end/fullback position. Similarly, we group interior and exterior defensive lineman. In the extensions section, we discuss the robustness of our findings to alternative positional groupings but use the above groupings for our main analysis. 9 Table 1 shows the average shares of games missed for NFL teams at each position due to injury and suspension over the course of our data panel.
Average Shares of Projected Starter Games Missed Due to Injury and Suspension, by Position, for NFL Teams in 2008, 2010, and 2012.
Note. For each position, we performed tests to determine whether the average number of games missed due to suspension and injury is statistically different from the average at all other positions. There are no statistically significant differences for any position at the 5% level, although the differences for cornerbacks (more likely to get injured) and tight ends/fullbacks (less likely to get injured) are statistically significant at the 10% level.
Empirical Strategy
To estimate positional WAR in the NFL, we estimate the following empirical model:
In Equation (1), W it is the number of wins for team i in year t. PW it is the preseason over-under win total for team i in year t, taken from a Las Vegas sportsbook. We include the “predicted wins” variable for the sole purpose of improving the predictive power of the model and thus the statistical precision of our estimates. 10 It is moderately valuable in this role—it explains 12% of the variation in win totals across the data panel. The variables QB it, RB it, TF it, WR it, IOL it, EOL it, CB it, S it, LB it, and DL it indicate the share of total starter games lost due to injury or suspension for team i in year t at each position (as listed in the previous section). These are the independent variables of interest. ∊it is the error term. We cluster the standard errors from Equation (1) at the team level.
The coefficients
Results
Table 2 shows results from the estimation of Equation (1), with and without the inclusion of the predicted-wins variable. The table shows that the positions where lost starters are most important are quarterback, tight end/fullback, wide receiver, and exterior offensive lineman. For the other positions, there are not statistically identifiable consequences associated with using replacement players in the face of an unexpected absence of a starter.
Estimated Effects on Total Wins of Games Missed Due to Injury/Suspension, by Position.
Note. Standard errors clustered at the team level are in given parentheses.
**Indicates statistical significance at the 1% level. *Indicates statistical significance at the 5% level.
We note two complications with interpreting the estimates in Table 2. First, the variables of interest are coded as the percentage of games missed at each position. Thus, the strict interpretation of the coefficients is that they indicate the effect of going from 0% to 100% games missed on total wins. However, interpreting the estimates in this way involves extrapolating well out of the range of variation in the data, as indicated by Table 1. Second, the percentage variables for games missed by position correspond to different numbers of total games missed. For example, losing 50% of the games at quarterback corresponds to missing eight games, whereas losing 50% of the games at wide receiver corresponds to missing almost 17 games (on average across teams—see Table 1).
To provide a more direct comparison of positional value based on our estimates, Table 3 shows the effect of losing exactly four games due to injury/suspension at each position based on our estimates in Table 2. We use the estimates from the model that includes the predicted-wins variable (the first set of estimates shown in Table 2). Table 3 shows that on a per-game basis, there is no position more valuable than quarterback. Players at wide receiver, tight end/fullback, and exterior offensive-line positions are all similarly valuable; and losing players to injury at other positions does not affect win totals. Note that for the positions where we do not find statistically significant WAR values, statistical power is not the issue. The point estimates for the effects of losing four games to injury for running backs, interior offensive linemen, and all defensive positions are quite small.
Estimated Effects on Total Wins of Missing Exactly Four Games Due to Injury/Suspension at Each Position.
Note. The effect of a four-game absence at each position is estimated based on the full model shown in Table 2.
**Indicates statistical significance at the 1% level. *Indicates statistical significance at the 5% level.
To put our estimates into context, consider that the average number of wins in the NFL is eight. When a quarterback is injured for four games, the team is expected to lose more than one additional game, on average, which is a large effect. We also reinforce the point that our estimates of positional WAR should be interpreted to indicate positional value on average, not the value of individual players. They do not preclude situations where an individual player at a position with low positional WAR is more valuable than an individual player at a position with higher positional WAR.
Sensitivity Analysis
Table 4 shows estimates from a model similar to Equation (1) except that we subdivide tight ends and fullbacks. 11 For brevity, we only show estimates from the full model that includes the predicted-wins variable. Recall from above that not all teams use a fullback—in the data, just over a third of the teams (37.5%) use a fullback in the starting lineup.
Estimated Effects on Total Wins of Games Missed Due to Injury/Suspension, by Position, With Expanded Positional Categories.
Note. Standard errors clustered at the team level are given in parentheses.
**Indicates statistical significance at the 1% level. *Indicates statistical significance at the 5% level.
The extended model shown in Table 4 suggests that between tight ends and fullbacks, it is the fullback position that is more important. Note that this is only a nominal result—put differently, we cannot statistically reject the null hypothesis that tight ends and fullbacks are equally valuable. Still, the fact that our combined tight-end/fullback estimate is not driven by tight-end runs counter to what one might expect given the rising prominence of the tight-end position in the NFL (Benoit, 2012).
We can only speculate as to why fullbacks are important for wins. One possibility is that the decline of the prominence of the position in the modern NFL is driven in part by the lack of availability of starter-quality fullbacks, which would make the fullbacks we do see more valuable. It may also be that the use of a fullback implies a dearth of talent on the roster, which would impact our positional WAR estimates driven by the difference in talent between starters and reserves—put differently, maybe all teams would rather start an additional player at a different position rather than a fullback, but some teams do not have a strong “next in line” tight end or wide receiver. Finally, some circumstances might add considerable value to fullbacks; for example, it may be that there are quarterbacks who perform better with backfield blocking, which is typically a role filled by the fullback.
Discussion
While it is beyond the scope of the present study to perform an in-depth analysis of the mechanisms that contribute to our positional WAR findings, in this section we briefly discuss three plausible explanations for why some positions have higher WAR values than others. In particular, we discuss (1) the importance of scheme in compensating for losses in player quality, (2) positional talent scarcity, and (3) constraints that affect franchises’ team-building decisions.
With regard to our findings on the defensive side of the football, one possible explanation for the lack of positive positional WAR is that defensive schemes can be adjusted to account for replacement players more easily than offensive schemes. For example, a downgrade at cornerback can be facilitated by more help from safeties and/or a broader adjustment to the coverage scheme. Losing one player on defense may not have a significant impact if the other players are able to continue functioning as a whole. This would not imply that defensive positions are insignificant; rather, it would imply that defenses overall function as a unit in which individual personnel are more easily interchangeable. It is also important to remember that our findings are estimated within the current investment equilibrium in professional football. It would be a mistake to interpret our findings to imply that a defense consisting entirely of backup players would perform no worse than a defense full of starters. Put differently, our finding that teams can fully compensate for the loss of an injured player on defense with their current personnel, on average, does not mean that a defense would not be affected by the simultaneous downgrade of numerous defensive positions at the same time. 12
Still, our findings for defensive positions are not consistent with the popular “defense wins championships” mantra, although our study is not the first to suggest that offensive efficacy is more important. A more thorough look at regular season success shows that teams with better offenses win more frequently than teams with the better defenses (Moskowitz, 2012). Further research has looked at the distribution of success of past offenses and defenses and found that exceptionally talented offenses are more successful than exceptionally talented defenses (Burke, 2008). Our results complement these previous studies by showing that the value of starting players at several offensive positions is greater than that of their defensive counterparts.
In addition to potentially driving our findings for defensive players, scheme may also be important on offense. Consider the comparison of exterior and interior offensive linemen as an example. It may be that teams can largely compensate for the loss of a starting interior offensive lineman by changing responsibilities within the line, while for an exterior lineman compensation might require diverting other players to help as blockers, such as a tight ends, fullbacks, or running backs.
A second explanation for our findings is differences across positions in the scarcity of talent. Because positional WAR inherently measures starter quality relative to the quality of replacements, if there is a high supply of talented players at a given position, then the drop off between the starter and the backup will be less and positional WAR will be small. Anecdotally, the position where talent is the most scarce is quarterback and unsurprisingly we find the largest positive WAR for the quarterback position by a wide margin.
A third explanation for our findings relates to the constraints that NFL franchises operate within to build their teams. These constraints include the salary cap, the cap on roster size, and limited draft picks, all of which force teams to make talent trade-offs throughout their rosters. For example, with regard to the salary cap, in 2008, the average quarterback salary was US$3.47 million dollars (median salary of US$1.6 million) and the average salary for a running back was US$1.67 million (median salary of US$755 thousand). Thus, on average, a team could sign two running backs for the same price as one quarterback. Similarly, the mean salary of a defensive cornerback was US$2.06 million (median of US$933 thousand) or about 41% less than a quarterback.
We use the following example to illustrate how these constraints might affect roster decisions. Consider two teams, A and B, where Team A has two quarterbacks that are of higher quality than the starting quarterback for Team B. The value marginal product (VMP) of second-best quarterback on Team A, serving as a backup, will be lower than the VMP of that same player as a starter on Team B. The divergence is caused by the expected number of snaps played, which per above is significantly higher for starters. Therefore, while Team A would prefer to carry both quarterbacks on their roster, the willingness to pay of teams who would use that player in a starting role will be higher (whether in terms of salary or other scarce resources, like draft picks). The differential value that the same player might provide to two separate teams is discussed by Leeds and Kowalewksi (2001) in the context of understanding how free agency can affect player salaries. With each team facing a budget and roster size constraint, it is not an equilibrium outcome (subject to contract rigidities) to have players serving as nonstarters when inferior players are starting for other teams. 13
Finally, we conclude with a brief examination of whether NFL teams act as if they have knowledge of positional WAR in making their personnel decisions. We focus on how decision makers use one of the scarce resources at their disposal—first-round draft picks. Specifically, we correlate the share of first-round draft picks devoted to different positions, weighted by the inverse of the positional share in the average NFL starting lineup, with positional WAR (using the four-game measures from Table 3). We do this for the five NFL drafts spanning our data panel: 2008-2012. The weighting by positional share is necessary because, for example, all else equal a team will draft more wide receivers than quarterbacks owing to the fact that wide receivers constitute a larger share of the starting lineup. 14
With the important caveat that the correlation between first-round draft share and positional WAR is calculated based on a small number of items (our 10 primary positional categories) and thus is imprecise, we estimate that it is positive at .33. Further investigation reveals that the positive correlation is driven entirely by the overrepresentation of quarterbacks as first-round draft picks given their high-positional WAR. Our correlational analysis offers suggestive evidence that NFL decision makers understand some but not all aspects of positional value. In football, like with other sports, it will be of interest in future research to monitor the responsiveness of personnel decisions to the rapidly increasing base of empirical information about player productivity.
Conclusion
The contribution of the present study has been to develop an empirical approach for estimating positional WAR in the NFL. Our identification strategy leverages the use of replacement players when starters miss games due to injury and suspension. Using data from all 32 NFL teams over three seasons, we find that games lost by starting quarterbacks are by far the most important in terms of affecting win totals. We also show that games lost by starting wide receivers, fullbacks/tight ends, and exterior offensive lineman are important. Games lost by starters at all other positions do not affect wins, on average.
Our study represents the first rigorous attempt to quantify positional value in the NFL of which we are aware and moves us toward an improved understanding how players contribute to team wins in football. Future research can advance this line of study in several ways. One direction would be replicate our approach in other sports where individual performance measures can be constructed that are less dependent on the performance of other players. Aggregating the individual performance measures in these sports to the position level should produce positional WAR measures that align with those estimated using our approach based on injuries and suspensions, and discrepancies would be worthy of investigation. Our approach can also be replicated at the college level, where, like in the NFL, football is a major business and generates significant revenue (Isidore, 2013). 15 In addition, one could imagine using more and better data in the future to identify the value of particular player attributes by position (e.g., speed, height, weight, etc.) and the interaction of player attributes on the field. Finally, our findings can help to inform the development of the theoretical literature on constrained team building (as in Leeds & Kawolewksi, 2001), and relatedly, to help NFL executives, aiming to maximize wins subject to constraints such as the salary cap and access to draft picks.
Footnotes
Appendix A
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
