Abstract
The study tries to predict the outcome of Twenty20 cricket matches based on two prime skills of the game: batting and bowling. The two different measures, namely batting performance (BP) and adjusted combined bowling rate (ACBR), developed by Lemmer (2011, European Journal of Sport Science, vol. 13, pp. 200–206) and Lemmer (2012, European Journal of Sports Science, vol. 14, pp. S191–S196), respectively, are used to quantify the overall batting and bowling performances of cricket teams. The rationale of choosing BP is that it takes into consideration the match situation against which the runs are scored, and ACBR considers the strength of a batsman using wicket weights and adjusted number of runs conceded by the bowler along with the match situation. Both the measures, BP and ACBR, are combined to get the overall strength of a cricket team prior to the knock-out or play-off stage of any given tournament. Thereafter, the overall strength of a team is used to quantify the team’s current form. If the current form of Team A (say) is considerably higher than that of Team B (say), then it is predicted that Team A will win the match and vice-versa.
Introduction
Innovation is a vital aspect of cricket’s continued development to remain relevant and to attract new fans and supporters. One suitable example of this is the Twenty20 format of the game, probably the most significant development in the twenty-first century in terms of any team game. Batting and bowling are the prime skills of any format of cricket. As each ball bowled in cricket generates enormous statistics, both batting and bowling performances of individual players can be quantified and aggregated overall to quantify a team’s performance. Usually in cricket, measures like batting average and strike rate are used to understand the performance of batsmen, and bowling average, economy rate and bowling strike rate for that of bowlers. However, most of the existing measures available in the scorecard are not judicious in sensing the true ability of players. For example, batting average helps us to understand how many runs a batsman can score on average before losing his wicket. That means higher the batting average, better the ability of the batsman to score runs. However, from this measure, we cannot understand how effective a batsman is at scoring quickly. Similarly, when we consider the economy rate, we understand the rate at which a bowler concedes runs but not his wicket-taking capability. Due to this fact, several performance measures have been proposed by different authors like Barr and Kantor (2004), Lemmer (2002, 2004, 2011), Lewis (2005), Damodaran (2006), Gerber and Sharp (2006), Lemmer (2011), etc. to quantify the batting and bowling performances of cricketers through combining the traditional performance statistics. Authors like Suleman and Saeed (2008), Gerber and Sharp (2006), Beaudoin and Swartz (2003), Saikia et al. (2013), etc. have made remarkable contributions on the issue of developing combined performance measures in cricket, which could be used to quantify the performance of a player combining all the skills of the game like batting, bowling, fielding and wicketkeeping. An analysis of team performance was performed by Douglas and Tam (2010) for ICC World Twenty20 World Cup 2009. They found that for success in Twenty20 cricket, wicket-taking and dot-ball-bowling capabilities of the bowling team, and 50-plus (50+) partnerships and boundary-hitting capabilities of the batting team, are momentous factors. However, most of the measures proposed by the above-mentioned authors are player-specific rather than team-specific. Thus, it is essential to extend some of the above-mentioned performance measures to quantify overall teams’ performances.
Apart from performance measurement in cricket, there are some studies that specifically focus on predicting the outcome of a match. In literature, it has been observed that match outcome prediction in Test and One-day International (ODI) cricket is determined by several authors through modelling. Bailey and Clarke (2006) predict the match outcome in ODI cricket matches while the game is in progress. They use the multiple linear regression model, and predictor variables are numerically weighted according to statistical significance and used to predict the match outcome. A linear model was used by Clarke and Allsopp (2001) to fit least squares ratings to margins of victory in the 1999 ODI World Cup. The margin of victory was recorded in difference between the runs scored by the teams. When the team batting second won the match, the margin of victory was recorded as the number of wickets in hand, irrespective of the number of balls left. In that case, the Duckworth–Lewis (D/L) rain rule method was used to determine the winning margin for that team. With data related to the outcome of Test matches, Scarf and Shi (2005) developed simple decision support tools. On the basis of these tools, match outcome probabilities were estimated using the multinomial logistic regression model. Considering these probabilities as a function of target aimed for and run rate helps a team determine the position of the match at a potential declaration point. A logistic regression model was applied by Bandulasiri (2008) to predict the winner of the 2007 ODI World Cup. To determine the relative batting and bowling strengths of teams in ODI and Test cricket, Allsopp and Clarke (2004) used the multiple regression model. Thereafter, multinomial logistic regression was applied to explore how different factors, along with relative batting and bowling strength, affect the outcome of ODI and Test matches. Artificial Neural Network (ANN) was used by Choudhury et al. (2007) for predicting the outcome of cricket tournaments. The match outcomes of ODI matches played by different teams in the past 10 years were used to train the ANN with various input variables. Lemmer (2012) predicts the names of the best teams in International Cricket Council (ICC) World Cup 2011 before the knockout phase of the tournament using the logistic regression model. A dynamic logistic regression model for forecasting the outcomes of ODI cricket matches is used by Asif and McHale (2016) while the game is in progress. The model is dynamic in the sense that the parameters of the underlying logistic regression are allowed to evolve slickly as the match progresses. A team’s ‘good form’ is used as one of the predictors in the said dynamic logistic regression model. Following the idea of measuring a team’s ‘good form’ proposed by Asif and McHale (2016), this study tries to quantify the current form of cricket teams by developing an index called ‘Current Form (CF)’ through strength of the cricket teams.
Since the game of cricket has two distinct phases, namely the batting phase and the bowling phase, a match outcome can be constructed as the combined effect of batting and bowling abilities of the teams (Allsopp & Clarke, 2004). In this study, we are trying to predict the outcome of a Twenty20 cricket match based on the current form of the teams. As the outcome of the game depends on the teams’ strength, the combined measure of batting and bowling performance can be considered as the strength of a team. This team strength is used to quantify the current form of the cricket teams and accordingly predict the outcome of the match.
Methodology
Quantifying the Current Form of the Cricket Teams
Asif and McHale (2016) defined a team’s form as:
where w(t,θ) = (1 – θ)t–1 and 0 < θ < 1.
yt is a binary variable that takes the value 1 if a team won the match played t matches ago (i.e., last 5 matches) and 0 otherwise. The function w(t, θ) is a discounting factor so that the most recent match receives the highest weight. The value of θ is estimated by fitting the dynamic logistic regression model. However, how to estimate the value of θ while fitting the model is not explained explicitly in their work.
Thus, with a slight modification in Equation (1), the following measure is defined to quantify the CF of the ith cricket team as
where 0 < ki < 1,
and ki represents the strength of the ith cricket team, which is calculated using Equation (2).
It may be noted that to compute a team’s current form, a weighted average of match outcomes over the last five games of the team is used. The CF value of a team lies between 0 and 1. A cricket team shall have CF = 1 if the team has won the five most recent matches and CF = 0 if none of the last five games were won by the team. In addition, for a given value of k, two cricket teams with the same number of wins in the last five matches would have different values of CF depending on the order of their wins.
Measuring the Strength of Cricket Teams
The strength of the ith team obtained by combining all the j measures (i.e., j = 1 indicates the BP of the team and j = 2 indicates the ACBR of the team) and is denoted by ki, given by:
Further, Yij = the normalized score of the BP and ACBR for the ith team and wj = weights associated with the normalized score of the BP and ACBR.
Since the measure BP is positively associated with the batting performance of the teams, it is normalized as:
However, the bowling performance measure, ACBR, is a reverse measure, which means a low value of ACBR is considered as good bowling performance of the team. Thus, it is normalized as:
This process of normalization helps limit the BP and ACBR values within an interval of 0 and 1 so that they are always non-negative. Furthermore, the conception of Iyenger and Sudarshan (1982) has been applied here (the weights vary inversely with the variation in the respective factors) to determine the weights associated with the BP and ACBR of the cricket teams. Therefore, the weights wj can be expressed as:
where w1 + w2 = 1 and C is a normalizing constant which follows:
The choice of weights in this manner implies that the large variation in the BPs of the teams would not unduly dominate the contribution of ACBR or vice versa.
Measuring the Batting Performance of the Teams
Let Rij be the runs scored by the ith player in the jth match and Bij be the balls faced by the ith player in the jth match. Then, the batting performance of a player in that match is defined by:
where Rij are the runs scored by the ith player in the jth match.
SR
ij
, the strike rate of the ith player in the jth match =
This BP measure takes into consideration the match situation in which the runs are scored. This is done by introducing MSRj in Equation (8), which is the strike rate of the entire match. To compute the batting performance of a team, the average BP values of the players across all the matches are considered.
Measuring the Bowling Performance (ACBR) of the Teams
Lemmer (2002) proposed a bowling performance measure called the combined bowling rate (CBR), which is the harmonic mean of three traditional bowling statistics, namely bowling average, economy rate and bowling strike rate. If R is the total number of runs conceded by a bowler, W is the total number of wickets taken by a bowler and B is the total number of balls bowled by a bowler in a particular series of matches, then the traditional bowling statistics can now be defined as:
To bring parity in the numerator of the above factors, a prerequisite of the harmonic mean, the bowling strike rate, was adjusted as:
Thus, the CBR is defined by
However, for the case of a small number of matches, the CBR had been adjusted to take the weights of the wickets taken by the bowler (Lemmer, 2005). Then, it was defined as CBR#. Later, in order to take the match situation into account, Lemmer (2012) improved the CBR# to an adjusted measure called the ACBR. The ACBR is more appropriate for quantifying the bowling performance for a small number of matches along with the match situation. The ACBR for the ith bowler is given by:
where Bi = number of balls bowled by the ith bowler;
Wi* = sum of weights of the wickets taken by the ith bowler; and
where RPB
ij
=
RPBM
j
=
Two issues regarding this measure are noteworthy. The factor (RPB ij / RPBM j ) considers the match situation in which the ith bowler delivered, and the factor Wi* refuses to give equal importance to all the wickets taken by the bowler but weighs them differently based on their batting position. The detailed discussion and the different values of Wi* are available in Lemmer (2005). It may be noted that the ACBR has a negative dimension, that is, lower the value, the better is the bowler. To compute the bowling performance of a team, the average ACBR values of the players across all the matches are considered.
Results and Discussion
The relevant data to validate the above-discussed methodology is obtained from Indian Premier League (IPL) 2018. In IPL 2018, eight teams participated, namely Chennai Super Kings (CSK), Sunrisers Hyderabad (SRH), Mumbai Indians (MI), Delhi Daredevils (DD), Royal Challengers Bangalore (RCB), Kolkata Knight Riders (KKR), Kings XI Punjab (KXIP) and Rajasthan Royals (RR). Each team had to play 14 matches before the playoffs stage of the tournament. Now, based on the information of these 14 matches for each team, the strength of the teams using Equation (2), through combining BP and ACBR, is measured. The weights (wj) associated with BP and ACBR are calculated using Equation (6), which are 0.51406 and 0.48594, respectively. The results can be seen in Table 1.
In the playoffs stage, only four matches had been played among the top four teams according to the point’s table, namely Qualifier 1, Eliminator 1, Qualifier 2 and Final. The top four teams according to the points table in IPL 2018 are SRH, CSK, KKR and RR. In IPL playoffs matches, the top four teams of the group stage are given more chances to qualify for the final match. Thus, the CF is calculated for these four teams, and the outcome of the matches played by these four teams in the playoffs stage of IPL 2018 is predicted based on the values of CF.
For example, let us consider the final match of IPL 2018, which was played between CSK and SRH at Wankhede Stadium, Mumbai on 27 May 2018. Before this final match, CSK had won three matches and SRH had won only one match out of its last five matches. With this information, Table 2 shows how to calculate the CF for the teams CSK and SRH. The values of ki for CSK and SRH are 0.79385 and 0.72197, respectively (see Table 1).
According to the measures of the current form, the team having the higher value of CF will win the match. From Table 2, it can be observed that the CF value for team CSK (0.96481) is greater than that of team SRH (0.72317); thus, team CSK won the final match of IPL 2018, despite SRH being at the top of the points table throughout the tournament. It may be noted that in Table 2, for the teams CSK and SRH, t = 1 represents the most recent or last match played by the teams before the final match, t = 2 represents the second most recent match played by the teams before the final match and so on. The CF values of the teams for all the four matches played in the playoffs stage of the IPL 2018 tournament is shown in Table 3.
In Table 3, except for the predicted outcome of the Qualifier 2 match, the other predicted match outcomes based on CF values are correctly predicted. It means that this measure when applied for actual validity gives 75 per cent correct results. This result is based on only four match outcomes. However, it has been observed that the measure of CF can be used effectively to predict the match outcome or winner of any other tournament for limited overs cricket formats like ODI and Twenty20.
Strength of the Cricket Teams
Current Form (CF) of CSK and SRH Before the Final Match of IPL 2018
CF Values of the Playoffs Stage Matches in IPL 2018
Conclusion
The individual batting and bowling performances of the players in the respective teams based on 14 league matches in IPL 2018 are measured and then combined to compute the strength of the cricket teams. Thereafter, using the strength of the cricket teams, a measure called CF is defined to quantify the current form of the teams. Then, the measure CF is used to predict the outcomes of the playoffs matches in IPL 2018 and then compared with the actual outcomes of the matches. It has been observed that CF is useful to correctly predict the outcome of three out of four matches. Though the performance of the players at the time of the actual match may control the outcome of the match, the CF method is expected to work better if the performances of the players do not vary much from their performances in earlier matches. Overall, the CF measure has the ability to predict the outcome of a cricket match on the basis of the current form of the teams. Most likely, it will be useful for spectators in making bets to predict the winner of a match before the match is played. Further scope of this study is in finding the optimum difference between teams’ current forms to predict the winner precisely.
Footnotes
Declaration of Conflicting Interests
The author declared no potential conflicts of interest with respect to the research, authorship and/or publication of this article.
Funding
The author received no financial support for the research, authorship and/or publication of this article.
