Showing posts with label Afghanistan. Show all posts
Showing posts with label Afghanistan. Show all posts

Monday, 17 June 2019

World Cup simulation update

The group stage of the World Cup is now roughly half way through, and there are 4 clear favourites to be the semi-finalists.

Afghanistan is the first team to be eliminated (they may have a mathematical possibility, but they don't have a statistical one). At this point, Sri Lanka are not far behind.

The rankings of the teams have remained fairly consistent, suggesting that the extra weighting for world cup matches is about right.
The fact that almost all the teams seem to have gone up is due to them all being relative to Afghanistan. Afghanistan do not seem to be quite as good as they were seeming to be and so they have dropped, but as they are set to 0, it's pushed everyone else up slightly.

The semi-final probability is the most interesting. 

I personally feel that this is underestimating the chances of South Africa, but we will see as the tournament progresses.

The key point on this graph is match 5, where Bangladesh overcame South Africa. If South Africa had won that match, they would be on about 40% and New Zealand and Australia would both be a lot lower.

The simulation also puts out the points for 4th, 5th and the difference between them. This suggests at the moment that there's only a fairly low chance that net run rate will come into play. However, one more rained out match, or a Bangladesh upset of Australia, and this could change dramatically. This makes the expected lines to be 9 points for 5th place, and 11 points for 4th place.

So far of the teams that I've had as favourite to win, 14 out of the 17 have won. Given the probabilities that the models assigned them, that's slightly higher than I would have expected - I would have expected there to have been 4 upsets rather than 3, but it's still telling me that my model is working quite well. That may be due to teams not always playing their best combinations in every match between the world cup, adding extra uncertainty to the results than exist inside a world cup.

It will be interesting to see if it continues to have the same success rate after the cup is finished.

Finally, applying the same system to find the probable winner gets the following results:
England are still favourites, but India are not far behind them.

Thursday, 30 May 2019

A simulation to see who will win the World Cup


One of the main purposes of statistics is to help inform decisions. Cricket statistics are often used when deciding on selection of players, or (more often) arguments about who is the best at a particular aspect. They can help decide which strategies are best, what an equivalent score is in a reduced match (with a particular case of Duckworth Lewis Stern) or which teams should automatically qualify for the World Cup (David Kendix). They are often also used by bookmakers (both the reputable, legal variety and the more dubious underworld version) to set odds about who is going to win.

I decided to attempt to build a model to calculate the probability of each team winning, based on their previous form. This was going to allow me (hopefully) to predict the probabilities of each outcome of the world cup, by using a simulation. It didn’t prove to be as easy as I had hoped.

My first thought was to look at each team’s net run rate in each match, adjust for home advantage, and then average it out. That seemed sensible, and the first attempt at doing that looked like it would be perfect. Most teams (all except Zimbabwe) had roughly symmetrical net run rates, and they fitted a normal curve really well. The only problem was that Afghanistan was miles ahead of everyone else. The fact that they had mostly played lower quality opponents in the past 4 years meant that they had recorded a lot more convincing wins than anyone else.

This was clearly a problem. India and England both had negative net run rates, while Afghanistan, Bangladesh and West Indies were all expected to win most of their matches.

I then tried a different approach, based off David Kendix’s approach of using each result to adjust a ranking. But rather than having a ranking that was based off wins, I based it off net run rate. So if a team had an expected net run rate of 0.5, and another had an expected net run rate of 0.6, the first team would have an expected net run rate of -0.1 for their match. If they did better than that, they went up, and if they did worse than that, they went down.

However, I found that some results ended up having too much bearing. If I made it sensitive to a change in the results, it ended up changing way too much based off one big loss/win. England dropped almost a whole net run per over based on the series in the West Indies. So this was clearly not a good option.

Next, I decided to try using logistic regression, and seeing how that turned out. Logistic regression is a way of determining probabilities of events happening if there are only two outcomes. To do that, I removed every tie or match with no result, and set to work building the models.

My initial results were exciting. By just using the team, opposition and home/away status, I was able to predict the results of the previous three world cups quite accurately using the data from the preceding 4 years. (I could not go back further than that, as they included teams making their ODI debut, and there was accordingly no data to use to build the model.

The results were really pleasing. I graphed them here, grouped to the nearest 0.2 (ie the point at 0.6 represents all matches that the model gave between 0.5 and 0.7 as the chance for a team to win), compared to the actual result for that match. It seems that they slightly overstate the chance of an upset (possibly due to upsets being more common outside world cups, where players tend to be rested against smaller nations), but overall they were fairly reliable, and (most importantly) the team that the model predicted would win, generally won.

I could then use this to give a ranking of each team that directly related to their likelihood of winning against each other. The model gave everything in relation to Afghanistan, with the being 0, and any number higher than 0 being how much more likely a team was to win against the same opponent as Afghanistan. (Afghanistan was the reference simply because they were first in the alphabet).







This turns out to be fairly close to the ICC rankings. So that was encouraging.

I tried adding a number of things to the model (ground types, continents, interactions, weighting the more recent matches more highly) but the added complexity did not result in better predictions when I tested them, so I stuck to a fairly simple model, only really controlling for home advantage.
Next I applied the probabilities to every match and found the probabilities of each team making the semi-finals.


The next step was to then extend the simulation past the group stage, and find the winner.

After running through the simulation a few more times, I came out with this:


A couple of points to remember here: every simulation is an estimate. The model is almost certainly going to estimate the probabilities incorrectly, but it will get them close, and they will be close enough to give a good estimate of the actual final probabilities. It is also likely to overstate Bangladesh’s ability due to their incredible home record; overstate Pakistan’s ability as a lot of neutral matches for them they have had a degree of home advantage in UAE; and understate West Indies, due to them having not played their best players in a lot of matches in the past 4 years. But these are not likely to make a massive difference to the semi-finalist predictions.



Given this, I’d suggest that if you are wanting to bet on the winner of the world cup, these are the odds that I would consider fair for each team:


I will try to update these probabilities periodically throughout the world cup, and report on their accuracy.

Wednesday, 13 June 2018

How well will Afghanistan's spinners go?

It's a fascinating question - at least two out of 4 spinners who have been very successful in limited over cricket are likely to make their test debut tomorrow. They will be playing in India, a country known for spinning tracks.

The Afghan captain, Ashgar Stanikzai is certainly bullish about their ability: "In my opinion, we have good spinners, better spinners than India."


That's a big call. India have some very high quality spinners in test cricket. But does he have a point?

There is no doubt that Afghanistan's spinners have been good in limited overs cricket, but does that mean anything?

Some spinners have excelled in test cricket and limited overs. Warne, Murali, Shakib, Swann all have excellent stats in all forms of the game.  But others haven't seen such correlation. Amit Mishra, Ish Sodhi, Abdur Razzak and Sunil Narine all have very strong limited overs stats, but have not converted that to test matches.

The comparison of Sodhi and Rashid is an interesting one. When they've played against similar opponents, they have had very similar stats. See for example this graph of combined IPL and Big Bash statistics. (This is the 8 players who played a reasonable number of matches in both tournaments over the past 2 years)

They have clearly been two of the stand out bowlers in the two major domestic T20 tournaments, and yet, Sodhi averages over 40 in test cricket with the ball.

Part of that will be that Sodhi has to play half his cricket in New Zealand, but part of it is also that it is not always possible to predict absolutely test success based on limited over success. They are sometimes linked, but not always.

I wondered if there was a general relationship in the numbers. Were Narine and Sodhi the outliers, or were Warne and Muralitheran the odd ones.

So I tried to construct a model, to see what happened. It turns out, to my surprise, that it is possible to predict, vaguely, the average and strike rate of a bowler in test cricket based on their ODI and T20 performances. I looked at the 37 players who had played as spinners in all 3 formats, had played at least 40 combined limited overs matches and had taken at least one test wicket.

The model is not very reliable, but it was better at predicting the bowling statistics of the players than just taking the average for the group. So it did provide some interesting numbers.

The equations that it came out with were as follows:
Only 4 of the 6 Afghani spinners had played enough limited overs cricket to have meaningful numbers, and here are their predicted results:


That would relate to the following results based on overs bowled:


Those wouldn't be particularly bad returns for a first ever test against India, but they also aren't really the results that Stanikzai is hoping for.

But the question comes up, how reliable is this anyway?

There is a degree of randomness in cricket results that means that any predictions are always quite unreliable. Looking at the graphs of the predicted vs actual for the career stats, it shows that there is quite a bit of variation about the trend.


I've added in the red lines manually, They show where roughly 95% of the data fits. Within 17 of the average and within 20 of the strike rate. Those are huge variations, which shows that it's very difficult to predict test performances based on limited overs performances.

I have added in the best realistic and worst realistic expected figures, based on these confidence bands:



It will be interesting to see which one of these the Afghani spinners actually get closest to.