How Should We Value the Masters and Premier Titles in the Bubble?

Tennis is back, but plenty of top players are still at home–or crashing out in the early rounds of their first tournament in months. While the ATP “Cincinnati” Masters event delivered the expected winner in Novak Djokovic, the Serb never had to face a top-ten opponent. The same was true of Victoria Azarenka, who won the WTA Premier tournament with the benefit of Naomi Osaka’s withdrawal in the final round, and without playing a top-tenner on her way there.

The tennis world’s “asterisk” talk has mostly focused on the US Open, since most people care about slams and don’t care about anything else. But judging from these easy paths to the two Cincinnati titles, should we be talking asterisk about the event just passed?

Novak’s 35th, but not (quite) his easiest

Last week, I explained why I thought the asterisk talk was premature, if not wrong. The field doesn’t matter, because the player who wins the title faces only a handful of players. The presence of, say, Rafael Nadal doesn’t have much to do with the difficulty of winning the title unless the eventual winner has to go through Rafa. If the champion’s opponents are very good, the path to the title is hard; if they are relatively weak, the path to the title is easy. Keep in mind I’m using the terms “good” and “weak” in theoretical terms. On paper, Djokovic was fortunate that his semi-final and final opponents were ranked 12th and 30th, respectively, and his title path was “easy.” As it happened, he was forced to work hard for both wins.

We now know that the title paths of the Cincinnati champions were relatively easy. But just how weak were they?

I calculate the difficulty of a path-to-the-title by determining the probability that the average Masters champion on that surface would beat the opponents that the champion faced. By using the “average Masters champion,” we are taking the skill level of the actual champ out of the equation, and looking only at the quality of his opposition. The resulting numbers vary wildly, from 2.5%–the odds that a typical Masters champion would have beaten the players that Jo Wilfried Tsonga defeated to win the 2014 Canada Masters–to 61.2%–the chances that an average titlist would have beaten the players that confronted Nikolay Davydenko at the 2006 Paris Masters.

Novak’s number this week was 40.5%. In other words, an average hard-court Masters champion would have a four-in-ten shot at beating the five guys that fate threw in Djokovic’s path. That’s the 11th easiest Masters title since 1990:

Title Odds  Tournament       Winner             
61.2%       2006 Paris       Nikolay Davydenko  
50.5%       2012 Paris       David Ferrer       
49.8%       2000 Paris       Marat Safin        
48.3%       2004 Paris       Marat Safin        
47.0%       1999 Paris       Andre Agassi       
44.5%       2013 Shanghai    Novak Djokovic     
43.3%       2002 Madrid      Andre Agassi       
42.9%       2005 Paris       Tomas Berdych      
41.4%       2009 Canada      Andy Murray        
41.3%       2017 Paris       Jack Sock          
40.5%       2020 Cincinnati  Novak Djokovic     
39.6%       2011 Shanghai    Andy Murray        
39.1%       2019 Canada      Rafael Nadal       
37.9%       2008 Rome        Novak Djokovic     
36.2%       2007 Cincinnati  Roger Federer

Unless we’re prepared to put a permanent asterisk next to the Paris Masters, we should hold off on cheapening this year’s Cincinnati title. Surprisingly, Djokovic’s path was even easier at the 2013 Shanghai Masters. He had to face two top-ten opponents in the final rounds (Tsonga and Juan Martin del Potro), but Elo didn’t think that highly of them at the time.

Azarenka: asterisk squared

Evaluating the WTA title is trickier. Part of the problem is the small number of “Premier Mandatory” events, and the fact that two of them (Indian Wells and Miami) have substantially larger draws, and are thus that much harder to win. The even bigger issue is how to think about Azarenka’s final-round walkover.

Let’s start with the numbers. If we consider the five opponents that Vika defeated on court and calculate the odds that an average WTA Premier (not just Premier Mandatory) champion would beat them, her path-to-the-title number is 20.7%. If we add Osaka to the mix, on the theory that Azarenka should get credit for beating her, the resulting number is 7.4%.

Compared to the ATP numbers above, those sound pretty good. But the devil lies in the tournament-category details–the average WTA Premier event is much weaker than a marquee (dare I say “premier”?) tour stop like Cincinnati. Here’s how the Cinci title-paths stack up for the last dozen years:

20.7%       2020  Victoria Azarenka  (W/O Osaka)  
7.4%        2020  Victoria Azarenka  (d. Osaka)   
7.3%        2016  Karolina Pliskova             
5.5%        2010  Kim Clijsters                 
5.5%        2012  Li Na                         
5.3%        2015  Serena Williams               
4.5%        2011  Maria Sharapova               
4.3%        2014  Serena Williams               
4.2%        2017  Garbine Muguruza              
3.9%        2019  Madison Keys                  
2.9%        2013  Victoria Azarenka             
2.0%        2009  Jelena Jankovic               
1.3%        2018  Kiki Bertens

20.7% is respectable for a run-of-the-mill Premier–in fact, Vika’s 2016 Brisbane title was almost exactly the same, at 20.8%. But Cincinnati reliably offers tougher competition. Even if we factor in the difficulty of beating Osaka, Azarenka’s path was (barely) the easiest at the event since the Premier-level designation came into being.

Yay, nay, meh

I’ll reiterate a main point from my last article about the US Open asterisk debate: There’s no simple yes or no answer when it comes to whether a title should “count.” (That’s assuming that you even think there are circumstances under which a title should be formally discounted.) Long before the COVID-19 pandemic messed with everything, there were titles–even at the grand slam level–that were a lot easier to win than others.

Djokovic’s championship falls squarely within the usual continuum, even if it will go down as one of his least challenging. Azarenka’s is tougher to define, but more because of Osaka’s withdrawal than because of the weakness of the field. The level of competition, despite missing many top players, was plenty good enough to offer Azarenka a path to the title that was comparable at least one recent Cinci championship, and plenty of other top-tier events.

With that in mind, I’ll leave you with a couple of predictions. First: the US Open champions will face relatively easy paths to their titles, but like Djokovic’s, they will fall on the established continuum. And second: by the end of the fortnight, you’ll hope to never hear the word “asterisk” again.

How Sports are (Analytically) Different in the Bubble

Most of the world’s major sports have resumed, or will pick up again soon, in some form or other. But a lot is different, with most leagues forming one or more bubbles, often excluding fans, limiting travel, and tweaking things like officiating rules to better maintain social distance.

Many of these changes have second-order effects. For instance, the “Cincinnati” tennis event requires that players fetch their own towels–which probably slows down play–but has no fans–which could accelerate it. We’ll soon have enough data to draw some preliminary conclusions about the overall effect of the new rules on pace of play.

Some of the issues that arise when a league moves into a bubble apply across sports, like home-court advantage. With that in mind, I’m gathering evidence of how sports are playing differently in our time of social distance. I’ll try to keep this post updated as we learn more. The comments are open, so you can contribute any demonstrated effects that I haven’t listed here. (Or similar effects in other sports.) You can also tweet at me.

Baseball

So far, home-field advantage is almost non-existent. Historically, home teams win about 54% of games.

Basketball

NBA offenses can’t stop scoring. Refs are calling more fouls, and fewer off-court distractions get in the way of making shots.

The WNBA is showing the effects of a league full of fresh legs, and has displayed a record-setting pace of play. And despite playing on the same court every night, there is a marked home-court advantage.

Hockey

Fighting is up! Lucky NHLers–most of us don’t go to work where it’s culturally acceptable to hit people.

Soccer

Home-field advantage is reduced, but it still exists, even behind closed doors. A recent paper (summary / PDF) notes that refs have been more lenient than usual toward away teams. That tallies with long-held conventional wisdom that home-advantage stems from officiating bias, which is driven by noisy, partisan crowds.

Speaking of officiating, refs were more likely to grant penalty kicks, but despite the quieter environment, penalties aren’t converted any more often.

For more detail on home-field advantage in various leagues since the restart, here is a valuable Twitter thread from @recspecs730.

Tennis

I’m keeping tabs on whether match results are less predictable than usual. (They are, but we haven’t really seen enough to be sure.) Other than that, it’s still speculation. We’ll know more after “Cincinnati,” and much more after the US Open.

US Open Asterisk Talk is Premature. It Might be Flat-Out Wrong.

Many high-profile players will be missing from the 2020 US Open. Rafael Nadal opted out of the abbreviated North American swing, and Roger Federer will miss the rest of the season due to injury. More than half of the WTA top ten is skipping Flushing Meadows as well. The thinned-out fields increase the odds that a few remaining favorites, such as Novak Djokovic and Serena Williams, add another major trophy to their collection.

As a result, pundits and fans are discussing whether the 2020 US Open deserves an “asterisk.” The idea is that, because of the depleted fields, this slam is worth less than others, so much so that the history books* should note the relative meaninglessness of this year’s titles.

* Nobody buys history books anymore, so we’re really talking** about a page on the US Open website, and a never-ending edit war on Wikipedia.

** Yes, I see the irony.

From what I’ve seen, people are thinking about this the wrong way. Yes, a weak field makes it easier–in theory–to win the tournament. It’s certainly true that the 2020 champions won’t have to go through Nadal or Ashleigh Barty to get their hardware. But the field isn’t what matters.

The field isn’t what matters

I repeated that on purpose, because it’s that important. The winner of a grand slam must get through seven matches. The difficulty of securing the title depends almost entirely on his or her opponents in those seven matches. Each main draw consists of 128 players, but 120 of them are mostly irrelevant.

I say “mostly” because I can foresee some objections. Sometimes a player can compete so hard in a loss that they weaken their opponent for the next round. Take the 2009 Madrid Masters, in which Nadal needed four hours to defeat Djokovic in the semi-final, then lost to Federer in the final. We could say that Djokovic’s presence was relevant, even though Federer won the title without playing him. That sort of thing happens, though probably not as much as you think. Even when it does, it needn’t be a top tier player who wears out their opponent in an early round.

Another objection is that a depleted field affects seedings. For instance, Serena’s current WTA ranking is 9th, an unenviable position going into most slams. The 9th seed lines up for a fourth-round match with a top-eight player, meaning that she could face four top-eight players en route to the title. But with all the absences, Williams will instead be seeded third, behind only Karolina Pliskova and Sofia Kenin.

I’m not dismissing these concerns out of hand. They do matter a bit. But they only matter insofar as they affect the way the tournament plays out. The difference between the difficulties facing the 3rd and 9th seeds could be enormous … or it could be nothing, especially if the draw is riddled with early upsets.

Difficulty is a continuum

Even if you grant some credence to the objections above (or others that I haven’t mentioned), I hope you’ll agree that the most meaningful obstacles standing between a player and a grand slam title are the seven opponents he or she will need to overcome.

If those seven opponents are, on average, very strong, we would say that the player faced a particularly tough path to a slam title. Take Stan Wawrinka’s 2014 Australian Open title: he beat both Djokovic and Nadal at a time when those two were dominating the game. If the collective skill level of the seven opponents doesn’t amount to much–at least by grand slam standards–we’d say it was an easy path. For example, Federer clinched the 2006 Australian Open despite facing only a single player ranked in the top 20, and none in the top four.

We can quantify path difficulty in a variety of ways. One approach that will be useful here is to calculate the odds that an average slam champion would beat those seven opponents. The difference between easy and hard championships is enormous. The typical major titlist (that is, someone with an Elo rating around 2100) would have had a 3.3% chance of beating the seven men that Wawrinka drew in Melbourne the year that he won. Only two slam paths have ever been tougher: Mats Wilander’s routes to the 1982 and 1985 French Open titles. By contrast, the average slam champion would have had a 51% chance of going 7-0 when faced by Federer’s 2006 Australian Open draw.

The extreme “easy” draw is fifteen times easier than the extreme “hard” draw. Fifteen times! You can find plenty of champions for any approximate level of difficulty in between those extremes. The typical slam champ would’ve had a 10% chance of doing what Djokovic did in progressing through seven rounds at the 2011 US Open. Same in New York in 2012. Andy Murray’s 2016 Wimbledon path would have given the average champion a 20% chance. The 2018 Roland Garros draw was manageable for Rafael Nadal, and a typical major titlist would have had a 30% chance of securing those seven match wins.

None of this is to say that any of those players did or didn’t “deserve” their titles. Federer didn’t choose his 2006 Melbourne opponents any more than Wawrinka selected his foes eight years later. The trophy is the same, and in many important ways, their achievements are the same–both of the Swiss stars swept away all of their opponents, who in turn were the best performers (at least during those fortnights) of the players who showed up.

Asterisks for everybody

Here’s another thing 2006 Roger and 2014 Stan had in common: Almost all of the best players in the world participated in the tournaments that they ultimately won. (I say “almost” because defending champion Marat Safin was injured and missed the 2006 Aussie Open.) The “field” was effectively the same, but to win the titles, one player cruised through a two-week cakewalk and the other needed to put together one of the most impressive final weeks of the modern era.

Tennis fans have collectively decided that each major title counts as “one.” It doesn’t have to be that way: We could give more “slam points” for achievements like Wawrinka’s and grant fewer for the easy ones. Most people don’t like this idea, and I admit that it sounds a bit weird. I’m not advocating it for general use, though it is an interesting concept that I’ve pursued in a number of earlier articles, showing that Djokovic’s majors are–on average–more impressive than Nadal’s, which in turn have been tougher than Federer’s. Weighting majors by difficulty results in some changes in the order of the all-time grand slam list, ensuring that fans of all players hate me because I wrote some code and played with some spreadsheets.*

* With, I admit, malice aforethought.

Adjusting slam counts for difficulty is, in a sense, asterisking every slam title. The tricky draws get an acknowledgement of their difficult, and the ones that opened up get tweaked to account for their ease. It’s a continuum, not a simple up-and-down decision between normal slams and abnormal slams.

The 2020 US Open champions will probably have title paths that sit in the easier half of that continuum. But even that modest claim is far from guaranteed.

Let’s say Venus Williams recaptures her vintage form and wins the title, beating 3rd seed Serena in the quarter-finals, 2nd seed Kenin in the semis, and top seed Pliskova in the title match. (It doesn’t matter if the surprise winner is Venus–it could be any lower-ranked player, though Venus seems more plausible than most.) An average slam champion would beat those three players in succession about 37% of the time. 37% is already lower odds than about 20% of women’s slam draws in the last 45 years. (Kenin’s Australian Open title rated 39%.)

37% for Venus’s hypothetical title isn’t even the whole story–four more rounds of journeywomen would knock the number down to around 26%–harder than one-third of women’s slam draws. Add in another tricky opponent or two–maybe Cori Gauff, or Petra Kvitova in the fourth round–and suddenly the path to the 2020 US Open women’s championship is just as hard as the typical slam.

It’s even easier to illustrate how the 2020 US Open men’s title could be as difficult as many other slams. By the numbers, simply upsetting Djokovic (simply! ha!) is more difficult than it was to defeat all seven of Federer’s opponents at the 2006 Australian Open. That’s right: Six withdrawals and one win over Novak wouldn’t be the easiest slam victory in the last 15 years. Tack on six actual wins, including a few against strong opponents, and the result is a seven-match path that stands up against the typical non-pandemic slam.

Ironically, the player who could win the title with the weakest possible draw is Djokovic. It would be odd to claim that any of Novak’s accomplishments should be asterisked, but it does make things much simpler when he doesn’t have to beat himself.

Masked competitiveness

Once again, the field doesn’t really matter. When we focus on the players who are in New York instead of the few dozen who aren’t, we see that the ingredients are in place for a couple of respectable path to US Open titles. Wilander’s and Wawrinka’s marks are probably safe, but it’s more than possible that the winners will have faced competition equivalent to that of the average slam champ.

At the very least, we don’t know any better until the tail end of the second week. Until then, asterisk talk is premature. After that, it will probably be moot.

The Post-Covid WTA is Drifting Back to Normal

In the two latest WTA events, we saw a mix of the expected and the unusual. Simona Halep, the heavy favorite in Prague, wound up with the title despite a couple of demanding three-setters in her first two rounds. The week’s other tournament, in Lexington, failed to follow the script. Serena Williams and Aryna Sabalenka, the big hitters at the top and bottom of the bracket, combined for three wins, with four unseeded players making up the semi-final field.

Last week I pointed out that Palermo–the tour’s initial comeback event–was so unpredictable that you would’ve been better off to treat each match as a coin flip than to use pre-layoff player strength ratings (such as Elo) to forecast outcomes. Such an upset-ridden event isn’t unheard of, even in pandemic-free times, but it is suggestive that the WTA rank-and-file haven’t quite returned to their usual form.

Prague and Lexington give us three times as much data to work with. Plus, we might theorize that Prague would be a little more predictable because so many players in that field also took part in the Palermo event, meaning that they have a little more recent match experience. While our sample of 93 main draw matches is still flimsy, it brings us a little closer to understanding how well traditional forecasts will handle this unusual time.

A thorny Brier patch

The metric I’m using to quantify predictability–or to put it another way, the validity of pre-layoff player ratings–is Brier Score, which takes into account both raw accuracy (did the forecast pick the right player to win?) and confidence level (was the forecast too strong, too weak, or just right?). Tour-level Brier Scores are usually in the range of 0.21, while a score of 0.25 means the predictions were no better than coin flips. A lower score represents more accurate predictions.

Here are the Brier Scores for Palermo, Lexington, and Prague, along with the average of the three, and the average of all WTA International events (on all surfaces) since 2017. (The scores are based on forecasts generated from my Elo ratings.) We might expect the first round to be different, since players are particularly rusty at that stage, so I’ve also broken out first round (“R32 Brier”) matches for each of the tournaments and averages in the table.

Tournament    Brier  R32 Brier  
Palermo       0.268      0.295  
Lexington     0.226      0.170  
Prague        0.212      0.247  
Comeback Avg  0.235      0.237  
Intl Avg      0.217      0.213

As we last week, the Palermo results truly defied expectations. More than half of the matches were upsets (according to my Elo ratings), with a particularly unpredictable first round.

That didn’t last. The Prague first round rated 0.247–just barely better than coin flips–but the messiness didn’t last beyond the first couple of days. The event’s overall Brier Score was 0.212, slightly better than the average WTA International. In other words, this group of 32 women, only recently returned from a months-long break, delivered results that were roughly as predictable as we would expect in the middle of a normal season.

The Lexington numbers are a bit more difficult to make sense of, but like Prague’s, they point to a post-coronavirus world that isn’t all that weird. The opening round closely followed the script, with a Brier Score of 0.170. Of the last 115 WTA International events, only 22 were more predictable. The forecast accuracy didn’t last, in large part because of Serena’s loss at the hands of Shelby Rogers. The rating for the entire tournament was 0.226, less predictable than usual, but much better than random guessing and closer to tour average than to the assumption-questioning Palermo numbers.

Revised estimates

We’re still early in the process of evaluating what to expect from players after the COVID-19 layoff. As more tournaments take place, we can identify whether players become more predictable with more matches under their belts. (Perhaps the Prague participants who skipped Palermo were more difficult to forecast, although Halep is an obvious counterexample.)

At this point, anything is possible. It could be that we will steadily drift back to business is usual. On the other hand, the new social-distancing-oriented rules–with few or no fans on site, nightlife limited to Netflix, players fetching their own towels, and new variations of on-court coaching–might work to the advantage of some women and the disadvantage of others. If that’s the case, Elo ratings will go through a novel period of adjustment as they shift to reflect which players thrive on the post-corona tour.

It’s too early to do much more than speculate about something as significant as that. But in the last week, we’ve seen forecasts go from wildly wrong (in Palermo) to not half bad (in Lexington and Prague). We’ve gained some confidence that for all the things that have obviously changed since March, our approach to player ratings may be one thing that largely remains the same.

Are Tournament Draws Giving Us Suspiciously Many Venus-Serena Clashes?

This week in Lexington, top seed Serena Williams faces her sister, Venus Williams, in the second round. They are both among the all-time greats, and they have played each other nine times in grand slam finals, so it’s always jarring to see them turn up in the same section of a draw and play on a Thursday.

Lately, their encounters seem to always happen long before the business end of a tournament. Their three matches between the 2017 Australian Open final and this week in Lexington all happened in the round of 32, including a planned 2019 Rome meeting from which Serena withdrew. Venus is usually unseeded, no longer the world-beater she once was, so it is at least possible that the Williams sisters would be bracket neighbors in any given week.

But should it happen quite so often? It is an understatement to say that Serena and Venus were not universally embraced upon arrival in the tennis world. If you’re conspiracy minded, every tournament draw is an opportunity to commit dastardly deeds. Perhaps early in the Williams era, it was the work of racist or otherwise misguided tournament officials who wanted to avoid all-Williams finals. Or nowadays, event honchos recognize that Venus is unlikely to reach the final, so they tinker with the bracket to make a headline-grabbing Williams-versus-Williams clash more likely.

I’m sure that most draws are conducted on the up-and-up, but the process is sufficiently opaque that it’s easy to get suspicious. It’s also easy to make mistaken generalizations from insufficient data. Let’s see what the numbers can tell us.

150 tournaments!

Lexington is the 150th tour event with both Serena and Venus in the field.*

* I think. My WTA data isn’t perfect for the early years of their careers, and there was an uncomfortable amount of manual tabulation involved in this post. Their TennisAbstract player pages are missing the 1999 Grand Slam Cup, but I’ve included it in all the numbers here. For the purposes of doing analytics, it doesn’t matter much if the total is 148 or 151, but if you’re printing a banner or making a cake, you should double-check.

Thursday’s match in Lexington will be their 31st, plus one withdrawal apiece. In 13 of the 150 events, the Williams sisters were either the top two seeds or the 3rd and 4th seeds, meaning that draw shenanigans were out of the question–they could not face each other until the final. 4 of those 13 times, that’s exactly what they did.

What are the odds?*

* Of me being able to use this sub-heading in any given blog post?

I went through the remaining 137 tournaments and identified the round in which they either did meet or could have met. For the purposes of analyzing draws, there isn’t really a difference. For instance, Serena and Venus have landed in the same half 73 out of a possible 137 times, a bit more than the 68 or 69 times that we would expect.

Because of their seeds, they had the chance of ending up in the same quarter 116 times, and that’s how it worked out 28 times, just under the 29 times that an exact one-in-four rate would’ve given them. The smaller the draw section, the fewer tournaments that Serena’s and Venus’s seeds made it possible for them to meet.

I counted the number of tournaments with a possible meeting on or before a certain round, and then the number of events in which the draw delivered that meeting, regardless of whether both Williamses got that far. Here are the results, along with the probability of that many or more actual meetings:

Section  Possible  Actual  Chance  
Half          137      73     25%  
Quarter       116      28     62%  
Eighth         85      17      3%  
16th           64       5     37%  
32nd           42       1     74%

There’s a one-in-four chance that Serena and Venus would’ve landed in the same half as many times as they have throughout their entire careers. That’s a bit of bad luck, but it’s hardly a smoking gun. The same is true for the same quarters, as well as very early meetings that would pit them against each other in the round of 32 or 64.

That leaves one eyebrow-raising number to discuss. On 85 occasions, at least one of the two women was seeded outside the top eight, making possible a meeting in the round of 16 or earlier. Given random draws, we’d expect 10 or 11 brackets in which they could face each other so early. Instead, we got 17.

A 3% chance of so many early encounters isn’t quite as bad as it sounds. I’ve tried to walk you through this process in the way I approached it. While I wondered if Serena and Venus have met more often than random draws would normally deliver, I didn’t have a particular round in mind. As you’ve seen, I generated a bunch of numbers, and one of the five looked suspicious. You might be able to construct a story that explains why the round of 16 is different from the others (such as my theory that tournament directors want mid-week headlines), but because we generated so many numbers, we were that much more likely to end up with an extreme percentage simply by chance.

The smoking (nerf) gun

Thus, we’re able to raise the possibilities that some draws weren’t random, but we can hardly prove it. One problem–one that we could’ve foreseen from the get-go–is that some draws are definitely not tampered with. Probably most draws. And even if they were, most tournaments wouldn’t have any reason to mess with Serena’s or Venus’s placement in the bracket. Or if they did, they might prefer an all-Williams final, and thus alter the bracket in the opposite direction of what we’re hunting for.

If you like conspiracy hunting, I’ve got a tiny sample for you. Since the beginning of 2018, Venus and Serena have played in the same tournament 15 times, and their seedings (or lack thereof) made it possible for them to be drawn in the same eighth 14 of those times. Of the 14, they were placed in position for a round-of-16 or earlier meeting 5 times. There’s only a 2% chance of that … if you set aside the fact that I’m checking all sorts of subsets of matches looking for (probably spurious) patterns. If nothing else, the 5-of-14 figure explains why it seems like Serena and Venus keep landing in the same draw sections lately. They do!

Broadly speaking, then, this is all much ado about nothing. (I don’t even know if these conspiracy theorists exist, so maybe I just invented a conspiracy and spent my evening debunking it. Hooray?) It’s possible that a few tournament directors are producing non-random draws … but it would take a very different kind of investigative work to prove it. Worst case scenario, we get a few more Serena-Venus matches. It may not be fair to the older sister, but it’s a pretty good deal for tennis fans.

Did Palermo Show the Signs of a Five-Month Pandemic Layoff?

Are tennis players tougher to predict when they haven’t played an official match for almost half a year? Last week’s WTA return-to-(sort-of)-normal in Palermo gave us a glimpse into that question. In a post last week I speculated that results would be tougher than usual to forecast for awhile, necessitating some tweaks to my Elo algorithm. The 31 main draw matches from Sicily allow us to run some preliminary tests.

At first glance, the results look a bit surprising. Only two of the eight seeds reached the semifinals, and the ultimate champion was the unseeded Fiona Ferro. Two wild cards reached the quarters. Is that notably weird for a WTA International-level event? It doesn’t seem that strange, so let’s establish a baseline.

Palermo the unpredictable

My go-to metric for “predictability” is Brier Score, which measures the accuracy of percentage forecasts. It’s nice to pick the winner, but it’s more important to assign the right level of probability. If you say that 100 matches are all 60/40 propositions, your favorites should win 60 of the 100 matches. If they win 90, you weren’t nearly confident enough; if they win 50, you would’ve been better off flipping a coin. Brier Score encapsulates those notions into a single number, the lower the better. Roughly speaking, my Elo forecasts for ATP and WTA matches hover a bit above 0.2.

From 2017 through March 2020, the 975 completed matches at clay-court WTA International events had a collective Brier Score of 0.223. First round matches were a tiny bit more predictable, with R32’s scoring 0.219.

Palermo was a roller-coaster by comparison. The 31 main-draw matches combined for a Brier Score of 0.268. Of the 32 other events I considered, only last year’s Prague tourney was higher, generating a 0.277 mark.

The first round was more unpredictable still, at 0.295. On the other hand, the combination of a smaller per-event sample and the wide variety of first-round fields means that several tournaments were wilder for the first few days. 9 of the 32 others had a first-round Brier Score above 0.250, with four of them scoring higher–that is, worse–than Palermo did.

The Brier Score of shame

I mentioned the 0.250 mark because it is a sort of Brier Score of shame. Let’s say you’re predicting the outcome of a series of coin flips. The smart pick is 50/50 every time. It’s boring, but forecasting something more extreme just means you’re even more wrong half the time. If you set your forecast at 50% for a series of random events with a 50/50 chance of occurring, your Brier Score will be … 0.250.

Another way to put it is this: If your Brier Score is higher than 0.250, you would’ve been better off predicting that every match was 50/50. All the fancy forecasting went to waste.

In Palermo, 17 of the 31 matches went the way of the underdog, at least according to my Elo formula. The Brier Scores were on the shameful side of the line. My earlier post–which advocated moderating all forecasts, at least a bit–didn’t go far enough. At least so far, the best course would’ve been to scrap the algorithm entirely and start flipping that coin.

Moderating the moderation

All that said, I’m not quite ready to throw away my Elo ratings. (At the moment, they pick Simona Halep and Aryna Sabalenka, my two favorite players, to win in Prague in Lexington. So there’s that.) 31 matches is small sample, far from adequate to judge the accuracy of a system designed to predict the outcome of thousands of matches each year. As I mentioned above, Elo failed even worse at Prague last year, but because that tournament didn’t follow several months of global shutdowns, it wouldn’t have even occurred to me to treat it as more than a blip.

This time, a week full of forecast-busting surprises could well be more than a blip. Treating players as if they have exactly the abilities they had in March is probably the wrong way to do things, and it could be a very wrong way of doing things. We’ll triple the size our sample in the next week, and expand it even more over the next month. It won’t help us pick winners right now, but soon we’ll have a better idea of just how unpredictable the post-COVID-19 tennis world really is.

Did Jimmy Connors Choke in the 1975 Wimbledon Final?

From our vantage point almost a half-century later, it’s easy to forget just how big an upset Arthur Ashe scored with his 1975 Wimbledon victory over Jimmy Connors. Connors was the top seed and defending champion, still riding high from a 1974 campaign that ranks among the best ever. Ashe was a few days short of his 32nd birthday, had a reputation of coming up short in finals, and had lost to Connors in their three previous meetings.

(For what it’s worth, my Elo algorithm thinks it was a much closer match than the bookies did at the time. It rated Ashe the second-best player in the tournament on grass courts, and gave the underdog a 39% chance of winning.)

Ashe ran away with the first two sets and held on to win in four, 6-1 6-1 5-7 6-4. Perhaps because the two men didn’t get along–apart from striking personality differences, Connors and his manager targeted Ashe with one of many lawsuits–the veteran was uncharacteristically critical of his opponent after the match. Ashe claimed that Connors missed many of his shots into the net (rather than long), a sign of choking.

Connors denied it, of course. It later came out that Jimmy was dealing with a foot problem which probably affected his play that day. In any case, fans and pundits surely had their fun debating whether Connors was a choker. I don’t know of anyone who took the question beyond simple speculation. No amount of statistical analysis can settle whether a player choked, but we can often answer adjacent questions to shed more light on the issue.

Counting errors

A couple of years ago I charted the Wimbledon final for the Match Charting Project, so we have a full count of errors–forced and unforced, serves and rallying shots, net and deep–for the entire match. We also have similar shot-by-shot stats for 25 other Connors matches for comparison. (Unfortunately, 24 of the 25 are chronologically later than the Ashe match, because there’s not much full-match footage from the early 70s.)

Here’s the tally: Excluding serves, Connors committed 13 unforced errors, 10 of them into the net. I recorded the type of error for 65 more forced errors: 32 into the net, 33 other. (Ashe was a netrusher, so many of Jimbo’s mistakes were failed passing shots.) On serve, he missed 29 first deliveries: 16 into the net, 13 otherwise. And his two second serve faults were split between one into the net and one elsewhere.

The unforced error split of 10-to-3 means that 77% of his UFEs were netted. That’s the most extreme of any of his charted matches; on average, his unforced errors were half nets, half others. While suggestive, that’s an awfully small sample from which to draw any conclusions.

Using larger samples that include forced errors and serves, the Wimbledon final doesn’t particularly stand out among other charted Connors matches. 54% of his non-serve errors (forced or unforced) in that match were netted, compared to 52% over the whole sample. 55% of his service faults against Ashe were hit into the net, versus 49% across the 26 matches. Altogether, Connors made 54% of his total errors and faults into the net in the Wimbledon final, compared to 51% in the broader sample.

Does it matter?

You’ve probably heard the tennis coaching conventional wisdom that it’s better to hit long than to hit into the net. Like most tennis shibboleths, this one has been around for a very long time. Ashe had surely heard it, which partly explains why he made the comment he did. Arthur didn’t have a printout with match stats generated by a consulting company with a gargantuan marketing budget, so he probably recalled a few key points and generalized from there.

If error types matter, we’d expect to see at least a mild correlation between results (say, percentage of points won) and error types. Let’s stay focused on the 26 charted Connors matches for today’s purposes. Here’s a version of the Ashe hypothesis, stripped of emotional content:

When Connors hits more errors than usual into the net, it’s a sign that he’s playing below his standard level.

It turns out that this theory is wrong–or, at best, possibly correct if narrowly defined. I considered five main stats as indicators of errors and faults going into the net:

  • Unforced errors (excluding double faults) into the net as a percentage of total unforced errors
  • Total rally errors (forced and unforced) into the net as a percentage of total errors
  • First serve faults into the net as a percentage of total first serve faults
  • All serve faults into the net as a percentage of all serve faults
  • All errors and faults into the net as a percentage of all errors and faults

The second (total rally errors) and last (all errors and faults) seem like the most valid of the five, because they give us a decent sample of error types for each match. There is almost exactly zero correlation between the last stat and total points won. And there is a very weak negative correlation (r^2 = 0.05) between the second stat and total points won.

In other words, the Ashe hypothesis might be on to something very minor if our focus in on rally shots. But the correlation is so weak that no human observer would ever notice it, unless they lucked into it by watching a few confirming key moments after being primed by the conventional wisdom.

He didn’t choke like that

I said above that statistical analysis couldn’t settle issues like whether a player choked. We can study what happened, but without machines hooked up to a player’s brain, we can’t tell what was going on inside their heads that might have caused it.

So we can’t say that Connors didn’t choke in the 1975 Wimbledon final. But we have seen that his percentage of into-the-net errors wasn’t that unusual for him (except for the small sample of unforced errors), and we’ve recognized that the number of mistakes he made into the net didn’t have much to say about his level of play that day. If Connors choked, then, it didn’t have anything to do with the low trajectory of his missed shots.

—

I learned of Ashe’s post-match comment in Raymond Arsenault’s excellent biography, Arthur Ashe: A Life.

Elo, Meet COVID-19

Tennis is back, and no one knows quite what to expect. Unpredictability is the new normal at both the macro level–will the US Open be a virus-ridden disaster?–and the micro level–which players will come back stronger or weaker? While I plead ignorance on the macro issues, estimating player abilities is more in my line.

Thanks to global shutdowns, every professional player has spent almost five months away from ATP, WTA, and ITF events–“official” tournaments. Some pros, such as those who didn’t play in the few weeks before the shutdowns began, or who are opting not to compete at the first possible opportunity, will have sat out seven or eight months by the time they return to court. Exhibition matches have filled some of the gap, but not for every player.

Half a year is a long time without any official matches. Or, from the analyst’s perspective: It’s tough to predict a player’s performance without any data from the last six months.

Increased uncertainty

Let’s start with the obvious. All this time off means that we know less about each player’s current ability level than we did before the shutdown, back when most pros were competing every week or two. Back in March, my Elo ratings put Dominic Thiem in 5th place, with a rating of ~2050, and David Goffin in 15th, with a rating of ~1900. Those numbers gave Thiem a 70% chance of winning a head-to-head.

What about now? Both men have played in exhibitions, but can we be confident that their levels are the same as they were in March? Or that they’ve risen or fallen roughly the same amount? To me, it’s obvious that we can’t be as sure. Whenever our confidence drops, our predictions should move toward the “naive” prediction of a 50/50 coin flip. A six-month coronavirus layoff isn’t that severe, so it doesn’t mean that Thiem is no longer the favorite against Goffin, but it does mean our prediction should be closer to 50% than it was before.

So, 60%? Maybe 65%? Or 69%? I can’t answer that–yet, anyway.

The (injury) layoff penalty

My Elo ratings already incorporate a layoff penalty, which I introduced here. The idea is that if a player misses a substantial amount of time (usually due to injury, but possibly because of suspension, pregnancy, or other reasons), they usually play worse when they come back. But it’s tough to predict how much worse, and players regain their form at different rates.

Thus, the tweak to the rating formula has two components:

  • A one-time penalty based on the amount of time missed (more time off = bigger penalty)
  • A temporarily increased k-factor (the part of the formula that determines how much each match increases or decreases a player’s rating) to account for the initial uncertainty. After an injury, the k-factor increases by a bit more than 50%, and steadily declines back to the typical k-factor over the next 20 matches.

Not an injury

A six-month coronavirus layoff is not an injury. (At least, not for players who haven’t lost practice time due to contracting COVID-19 or picking up other maladies.) So the injury-penalty algorithm can’t be applied as-is. But we can take away two ideas from the injury penalty:

  • If we generate those closer-to-50% forecasts by shifting certain players’ ratings downward, the penalty should be less than the injury penalty. (The minimum injury penalty is 100 Elo points for a non-offseason layoff of eight or nine weeks.)
  • The temporarily increased k-factor is a useful tool to handle the type of uncertainty that surrounds a player’s ability level after a layoff.

The injury-penalty framework is useful because it has been validated by data. We can look at hundreds of injury (and other) layoffs in modern tennis history and see how players fared upon return. And the numbers I use in the Elo formula are based on exactly that. We don’t have the same luxury with the last six months, because it is so unprecedented.

Not an offseason, but…

The closest thing we have to a half-year shutdown in existing tennis data is the offseason. The sport’s winter break is much shorter, and it isn’t the same for every player. Yet some of the dynamics are the same: Many players fill their time with exhibitions, others sit on the beach, some let injuries recover, others work particularly hard to improve their games, and so on.

Here’s a theory, then: The first few weeks of each season should be less predictable than average.

Fact check: False! For the years 2010-19, I labeled each match according to how many previous matches the two players had contested that year. If it was both players’ first match, the label was 1. If it was one player’s 15th match and the other’s 21st, the label was the average, 18. Then, I calculated the Brier Score–a measure of prediction accuracy–of the Elo-generated predictions for the matches with each label.

The lower the Brier Score, the better. If my theory were right, we would see the highest Brier Scores for the first few matches of the season, followed by a decrease. Not exactly!

The jagged blue line shows the Brier Scores for each individual label (match 1, match 2, match 23, etc), while the orange line is a 5-match moving average that aims to represent the overall trend.

There’s not a huge difference throughout the season (which is reassuring), but the early-season trend is the opposite of what I predicted. Maybe the women, with their slightly longer offseason, will make me feel better?

No such luck. Again, the match-to-match variation in prediction accuracy is very small, and there’s no sign of early-season uncertainty.

I will not be denied

Despite disproving my own theory, I still expect to see an unpredictable couple of post-pandemic months. The regular offseason is something that players are accustomed to, and there is conventional wisdom in the game surrounding how to best use that time. And it’s two months, not five to seven. In addition, there are many other things that will make tour life more challenging–or different, at the very least–in 2020, such as limited crowds, social distancing protocols, and scheduling uncertainty. Some players will better handle those challenges than others, but it won’t necessarily be the strongest players who respond the best.

So my Elo ratings will, for the time being, incorporate a small penalty and a temporarily increased k-factor. (Something more like 69% for Thiem-Goffin, not 60%.) I haven’t finished the code yet, in large part because handling the two different types of layoffs–coronavirus and the usual injuries, etc–makes things very complicated. If you’re watching closely, you’ll see some minor tweaks to the numbers before the “Cincinnati” tournament in a few weeks.

There is a right answer

It’s clear from what I’ve written so far that any attempt to adjust Elo ratings for the COVID-19 layoff is a bit of a guessing game. But it won’t always be that way!

By the end of the year, we’ll know the right answer: just how unpredictable results turned out to be in the early going. Just as I’ve calculated penalties and k-factor adjustments for injury layoffs based on historical data, we will be able to do the same with match results from the second half of 2020. To be more precise, we’ll be able to work out a class of right answers, because one adjustment to the Elo formula will give us the best Brier Score, while another will best represent the gap between Novak Djokovic and Rafael Nadal, while others could target different goals.

The ultimate after-the-fact COVID-19 Elo-formula adjustment won’t help you win more money betting on tennis, but it will give us more insight into how the coronavirus layoff affected players after so much time off, and how quickly they returned to pre-layoff form. We’ll understand a little bit more about the game, even if we desperately hope never to have reason to apply the newly-won knowledge.

The Rarity of Winning Two Titles at One Tournament

This is a guest post by Peter Wetz.

With all the drama in the tennis world right now–paradoxically despite the lack of official match results–a dry analytical article might be just what you need. And what better opportunity than quarantine to work through my long list of articles to write?

In June 2019, Feliciano Lopez had to complete five matches in two days. Not because he had to hop between tournaments as a 22-year-old Jo-Wilfried Tsonga did in 2007, but because Lopez went deep into both the singles and doubles draws on the grass courts at Queen’s Club, ultimately winning both titles.

Lopez won all four of his singles matches in the deciding set, and there was not much time to celebrate and recover after the final, because the doubles title match awaited. Partnering a rehabilitating Andy Murray seems to have been a sensible decision based on the fact that Murray’s most lopsided head-to-head of 11-0 is against Lopez. By doing so, Lopez could be guaranteed to avoid facing Murray in the doubles draw. An unusual strategy–and probably not his top consideration in choosing a partner–but it worked.

Lifting two trophies on finals day happens quite often at the Challenger tour, but is unusual on the main tour, where the best singles players often skip the doubles draw entirely. But how rare is it? And has it changed over the years? Longtime fans will immediately think of John McEnroe and his nearly equal tally of doubles titles (78) and singles titles (77). The modest title counts of Roger Federer (6) and Rafael Nadal (11) pale in comparison, even though the Spaniard is an exceptional doubles player.

Let’s take a look at the instances when a player won both trophies at the same tournament since 2005.

Year	Tournament	Player (Partner)
2005	Dusseldorf	Tommy Haas (Alexander Waske)
2005	Halle		Roger Federer (Yves Allegro)
2005	Basel		Fernando Gonzalez (Agustin Calleri)
2006	Vina del Mar	Jose Acasuso (Sebastian Prieto)
2007	Chennai		Xavier Malisse (Dick Norman)
2007	Delray Beach	Xavier Malisse (Hugo Armando)
2007	Munich		Philipp Kohlschreiber (Mikhail Youzhny)
2007	Dusseldorf	Agustin Calleri (Juan Ignacio Chela)
2008	Monte Carlo	Rafael Nadal (Tommy Robredo)
2008	Dusseldorf	Robin Soderling (Robert Lindstedt)
2009	Costa Do Sauipe	Tommy Robredo (Marcel Granollers)
2009	San Jose	Radek Stepanek (Tommy Haas)
2009	Newport		Rajeev Ram (Jordan Kerr)
2010	Memphis		Sam Querrey (John Isner)
2010	Marseille	Michael Llodra (Julien Benneteau)
2010	Bucharest	Juan Ignacio Chela (Lukasz Kubot)
2011	Tokyo		Andy Murray (Jamie Murray)
2012	Zagreb		Mikhail Youzhny (Marcos Baghdatis)
2013	Newport		Nicolas Mahut (Edouard Roger Vasselin)
2014	Newport		Lleyton Hewitt (Chris Guccione)
2017	Montpellier	Alexander Zverev (Mischa Zverev)
2018	Gstaad		Matteo Berrettini (Daniele Bracciali)
2019	London		Feliciano Lopez (Andy Murray)

Two things may catch one’s eye when looking at the list: First, since 2011 the double-title feat occurred slightly less than once per year. But before that it happened several times a year with the sole exception of 2006. Second, the only player who managed to win both titles at a Masters event is Nadal at Monte Carlo in 2008.

It is obvious, and a frequent topic of tennis hipster talk, that top singles players do not care as much about doubles anymore, certainly not as much as McEnroe and his peers did. One line of argument is that the way that modern doubles tennis has evolved to become more and more different from the singles game. In order to keep up with that, singles players would need to adapt their practice routine, which might detract from potential singles success. Long story short, the argument is that doubles became too “difficult” for singles players.

But let’s look at the numbers. The following graphs show the composition of draws since the year 2000. We see the percentage of players in singles draws, who also entered the doubles draw of the same tournament for three different categories (A = All, M = Masters, G = Grand Slams). The first graph shows the numbers for top 50 singles players and the second graph for top 10 singles players.

Percentage of top 50 players entering doubles draws per 5 years
Percentage of top 10 players entering doubles draws per 5 years

The first graph is not very dramatic, but it establishes that the habits of top 50 singles players have been quite steady over the past 20 years among all tournament categories. Since the year 2000, irrespective of event categories, between 41 and 47 percent of top 50 players entering a singles draw also entered the doubles draw of the same tournament.*

The second graph shows us that the numbers for top 10 players are a different story entirely. Ignoring tournament categories, the number of top 10 players participating in doubles draws has plummeted from 35 to 22 percent. While the numbers also decreased if we only look at Masters tournaments, it is interesting that it remains higher than the overall number. This can likely be explained by the fact that the prize money for doubles at Masters events is significantly higher than at regular tour events. Often the organizers of these tournaments also have the financial power to persuade top players to play doubles in order to–I am hypothesizing here–increase ticket sales or attendance in the early days of a tournament. See the Indian Wells Masters for instance, which is known for its stellar doubles draw every year.

The most drastic decline in doubles attendance by top 10 singles players can be seen at the Grand Slams, however. While in the period between the years 2000 and 2004 every fifth singles player took part in the doubles, in the past five years only one out of 183 singles entries also appeared in the doubles draw. The sole exception (of course!) was Dominic Thiem, who entered the 2016 US Open doubles competition ranked number 10 in singles with his countryman Tristan Samuel Weissborn.

As with many analyses it is difficult to provide a definitive answer to the question at hand. But the numbers help us to see the size of the effects and theorize about its causes. That doubles competition has become more and more specialized certainly has its validity. At the same time, the numbers also suggest that top singles players simply optimize for prize money, which means focusing on singles, not doubles. If there was a McEnroe-esque player on tour today (as Rafa might be), he just wouldn’t play enough doubles to win nearly 80 titles.

However, it is hard to tell which was first: The decline of singles players playing doubles due to reasons such as financial motivation (among possibly many others), or the players’ realization that they simply cannot keep up with the elite doubles competition? One thing may be for sure though: Had TennisTV already existed a few decades ago, it would have shown a lot more doubles than it does now.

—

* Note that there is the possibility that a few singles players might have been willing to enter the doubles draw of a tournament, but couldn’t, because their ranking was too low among other reasons. However, I think this affects the analysis only marginally, if at all.

—

Peter Wetz is a computer scientist interested in racket sports and data analytics based in Vienna, Austria.

Podcast Episode 83: Is the Practice Court Broken?

Episode 83 of the Tennis Abstract Podcast features co-host Carl Bialik, of the Thirty Love podcast, and guest Jeff McFarland of Hidden Game of Tennis. This week we dip our collective toe into a debate in the tennis coaching world.

With rallies short and aggressive, should players be using practice time differently? What types of skills can still be improved, once a player has reached the top? What tactics can a coach teach their charges, and which ones are too deeply ingrained in the physical nature of hitting the shots? The line between technique and tactics may not be a clear-cut as we think.

Is a 3- or 4-shot rally qualitatively different from a 5- or more-shot rally? How would you teach Madison Keys to retain the positives of her aggressive style while dialing back the aggression a bit? We offer more questions than answers, which seems appropriate for a topic that is far from settled, and is likely to remain controversial for years to come.

Thanks for listening!

(Note: this week’s episode is about 67 minutes long; in some browsers the audio player may display a different length. Sorry about that!)

Click to listen, subscribe on iTunes, or use our feed to get updates on your favorite podcast software.