Across 116,382 completed best-of-three singles matches played between 1 January 2023 and 14 September 2026, the player who lost the first set went on to win 16.52% of the time. But that single number hides the finding: the rate falls steadily as you move down the tours, from 18.21% on the ATP tour to 15.18% in ITF men's singles. The surface you play on barely moves it. The tier you play in does.
But the strongest predictor is neither the tour nor the surface: it is how close the first set was. A set lost 7-6 is recovered 21.92% of the time; a set lost 6-0 is recovered 7.18% of the time. That spread — 14.74 points — is five times the difference between tours and twenty times the difference between surfaces.
These are counted figures from our own point-by-point tape, not estimates, and the method is stated below so you can disagree with it.
Does the first-set score predict the comeback?
More than anything else measured here. Same 116,382 matches, grouped by the score of the opening set:
| First set | Margin | Matches | Comeback rate |
|---|---|---|---|
| 7-6 | 1 game | 14,265 | 21.92% |
| 6-4 | 2 games | 25,546 | 19.41% |
| 7-5 | 2 games | 10,252 | 18.89% |
| 6-3 | 3 games | 25,970 | 17.39% |
| 6-2 | 4 games | 19,739 | 13.30% |
| 6-1 | 5 games | 14,269 | 11.06% |
| 6-0 | 6 games | 5,833 | 7.18% |
The ordering is monotonic in the game margin, and the detail that makes it convincing is that 6-4 and 7-5 land together — 19.41% and 18.89%. Both are two-game margins, and they behave like two-game margins despite looking like different scorelines. What predicts recovery is the size of the gap, not the shape of the scoreboard.
So a player who drops a first-set tiebreak is roughly three times more likely to win the match than one who is bagelled, and that single split matters far more than knowing which tour they are on or what they are playing on.
How often is a one-set deficit recovered?
| Tour | Matches | Comeback rate | 95% CI |
|---|---|---|---|
| ATP | 10,945 | 18.21% | ±0.72pp |
| Challenger (men) | 30,406 | 17.32% | ±0.43pp |
| Challenger (women) | 4,731 | 17.25% | ±1.08pp |
| WTA | 13,543 | 17.04% | ±0.63pp |
| ITF (women) | 29,331 | 15.97% | ±0.42pp |
| ITF (men) | 27,415 | 15.18% | ±0.42pp |
The gap between the top and bottom rows is 3.03 percentage points. On a two-proportion test that is z = 7.08, p ≈ 1.4 × 10⁻¹², so it is not sampling noise, and the confidence intervals for ATP and both ITF tiers do not overlap.
Best-of-five changes the picture as you would expect. Over 1,597 completed best-of-five ATP matches the comeback rate is 23.98% — more sets to recover in, and roughly a third more recoveries.
Does the surface matter?
Much less than the tour does.
| Surface | Matches | Comeback rate |
|---|---|---|
| Grass | 3,593 | 17.20% |
| Clay | 47,944 | 16.79% |
| Hard | 61,751 | 16.47% |
That is a spread of 0.73 percentage points, against 3.03 across tiers — the tier accounts for about four times as much variation.
Surface and tier are confounded, though, because tours do not play the same surface mix. So the honest check is whether each effect survives the other. It does: on clay the order runs ATP 19.42% → Challenger men 17.38% → ITF men 15.21%, and on hard it runs ATP 18.91% → Challenger men 17.22% → ITF men 15.22%. ATP is highest on both surfaces and ITF men lowest on both, while within any single tier the surfaces sit within about half a point of each other.
From which second-set scores is the match still winnable?
The rate above is measured at the moment the first set ends. It moves fast. Taking every distinct second-set game score reached in those matches — counted once per match, so a long match does not vote more than once — the chaser's eventual win rate is:
| Second set (chaser first) | Matches | Chaser wins match |
|---|---|---|
| 6-1 | 3,682 | 58.53% |
| 6-2 | 5,044 | 56.38% |
| 6-3 | 8,943 | 52.19% |
| 6-4 | 8,132 | 50.16% |
| 3-0 | 8,755 | 45.22% |
| 3-1 | 18,412 | 39.02% |
| 2-0 | 16,660 | 36.80% |
| 6-5 | 11,369 | 31.93% |
| 3-2 | 30,218 | 27.31% |
| 2-1 | 38,611 | 25.38% |
| 1-0 | 53,887 | 23.64% |
| 5-5 | 22,506 | 21.13% |
| 3-3 | 33,832 | 19.47% |
| 1-1 | 63,405 | 17.66% |
| 0-0 | — | 16.52% |
| 1-2 | 47,102 | 10.42% |
| 0-1 | 61,626 | 10.36% |
| 2-3 | 37,473 | 10.19% |
| 3-4 | 31,243 | 9.51% |
| 4-5 | 26,514 | 8.17% |
One game decides more than the whole first set did. At 0-0 in the second the chaser is on 16.52%. Win the opening game and it is 23.64%; lose it and it is 10.36%. A single game more than doubles the spread between the two branches.
The table validates itself at one row. A chaser who takes the second set 6-4 is level at one set all with a deciding set to play — and the table, built only from counted outcomes with no model behind it, puts them at 50.16%. Nothing forced that to land on a coin flip. It is the strongest reason to trust the rest of the rows.
Two further patterns worth reading off it. Level scores drift upward as the set goes on — 1-1 is 17.66%, 3-3 is 19.47%, 5-5 is 21.13% — because a chaser still level late has been holding serve and a deciding set is getting closer. And trailing scores drift downward — 0-1 is 10.36% but 4-5 is 8.17% — because the same deficit with fewer games left is worth less. The exception is 5-6 at 10.11%, which is higher than 4-5 because the chaser can still hold to force a tiebreak.
Where the surface does show up: tiebreaks
The same corpus, counting how many completed sets reached a tiebreak:
| Surface | Sets reaching a tiebreak | Matches with at least one |
|---|---|---|
| Grass | 15.91% | 31.73% |
| Hard | 12.31% | 25.28% |
| Clay | 10.72% | 22.52% |
Grass produces roughly half again as many tiebreaks per set as clay. That ordering is what anyone who watches tennis would predict, which is part of why we trust the pipeline that produced the comeback numbers: the method reproduces a known result on data it was not tuned for.
How this was measured
- Corpus. Completed singles matches on ATP, WTA, Challenger and ITF, 2023-01-01 to 2026-09-14, counted 14 September 2026.
- Retirements and walkovers are excluded. Only matches the feed marks
Finishedare counted — 5,173 retirements and 84 walkovers inside this corpus were dropped, because a match the opponent abandoned is not a comeback. - The first-set winner is read from the first tape row where the set count first totals one. The match winner is read from the final tape row.
- Truncated tapes are excluded. 1,514 tapes end with neither player on two sets, plus 43 with a set count impossible for the format. They are dropped rather than guessed at.
- A caveat we will not bury: the
winner_sidecolumn is unpopulated for almost every historical match, so the winner is derived from the tape's own final state rather than from an independent field. The two therefore cannot disagree, and this study does not claim they were cross-checked. - One structural skew. Player one wins 58.7% of matches in this corpus, because feeds conventionally list the stronger or seeded player first. A comeback rate is player-agnostic — it asks only whether the first-set loser won — so the skew does not bias it, but it would badly bias any statistic tied to player order.
Are these points real, or reconstructed?
Both, and the distinction is worth stating. Of 27,381,531 point-state rows, about 85% carry the provenance reconstructed and 15% observed. Reconstructed does not mean invented: those are the vendor's own recorded point-by-point sequences for finished matches, expanded into score states. What reconstructed rows deliberately lack is a per-point wall clock and any model output — both are left NULL rather than synthesised, because inventing a timestamp would make the tape lie about when a thing was known.
For this study that is the right corpus, because it uses only the sequence of sets, which is real in both. Any analysis that depends on timing between points, or on model win-probability, must restrict itself to observed rows — and the data makes that easy, because the columns are simply empty otherwise.
Reproducing it
Every figure here comes from two endpoints on the Basic ($9.99/mo) tier. List completed matches, then pull each tape:
import os, requests
BASE = "https://api.livetennisapi.com/api/public/v1"
H = {"X-API-Key": os.environ["LIVE_TENNIS_API_KEY"]}
r = requests.get(f"{BASE}/history/matches",
params={"tour": "atp", "limit": 100}, headers=H, timeout=30)
r.raise_for_status()
for m in r.json().get("matches", []):
if not m.get("has_tape"):
continue # coverage is 96%, not 100% — skip, never assume
tape = requests.get(f"{BASE}/history/matches/{m['id']}", headers=H, timeout=30).json()
# each state carries sets, games, points and the server: the first row where the
# set count totals one tells you who took the opening set
Check GET /history/coverage before treating any slice as complete — it returns measured completeness per tour with its own as_of date.
Full endpoint detail is on the point-by-point history page and the full reference. Tiers and limits are on the pricing page.