Across 238,716 player-matches from 2018 to 2026, players returning from a layoff win less often than their ELO expects, on a monotone curve running from 0.9 percentage points below expectation after 8 to 14 days out to 9.8 points below after more than a year. The deficit fades over four to six matches back.
Every preview written about a returning player says the same thing. He has been out for four months. She has not played since March. Watch for rust.
Nobody says how much rust. The claim is universal, it is never quantified, and it is never checked against the rating the returning player carries into the match. So we measured it on every player-match in our database from 2018 to 2026, 238,716 observations, and we shipped the curve into our own model.
Rating systems have known about inactivity for three decades, but they encode it as uncertainty rather than as a deficit. Glickman's Glicko system extends ELO with a second number per player, the Rating Deviation, and the technical note is explicit that "time passing without competing in rated games always increases a player's RD" [1], [2]. The prediction step then attenuates the rating gap by a factor g(RD), so a wide RD on either side pulls the predicted probability toward 50%. That is a statement about how much we know, not about how well the player will play.
The applied tennis literature does carry a layoff-adjacent feature, and it is worth being precise about what it publishes. Sipko's Imperial College thesis, a widely cited feature-engineering treatment in this space, extracts 22 features including FATIGUE, games played in the past three days discounted at 0.75 per day (our own measurement of that short-window effect is the fatigue study), and RETIRED, a binary flag for a player's first match since retiring from one [3]. The only magnitude he publishes for RETIRED is a raw split, on his 2004-2010 ATP training sample: 48.5% won, 51.5% lost. That is an unadjusted win rate, not a residual against a rating, so it cannot say how much a rating overstates a returning player.
More interesting is the feature he considered and dropped. He initially looked at time since a retirement as a severity measure, then rejected it because "a player that has not competed for longer has also had more time to recover, so the relationship is unclear" [3]. As far as we can find, that sentence is why the layoff curve does not exist in the published literature. The sign was assumed ambiguous, so nobody drew it.
It is not ambiguous. It is monotone. And none of this is a niche correction: in the most cited head-to-head comparison of tennis forecasting methods, Kovalchik tested 11 published models on a full ATP season and found ELO-based methods among the most accurate available [4], so a systematic error in an ELO rating is an error in the strongest published baseline.
Every scored observation is one player-match pair, and the unit is the residual: won (0 or 1) minus the ELO-expected win probability, in percentage points. Raw win rate is useless here, because it only rediscovers that better players win. Our rating table stores the post-match rating plus the delta applied, so the as-of rating is the stored rating minus that delta. Nothing after the match is ever read.
The headline view takes one observation per match, the rustier of the two sides, and drops matches where both players sit in the same layoff bucket. That matters more than it sounds: the two residuals in a match are exact negatives, so averaging over all player-matches gives identically zero, and same-bucket pairs cancel inside their own bucket. Taking the rustier side keeps one clean observation per match.
The layoff clock runs over every match in the database, including exhibitions and juniors, because any time on court is anti-rust. Walkovers are skipped, since nobody played.
| Layoff | n | Observed minus expected (pp) | 95% CI | Implied ELO overstatement |
|---|---|---|---|---|
| 8-14 days | 24,088 | −0.89 | [−1.46, −0.31] | +6 |
| 15-28 days | 21,993 | −2.41 | [−3.01, −1.81] | +17 |
| 29-56 days | 13,556 | −4.20 | [−4.95, −3.44] | +30 |
| 57-90 days | 6,163 | −4.00 | [−5.11, −2.89] | +29 |
| 91-180 days | 4,591 | −4.20 | [−5.47, −2.93] | +31 |
| 181-365 days | 2,518 | −4.88 | [−6.62, −3.13] | +36 |
| 366+ days | 768 | −9.75 | [−12.85, −6.64] | +73 |
The right-hand column is the number we actually use. It inverts the logistic at each bucket's mean expected probability and answers one question: how many rating points was this player's ELO overstated by? Two weeks off is worth about 6 points, which is nothing. A month is worth 30. A year is worth 73, which is roughly the gap between a solid Challenger regular and the man two rungs above him. The curve replicates separately on ATP and WTA, on the post-2023 window, and on the surface rating book.
A residual method can manufacture effects. Our ELO is not perfectly calibrated across the probability range, and returning players skew hard to one side of the board, so any bucket that concentrates favourites or underdogs inherits that local bias and prints it as if it were a layoff effect.
So we ran the same measurement on players with no layoff at all: 0 to 7 days off, against an opponent who had themselves played within 14 days. If the instrument invents effects, this is where it shows. It printed +0.06pp, CI [−0.09, +0.22]. Correctly null.
That is also why the shipped curve is exactly zero below 8 days rather than a smoothed tail into the origin. It is a measured floor, not a modelling convenience.
Permanent damage persists. A rating that has not caught up dies within a few matches as ELO re-converges. Splitting returns from 91 days or more by matches played since coming back:
| Match back | 1st | 2nd | 3rd | 4th to 6th | 7th to 12th | 13th to 30th |
|---|---|---|---|---|---|---|
| Residual (pp) | −3.11 | −2.77 | −2.34 | −0.91 | null | +0.54 |
It re-converges over roughly four to six matches. That is the signature of a stale rating with a genuine short-lived rust component on top, not of a permanently worse player.
The last cell is a mild positive overshoot. It is one bucket out of eight in a study that applies no multiplicity correction, where roughly 0.4 false positives at p<0.05 are expected by chance alone. It sits inside that budget, so we did not encode it. Our curve never turns into a bonus.
A player whose previous match ended in their own retirement underperforms by a further −1.85pp, CI [−2.86, −0.84], n=8,284, measured across layoff bands and therefore additive to the gap penalty.
The mirror case is the check that makes it meaningful. If this were really about having played an unusually short previous match, the player whose OPPONENT retired should show the same deficit. That cell is null: −0.59pp, CI [−1.53, +0.35]. The effect attaches to the body that broke down, not to the scoreline. Only the own-retirement side is encoded.
Candidate explanations predict different signatures, and they are separable. We split the rustier side's residual by whether that side was the favourite, 119,692 observations since 2020:
| Layoff | Rusty favourite (pp) | Rusty underdog (pp) |
|---|---|---|
| 8-14 days | +0.37 | +0.51 |
| 15-28 days | −1.90 | +0.02 |
| 29-56 days | −4.63 | −0.99 |
| 57-90 days | −7.38 | −2.54 |
| 91-180 days | −6.71 | −3.03 |
| 181-365 days | −9.01 | −3.92 |
| 366+ days | −14.33 | −2.51 |
| Pooled | −3.93 [−4.34, −3.52] | −1.48 [−1.81, −1.16] |
Symmetric Glicko attenuation is refuted by the underdog column. A g(RD) factor shrinks the rating gap for both sides equally, which pushes a rusty underdog's win probability UP toward 0.5, so their residual should come out positive. Rusty underdogs are negative in every bucket from 15-28 days onward. The effect is directional: the rusty side underperforms whether it is favoured or not, and uncertainty alone does not describe that.
A constant ELO subtraction does not explain the table either, and we should say so rather than declare victory. At the observed mean rating gaps of +129 for favourites and −163 for underdogs, a flat penalty predicts the two pooled residuals in a ratio of about 1.08. The observed ratio is about 2.7. The players with the most rating to lose lose the most, which sharpens the stale-rating reading, but we cannot derive that 2.7 from first principles. It is open, so the shipped curve is the pooled measurement and the split is recorded rather than fitted.
The curve went through our backtest gate as a variant we call "pen", on a 2023-01-01 cutoff with full-history rating state and a paired bootstrap interval against the unmodified baseline. In cells with n of at least 500 it is positive in 13 of 15, significant in 10, and positive on all three surfaces. Sweeping the penalty scale at 0.5, 1.0 and 1.5 wins at every scale, which rules out the result being an artifact of one chosen magnitude. One structural caveat belonged right here rather than in a footnote: the gate's validation window (2023 onward) sits inside the 2018-2026 window the curve was fitted on, so that first pass was weaker evidence than an out-of-sample pass would be. We have since closed it: refitting the curve on data through 2024 only and re-gating on a 2025 cutoff (a window the fit never saw) gives 14 of 15 cells positive, 9 significant, positive on all three surfaces, with the refit curve within a few ELO points of the full-sample one in every bucket. The effect is confirmed out-of-sample, not just indicated.
The improvement is small: roughly 0.0003 to 0.0012 Brier, on a base of about 0.22, so 0.15% to 0.5% relative. It is real, consistently signed and significant, and nobody should expect it to move returns on its own. We would rather print that number than a percentage that flatters it.
The version we did not ship is the more useful disclosure. Adding a K-factor inflation on top of the penalty, so a stale rating also re-converges faster, scored HIGHER than the plain penalty on hard and clay, in several cells nearly twice as high. It was negative in all three grass cells, and K inflation alone was the weakest and least consistent of the three variants. Our gate requires the same sign on at least two of three surfaces, and a variant that helps on two surfaces and hurts on the third is a variant we do not understand. We shipped the smaller number and left the better headline on the table, the same publish-then-encode policy as our surface-ELO change.
We found a large, monotone, replicating effect that the literature had declined to measure because one influential thesis judged its direction unclear. We shipped it. It improves our Brier score by about a third of one percent, and we still do not know whether the market has been pricing it correctly since 2018.
Both sentences are true at once. We publish every pick before the match and grade the wins and the losses on one public board, so we would rather tell you the size of the win than the size of the discovery.
Method: 238,716 player-match observations, matches 2018 through 2026, our own ELO system (main tour through Challenger and ITF), pre-match ratings only: the as-of rating is the stored post-match rating minus the update just applied, so nothing after the match is ever read. Each observation scores won (0/1) minus the ELO-expected win probability, in percentage points; the headline view keeps the rustier side of each match and drops matches where both players sit in the same layoff bucket. Layoff is calendar days since the player's previous played match, cross-surface by design, counted over every match in the database including exhibitions and juniors; walkovers skipped. The shipped penalty is keyed on the gap that STARTED the return, not the gap since the immediately preceding match: once a 200-day returnee plays once their current gap is three days, but their second match back is still −2.77pp, so a current-gap penalty would collapse to zero exactly where the effect is largest. Confidence intervals are match-clustered; the decay split is measured on returns of 91 days or more and applied only to that population. The favourite/underdog split is 119,692 observations since 2020. The backtest gate is variant pen on a 2023-01-01 cutoff with full-history rating state and paired bootstrap intervals against the unmodified baseline; the first gate window sat inside the fit window; the out-of-sample check (curve refit on pre-2025 data, gate scored 2025 onward) confirms it: 14 of 15 cells positive, 9 significant, all three surfaces positive.
References:
[1] M. E. Glickman, "The Glicko system," Harvard University. https://www.glicko.net/glicko/glicko.pdf (accessed Aug. 12, 2026).
[2] M. E. Glickman, "Parameter estimation in large dynamic paired comparison experiments," Journal of the Royal Statistical Society: Series C (Applied Statistics), vol. 48, no. 3, pp. 377–394, 1999.
[3] M. Sipko, "Machine learning for the prediction of professional tennis matches," MEng Computing final year project, Imperial College London, London, U.K., Jun. 2015. https://www.doc.ic.ac.uk/teaching/distinguished-projects/2015/m.sipko.pdf (accessed Aug. 12, 2026).
[4] S. A. Kovalchik, "Searching for the GOAT of tennis win prediction," Journal of Quantitative Analysis in Sports, vol. 12, no. 3, pp. 127–138, 2016, doi: 10.1515/jqas-2015-0059.
On 238,716 player-matches from 2018 to 2026, a returning player wins less than their ELO rating expects on a monotone curve: about 0.9 percentage points below expectation after 8-14 days out, 4.2 points after one to two months, and 9.8 points after more than a year, corresponding to a rating overstated by roughly 6, 30 and 73 ELO points. The deficit fades over roughly four to six matches back, which is the signature of a stale rating plus a short-lived rust component.
No, and that was the surprise. Glicko treats inactivity as uncertainty, which pulls predictions toward 50% and therefore predicts a rusty UNDERDOG should overperform their rating. In our data rusty underdogs underperform in every layoff bucket from 15-28 days onward (pooled -1.48pp, favourites -3.93pp). The effect is directional, the rusty side plays worse whether favoured or not, so we encode a penalty rather than an uncertainty term.
Unproven, and we say so in the article. The curve is measured against our own ELO rating, not against betting markets, so it shows that ratings overstate returning players, not that the market does. Encoding it improved our model's Brier score by roughly 0.15 to 0.5 percent relative, a real but small accuracy gain, confirmed out-of-sample by refitting the curve on pre-2025 data and validating on 2025 onward. Until it is tested against closing prices it is model calibration, not a betting edge.
See today's picks — published before the match, graded in public →