Peakfy

Can you trust your readiness score?

Updated 2026-08-17 · The Peakfy coaching team

Cyclist sitting on the edge of a bed in early morning light, kit half on, deciding whether to ride
The number is an opinion about your morning. You still have to decide.

The measurement underneath is often better than people assume; the score built on top of it is the part with no published validation. A 2025 review of fourteen readiness, recovery and strain scores across ten manufacturers found that none of them had published independent, peer-reviewed validation of their exact scoring formulas. Treat the number as one opinion among several, not as a verdict.

What the score is trying to do

A readiness or recovery score takes several overnight signals and compresses them into one number that answers a question you would otherwise have to answer yourself: how hard should today be?

Across the field the ingredients are fairly consistent. In that 2025 review, heart-rate variability appeared in 86% of the scores examined, resting heart rate in 79%, physical activity in 71% and sleep duration in 71%. So most scores are looking at broadly the same things. What differs is how they weigh them, over what window, and by what rule — and that is exactly the part nobody publishes.

The measurement is probably not your problem

This is the distinction that gets lost in the argument, and it matters: measuring your overnight heart rate and heart-rate variability is a different job from scoring you.

A 2025 validation study put five consumer devices against an electrocardiogram across 536 nights. The best performers tracked resting heart rate almost exactly — agreement of 0.97 to 0.98 on a concordance scale that tops out at 1, with average errors under 2%. For heart-rate variability the best device reached 0.99 agreement with roughly 6% error. The weakest device in the group averaged over 16% error on variability, with a spread wide enough that individual nights could be far off.

So the hardware is not uniformly bad, and on the better devices the raw numbers are good enough to act on. The question is what happens after those numbers go into the box.

What the review of ten manufacturers actually found

In 2025 a group of researchers catalogued fourteen composite scores across ten manufacturers — readiness, recovery, strain, energy, body resources, whatever each company calls it — and looked for the evidence behind them.

Their central finding is short: none of the manufacturers had published independent, peer-reviewed validation of their exact scoring formulas. They also found substantial differences between products in the time windows used, in how the inputs are weighted, and in the scoring method itself. And they noted that wearable-derived stress metrics did not consistently line up with what users themselves reported feeling.

That last point is the one to sit with. Two products can read the same night off the same wrist and disagree, and neither can show you why.

Why one number loses information on purpose

Compression is the whole point of a score, and it is also its weakness. Consider two mornings that both come out at 62.

On the first, you slept badly before a stressful day at work, but you are two days into a recovery week and your legs are fine. On the second, you slept normally but you are at the end of a heavy block and your resting heart rate has been drifting up for four days. Same number. Completely different sessions are correct.

A score cannot tell you which morning you are in, because the information that distinguishes them was thrown away in the compression. That is not a flaw in the implementation. It is what a single number is.

And acting on the signal is not automatically better

Here is the part that rarely makes it into the marketing. Adjusting your training according to heart-rate variability has been tested against simply following a plan, and the results are more modest than you would expect.

Pooled analyses find that variability-guided training is not superior for raising maximal oxygen uptake. It shows a medium-sized effect on submaximal physiological markers and only a small, statistically unconvincing effect on performance itself. There is also a quiet detail in the data: the variability-guided groups often ended up doing fewer hard sessions than the fixed-plan groups. Some of the benefit, where it appears, may simply be from backing off more often rather than from the measurement being clever.

The honest caveats: few studies, most shorter than eight weeks, and not many of them on well-trained athletes.

How to use a score without being ruled by it

None of this makes the number useless. It makes it one input.

Use it as a prompt, not a verdict. A low score is a good reason to check in with yourself before you decide, and a bad reason to skip a session on its own.

Look for agreement, not for a threshold. One signal is noise. When the score, your resting heart rate and how the first ten minutes feel all point the same way, that is information.

Ask what moved it. If a product can tell you which input pushed the number down, the number becomes useful. If it cannot, you are being asked to trust an oracle.

Watch the direction, not the digit. Your own trend over a week beats today's absolute value, and it beats comparing yourself with anyone else. That is the same argument as in reading your morning signals — and it is why a single reading was never the point.

Your numbers, explained the moment you ask.

Peakfy shows a readiness number too, so it is fair to ask why that is any different. The difference is that you can open it: the number sits next to the things that produced it, and you can ask the coach what moved it and what that changes about today's session. When a value has not actually been computed, nothing is shown — Peakfy does not fill the gap with a plausible-looking figure.

A score you can interrogate is a summary. A score you cannot is a black box that happens to be numeric. We would rather be the first thing, and we would rather you argue with it.

Questions riders ask

Accuracy is two questions. The overnight measurements underneath — resting heart rate and heart-rate variability — are reasonably accurate on the better devices, with agreement above 0.97 against an electrocardiogram in one 2025 study. The composite score built from them is a different matter: a review of fourteen such scores across ten manufacturers found none with published independent validation of its formula.
Not on the score alone. Use it as a reason to check other signals — how the first ten minutes feel, whether your resting heart rate has been drifting, how motivated you are. If two or three point the same way, adjust. If only the score is low and everything else is normal, start the session and decide after the warm-up.
Yes, as a trend against your own baseline rather than as a daily verdict. But keep expectations realistic: pooled research finds that guiding training by variability is not clearly better than following a well-built plan for raising maximal oxygen uptake, and part of the benefit that does appear may come from simply doing fewer hard sessions.

General sport-science guidance for healthy riders. Not medical advice.


← All articles

Your numbers, explained the moment you ask.

Peakfy computes your training load, readiness, zones and FTP estimate from your own rides — then a coach you can talk to explains what each one means for tomorrow. iOS and Android, six languages, built in Europe.

See how Peakfy works