Peakfy

Can you trust your AI cycling coach?

Updated 2026-08-17 · The Peakfy coaching team

Cyclist at a kitchen table in the morning, phone face down beside a mug, looking out of the window
The question is not whether it sounds confident. It is whether you can check it.

Trust the ones that can show their work, and stay sceptical of the rest — because the evidence that any of them work over time does not yet exist. A 2025 scoping review of twenty studies on AI coaching models found that none had evaluated long-term behaviour change or sustained physical improvement, and rated the median study 2.5 out of 5 for evaluation rigour. That makes explainability the only thing you can actually check yourself.

Everything in the category now describes itself the same way

Open any cycling training app in 2026 and you will meet the same three words: adaptive, personalised, AI-powered. That was a meaningful description about four years ago. Now it is close to universal — which means it has stopped carrying information.

When every product in a category makes the same claim, the claim cannot help you choose. So the useful question is not does this app use AI. It is: when this thing tells me to do something, can I find out why?

That question turns out to have a research literature behind it, and the literature is not flattering.

What has actually been tested — and what has not

In 2025 a group of researchers went looking for every study that evaluated an AI-based exercise or health coaching model. They found twenty, published between March 2023 and July 2025, and then assessed how well those studies were done.

The findings are worth reading slowly. Eight in ten relied on people rating outputs — experts scoring a plan, users filling in a survey. Only 40% used real user data, and only 40% evaluated anything in a real or simulated usage context. Fewer than half reported whether their raters agreed with each other. The median study scored 2.5 out of 5 on the reviewers' rigour scale; 55% were classed as low-rigour, and only 10% reached the top mark.

And then the finding that should change how you read every marketing page in this category: not one of the twenty studies evaluated long-term behaviour change or sustained physical improvement. Nobody has published evidence that using one of these tools makes you fitter over a season. That is not the same as saying they do not work. It means the question has not been asked properly yet.

The reviewers also noted that transparency and explainability received minimal attention across the whole literature. Their own summary of the field's state is blunt: the evaluation is "fragmented and methodologically weak".

Ask the same question twice

Here is a test you can run yourself, and it comes straight out of a 2026 preprint that did exactly this at scale.

The researchers took six scenarios and asked a large language model to produce an exercise prescription for each — then repeated the identical request twenty times per scenario, for 120 outputs in total, changing nothing at all between runs.

The wording held up well. Semantic similarity across repeats came out between 0.879 and 0.939, so each answer read like the same answer. The numbers were a different story: the quantitative parts moved between runs, and intensity moved most. In the resistance-training prescriptions, between 10% and 25% of outputs expressed intensity in a way that could not even be classified.

Two honest caveats, because they matter: this was a preprint rather than a peer-reviewed paper, and it tested exercise prescription in a clinical framing, not cycling training plans. But the mechanism it exposes is not specific to any field. Confident prose and a stable number are two different things, and a system can produce the first without the second.

So: ask your coach the same question on two quiet days, with nothing changed in your data. If the recommendation swings and nothing explains the swing, you have learned something useful for the price of two minutes.

Where the output tends to be weakest

A separate study handed AI-generated training plans to ten experienced coaches — an average of 9.6 years in the job — and had them rate the plans against 27 criteria on a five-point scale.

The verdict was middling rather than damning. Across three models, only five of the criteria ever reached 4.5 or higher. The consistent weak point was specific: the plans were vague about intensity — how hard, how close to the limit, how much load. The authors' reading was that the models leaned towards being safe rather than being effective, and that what they produce is best treated as a template that still needs adjusting.

That study looked at strength training, not cycling, so do not import its scores. Import the failure mode. Vagueness clusters exactly where the training effect lives. A plan can name the right session, the right day and the right duration, and still be useless if it will not commit to how hard.

A ten-minute check on any AI coach

None of the above requires you to trust a review, including this one. Each point below is something you can test in a free trial.

Ask where a number came from. Not what it means — where it came from. A good answer names your data: this ride, that week, this trend. A bad answer restates the number in longer words.

Ask the same question twice. Two days apart, data unchanged. Stable is good. Different is fine if something changed and it can tell you what.

Push back once. Say "this session does not fit me" and watch what moves. If the plan changes, you have a coach. If only the tone changes, you have a chat window in front of a fixed plan.

Give it an incomplete day. Ride without a heart-rate strap, or cut a session short. A tool that says "I cannot compute that today" is more trustworthy than one that always has a figure ready.

Ask it what it does not know. Anything that never expresses uncertainty is not being careful, and you will find out which the hard way.

Why an unexplained change costs more than a wrong one

Riders forgive a plan that gets it wrong far more readily than a plan that changes without saying why — and that instinct is correct.

A recommendation you can follow the reasoning of is a recommendation you can argue with. You can notice that it weighted a bad night's sleep too heavily, say so, and be right. That exchange is worth something even when the software was wrong, because you end up understanding your own training better.

A recommendation that simply appears teaches you nothing, and worse, it removes your ability to tell a sensible adaptation from a glitch. Both look identical from the outside: the number moved. That is the same argument as the one about trusting a readiness score — a figure you cannot open is not a summary, it is an oracle.

What good looks like from here

The honest position in 2026 is neither dismissal nor enthusiasm. These tools are genuinely useful at the things software is good at: being available at six on a Sunday morning, computing the same way every time, never getting tired of your questions. That case is made properly in our comparison of an AI coach and a human coach.

What has not been earned yet is the right to be followed without question. Until somebody publishes evidence that one of these tools changes what a rider can do over a season, an AI coach should be a very well-informed second opinion that supports your own judgement — not a replacement for it. You still know things about your legs, your week and your life that no dataset contains.

The standard to hold the category to is simple: coaching that shows its work. Everything else is a slogan that every product already uses.

Your numbers, explained the moment you ask.

This is the standard we built Peakfy to meet, so it is fair to hold us to it. Every recommendation in Peakfy arrives with the data it came from and a one-sentence reason in plain language — not a paragraph of hedging, one sentence you can check against your own ride. And when you tell the coach "this does not fit me", the plan changes; the objection is an input, not a complaint the software absorbs politely.

Where a value has not actually been computed, nothing is shown. We would rather leave a gap than fill it with a plausible-looking figure, because a number you cannot trace is exactly the thing this article is about.

Questions riders ask

Treat them as a well-informed starting point rather than an instruction. A 2025 scoping review of twenty studies on AI exercise and health coaching models found that none had evaluated long-term behaviour change or sustained physical improvement, and rated the median study 2.5 out of 5 for evaluation rigour. The evidence that these tools work over a season has not been produced yet, so the practical test is whether a tool can show you where its recommendations come from.
Run four checks during a free trial. Ask where a number came from and see whether it names your own data. Ask the same question twice on two days with nothing changed, and see whether the answer holds. Tell it a session does not suit you and see whether the plan actually changes. And give it a day with incomplete data — a tool that admits it cannot compute something is more trustworthy than one that always has a figure ready.
Not necessarily. A 2026 preprint repeated an identical exercise-prescription request twenty times across six scenarios and found the wording highly consistent — semantic similarity of 0.879 to 0.939 — while the quantitative parts varied, with intensity varying most. In some cases the intensity given could not even be classified. The study was not cycling-specific, but the lesson generalises: text that reads consistently is not the same as a number that is stable.

General sport-science guidance for healthy riders. Not medical advice.


← All articles

Your numbers, explained the moment you ask.

Peakfy computes your training load, readiness, zones and FTP estimate from your own rides — then a coach you can talk to explains what each one means for tomorrow. iOS and Android, six languages, built in Europe.

See how Peakfy works