Trust the ones that can show their work, and stay sceptical of the rest — because the evidence that any of them work over time does not yet exist. A 2025 scoping review of twenty studies on AI coaching models found that none had evaluated long-term behaviour change or sustained physical improvement, and rated the median study 2.5 out of 5 for evaluation rigour. That makes explainability the only thing you can actually check yourself.
Everything in the category now describes itself the same way
Open any cycling training app in 2026 and you will meet the same three words: adaptive, personalised, AI-powered. That was a meaningful description about four years ago. Now it is close to universal — which means it has stopped carrying information.
When every product in a category makes the same claim, the claim cannot help you choose. So the useful question is not does this app use AI. It is: when this thing tells me to do something, can I find out why?
That question turns out to have a research literature behind it, and the literature is not flattering.
What has actually been tested — and what has not
In 2025 a group of researchers went looking for every study that evaluated an AI-based exercise or health coaching model. They found twenty, published between March 2023 and July 2025, and then assessed how well those studies were done.
The findings are worth reading slowly. Eight in ten relied on people rating outputs — experts scoring a plan, users filling in a survey. Only 40% used real user data, and only 40% evaluated anything in a real or simulated usage context. Fewer than half reported whether their raters agreed with each other. The median study scored 2.5 out of 5 on the reviewers' rigour scale; 55% were classed as low-rigour, and only 10% reached the top mark.
And then the finding that should change how you read every marketing page in this category: not one of the twenty studies evaluated long-term behaviour change or sustained physical improvement. Nobody has published evidence that using one of these tools makes you fitter over a season. That is not the same as saying they do not work. It means the question has not been asked properly yet.
The reviewers also noted that transparency and explainability received minimal attention across the whole literature. Their own summary of the field's state is blunt: the evaluation is "fragmented and methodologically weak".
Ask the same question twice
Here is a test you can run yourself, and it comes straight out of a 2026 preprint that did exactly this at scale.
The researchers took six scenarios and asked a large language model to produce an exercise prescription for each — then repeated the identical request twenty times per scenario, for 120 outputs in total, changing nothing at all between runs.
The wording held up well. Semantic similarity across repeats came out between 0.879 and 0.939, so each answer read like the same answer. The numbers were a different story: the quantitative parts moved between runs, and intensity moved most. In the resistance-training prescriptions, between 10% and 25% of outputs expressed intensity in a way that could not even be classified.
Two honest caveats, because they matter: this was a preprint rather than a peer-reviewed paper, and it tested exercise prescription in a clinical framing, not cycling training plans. But the mechanism it exposes is not specific to any field. Confident prose and a stable number are two different things, and a system can produce the first without the second.
So: ask your coach the same question on two quiet days, with nothing changed in your data. If the recommendation swings and nothing explains the swing, you have learned something useful for the price of two minutes.
Where the output tends to be weakest
A separate study handed AI-generated training plans to ten experienced coaches — an average of 9.6 years in the job — and had them rate the plans against 27 criteria on a five-point scale.
The verdict was middling rather than damning. Across three models, only five of the criteria ever reached 4.5 or higher. The consistent weak point was specific: the plans were vague about intensity — how hard, how close to the limit, how much load. The authors' reading was that the models leaned towards being safe rather than being effective, and that what they produce is best treated as a template that still needs adjusting.
That study looked at strength training, not cycling, so do not import its scores. Import the failure mode. Vagueness clusters exactly where the training effect lives. A plan can name the right session, the right day and the right duration, and still be useless if it will not commit to how hard.
A ten-minute check on any AI coach
None of the above requires you to trust a review, including this one. Each point below is something you can test in a free trial.
Ask where a number came from. Not what it means — where it came from. A good answer names your data: this ride, that week, this trend. A bad answer restates the number in longer words.
Ask the same question twice. Two days apart, data unchanged. Stable is good. Different is fine if something changed and it can tell you what.
Push back once. Say "this session does not fit me" and watch what moves. If the plan changes, you have a coach. If only the tone changes, you have a chat window in front of a fixed plan.
Give it an incomplete day. Ride without a heart-rate strap, or cut a session short. A tool that says "I cannot compute that today" is more trustworthy than one that always has a figure ready.
Ask it what it does not know. Anything that never expresses uncertainty is not being careful, and you will find out which the hard way.
Why an unexplained change costs more than a wrong one
Riders forgive a plan that gets it wrong far more readily than a plan that changes without saying why — and that instinct is correct.
A recommendation you can follow the reasoning of is a recommendation you can argue with. You can notice that it weighted a bad night's sleep too heavily, say so, and be right. That exchange is worth something even when the software was wrong, because you end up understanding your own training better.
A recommendation that simply appears teaches you nothing, and worse, it removes your ability to tell a sensible adaptation from a glitch. Both look identical from the outside: the number moved. That is the same argument as the one about trusting a readiness score — a figure you cannot open is not a summary, it is an oracle.
What good looks like from here
The honest position in 2026 is neither dismissal nor enthusiasm. These tools are genuinely useful at the things software is good at: being available at six on a Sunday morning, computing the same way every time, never getting tired of your questions. That case is made properly in our comparison of an AI coach and a human coach.
What has not been earned yet is the right to be followed without question. Until somebody publishes evidence that one of these tools changes what a rider can do over a season, an AI coach should be a very well-informed second opinion that supports your own judgement — not a replacement for it. You still know things about your legs, your week and your life that no dataset contains.
The standard to hold the category to is simple: coaching that shows its work. Everything else is a slogan that every product already uses.
Your numbers, explained the moment you ask.
This is the standard we built Peakfy to meet, so it is fair to hold us to it. Every recommendation in Peakfy arrives with the data it came from and a one-sentence reason in plain language — not a paragraph of hedging, one sentence you can check against your own ride. And when you tell the coach "this does not fit me", the plan changes; the objection is an input, not a complaint the software absorbs politely.
Where a value has not actually been computed, nothing is shown. We would rather leave a gap than fill it with a plausible-looking figure, because a number you cannot trace is exactly the thing this article is about.
Questions riders ask
Keep reading
- AI cycling coach vs human coach: an honest comparison
- Can you trust your readiness score? What the research says
- How to increase your FTP: what actually moves the number