Something has quietly shifted in how people in the UK engage with their health before they ever book a GP appointment. Rather than reaching for NHS 111 or typing symptoms into a search engine, a growing number of patients are having full-length, back-and-forth conversations with large language models, asking things like “could this be lupus?” or “should I be worried about this lump?” The results of those conversations are shaping what people believe, how urgently they act, and in some cases, what they tell their doctors when they eventually do get a consultation.
AI self-diagnosis in the UK is not a fringe behaviour any more. A 2024 survey by the Nuffield Trust found that digital health tools, including AI-driven ones, were being used by patients in ways that outpaced any formal guidance on appropriate use. The question worth asking plainly is: how accurate is AI when it comes to interpreting human symptoms, and who is responsible when it gets it wrong?

What the research actually shows about chatbot diagnostic accuracy
The headline findings on accuracy are mixed, and I’d argue the nuances matter more than the averages. A 2023 study published in JAMA Internal Medicine found that ChatGPT performed reasonably well on standardised clinical vignettes, ranking the correct diagnosis in its top three suggestions around 72% of the time. Sounds reassuring until you consider what the other 28% of cases looked like, and that real patients presenting to chatbots rarely match the clean framing of a clinical vignette.
The more concerning pattern is directional confidence. Large language models, by design, generate fluent, authoritative-sounding text. A patient asking about chest tightness and fatigue might receive a response that confidently prioritises anxiety or acid reflux, without adequately surfacing the possibility of cardiac involvement. The model is not lying. It is doing exactly what it was trained to do: produce plausible, coherent output weighted by probability. But medicine is full of low-probability, high-consequence events that probability-weighted systems systematically underweigh.
Research from University College London published in early 2025 looked specifically at patients who had used AI tools to interpret symptoms before a GP visit. Around 34% had formed a strong prior belief about their diagnosis before the consultation, and GPs reported that challenging those beliefs added meaningful time pressure to already stretched appointments. This is a practical problem with systemic consequences.
Why certain symptom types carry higher risk
Not all AI symptom checking carries equal risk. For common, self-limiting conditions, a chatbot telling someone they likely have a cold or mild gastroenteritis is probably harmless, and might even reduce unnecessary demand on NHS services. The risk concentrates in two areas.
First, rare or atypical presentations of serious conditions. Symptoms of conditions like ME/CFS are notoriously diffuse and overlap with dozens of other diagnoses. An AI trained predominantly on mainstream clinical literature may push patients toward more common explanations and away from the right one. Second, mental health presentations. A patient describing low mood, cognitive fog, and fatigue might be told by a model that they sound stressed or sleep-deprived, when what they are experiencing is a prodrome of something requiring clinical assessment.

There is also a particular problem with supplement and lifestyle recommendations that chatbots often attach to their diagnostic suggestions. A patient convinced they have a magnesium deficiency based on an AI conversation might spend weeks self-treating before discovering the real cause of their symptoms. The information itself is not necessarily wrong; the sequencing and framing almost always is.
Where UK regulation currently stands
This is where things get genuinely complicated. The MHRA’s updated framework for software as a medical device, which I’ve covered in the context of AI-driven health apps, applies to tools that are intended to be used for a medical purpose. The word “intended” is doing enormous work here. A general-purpose large language model, such as ChatGPT or Google’s Gemini, is not marketed as a diagnostic tool. Its developers explicitly disclaim medical use. That puts it largely outside the MHRA’s current regulatory perimeter, even when patients are using it for exactly that purpose.
The Care Quality Commission, which regulates health services in England, has similarly limited reach here. CQC oversight applies to registered providers. A software product that refuses to call itself a health service sits in a gap that existing frameworks were never designed to address. The CQC’s 2025 annual report on innovation in health acknowledged the issue without offering a concrete regulatory pathway.
NHS England has issued informal guidance encouraging patients to use NHS 111 or the NHS App as first points of contact for health concerns, and to treat AI chatbots as general information tools rather than diagnostic ones. That guidance is reasonable. The gap between reasonable guidance and actual patient behaviour is, as usual, wide. You can find the NHS’s current position on digital health tools at nhs.uk.
What this means for patients navigating it now
I’d be dishonest if I said AI chatbots have no value in a health context. They can help patients organise and articulate symptoms more clearly before an appointment. They can provide useful background on how a diagnosis works, what questions to ask a GP, or what a prescribed medication does. Used in that framing, as a preparation tool rather than a diagnostic oracle, they are genuinely useful.
The problem is that the interface encourages a different behaviour. Asking a chatbot “what is causing my headaches?” and receiving a structured, confident three-paragraph response does not feel like reading a general information article. It feels like receiving a personalised assessment. That experiential difference matters, because it shapes how strongly a person holds the belief that follows.
My honest take: treat any chatbot response about personal symptoms the same way you’d treat advice from a well-read friend with no medical training. Useful context, possibly, but not a clinical opinion. If you are researching health information online more broadly and want better tools for evaluating what comes up in search results, resources like expert free SEO tools can at least help you understand which sources rank highly and why, though that is a separate skill from evaluating medical credibility.
The harder structural question
Underneath the accuracy debate is a more uncomfortable truth: patients in the UK are turning to AI in part because accessing a GP in 2026 is genuinely difficult. NHS England data shows average waiting times for a routine GP appointment still exceeding two weeks in many areas. When someone is worried and can’t get timely access to a clinician, a chatbot that responds immediately feels better than nothing. That is not a technology failure. That is a capacity failure that technology is filling imperfectly.
Regulatory frameworks that treat AI symptom use as purely a tech problem to be policed will miss this. The more productive frame is: how do we design AI health tools that are accurate about their own limitations, that actively route patients toward appropriate care rather than substituting for it, and that work within NHS pathways rather than around them? That design challenge is solvable. The regulatory appetite to mandate it is still catching up.

Leave a Reply