← All articles

AI in medicine: where it helps doctors, where it fails and how to read claims about its accuracy

Topics: AI, Medicine

A doctor's office in the evening: a chest X-ray on a monitor with semi-transparent marking frames from a program, one frame highlighted in yellow, a second monitor with two smooth curves, a stethoscope, a tablet and a cup

Short answer: the US FDA list as of September 4, 2026 contains 1,614 AI-enabled medical devices, and 1,230 of them (76%) belong to radiology, for example they work with X-rays, CT, MRI and mammography. In Moscow, according to a statement by the city government on June 3, 2026, computer vision services help doctors, and the final decision stays with the doctor. The list is incomplete, and it does not show how often these devices are used. The strongest evidence of benefit I found is the Swedish mammography trial MASAI with 106 thousand women: with AI, more cancers were found, doctors had to read 44% fewer images, and the share of false alarms did not grow. Three kinds of failure are also documented, and for anyone planning to adopt AI they matter just as much. A model may work worse in another hospital than it did for its developer (this happened with a sepsis model in Michigan). A doctor may follow a wrong suggestion or, over time, do worse without it. A language model answers well on its own, but in two randomized studies it did not noticeably help the people who used it: for doctors, a gain of 2 percentage points was not statistically confirmed, and ordinary people did no better than those who used a source of their own choice. Below I go through each point with numbers, explain why the word "accuracy" means almost nothing without the share of sick people, and give a checklist for a patient, for a clinic owner and for a company outside medicine.

Where AI already works and how many AI medical devices are approved

USA. The US Food and Drug Administration (FDA) keeps a public list of AI-enabled medical devices. On the page updated on September 4, 2026, I counted 1,614 rows. Of these, 1,230 belong to radiology (X-ray, CT, MRI, mammography), 154 to cardiology and 73 to neurology. The list contains 335 authorization decisions dated 2025. The FDA states plainly that the list is incomplete: devices get onto it when AI is mentioned in their description.

Russia. Under Article 38 of Russia's federal law on the fundamentals of public health protection (Law No. 323-FZ), software that its manufacturer intends for diagnosis or treatment counts as a medical device, and medical devices may be used in Russia only after state registration. In November 2025, Russia's Minister of Health Mikhail Murashko said that 48 AI programs were listed in the register of Roszdravnadzor, the Russian healthcare regulator. According to GxP News as of September 2025, 43 of the 48 registered AI devices were developed in Russia.

Moscow. According to a statement by the city government on June 3, 2026, Moscow healthcare uses more than 60 computer vision services (programs that look for signs of disease in medical images) across 45 clinical areas, they have processed more than 30 million examinations, and the average time doctors spend reading images and writing reports has fallen by about 30%. The same statement says that the final decision stays with the doctor. These figures come from the city itself; I found no independent verification of them.

The overall picture from this data: most AI-enabled devices on the FDA list work with medical images. How often they are used in clinics cannot be seen from the lists; there are only separate statements like the Moscow one.

What has the strongest evidence: the Swedish mammography trial

The simple kind of evidence looks like this: a model, on a set of old images, found a disease as well as doctors did. This evidence is weak: an archive has no live flow of patients, no tired doctor and no answer to what happened to the people afterwards. Strong evidence comes from a randomized trial: people are randomly split into two groups, one is examined with AI, the other as usual, and the outcomes are compared.

Such a trial was run in Sweden. MASAI randomly assigned 105,934 women who were going through routine screening (a planned examination of healthy people) for breast cancer. In the usual group, each image was read by two doctors, as is standard in Swedish screening. In the AI group, the program assessed the risk for each image: low-risk images were read by one doctor, high-risk images by two, and the program marked suspicious areas. The results were published in The Lancet on January 30, 2026.

What happened:

  • More cancer was found. According to Lund University, screening with AI detected 29% more breast cancers.
  • The share of false alarms did not grow. Sensitivity (the share of sick people in whom the examination found cancer) rose from 73.8% to 80.5%. Specificity (the share of healthy people whom the examination correctly found healthy) was the same in both groups, 98.5%.
  • Doctors read fewer images. The number of image readings by doctors fell by 44%. This measures image reading specifically, not all of a doctor's work.

The main measure of the trial was interval cancer: cancer found in the period before the next planned examination. It can be a tumour that was missed on the image or a tumour that grew quickly. With AI there were 82 such cases, without AI 93, that is 1.55 versus 1.76 per thousand women. Lund University, in its statement on the results, writes about 12% fewer interval cancers.

This is where the difficulty starts, and it is the reason to read studies rather than headlines. The trial was designed to prove that AI is not worse, and that was proven. But the 12% difference is not statistically confirmed: the ratio of rates is 0.88, but given random variation (the 95% confidence interval) the true value may lie between 0.65 and 1.18, so the effect could turn out to be a noticeable decrease or a small increase (p = 0.41). Also, 12% is a relative difference. In absolute numbers it is 0.21 cases per thousand women, about two cases per ten thousand women screened.

How to read this: the trial showed that screening with AI finds more cancer (the difference in sensitivity is statistically confirmed) with the same share of false alarms and takes almost half of the readings off doctors, with no more interval cancers. That the additional findings save lives, this trial has not shown yet; that takes many years of follow-up. And this is the result of one country with its own screening procedure: transferring it to a different procedure and different equipment needs its own check.

Why "95% accuracy" says almost nothing

News about medical AI uses the word "accuracy". For example, the same city government statement of June 3, 2026 says that in some areas the accuracy of AI services exceeds 95%, but does not say which measure is meant. There are several measures, and for a patient and for a clinic they mean different things.

Three measures worth knowing:

  • Sensitivity: of all sick people, what share the system found. Low sensitivity means misses.
  • Specificity: of all healthy people, what share the system correctly left alone. Low specificity means false alarms.
  • Share of true alarms (in medicine it is called positive predictive value): of all people the system flagged, what share are actually sick. This is what a doctor who receives an alarm sees.

The third measure depends on how many sick people there are among those being checked. Here is a hypothetical example. Take a system with 90% sensitivity and 90% specificity and use it to check 10,000 people in two places.

1% sick (screening healthy people) 20% sick (people with symptoms)
Sick people out of 10,000 100 2,000
Sick people the system found 90 1,800
False alarms in healthy people 990 800
Share of true alarms 8% 69%

The same system with the same "accuracy" in the first case raises 12 alarms for every truly sick person, while in the second almost seven out of ten alarms are true. So a vendor's figure says nothing about your clinic until you know the share of sick people it was obtained on and the share of sick people you have.

A real example of this arithmetic comes below, in the sepsis story: there, 12 out of every 100 alarms were true.

What happens when a model leaves the lab

Sepsis in Michigan. Sepsis is a life-threatening reaction of the body to an infection, and it is important to spot it early. Epic, a maker of an electronic medical records system, built into it a model that uses patient data to warn about the risk of sepsis. According to the study authors, by 2021 it was used in hundreds of US hospitals. The developer reported an AUC of 0.76 to 0.83 for it. AUC is the probability that the model gives a randomly chosen sick person a higher score than a randomly chosen healthy person: 0.5 corresponds to a coin toss, 1 to perfection.

Doctors at the University of Michigan checked the model on their own patients: 27,697 people, 38,455 hospitalizations from December 2018 to October 2019, with sepsis in 7% of hospitalizations. The result was published in JAMA Internal Medicine in June 2021:

  • AUC in their hospital was 0.63, lower than the developer claimed;
  • the model missed 67% of patients with sepsis (1,709 out of 2,552);
  • it raised an alarm for 18% of all hospitalizations, and 12% of the alarms were true: to find one future sepsis patient, a doctor had to check eight patients.

The lesson of this story is wider than one model. When a model is moved to another hospital with different patients, different data recording and different treatment routines, its quality may change, as it did here. External validation means checking on patients from other hospitals or another period than the data the model was built and tuned on. Here it was done by doctors at the University of Michigan, not by the developer. Without it, the developer's figure remains a promise.

A human next to AI is not always a safeguard

The answer to the risks of medical AI is the phrase "the final decision is up to the doctor"; the Moscow city government statement says this too. In the studies below, the "doctor plus AI" combination behaved differently from the doctor and the AI separately, so it has to be checked separately.

Blind trust in a suggestion. At the University of Cologne, 27 radiologists assessed 50 mammograms with an AI suggestion. In 12 cases the suggestion was deliberately made wrong. The result was published in the journal Radiology in May 2023. When the suggestion was correct, inexperienced doctors assessed 79.7% of images correctly; when it was wrong, 19.8%. For the most experienced doctors it was 82.3% and 45.5%. This is an experiment, not clinical work, but it shows that a wrong suggestion can mislead even an experienced doctor.

Loss of skill. In Poland, at four endoscopy centres taking part in the ACCEPT trial, AI that highlights polyps during colonoscopy (an examination of the bowel with a camera) was introduced at the end of 2021. The researchers compared colonoscopies that the same centres performed without AI in the three months before it appeared and in the three months after. The measure was the adenoma detection rate: the share of colonoscopies in which the doctor found at least one adenoma, that is, a polyp that can turn into cancer. It fell from 28.4% to 22.4%. The study was published in The Lancet Gastroenterology & Hepatology in 2025. It is observational, so it does not directly prove the cause, but the authors write that constant work with AI may reduce a doctor's results when AI is not there.

Language models for doctors. In a US randomized study, 50 doctors (26 experienced doctors and 24 residents) worked through descriptions of clinical cases in November and December 2023. Half could use GPT-4, half used ordinary reference resources. The median diagnostic reasoning score (half of the doctors scored higher, half lower): 76% with the model and 74% without it; the difference was not statistically confirmed. The model alone, without a doctor, scored a median of 92%, 16 percentage points more than the group with reference resources. The result was published in JAMA Network Open in October 2024. The model alone did better, but doctors gained almost nothing from it.

Language models for ordinary people. Oxford researchers gave 1,298 participants ten medical scenarios: work out what is wrong with the person and decide where to go and how urgently. Some used a language model (GPT-4o, Llama 3 or Command R+), others any source of their choice. The models alone, without people, named a relevant condition in 94.9% of cases. People using these same models named it in fewer than 34.5% of cases and chose the right action in fewer than 44.2%, and this is no better than those who used their own source. The paper appeared in Nature Medicine on February 9, 2026. The authors write that standard tests of medical knowledge do not predict such failures, and recommend testing models with real people before releasing them to patients.

The models in both studies are no longer the newest: GPT-4 in the late-2023 trial; GPT-4o, Llama 3 and Command R+ in the Oxford one. But the conclusion is not about a specific model, it is about how to check: a good answer from the model and a good decision by the person using it are two different measurements.

Who watches AI after approval

Medicines are monitored for side effects after they reach the market, and expert Boris Zingerman, in the GxP News article I cite below, compared the oversight that medical AI needs to that. In Russia, this kind of monitoring of AI began to be built in 2025-2026.

In September 2025, according to the industry outlet GxP News, Roszdravnadzor planned to automate the collection of data on how medical AI systems work: how often they are used, where they fail and deviate from the norm. Experts interviewed by the outlet noted that the agency did not have a platform for this at the time. On February 10, 2026, Roszdravnadzor issued order No. 123: manufacturers of AI software that counts as a medical device and can transmit data automatically must send information about the device's data and results of its work to the agency's information system. The order is in force from May 23, 2026 to December 31, 2027. How this monitoring works in practice, the sources I found do not yet say.

For the reader this leads to a simple point: approval of a device does not tell you whether a trial like the Swedish one, which compared patient outcomes, was done, or how the device works in your clinic. You have to ask about and check this separately.

What to do

For a patient

  1. A chatbot is fine for preparation, not for decisions. In my view, it is handy for making sense of unfamiliar words in a medical report or for drawing up a list of questions for the doctor. There is no need to send the bot the whole report with your name, date of birth and insurance number: for a question, the term itself or a phrase without personal data is enough. In the Oxford study above, people with a chatbot worked out what was wrong and what to do no better than those who used a source of their choice.
  2. Do not discuss danger signs with a bot. Chest pain, sudden weakness or numbness in an arm or leg, a drooping face, speech problems, severe difficulty breathing: call 103 or 112 (Russian emergency numbers). Time spent chatting with a bot in such cases is lost.
  3. If AI looked at your image, ask who signed the report. Moscow says that in its services the final decision is up to the doctor. If you are given a report without a doctor's signature, that is a reason to ask.

For a private clinic owner before buying AI

  1. Check the registration. Ask the vendor for the number of the medical device registration certificate and find it in the medical device register on the Roszdravnadzor website.
    • Who does it: the clinic manager or a lawyer.
    • How to check: the register has an entry with the same name and the same intended use you are buying it for (for example, mammography, not lung CT).
  2. Ask for three figures and recalculate them for yourself. Sensitivity, specificity and which patients they were obtained on, with separate data from checks at clinics where the system was not built. Then calculate the share of true alarms for your share of sick patients, as in the table above.
    • Who does it: the chief physician together with a doctor of the relevant specialty.
    • How to check: the calculation shows how many alarms a day a doctor will get and how many of them will be real. If the vendor names a single "accuracy" figure and cannot break it down into these measures, that is a reason not to buy until they do.
  3. A pilot on your own patient flow. For two to three months, record for each examination: what the AI said, what the doctor said, and what was confirmed later (by biopsy, a repeat examination, another doctor). Separately, time how long it takes to read one examination and write the report, before the pilot and during it.
    • Who does it: the head of the department.
    • How to check: the number of alarms per day, the share of true alarms, the number of findings the doctor missed without AI, and the minutes per examination.
  4. Protection against blind trust and loss of skill. My advice based on the studies above: have the doctor write their own report first, then look at the suggestion and record it if they changed their mind. Once a quarter it is useful to compare the doctor's own results on examinations without AI with what they were before AI was introduced.
    • Who does it: the head of the department.
    • How to check: the share of cases where the doctor changed their mind after the suggestion and how many of these changes turned out to be correct; the doctor's results without AI do not fall.
  5. Patient data. Images and reports are health information. If the service sends them to someone else's server, show the contract and the data transfer scheme to a lawyer before launch: they will tell you what consents and protection measures are needed.

How to calculate payback. A doctor's freed hours bring money in one of two ways, and you should count one of them, not both at once. First: you pay for fewer hours, and then benefit = hours saved × cost of an hour including taxes. Second: the doctor does additional examinations in those hours, and then benefit = number of additional examinations × the clinic's income per examination minus its costs, excluding the doctor's salary (it is already paid). Costs = the monthly service fee, including the fee for additional examinations, plus one-off staff hours for setup and the pilot. If the doctor simply finishes earlier, that is spare time, not savings.

An example with hypothetical numbers; plug in your own. A doctor reads 1,500 examinations a month and writes a report on each, at 6 minutes per examination, which is 150 hours. If this reading and reporting time falls by 30% (this is the average Moscow reports for its services; yours may differ), 45 hours are freed. The service costs 50 rubles per examination, that is 75,000 rubles a month.

  • First way: at a doctor's hourly cost of 1,500 rubles including taxes, the 45 hours saved give 67,500 rubles a month. That is less than the service fee, so it does not pay off.
  • Second way: with a 30% reduction, reading and reporting on one examination with AI takes 4.2 minutes, and in the 45 freed hours the doctor reports on about 640 more examinations. If the clinic earns 300 rubles on each above its costs (excluding the doctor's salary), that is 192,000 rubles. The service fee is now paid for 2,140 examinations, 107,000 rubles. That leaves 85,000 rubles a month. If setup and the pilot took 40 staff hours at 1,500 rubles, the one-off 60,000 rubles are recovered in the very first month.
  • With income of 150 rubles per examination, the second way gives 96,000 rubles against a fee of 107,000 rubles, and the service no longer pays off.
  • The second way works if there are patients for the additional examinations. If there are none, only the first way remains.

The benefit of diseases that would have been missed without AI cannot honestly be converted into rubles, but it is worth recording as a separate line after the pilot.

For a company outside medicine

In medicine, AI errors are analysed in published studies with numbers. In my view, four conclusions from them also apply to a company adopting AI in its own work.

  1. Measure the "human plus AI" combination, not AI on its own. Doctors with GPT-4 got a gain of 2 percentage points that was not statistically confirmed, and ordinary people with chatbots did no better than the rest, even though the models themselves answered well. In your own pilot, compare the results of employees with and without AI on the same tasks.
  2. Re-check vendor figures on your own data. The sepsis model in another hospital showed an AUC of 0.63 instead of the claimed 0.76-0.83. A demo on someone else's examples does not replace a trial on yours.
  3. Count the alarms people will have time to handle. The sepsis model raised an alarm for almost every fifth hospitalization, and for each truly sick patient eight patients had to be checked. Handling alerts becomes a separate job. Before launch, decide who handles the alerts and how many a day they can get through.
  4. Watch your employees' skills. If AI takes over part of the work, check once a quarter how people manage without it, especially on tasks where the AI may fail.

How to measure the benefit of AI in hours before and after, rather than in the number of closed tasks, I covered in the article Developers and analysts working with AI: how their work differs and how to check the results. For where AI has already made discoveries in science and where its results are still in question, see the article AI in science in 2026.

Summary

As of autumn 2026, three quarters of the AI-enabled medical devices on the FDA list belong to radiology. In Moscow, according to the city government as of June 2026, computer vision services have processed more than 30 million examinations, and there, according to the city government, the final decision stays with the doctor. This data does not show how often and which systems are used in clinics in general. The strongest evidence of benefit I found comes from the Swedish mammography trial: more cancer found, the same share of false alarms and 44% fewer image readings by doctors, while the 12% reduction in interval cancers is not statistically confirmed. The failures are documented just as thoroughly. The sepsis model performed worse than claimed in another hospital, and doctors in an experiment followed wrong suggestions. In the Polish centres, doctors' results without AI fell after it was introduced (the study is observational and does not directly prove the cause), and good answers from language models did not lead to noticeably better decisions by people in two studies. In my view, the main question to ask about any medical AI is this: what changed for patients when real doctors used the system on a real patient flow, and on whom was this checked.

I work on AI agents and automation. If you would like to see my projects or discuss your own task, take a look at my portfolio.

Sources