AI in science in 2026: what it has already discovered and where it still gets things wrong

The short answer: in several fields AI has already produced recognised results. Predicting the shape of proteins earned the 2024 Nobel Prize in Chemistry, the European weather forecasting centre has been issuing a forecast from a machine learning model since 2025, a drug found with the help of AI has reached phase three of trials, and models have started solving open mathematical problems. In a survey of scientists in the US and the UK in the summer of 2026, almost half said they use AI every day. At the same time, a headline discovery by an AI agent that Anthropic announced in September is disputed and was not reproduced in repeat runs, and checking what AI produces, according to the scientists, eats up a noticeable part of the time saved. Below I go field by field through what has been confirmed and what so far has only been claimed.
Where AI is already recognised: proteins and weather
Proteins. A protein works in the cell thanks to its three-dimensional shape, and determining that shape through experiments is slow and expensive. In 2020 Demis Hassabis and John Jumper of Google DeepMind presented AlphaFold2, a model that predicts this shape from the protein's composition. According to the Nobel Committee, it has been used to predict the structure of virtually all 200 million proteins known to researchers, and as of October 2024 more than two million people from 190 countries had used the model. In 2024 Hassabis and Jumper received the Nobel Prize in Chemistry together with David Baker, who was recognised for computational protein design. The Physics prize that same year went to John Hopfield and Geoffrey Hinton for discoveries on which machine learning with neural networks is built.
Weather. On 25 February 2025 the European Centre for Medium-Range Weather Forecasts (ECMWF) brought into operational use AIFS, a system trained on historical weather data. According to the centre, it outperforms the best physics-based models on many measures, the gain on tropical cyclone tracks reaches 20%, and a single forecast uses roughly a thousand times less energy. The centre has not switched off its physics-based model: both run in parallel, and the physics-based one computes the weather in more detail, on a grid of points 9 km apart versus 28 km for AIFS.
Drugs: an AI-derived drug has reached phase three
Using its AI systems, the company Insilico Medicine chose a target for treating idiopathic pulmonary fibrosis and generated a molecule for it. A target is a protein in the body that a drug is meant to act on, in this case the protein TNIK. Idiopathic pulmonary fibrosis is a disease in which lung tissue becomes scarred and breathing gets harder and harder, and the cause is unknown. The drug was named rentosertib.
The phase two trial involved 71 patients in 22 clinics in China, treatment lasted 12 weeks, and the results were published in Nature Medicine in 2025. In the group on a dose of 60 mg a day, the volume of air a person can breathe out (forced vital capacity) rose by 98.4 ml on average, while in the placebo group it fell by 20.3 ml. The main goal of this phase was safety, and the lung measures were secondary.
On 7 July 2026 the company announced the start of phase three. According to its plan, the study will enrol 320 patients at 47 centres in China and follow them for 52 weeks, and the main measure will be how quickly their exhaled volume declines. There are no phase three results yet. Rentosertib has not become an approved drug: even a successful phase three only provides data, and the decision to authorise a drug is made by the regulator. AI took part in finding the target and the molecule, while trials in humans proceed in the usual way, phase by phase.
Mathematics: AI has started solving Erdős problems
The Hungarian mathematician Paul Erdős left behind more than a thousand problems. According to Quanta Magazine as of 3 August 2026, the catalogue of his problems maintained by the British mathematician Thomas Bloom lists 565 solved and 652 open ones, and in 2026 they began to be used to test what models can do.
- On 20 May 2026 OpenAI reported that its internal model had found a counterexample to Erdős's 1946 unit distance conjecture. The conjecture concerns how many pairs of points on a plane can be the same distance apart, and mathematicians generally considered it true. The proof was checked by nine leading mathematicians. Timothy Gowers said that if a human had been the author, he would have recommended the work for publication in the Annals of Mathematics without hesitation.
- In May 2026 a Google DeepMind agent independently solved 9 of the 353 open Erdős problems it attempted. According to the team, each problem cost a few hundred dollars.
There are caveats here too. The problems vary greatly in difficulty and significance. One problem (number 333) was announced by a participant on the catalogue's website as solved with the help of AI, but a few hours later another participant pointed out that Erdős himself had solved it in a 1977 paper. Bloom says that AI is being used heavily by people who are not mathematicians, and there are more papers of 100-200 pages that no human has read.
AI agents as researchers: first discoveries and first disputes
An AI agent is a program that carries out a chain of steps on its own: it searches for data, runs calculations, writes a report. In 2026 companies began to show discoveries made by agents with almost no human involvement.
On 23 September 2026 Anthropic reported that its Claude model, on assignment from researchers, searched a large database of DNA sequences and found an undescribed system of genes, mainly in viruses that infect bacteria. In its structure it resembles CRISPR, the system from which gene editing tools grew. According to the company, about 950 agent sessions over 21.5 hours went through roughly 200 thousand groups of enzymes and handed people 19 reports. The laboratory experiments on this discovery were done by people. There are many limitations:
- what the system found actually does is not yet known, and the work has not been peer reviewed;
- the search was repeated 10 times, and in none of the repeats did the agents find this system again;
- the researcher Mario Rodríguez Mestre stated that he had previously discussed his unpublished results on this enzyme with Claude, so the independence of the discovery is disputed. Anthropic rejects this: according to the company, user conversations were not used to train the model that ran the search, and its team had no access to Mestre's conversations with Claude. Whether those conversations influenced the result cannot be established from public data.
Agents are also useful for checking other people's work. In the summer of 2026, at an open competition (hackathon) run by Hugging Face, 1221 participants used AI agents to try to reproduce 2226 papers from the ICML 2026 machine learning conference. At least one claim was confirmed in 51% of the papers, and at least one was refuted or disputed in 23%. The organisers write that it worked most reliably when a person directed the agents and checked their assumptions.
How many scientists use AI and how it helps them
The most recent figures come from a survey of 637 scientists in the US and the UK, conducted in July and August 2026 by the research organisation More in Common for Google's study "AI in Science: Early Insights" (September 2026). Bear in mind that the study was commissioned by a company that itself sells AI, and the answers are the scientists' own self-assessment.
| What was asked | Answer |
|---|---|
| Use AI every day | about 47%, another 31% weekly |
| How much time they save | about 6.9 hours a week on average |
| Spend more than a quarter of the time saved checking AI answers | about 46% of those who save time |
| The biggest bottleneck in their work has shifted to later stages (experiments, verification, preparing a paper) | about 44% |
| The queue of untested hypotheses is growing | about 41% (shrinking for 25%) |
| AI pushes them towards safer, more predictable projects | 49% (towards riskier ones 28%) |
| There are more weak papers in their field | 40% (fewer 35%) |
The study's authors offer a possible explanation but do not prove it: AI speeds up the production of hypotheses and calculations, while laboratories cannot test them experimentally at the same pace, and scientists themselves cannot keep up with checking what AI produces. The survey shows that this is what the scientists themselves think; it does not measure the cause.
The flip side: reviews and references written by AI
In science, before publication a paper is usually read and assessed by other specialists; this is called peer review. An example of how AI has entered this check came from the ICLR 2026 machine learning conference.
19,525 papers were submitted to it. Pangram, a company that makes an AI text detector, ran about 70 thousand reviews of these papers through its program in November 2025 and estimated that 21% were written entirely by AI, and more than half showed traces of AI. This is an estimate by the detector's own developer, and such programs make mistakes. The ICLR organisers ran the reviews through two different detectors and passed the flagged ones to the area chairs. Papers found to contain made-up references to non-existent articles were rejected without review, while the authors kept the option to contest the decision.
What this means beyond science
In science a result is expected to be checked independently, so the lessons of AI in science are useful to anyone introducing AI in their own work. Three conclusions I draw from the facts above:
- Budget time for checking. In the survey, almost half of the scientists who save time with AI spend more than a quarter of what they save checking its answers. In a company this means deciding in advance who checks the AI's output and how they will know it is correct.
- Look for where a new queue will appear. About 44% of the scientists surveyed believe that the biggest bottleneck in their work has shifted to later stages, and the survey's authors suggest that AI has sped up the production of ideas faster than they can be checked. If AI writes commercial proposals faster, the queue may move to whoever approves them.
- Demand sources and check them. At ICLR, papers with made-up references had to be rejected. When an AI text cites a law, a price or a study, it is worth opening that source and checking that it says what the AI claims.
What the successes in this article have in common, in my view, is one thing: the result can be checked independently. A protein's shape is compared with experiment, a weather forecast with the actual weather, a drug goes through trials, a proof is read by mathematicians. If you are introducing AI, start with tasks where it is equally clear how to check the answer. I wrote about what AI can do in attacks on websites and how to defend against it in the article Can AI hack a server or a website.
Summary
As of autumn 2026, the role of AI in science looks like this: in narrow tasks with a verifiable answer it is already producing recognised results, from a Nobel Prize for proteins to solved Erdős problems and a drug in phase three trials. Almost half of the scientists surveyed use it every day. The autonomous AI researcher is still only a claim: the headline discovery by the Claude agent is disputed and was not reproduced in repeat runs, while verification, experiments and peer review in the examples above rest with people.
I work on AI agents and automation. If you would like to see my projects or discuss your own task, take a look at my portfolio.
Sources
- Nobel Committee, 2024: press release on the Chemistry prize (Baker, Hassabis, Jumper). https://www.nobelprize.org/prizes/chemistry/2024/press-release/
- Nobel Committee, 2024: press release on the Physics prize (Hopfield, Hinton). https://www.nobelprize.org/prizes/physics/2024/press-release/
- ECMWF, 25.02.2025: ECMWF's AI forecasts become operational. https://www.ecmwf.int/en/about/media-centre/news/2025/ecmwfs-ai-forecasts-become-operational
- Insilico Medicine, 2026: start of the phase three trial of rentosertib and phase two results. https://insilico.com/news/xmjsn4l091-insilico-initiates-phase-iii-clinical-tr
- Quanta Magazine, 03.08.2026: Why the Legendary Erdős Problems Are Falling to AI. https://www.quantamagazine.org/why-the-legendary-erdos-problems-are-falling-to-ai-20260803/
- The Next Web, 2026: Anthropic says Claude found a new enzyme system with CRISPR-like repeats. https://thenextweb.com/news/anthropic-claude-enzyme-system-crispr-like-repeats
- Seoul Economic Daily, 29.09.2026: the dispute over the independence of Claude's discovery. https://en.sedaily.com/international/2026/09/29/can-ai-be-a-scientist-anthropic-discovery-claim-sparks
- Hugging Face, 2026: results of the hackathon on reproducing ICML 2026 papers. https://huggingface.co/blog/icml-2026-open-reproductions
- Google, September 2026: AI in Science: Early Insights (survey of 637 scientists, More in Common). https://ai.google/static/documents/AI-in-Science.pdf
- Pangram, 18.11.2025: estimate of the share of ICLR 2026 reviews written by AI. https://www.pangram.com/blog/pangram-predicts-21-of-iclr-reviews-are-ai-generated
- ICLR, 31.03.2026: A retrospective on the ICLR 2026 review process. https://blog.iclr.cc/2026/03/31/a-retrospective-on-the-iclr-2026-review-process/