The virtual cell: what it is, what makes up the $1.8 billion behind it and what not to expect from it yet

Short answer: a virtual cell is a computer model that, as its authors intend, should predict how a living cell will respond to a drug or another intervention. Then some experiments could first be run on a computer, and the most promising ideas tested in the lab. On October 7, 2026, the nonprofit institute Biohub, the US Department of Energy, the US National Institutes of Health (NIH) and new partners announced that they are putting $1.8 billion into this work, in the form of funding, data, computation and new measurement technology. Of this, Google DeepMind, Isomorphic Labs and Meta are together investing $300 million. Two things are important to understand. The announced investment goes not to a finished model but to what is needed to create one: measurements of how cells of different types respond to interventions, new technology, computation and research. And more than $500 million of the sum is not new money but the amount of past government investment through which the data now being contributed was created. The participants give no timeframe for when the model will work, and they describe its benefits with the words "may" and "could". Below is what a cell model is, what makes up the sum, what is already openly available and how to read similar news.
What a virtual cell is
A cell is the smallest living unit of an organism. A drug or a disease changes what happens inside a cell. Today, to find out how a cell will react, an experiment is run in the lab. Experiments are expensive and slow, and there are many possibilities to test.
The idea of a virtual cell is to train an AI model on a large number of such measurements so that it learns to predict a cell's response without an experiment. NIH Deputy Director Nicole Kleinstreuer describes the goal as universal cell models "with sufficient biological complexity to predict how any cell responds to an intervention". Biohub's Head of Science, Alex Rives, says in the joint press release that an accurate predictive model of biology "could dramatically accelerate scientific discovery by enabling scientists to perform experiments digitally". He calls the creation of a virtual cell one of the most important challenges for the next era of science.
Note the wording: "could", "models can". This describes a goal, not a finished result.
Why it all comes down to data
For a model to predict a cell's response, it needs to see many examples: a cell of such a type, in such a state, after such an intervention, responded in such a way. According to NIH, an enormous amount of such data is needed, and training the models requires measurements of cell responses across many more cell types and states than have been studied so far, as well as technology that studies cells at scales and speeds that current instruments cannot capture. Biohub says the same: the programme should deliver measurements of cell responses to interventions across far more cell types and conditions, and new technology to study cells faster, more accurately and at greater scale.
That is why the announced investment is directed at generating and preparing data, new measurement technology, modelling, computation and research. The share of each area is not given in the sources. The phrase "AI-ready data" here means measurements collected and described according to shared rules, so that a model can learn from data from different labs at once. Biohub says it is building shared standards, common identifiers and a single point of access for this.
What makes up the $1.8 billion
According to Biohub's press release of October 7, 2026:
| Participant | What it contributes | New commitment or past investment |
|---|---|---|
| US Department of Energy | more than $500 million over five years for lab measurement, modelling and computation | new commitment over five years |
| NIH | coordinates the contribution of datasets, repositories and knowledge bases created through more than $500 million in past government investment | past investment through which the data was created |
| Google DeepMind, Isomorphic Labs and Meta | $300 million together in the Virtual Biology Initiative | new commitment |
| Biohub | the programme's founding $500 million, announced in April 2026: $400 million for new measurement technology and $100 million for research outside Biohub | commitment announced in April |
The press release itself does not break the sum down by line. If we assume that Biohub's founding $500 million is included in the announced $1.8 billion (the release does not say so directly), the four lines add up to roughly $1.8 billion. This is my approximate reconstruction: two lines are given as "more than $500 million". Regardless of it, it is clear that more than $500 million of the sum is past government investment through which the data contributed by NIH was created, not money allocated now.
Biohub calls this the largest coordinated commitment to generating AI-ready biological data to date. That is Biohub's own assessment.
Who Biohub is and who else is involved
Biohub is a nonprofit research institute founded in 2016 by Mark Zuckerberg and his wife Priscilla Chan, The Verge writes. The institute describes its ultimate mission as "to cure or prevent all disease", and it works at the intersection of AI and biology.
Besides money, the programme has been joined by scientific institutions and consortia that, according to Biohub, have experience in organizing major international scientific collaborations, from the Human Genome Project onwards: the Allen Institute, the Broad Institute, the Gladstone Institutes, the Human Cell Atlas and Human Protein Atlas projects, and the Wellcome Sanger Institute. NVIDIA will support the programme with computing infrastructure and software, and Renaissance Philanthropy is helping to raise funding for data generation. The Department of Energy provides supercomputers and facilities of the national laboratories, including for electron microscopy.
What you can already look at
The programme's result lies in the future. But in the same press release, Biohub lists projects with open data that it has led: Tabula Sapiens, OpenCell and Zebrahub, as well as the shared data repositories CELLxGENE and the CryoET Data Portal, which it built and maintains. These are not a virtual cell but existing resources with biological data. Links to them are in the list of sources.
If you work with biological data or are looking for it for your own project, these resources are a good place to start. Each has its own terms of use, which you need to read on the resource's own website before building anything commercial on the data.
What not to expect
- A finished model by a specific date. The sources give no timeframe for when the virtual cell will work, so there is no basis for expecting it by any particular date.
- New drugs from this news. NIH writes that such models could ultimately help find drug targets, that is, what in the body a drug is supposed to act on, and choose ideas for laboratory and clinical testing. This is a possible future benefit, not a promise of drugs.
- An end to lab experiments. Even in the description of the goal, the model helps choose what to test in the lab rather than replacing the testing.
I wrote about how to read figures on the accuracy of medical AI in the article AI in medicine: where it helps doctors, where it gets things wrong and how to read figures on its accuracy. What AI has already done in science is covered in the article AI in science in 2026: what it has already discovered and where it still gets things wrong.
How to read similar news
Four questions that helped to take this news apart and that work for other announcements about science and AI. Readers can ask them themselves by opening the primary source.
- Is this new money or past investment? Here, more than $500 million of the $1.8 billion is past government investment through which the data contributed by NIH was created.
- How to check: in the primary source, for each sum it is clear whether it is a new commitment by a participant ("is investing", "will invest") or money spent earlier ("past investment", "created previously"). Not every new commitment has a timeframe attached, and the absence of a timeframe does not by itself make an announcement doubtful.
- What exactly is the money for: a result or preparation for one? Here, for data, technology, computation and research, not a finished model.
- How to check: the text says the result has already been achieved, and there is evidence you can open yourself: a publication, open data or a working tool. Until there is such evidence, treat the result as unconfirmed and check separately whether it is announced as already achieved or only planned.
- Which verbs are used. "May", "could", "ultimately" usually describe a hope. "Showed", "measured", "published" usually describe something done.
- How to check: find at least one statement in the text about a result already achieved, and its source.
- Who is making the assessment. "The largest", "one of the most important challenges" in a press release are the participants' own assessments.
- How to check: whether someone who is not involved in the project gives the same assessment.
Summary
On October 7, 2026, Biohub, the US Department of Energy, NIH, Google DeepMind, Isomorphic Labs and Meta announced $1.8 billion for data for AI models of the cell. The virtual cell is a goal: a model that would predict a cell's response to a drug or another intervention, so that some experiments could first be run on a computer. For now, the investment goes to measurements, technology, computation, research and shared data standards, and more than $500 million of the sum is past government investment through which the data contributed by NIH was created. No timeframe for the model is given, and the benefits are described as possible. The news is important, but it is about the start of a large data effort, not about a finished tool.
I work on AI agents and automation. If you would like to see my projects or discuss your own task, take a look at my portfolio.
Sources
- Biohub, 07.10.2026: press release on the expansion of the Virtual Biology Initiative. https://biohub.org/news/virtual-biology-initiative-expansion/
- National Institutes of Health, 07.10.2026: NIH joins effort to build SI-ready data for predictive models of human biology. https://www.nih.gov/news-events/news-releases/nih-joins-effort-build-si-ready-data-predictive-models-human-biology
- Biohub's open resources named in its press release: Tabula Sapiens https://tabula-sapiens.sf.czbiohub.org/ , OpenCell https://opencell.sf.czbiohub.org/ , Zebrahub https://zebrahub.sf.czbiohub.org/ , CELLxGENE https://cellxgene.cziscience.com/ , CryoET Data Portal https://cryoetdataportal.czscience.com/
- The Verge, 07.10.2026: Google invests millions in Mark Zuckerberg's efforts to create a 'virtual cell'. https://www.theverge.com/tech/1006766/google-meta-biohub-investment-virtual-cell