Scientists Are Building an AI “Virtual Cell” — Could It Speed Up New Medicines?
A $1.8 billion initiative involving Biohub, the U.S. government, Google DeepMind and Meta aims to build AI models that predict how living cells respond to drugs and genetic changes.
Contents
- What is a virtual cell?
- Why is biology so difficult for AI?
- Where does the $1.8 billion come from?
- Why are Google DeepMind and Meta involved?
- The real project is about data
- Could scientists test drugs inside a computer?
- Why could this matter for drug discovery?
- When could a usable virtual cell exist?
- Why open data matters
- What technologies will generate the data?
- Can AI really predict a whole living cell?
- The biggest risk: an AI model that is confidently wrong
- Could virtual cells reduce animal experiments?
- Is this another AlphaFold moment?
- Why the project could still succeed even if the full virtual-cell goal takes longer
- When could patients benefit?
- Why this is a major AI story
- Frequently asked questions
- Has a complete AI virtual cell already been built?
- How much is being invested?
- Who is involved?
- Could medicines eventually be tested entirely on computers?
- Will the datasets be public?
- The bottom line
- Sources

Last reviewed: October 8, 2026. Artificial intelligence has already learned to generate text, images and code, and it has transformed areas such as protein-structure prediction. Now researchers are attempting something more ambitious: building AI models that can predict how living cells respond when scientists change a gene, add a drug or alter the cell’s environment.
On October 7, Biohub announced a major expansion of its Virtual Biology Initiative, bringing together the U.S. Department of Energy, the National Institutes of Health, Google DeepMind, Isomorphic Labs, Meta and other research organizations in a coordinated effort valued at about $1.8 billion.
The goal is not to create a perfect digital human cell overnight. The project is trying to generate the enormous amount of standardized experimental data required to train predictive models of biology — models that could eventually let researchers test some ideas on a computer before deciding which experiments are worth doing in the lab.
What is a virtual cell?
A virtual cell is the idea of an AI model that can simulate at least part of a living cell’s behavior. Instead of simply describing what scientists already know, the model would try to predict what happens next.
Imagine a researcher considering a new cancer drug. Today, the scientist might expose cells to the compound, wait for the result, measure how genes and proteins change, adjust the experiment and repeat the process many times.
A sufficiently accurate virtual-cell model could let the researcher ask questions first:
- What happens if this drug is added to this specific cell type?
- Which genes become more or less active?
- Does the cell die, recover or become resistant?
- What happens if a particular gene is switched off?
- Which experiment is most likely to produce useful information?
The physical experiment would still matter. The difference is that AI could help researchers choose better experiments instead of testing every possibility blindly.
Why is biology so difficult for AI?
Modern language models became powerful partly because the internet contains enormous amounts of text. Biology does not have an equivalent dataset.
There are huge biomedical databases, but the underlying experiments are often performed in different laboratories, under different conditions, with different instruments and measurement methods. That makes the data difficult to combine cleanly.
Reuters described the shortage of standardized biological data as a major obstacle preventing AI from reaching its full potential in drug discovery. A model cannot reliably learn the rules of cellular behavior if the training data are fragmented, inconsistent or too narrow.
That is why the Virtual Biology Initiative is focusing so heavily on generating new experimental data rather than simply building a bigger neural network.
Where does the $1.8 billion come from?
The headline figure is a combined commitment of funding, existing datasets, computing resources and new measurement technology — not a single $1.8 billion cash payment.
Biohub says its original commitment to the initiative was about $500 million. The U.S. Department of Energy plans to contribute more than $500 million over five years in measurement, modeling and computing.
The NIH will coordinate datasets and resources that grew out of more than $500 million in previous federal investment. Google DeepMind, Isomorphic Labs and Meta are collectively investing another $300 million into the Virtual Biology Initiative.
Together, those pieces bring the coordinated effort to roughly $1.8 billion.
Why are Google DeepMind and Meta involved?
AI companies increasingly see biology as one of the next major frontiers for machine learning.
Google DeepMind’s AlphaFold systems showed that AI could make a major scientific contribution by predicting protein structures at scale. But predicting a whole living cell is far more difficult than predicting the shape of one protein.
A cell contains thousands of interacting genes, proteins, signaling pathways, membranes, energy systems and molecular processes. Those systems change over time and react differently depending on the cell’s environment.
Meta has also invested heavily in biological AI, while Isomorphic Labs is focused on applying AI to drug discovery. A model capable of predicting cellular responses could become foundational technology for many areas of biotechnology.
The real project is about data
The most important part of the announcement may not be the AI model at all. It is the plan to create enormous, standardized biological datasets designed specifically for machine learning.
Biohub says the initiative will expand measurements of how cells respond to interventions across far more cell types and conditions than have previously been studied.
The effort is expected to combine technologies such as advanced microscopy, cryo-electron tomography, molecular measurements, cellular engineering, autonomous laboratories and exascale computing.
The basic loop is straightforward: scientists deliberately change a biological system, measure what happens in detail, then use those input-output relationships to train AI models.
Could scientists test drugs inside a computer?
That is one of the most exciting possibilities, but it needs careful wording.
A virtual-cell model would not immediately replace clinical trials, animal studies or laboratory validation. Human biology is vastly more complicated than a single isolated cell.
But even an imperfect predictive model could still be useful. Suppose researchers have 100,000 possible molecules that might affect a disease pathway. AI might help identify the few hundred that are most promising, allowing scientists to spend laboratory time on a much smaller and better-targeted set.
That would not eliminate experimentation. It would change which experiments get performed.
Why could this matter for drug discovery?
Drug development is slow and expensive partly because many apparently promising ideas fail.
A molecule may look useful in an early experiment but later fail because researchers did not fully understand how different cells would respond.
Better predictive models could help researchers ask more precise questions, such as:
- Which patients might respond differently to the same treatment?
- What happens when two drugs are combined?
- Which mutation makes a tumor resistant?
- Which cellular pathway should researchers target next?
- Which experiment is likely to fail before it is performed?
If AI can eliminate poor candidates earlier, it could reduce wasted experiments and shorten some parts of the drug-development process.
When could a usable virtual cell exist?
Biohub’s timeline is ambitious. Reuters reported that the first large-scale dataset is expected within roughly a year, while more functional predictive models are a goal over the initiative’s five-year horizon.
That does not mean scientists expect a perfect digital human cell next year. The first milestone is better data. The harder milestone is proving that AI predictions are reliable enough to guide real scientific decisions.
The initiative is essentially trying to compress progress that could otherwise take decades into about five years.
Why open data matters
Biohub says the resulting resource is intended to be open to the broader scientific community.
The initiative includes organizations such as the Allen Institute, Broad Institute, Human Cell Atlas, Human Protein Atlas and Wellcome Sanger Institute. The aim is to create shared standards and identifiers so datasets produced in different places can work together.
That could be as important as the AI itself. A powerful model trained on incompatible datasets will still struggle to generalize.
What technologies will generate the data?
Biohub’s announcement describes several major technologies that will contribute to the project:
- Cryo-electron tomography for near-atomic views inside cells.
- Microscopy systems designed to image enormous numbers of cells.
- Tools that perturb genes, molecules, cells and tissues in controlled ways.
- Exascale supercomputing and advanced AI analytics.
- Autonomous laboratories and high-throughput measurement systems.
- Standardized datasets linking interventions to cellular responses.
The goal is to create training data that capture not just what cells look like, but how they change when something is done to them.
Can AI really predict a whole living cell?
That remains an open scientific question.
Protein-structure prediction was a narrower problem than cell simulation. A cell contains interacting systems that unfold over time and are influenced by the environment.
Researchers still do not know how far today’s AI approaches can scale in biological prediction. The initiative is therefore building both the data and the validation systems needed to test whether the models are actually useful.
The biggest risk: an AI model that is confidently wrong
A model can produce a prediction that looks scientifically plausible and still be incorrect.
That is inconvenient in ordinary AI applications. In drug research, it can waste months of work or direct resources toward the wrong biological mechanism.
A useful virtual-cell model therefore needs more than an impressive benchmark score. Researchers will need to test whether its predictions hold up on new cell types, new drugs and new biological conditions that were not included in training.
Validation will be just as important as prediction.
Could virtual cells reduce animal experiments?
Potentially, but they are unlikely to eliminate them anytime soon.
If AI can rule out thousands of weak drug candidates before they reach animal testing, fewer unnecessary experiments may be required. That would be valuable scientifically and ethically.
But virtual predictions still need to be checked against real biology. Laboratory cells, animal models and human clinical trials will remain necessary for proving safety and effectiveness.
Is this another AlphaFold moment?
Possibly — but it is too early to know.
AlphaFold transformed structural biology because protein structure turned out to be highly suitable for machine learning. A predictive cell model is a broader and more dynamic challenge.
If scientists eventually succeed, the impact could be even wider because a virtual cell would not only describe structure; it would try to predict biological response.
The challenge is correspondingly larger.
Why the project could still succeed even if the full virtual-cell goal takes longer
The initiative could produce valuable science even if researchers do not achieve a universal digital cell within five years.
A massive standardized dataset showing how different cells respond to drugs, genetic changes and environmental conditions would be useful on its own.
Scientists could use those data to identify biomarkers, study disease mechanisms, train smaller models and discover patterns that are difficult to see in conventional experiments.
That means the data infrastructure could remain valuable even if the most ambitious AI predictions take longer than hoped.
When could patients benefit?
This is not a near-term medical treatment.
The first step is building data and predictive models. Any drug or clinical tool developed using those models would still need laboratory testing, safety studies and clinical trials.
The more realistic near-term benefit is better research efficiency: choosing more promising experiments, identifying better drug targets and eliminating weak candidates earlier.
Over time, those improvements could shorten some drug-development programs.
Why this is a major AI story
Most AI headlines focus on chatbots, image generation or workplace automation. Virtual biology points in a different direction: AI as a scientific instrument.
Instead of asking a model to summarize knowledge that humans already have, researchers want it to predict experiments that have never been performed.
If that works, AI would move from organizing scientific knowledge to helping generate new scientific knowledge.
That is why governments, research institutes and major technology companies are willing to invest at this scale.
For another example of biology producing unexpected clues about health and aging, see our article on Jonathan the 194-year-old tortoise and what scientists found in his DNA.
Frequently asked questions
Has a complete AI virtual cell already been built?
No. Researchers are building the datasets, technologies and predictive models needed to work toward that goal.
How much is being invested?
The expanded initiative represents roughly $1.8 billion in combined funding, existing data resources, computing and measurement technology.
Who is involved?
Participants include Biohub, the U.S. Department of Energy, NIH, Google DeepMind, Isomorphic Labs, Meta and several major research organizations. NVIDIA is also supporting the effort with accelerated-computing expertise.
Could medicines eventually be tested entirely on computers?
Not with current technology. Virtual models may help prioritize experiments and drug candidates, but physical experiments and clinical trials remain necessary.
Will the datasets be public?
Biohub says the initiative is being developed as an open scientific resource for the broader research community.
The bottom line
The Virtual Biology Initiative is not about producing a perfect digital human cell next year. It is an attempt to build something biology has never had at this scale: standardized experimental datasets created specifically to train AI models to predict how living cells respond to change.
If those models become reliable, researchers could test some ideas digitally before spending time and money in the laboratory. That could speed up drug discovery and reveal biological mechanisms that are difficult to uncover today.
But the hardest part is still ahead: proving that AI can predict living biology accurately enough to be trusted.
If researchers succeed, one of the next major leaps in artificial intelligence may not happen inside a chatbot at all. It may happen inside a simulated cell.
Sources
Biohub — $1.8 billion expansion of the Virtual Biology Initiative.
Reuters — U.S. government and Google join Biohub in $1.8 billion AI biology push.
