Lonza - PBMCs

Foundation Models in Drug Discovery: The Next Frontier of Pharmaceutical Research

Lakshmi, Editorial Team, Pharma Focus America

Foundation models pretrained on vast biological corpora are replacing the one-model-per-task economics that has long limited computational drug discovery. This article explains what makes these systems different, maps where they are genuinely mature and where they are not, examines the gap between them and current FDA expectations, and sets out the build-or-buy and data investment decisions now facing American research leaders.

Introduction: The End of the One-Model-Per-Task Era

For most of the past decade, computational drug discovery ran on bespoke machinery. One model per assay. One model per target class. One model per endpoint. Each required its own labeled dataset, its own validation exercise, its own maintenance schedule, and its own team of people who understood what it did. In many organizations the portfolio of models grew considerably faster than the portfolio of drugs, and the cost of keeping them all honest quietly consumed the productivity they were built to deliver.

Foundation models break that pattern. Pretrained once on enormous quantities of unlabeled data — billions of protein sequences drawn from across the tree of life, hundreds of millions of single-cell measurements, entire genomes, the accumulated text of the biomedical literature — they are then adapted to specific tasks with comparatively little additional data. That is not an incremental gain in predictive accuracy. It is a change in the unit economics of computational biology, and therefore a change in what an American research organization can afford to attempt. Executives who read it as a faster version of what they already have will misjudge both the opportunity and the risk.

What Actually Makes a Model “Foundational” — and Why It Belongs on the P&L

Three properties separate a foundation model from the predictive tools that preceded it. The first is the scale and nature of pretraining: these systems learn from unlabeled data, which removes the historic bottleneck of expensive expert annotation. The second is generality — a single base model supports many downstream tasks rather than one. The third, and the most commercially interesting, is transfer: performance on problems the model was never explicitly trained to solve, because it has absorbed something closer to the underlying grammar of proteins, cells or molecules than to any particular endpoint.

The modality coverage has widened quickly. Base models now exist for protein sequence and structure, for genomic sequence, for chemical structure, for cellular imaging, for clinical and scientific text, and — most recently and most consequentially — for how cells respond to perturbation. Releases over the past eighteen months include protein models trained on billions of sequences and multi-billion-parameter systems built specifically to predict cellular response to genetic perturbation.

The financial translation is straightforward. In a bespoke world, every new prediction task is a project with a budget, a timeline and a hiring plan. In a foundation-model world, most new tasks are a fine-tuning run against an existing base. The marginal cost of asking a new computational question collapses — and when the marginal cost of a question falls far enough, organizations start asking questions they previously could not justify.

Figure 1: The foundation model stack. The base model is increasingly a commodity; the adaptation layer, built on proprietary assay data, is where durable competitive advantage is created.

The Data-Efficiency Dividend Is the Real Story

The binding constraint on machine learning in pharmaceutical research has never been algorithms. It has been labeled data. Assays are expensive, proprietary, inconsistently run across sites and years, and small — most internal datasets contain hundreds of well-characterized examples, not the tens of thousands that conventional supervised learning demands.

Foundation models change the shape of that curve rather than shifting a point along it. Because the base model arrives already carrying a representation of chemical and biological structure, a downstream task that once required tens of thousands of labeled examples may become tractable with a few hundred. The practical consequence is that a great deal of proprietary data which was previously too small to be useful suddenly becomes usable.

Three strategic implications follow. Proprietary data becomes more valuable per data point, not less, because it is the input to the differentiating adaptation layer rather than to a commodity base. Rare diseases and understudied target classes become computationally tractable, since they were excluded largely by data scarcity. And the competitive question shifts from who holds the largest dataset to who holds the most informative one for the specific problems they intend to solve.

One warning belongs alongside this. Published benchmarks flatter these systems. Independent evaluations have repeatedly found that impressive zero-shot performance on curated public datasets degrades under distribution shift — different cell types, different donors, different assay protocols, different chemical series — and that carefully tuned simple baselines remain competitive on several widely reported tasks. No vendor number should be believed until it has been reproduced on an internal held-out set that was never touched during adaptation.

Figure 2: Task performance against the volume of labeled data available. The shaded area is the practical value of pretraining — and it is widest precisely in the data range where most proprietary pharmaceutical assays sit.

A Candid Map of What Works Today — and What Doesn’t

Capability is not uniform across the discovery cascade, and the pattern is unusually consistent: performance tracks the density and quality of the underlying training data almost mechanically. Protein structure prediction rests on hundreds of thousands of experimentally solved structures and is now routine. Binder and protein design, molecular property and ADMET prediction, and literature synthesis are all in productive use.

Cellular perturbation modeling and target identification sit a tier below — genuinely promising, actively improving, and not yet reliable enough to carry a program decision unaided. Translational prediction from cell to animal to human, and prediction of clinical efficacy, remain unproven. That is not a failure of engineering. For any given target, the number of well-characterized human clinical outcomes is measured in single digits, and no architecture compensates for the absence of the observations it would need to learn from.

The executive implication is a two-handed one. Deploy foundation models aggressively where they are mature, because the cost and cycle-time savings are available immediately and competitors are already taking them. Simultaneously, fund data generation where the models are weak. For many American organizations, the highest-return investment available is not a larger model but a large, well-controlled perturbation dataset that nobody else possesses — because that is the asset a commoditized base model cannot replicate.

Figure 3: Demonstrated maturity of foundation model capabilities across the discovery cascade. The gradient follows data density, not research effort.

Washington Wrote the Rules for a Different Kind of Model

American sponsors operate under a regulator that has been unusually forward-leaning on this subject. FDA’s drug center received more than 500 submissions containing AI components between 2016 and 2023, and the agency has authorized well over a thousand AI-enabled devices. In January 2025 it issued draft guidance establishing a risk-based credibility assessment framework for AI used to support regulatory decisions, anchored on the concept of context of use — the principle that identical models can require very different evidence depending on the decision they inform. In January 2026 the agency joined its European counterpart in publishing ten guiding principles for good AI practice across the drug lifecycle.

The difficulty is that this framework was designed when models were narrow and purpose-built, and foundation models fit it awkwardly. Context of use presumes a defined purpose, whereas a pretrained base model has thousands. Credibility assessment leans heavily on training data provenance, whereas base models are pretrained on heterogeneous public corpora whose composition is sometimes only partially documented. And version control becomes a live regulatory question: when a supplier updates model weights, it is not obvious whether the credibility assessment travels with it.

Waiting for a final guidance to resolve this would be a mistake. Four practices can be established now, cost little, and will almost certainly be asked for later. Maintain documentation for every base model in use, including the specific checkpoint identifier. Define context of use per downstream application rather than per model. Freeze checkpoints for anything that will touch a submission, which in practice argues against renting critical capability through an interface you do not control. And hold back internal validation sets that are never used in adaptation. Sponsors who can answer “which version of which model produced this number, and what was it trained on” will move through review measurably faster than those who cannot.

Figure 4: A practical reading of context of use. The depth of credibility evidence required scales with how consequential the decision is and how much the model influences it — not with the sophistication of the model itself.

Build, Buy, or Rent: The Capital Question With No Clean Answer

Three postures are available, and most organizations should hold all three deliberately rather than drifting into one by default. Building a proprietary base model is justified only where a company holds a genuinely unique data asset at scale and can sustain the compute budget; costs run into the tens of millions and the resulting asset depreciates quickly as open releases catch up. Adapting openly released weights has become the pragmatic default, particularly for protein and single-cell modeling, where open base models have improved rapidly and carry workable licenses. Renting capability through a vendor interface is the fastest route to value and the one with the least regulatory optionality, because a checkpoint you do not control cannot be frozen and a training corpus you cannot inspect cannot be documented.

The sensible discipline is to tier applications by regulatory criticality and match the posture to the tier. Rent for exploratory and internal-productivity work. Adapt open weights for the discovery engine. Reserve building for the one or two areas where proprietary data genuinely justifies it. What organizations should avoid is arriving at a vendor-dependent architecture because procurement moved faster than strategy — a common outcome, and an expensive one to unwind two years into a program.

Five questions make the tiering concrete, and they are the ones a research board should be putting to its own organization. Which of our proprietary datasets would become useful for the first time under a data-efficient adaptation approach? Have we reproduced any vendor performance claim on our own held-out data, or are we buying against published benchmarks? For every model that could touch a regulatory submission, can we name the checkpoint and describe its training corpus? Where are we renting capability that we will later need to freeze, document, or defend to a reviewer? And if base models are commoditizing, what exactly is our durable advantage — and are we funding it?

Conclusion: The Advantage Is Moving to the Adaptation Layer

Foundation models will not announce themselves through a single dramatic approval. What they change is subtler and, over a decade, more consequential: they reduce the price of asking a scientific question. Pharmaceutical research has always been constrained less by the ambition of its hypotheses than by the cost of testing them, and a technology that lowers that cost by an order of magnitude alters which hypotheses ever get tested at all.

The strategic risk for American companies is a specific one, and it is not falling behind on models. Base models are commoditizing quickly, and open weights travel across borders instantly and at no cost. What does not travel is proprietary experimental data, the laboratory capacity to generate more of it, the regulatory documentation discipline to defend what the models produce, and the organizational judgment to know which of their outputs to believe. Those are the assets that compound.

The United States retains real structural advantages here — deep capital markets, concentrated compute capacity, and a regulator engaging with the technology earlier and more openly than most. None of those advantages is permanent, and none of them substitutes for the unglamorous work of building a data asset worth adapting a model to. Boards that fund the base model and neglect the adaptation layer will have bought the part of the stack their competitors can download.

Lakshmi

Lakshmi is a science writer with a foundation in the laboratory. She earned her master's in biotechnology and trained through research internships at ICGEB (JNU) and DIPAS, DRDO, with her work appearing in the Egyptian Journal of Veterinary Sciences. Now APCRM-certified and part of the editorial team at Pharma Focus America and Pharma Focus Europe, she reports on pharmaceutical technology, research, and innovation — giving complex science a clear and confident voice for industry leaders.