Twice in eleven weeks, the market marked down drug-discovery companies after frontier AI labs launched new science capabilities.

When OpenAI released GPT-Rosalind and Rosalind Workbench on 16 April, Recursion and Schrödinger each fell more than 5%, IQVIA by as much as 3.2% and Charles River by 2.6%. The sell-off ran wider than software, catching contract research organisations on the reasoning that more capable in-house teams would outsource less work.

When Anthropic launched Claude Science on 30 June, Schrödinger fell 6.6% on the day and Recursion traded heavily before recovering some of the decline.

Twice, the market reached essentially the same conclusion: if frontier models can do drug discovery, specialist drug-discovery software must be worth less.

I think that confuses two different things.

A frontier model can read, infer, propose and explain. Increasingly, an agent wrapped around that model can also use scientific tools, execute workflows and keep working after you close your laptop.

A drug-discovery platform has to do something else. It has to remember the project: every compound, every assay, every model prediction, every experiment; what was tried, what failed, what was discounted and why. It has to connect a hypothesis to a calculation, a calculation to a molecule, a molecule to an experiment, and the result of that experiment to the next decision.

The reasoning layer is what happens inside an agent. The platform is what persists between agents, sessions, experiments and people.

The distinction matters because OpenAI, Anthropic and Google have now made remarkably similar moves into science. And none of them has built the thing the market appears to think they have replaced.

Five months ago, I called this Gen 3

In my previous article, I described AI drug-design platforms as having evolved through three generations:

  • Gen 1 (plumbing): the lakehouse, ML platform, scalable molecular scoring, workflow infrastructure and elastic compute required to run computational drug discovery at scale

  • Gen 2 (ecosystem): integrated workflows, compound stores, traceability, laboratory integration and the infrastructure required to close the design-make-test-learn loop

  • Gen 3 (agentic): context engineering, agentic interfaces, tool registries, knowledge mining, durable agentic workflows and reasoning across the whole system

At the time, my argument was that Gen 3 was additive. Agents did not make Gen 1 and Gen 2 obsolete; they made them more important, because an agent is only as useful as the context and capabilities beneath it.

Five months later, I would revise one part of that argument.

The frontier labs are commoditising parts of Gen 3 much faster than I expected.

Tool discovery, orchestration, specialist agents, scientific reasoning, long-running tasks and increasingly sophisticated scientific interfaces are becoming general-purpose infrastructure.

But the frontier labs have approached the architecture from the opposite direction to drug-discovery platforms.

They started with the agent and moved down into science. Drug-discovery platforms started with the project and moved up towards agents.

Those architectures are converging, but important gaps remain in project data and integration.

What they actually built

The three offerings differ substantially in implementation, but they occupy remarkably similar architectural territory.

  • Anthropic: Claude Science. A local-first scientific workbench that can drive Python and R kernels locally, use Slurm clusters over SSH or external compute, and combine scientific connectors and skills across genomics, single-cell analysis, proteomics, structural biology and cheminformatics. A coordinating agent can delegate work to specialists, while reviewer capabilities check outputs such as citations and figures.

  • OpenAI: GPT-Rosalind and Rosalind Workbench. GPT-Rosalind is a specialised model for life sciences, with capabilities spanning genomics, wet-lab troubleshooting and medicinal chemistry. Rosalind Workbench provides the workspace around it, with scientific viewers, guided workflows and access to a Life Sciences Research plugin connecting the agent to dozens of public scientific databases and tools.

  • Google DeepMind: Science Skills for Antigravity. Open skill files that teach Google’s general-purpose agent environment how to work with AlphaGenome, AlphaFold, UniProt, PubChem, ChEMBL and dozens of other scientific databases and tools.

The implementations differ. Architecturally, however, all three have made essentially the same move:

General-purpose model → agent → scientific tools → scientific workflows

These offerings cover a substantial part of Gen 3. Gen 1 and Gen 2 still need to be provided by the surrounding platform.

Three labs, one architectural move

Look past the branding and each frontier lab has taken an agent product it had already built, taught it about scientific databases and domain tools, and packaged the result for scientists.

That is exactly what a horizontal AI vendor should do. A scientific workbench can ship once and be useful to research organisations around the world.

A drug-discovery platform cannot.

It has to meet each organisation’s compounds, assays, models, registry, workflows, compute infrastructure and laboratory. It has to understand that organisation’s representation of a compound and preserve that identity across computational and experimental systems.

That is deep integration work and a very different business to selling a general-purpose assistant.

None of the three frontier offerings therefore provides the system of record for a drug-discovery project:

  • Compound registry

  • Integrated assay stack

  • Project-wide design history

  • Persistent project-level provenance

  • Closed design-make-test-learn loop

These products focus on the agent workspace. Project-wide records and integration remain responsibilities of the surrounding platform.

A catalogue of tools is not a capability

The chemistry illustrates the distinction particularly well.

The tool lists initially look formidable. Claude Science connects agents to chemical databases and exposes capabilities around structure prediction, docking and molecular modelling. OpenAI’s Life Sciences Research plugin reaches ChEMBL, PubChem and dozens of other scientific sources. Google’s skills similarly expose agents to chemistry databases and computational tools.

This is useful. But access to a tool is not the same thing as having a drug-design capability.

A tool registry is a Gen 3 capability. A drug-design capability appears when that registry is connected to Gen 1 and Gen 2.

Somebody still has to define what good means for this project:

  • Which properties matter?

  • How should potency, selectivity, permeability, metabolic stability, novelty and synthesizability trade against one another?

  • Which cheap filters should run before expensive calculations?

  • Which scoring methods should operate over ten molecules, and which need to operate over ten million?

  • What constitutes an improvement over the current series?

The tools also have to operate on your compounds and your measurements, alongside public scientific data.

That means compounds need identities. Assays need conditions. Predictions need model versions. Results need provenance. Decisions need to survive the session in which they were made.

Together, these platform capabilities turn a catalogue of tools into a design cycle.

Reasoning about a molecule is not running a design campaign

GPT-Rosalind’s medicinal-chemistry results illustrate both how far the reasoning has come and where the distinction remains.

OpenAI reports GPT-Rosalind scoring 27.5% on its medicinal-chemistry benchmark against 25.1% for GPT-5.5. That is meaningful progress on a hard problem.

But the benchmark measures reasoning about chemistry. It does not measure the operation of a drug-design campaign.

There is a large architectural distance between discussing a molecule intelligently and generating and scoring a million molecules against a project’s objectives.

All three frontier offerings can reason about structures. Some can invoke docking, co-folding and other scientific models.

What they do not ship is the machinery to:

  • Enumerate a virtual library

  • Run project-specific filters

  • Distribute expensive scoring across elastic compute

  • Combine calculated and experimental properties

  • Rank the resulting molecules against a project’s optimisation profile

  • Hand a medicinal chemist the most promising candidates

Operating at that scale requires a compute and data platform.

The distinction also depends on the science. If the problem is target discovery or validation from omics data, these new scientific agents are much closer to being sufficient. For iterative small-molecule optimisation, the gap is considerably larger: a chemical series must improve ADMET without sacrificing potency or selectivity.

I do not assume that will remain true indefinitely. But much of the gap is architectural rather than a deficit in model intelligence.

Those are different things to bet against.

Where your data actually goes

The old argument against frontier models was confidentiality: proprietary science should not go anywhere near a third-party model.

At the enterprise level, that argument has weakened substantially. Enterprise agreements can exclude customer data from model training, regulated environments exist and vendors increasingly provide the contractual and technical controls required by large research organisations.

Model-provider training, taken on its own, is no longer a compelling reason to build your own platform.

Data also crosses organisational boundaries through tool calls.

An agent querying UniProt, PubChem, ChEMBL, NCBI or a public sequence-alignment service is communicating with infrastructure outside your organisation. Your agreement with the model provider does not govern what happens at that endpoint.

And the query itself can be sensitive. A search against a public endpoint can disclose information about the scaffold, target or sequence being investigated to infrastructure outside the organisation. Put a proprietary sequence through a public alignment service and you have done something many organisations expressly prohibit.

Pharmaceutical companies have long addressed this problem through internal mirrors of public scientific databases and tools such as local BLAST.

You ingest public scientific data into your own environment so that it can be queried privately and joined to project data.

That requires the lakehouse and ETL capabilities from Gen 1. Private access to scientific data depends on those foundations.

Session memory is not project memory

The frontier offerings already provide useful execution and session-management capabilities.

Claude Science can preserve computational environments and produce reproducible artefacts. Rosalind can save workflows that teams inspect, share and rerun. Antigravity can orchestrate agents, run background tasks and schedule work.

These capabilities preserve session state. A platform also needs to preserve project state across sessions, people and experiments.

A drug-discovery project can run for years across dozens or hundreds of scientists. Somewhere, something has to remember every compound, every calculation, every assay and every decision.

The same applies to provenance. A reproducible agent session is valuable. An auditable workflow run is valuable. Project-level provenance asks a different set of questions:

  • Which version of which model produced this score?

  • Against which structure and using which parameters?

  • Which assay batch contradicted it?

  • Why was the result discounted?

  • Who approved synthesis?

  • What happened next?

Patent attorneys ask questions like these. Regulators can ask them. Scientists inheriting somebody else’s project certainly ask them.

Forty individually reproducible agent sessions do not, by themselves, provide the answer.

And then there is the laboratory

None of these systems owns the physical loop.

An agent may interpret an NMR spectrum, analyse LC-MS data or troubleshoot a protocol. Those are useful scientific capabilities.

They are not a compound-registration system, an assay stack, a synthesis queue or laboratory integration.

Drug discovery compounds value through a loop:

Design → Make → Test → Learn → Design again

The model can increasingly reason at every point in that loop.

But the loop itself has to exist somewhere.

If the experimental result does not flow back into the same project state from which the next design decision is made, the loop has not closed.

The strongest argument against the platform

Perhaps the frontier labs don’t need to build the whole platform. Extensible agents could work with infrastructure a research organisation already owns.

All three are doing exactly that. Claude Science supports connectors and skills. OpenAI exposes plugins, Codex and MCP. Antigravity exposes skills and agent orchestration.

A pharmaceutical company could therefore connect a frontier agent to infrastructure it already owns:

  • Compound registry

  • Assay systems

  • Lakehouse

  • Models

  • Scoring services

You potentially get an agentic layer over the existing estate without buying another integrated AI platform.

I think this is the biggest architectural change platform vendors need to respond to.

General-purpose vendors are increasingly commoditising the parts of Gen 3 that decompose cleanly. Platform teams should take advantage of that rather than rebuilding everything themselves.

The question moves down a level:

Do you already have infrastructure worth orchestrating?

For those organisations, frontier agents could provide useful capabilities over the existing estate. Integration still needs careful design.

A connector gives an agent a way to call a scoring service. It does not automatically give that service the same definition of a compound as the assay system. It does not make experimental conditions part of the context. It does not preserve why a number was discounted during the previous design cycle. It does not ensure the agent writes a revised hypothesis back into the same project record from which the next scientist works.

You can connect systems without making them one system.

And in drug discovery, much of the value lives at those joins.

The platform has moved

This changes the build-or-buy question. It also changes what a drug-discovery platform team should build.

If frontier labs can provide increasingly capable reasoning, tool use, orchestration, scientific interfaces and specialist agents, rebuilding generic versions of those capabilities is unlikely to be where a specialist platform creates the most value.

The differentiated platform layer increasingly becomes:

  • Compound identity: registration, identity resolution and series relationships

  • Assay context: conditions, results and experimental metadata

  • Project state: hypotheses, decisions and project history

  • Scientific provenance: the chain from intent to calculation to experiment

  • Scoring infrastructure: project-specific models operating at the required scale

  • Workflow execution: reliable, reproducible computational pipelines

  • Experimental integration: connecting design to synthesis and assay results

  • Approval and governance: keeping humans appropriately inside the loop

The joins between these capabilities let an agent work with a coherent drug-discovery project. Those project-specific relationships are harder to commoditise than general-purpose agent capabilities.

So do you still need a platform?

When I wrote about the third generation of AI drug-design architecture, I described it as additive: Gen 3 sits on Gen 2, which sits on Gen 1.

I still think that is true, but five months of frontier-model development have made the implication clearer.

Some of Gen 3 is becoming commodity infrastructure.

Frontier labs can build general-purpose agent orchestration, tool use, scientific reasoning and interfaces faster than almost any specialist platform team.

We should use that.

What they cannot generalise so easily is the project underneath it:

  • Your compounds and assays

  • Your models and scoring objectives

  • Your experimental history

  • The decisions that killed one series and advanced another

  • The provenance connecting a prediction to an assay result and ultimately to the next molecule somebody chooses to make

That is where Gen 1 and Gen 2 stop being plumbing and become context for Gen 3.

The three generations can live in different codebases, but they need to behave as one system.

The agent has to reason over the same project state that the workflow engine changes. The scoring service has to operate on the same compound identity the laboratory uses. Experimental results have to flow back into the context that determines the next decision. Provenance has to run unbroken from scientific intent, through reasoning and computation, to molecule and measurement.

Break those joins and you have an extraordinarily capable scientific assistant sitting beside your drug-design platform.

Preserve them and you have an agentic AI drug-design platform.

OpenAI, Anthropic and Google have moved much further into the first of those than I expected this quickly. They have also made parts of Gen 3 dramatically cheaper to build.

Scientific platforms still need to own project identity, evidence and integration, even as more of the agent infrastructure becomes available from frontier labs.