The real impact of AI drug design technology does not come when a new scientific algorithm is released. It comes when the software architecture is built around that algorithm so it can actually be used. So what do you need to build to get there?
Three generations of platform architecture, each forced into existence by a different constraint.
The 1st Generation: Plumbing
The constraint was access. The science existed, but the data was trapped: in HPC file systems, in scientists’ home directories, in instrument exports that nobody could query. Nothing could be industrialised until that was fixed. It is a heavy lift and not to be underestimated. What to build:
-
Molecular generation UI: enabling users to configure the complex reinforcement learning algorithms
-
Data lakehouse: for all drug discovery data, so that it can be used for ML
-
ETL/ELT workflows: to get all of the drug discovery data into the above
-
Machine learning platform: automated data extraction, model build, versioning, benchmarking and scalable inference
-
Modular molecular scoring: scalable web services that can score anything from 1 to 1 million molecules across hundreds of potential scoring methods
-
Computational chemistry integration: traditional techniques working alongside, and inside, the AI algorithms
-
Containers: taken for granted now, but not long ago everything was still a Python import on an HPC platform
-
Cloud and Kubernetes: why would you build a modular platform on anything else?
-
Continuous integration and deployment for everything: also novel in drug discovery software until recently
The 2nd Generation: Ecosystem
The constraint was headcount. A gen 1 platform is powerful, but only a handful of experts can drive it, which means the value of the platform is capped by the number of people who understand it. Gen 2 removes that cap. You need everything you built (or licensed) in gen 1, and additionally:
-
Fully integrated AI design UI: the gen 1 UIs were really only for a few expert users; at this point you want a much broader userbase
-
Integrated workflow and compute engine: all compute runs on one elastic, scalable platform, where algorithms can be combined into complex workflows
-
Cost control: the cost of running workflows needs to be surfaced to users, with limits
-
Abstracted data management: users no longer need to think about where data lives. This comes as a pair with traceability
-
Generative design compound store: maintain knowledge about every compound you generate and score
-
Traceability database: link the workflows, tools, data and models with compound generation and scoring, so that you have complete molecular provenance
-
Lab in the loop: link design with experimental platforms so that experimental results return to the design process
The 3rd and Current Generation: Agentic
Until a couple of years ago, if you had all of the above you were mostly done. Even the newest co-folding models like Boltz-2 still deploy as a web service running on a GPU: a hard engineering problem, but a familiar one.
Then the thinking models arrived, and the constraint changed again. This time it is judgement. A gen 2 platform will run anything you ask of it, but a person still has to decide what to ask: which series to push, which of a thousand possible next experiments is worth the compute and the chemistry. That decision is the slow step, and it does not get faster by adding more platform.
A model that can reason over a project can now propose that next step itself. That breaks an assumption buried in every workflow engine we ever built: that the shape of a workflow is known before it starts. It is what makes this generation agentic, and it requires:
-
Agentic UI: provide rich information to a user who can then work collaboratively with the agents to guide them
-
Context driven drug design: building the right context around project data so that the LLMs can work out what to do next
-
Agentic tool registry: tools need to be explained to the LLM, with guidance on applicability and usage
-
Knowledge miner: using LLMs to mine research papers so that the knowledge can be used alongside project data
-
Agent memory: what an agent carries forward between sessions, and what it is allowed to forget, is a design decision, not a side effect
-
Agentic-enabled workflow engines: running agents as workflows that call tools as workflows is really powerful with long running drug design processes
-
Durable workflows: to maintain robustness across the expanding complexity
-
Agentic traceability: tracking how the intent translated to decisions and tool calls
-
Approval gates: the points where an agent must stop and ask. Ordering compounds, committing lab capacity, closing off a series: these need a human signature, and the architecture has to know where they are
-
Agent evaluation: you benchmark models in gen 1; you now need the equivalent for agent decisions. Knowing whether an agent chose the right next experiment is harder than knowing whether a model predicts well, and it is the least solved problem on this list
-
Context confidentiality: project data is now prompt content. The platform architecture must govern which model sees which data and under what contractual terms
-
Reasoning cost control: gen 2 cost control assumed compute you could predict from the job. A reasoning loop decides its own depth, so cost has to be governed at run time
-
Agentic visualisation: created for a user based on the data
Gen 3 depends on the layers underneath it. An agent can only call a tool that has a machine-readable contract, and can only be trusted with a result that carries its own provenance. Both of those are gen 1 and gen 2 problems.
Why Context Engineering is the Hard Part
The new challenge is providing the optimal context to the thinking LLMs running in an AI agent native platform. In most domains that means retrieval and a prompt template. Drug design makes it harder in three specific ways.
-
The context is multi-modal: a useful picture of a project is structures, assay results, synthesis history, ADMET predictions, safety flags and the literature, and no two of those share a representation
-
The context has to carry its own provenance: it is not enough for the model to know that a compound scored 8.2. It needs to know which model version produced that, against which assay, and with what confidence, or it will reason confidently from a number that a chemist would have discounted on sight
-
The unit of work is the campaign, not the conversation: a design cycle runs for months; the context window does not. Everything that matters has to survive being compressed, dropped and reconstructed
So What Do You Need to Build?
Building gen 3 requires the data, compute and integration capabilities of gen 1 and gen 2.
Most of the current attention is on the top layer: the agent, the conversational UI, the demo where a model proposes a molecule. But an agent is only as good as the tools it can call and the data it can trust. A tool registry is worthless without modular, callable tools behind it. Context engineering is worthless without a lakehouse and a traceability database to build the context from.
A useful starting question is “what did I build over the last ten years, and can an agent reach it?” If scoring still runs as a script on a shared drive, and provenance still lives in a chemist’s notebook, no amount of LLM will paper over that. You are still going to have to build the plumbing.
Teams with strong data, compute and integration foundations are better placed to build useful agents.
