Agents change who decides what happens next. They do not remove the need to execute those decisions reliably, safely and with scientific provenance.
Over the past 20 years, I have used a wide range of workflow engines, including Pipeline Pilot, KNIME, Argo Workflows, AWS Step Functions, Prefect and Temporal. I spent many years at BIOVIA, which sells Pipeline Pilot.
In my previous two articles, I explored agents, context engineering and where lasting value sits in scientific platforms. Project memory, provenance and connections between scientific systems were central to that argument. This raises the next architectural question: as agents choose tools and adapt their plans, what role remains for the workflow engine?
KNIME makes a scientific method visible. A scientist or informatician connects nodes into a workflow that retrieves data, applies scientific algorithms and produces a result.
An agent does not need every step to be defined in advance. It can interpret an objective, select tools, delegate work to specialist agents, inspect the results and revise its plan. The graph can emerge while the work is happening.
In science, the consequences are concrete. A duplicated calculation wastes compute; a duplicated laboratory request may consume a physical sample. Losing the link between an action and its evidence can make a result difficult to interpret or reproduce.
What agentic orchestration changes
Agentic orchestration moves part of the control flow into the reasoning system.
An agent can create a plan from a scientific objective, select tools, delegate to specialists and run investigations in parallel. It can inspect intermediate results, revise an unpromising route and ask a scientist for clarification or approval.
Anthropic has described its Research system as an orchestrator-worker architecture in which a lead agent creates specialised subagents dynamically. OpenAI's agent tooling similarly supports handoffs and manager-style patterns in which one agent calls other agents as bounded capabilities. Agent runtimes increasingly coordinate the work around model calls. Anthropic's multi-agent architecture OpenAI orchestration and handoffs
The overlap is becoming more direct. OpenAI's Agents API now describes itself as managing sessions, orchestration, context compaction and recovery while the application supplies tools and chooses the execution environment. Those are capabilities that would previously have sat clearly outside the model layer. OpenAI Agents API
It is therefore reasonable to ask whether the agent runtime is becoming the workflow engine.
For bounded work inside one agent session, it sometimes is.
For an end-to-end scientific process, the test is whether those capabilities cover the whole investigation, including external jobs, approvals and laboratory work.
Agentic orchestration and workflow orchestration solve different problems
Agentic orchestration answers a cognitive question:
Given the objective, the available context and the result so far, what should I do next?
Workflow orchestration answers an operational question:
Given what we have decided to do, how do we ensure that it happens reliably and within the rules?
The agentic loop concerns reasoning, planning, tool selection and adaptation. The operational layer preserves execution state and coordinates retries, timeouts and recovery. Platform code uses that layer to apply approval rules, permissions and budgets.
These responsibilities can sit within a single runtime or be divided across cooperating systems.
| Agentic orchestration | Workflow orchestration |
|---|---|
| Selects the next action | Preserves the state of the process |
| Adapts the plan to new evidence | Applies retry and timeout policy |
| Chooses tools and specialists | Coordinates platform approval checks |
| Handles ambiguity | Supports recovery from infrastructure failure |
| Pursues an objective | Manages lifecycle and coordinates budget checks |
| Can make probabilistic decisions | Records execution outcomes for recovery |
Modern workflow code can contain loops, branches and external events. The useful distinction is where judgement sits and which guarantees each layer provides.
The scientific workflow landscape
Scientific workflow engines already serve different responsibilities. This landscape includes both science-focused systems and general-purpose orchestrators used within scientific platforms. The following grouping gives representative examples, not an exhaustive catalogue or a ranking.
| Main role | Representative systems | Primary focus |
|---|---|---|
| Visual analytical workflows | Pipeline Pilot, KNIME, Galaxy | Compose scientific methods from reusable analytical components. |
| Portable analysis pipelines | Nextflow, Snakemake, Cromwell, Toil | Connect computational steps and data dependencies across execution environments. |
| Distributed scientific computing | FireWorks, Parsl, Pegasus, AiiDA | Coordinate high-throughput scientific calculations across clusters and other distributed resources, with scientific provenance capabilities that vary by engine. |
| Data and machine learning workflows | Apache Airflow, Dagster, Prefect, Flyte, Kubeflow Pipelines | Coordinate data processing and model workflows, including scheduling, execution tracking and recovery. |
| Kubernetes task orchestration | Argo Workflows | Execute sequences and graphs of containerised tasks on Kubernetes. |
| Durable application workflows | Temporal | Preserve process state across failures, external events and long waits. |
These categories overlap, and a platform can combine several engines. Systems grouped together can differ in their execution models, durability guarantees and scientific provenance capabilities. The useful distinction is which responsibilities each system provides, not which label it carries.
Dynamic execution predates agents: FireWorks can change workflows during execution, while Parsl builds dynamic graphs from Python tasks. Agents add model judgement about the next action. FireWorks dynamic workflows Parsl dataflow
Scientific workflows as agent tools
An agent does not need to invent every analytical procedure. It can select and parameterise an established workflow: a KNIME workflow, a Nextflow pipeline or a Flyte workflow, for example. The method remains explicit while the agent decides when to use it.
An agent might choose a Nextflow pipeline for sequence analysis, a Flyte workflow for model training or a FireWorks workflow for a set of materials calculations. In a layered design, the outer workflow applies platform policy and tracks the requested run; the selected engine manages its computational steps and returns outputs with provenance.
Established workflows capture expertise that can be tested, reviewed and reused. The platform should reuse the selected engine's execution guarantees. A separate durable layer becomes useful when the wider investigation also crosses people, laboratories and other systems beyond that engine's remit.
Temporal beneath an agentic platform
Temporal provides one example of a durable execution backend for an agentic platform. One useful capability is that an Activity can perform compute itself or submit a job to another execution system.
A Temporal Workflow holds the durable control logic. Activities perform external operations, including service calls, calculations and job submissions. This allows the same process to coordinate a Kubernetes Job, an HPC calculation or a scientific pipeline without requiring a single execution environment. Temporal Activities
For example, a workflow can submit a GPU calculation and continue when the result arrives, or wait for scientific approval and laboratory results. The wait does not require the original agent or submitting worker to remain running.
Temporal preserves workflow state across failures and supports processes that last from seconds to years. External results can return through callbacks or polling, allowing the workflow to resume after a long wait. Temporal Workflow Execution External Activity completion
The diagram below shows how the workflow connects an agentic investigation to scientific infrastructure. Activities call models and scientific tools, while platform policy and any required scientific approval govern which actions may execute.
Loading architecture diagram…
Diagram source
flowchart TD
S["Scientist"] -->|"Objective and approvals"| W["Temporal Workflow and agent loop"]
W -->|"Approval requests"| S
W <-->|"Model Activities"| M["Model services"]
W <-->|"Authorised tool Activities"| T["Scientific tools and pipelines"]
T <-->|"Jobs and results"| C["Kubernetes, HPC and services"]
T -->|"Outputs and provenance"| R["Scientific systems of record"]
R -->|"Evidence"| WTemporal preserves process state and coordinates execution. The agent interprets evidence and proposes the next step. Scientific tools perform the selected method, while systems of record retain the resulting evidence. Platform code implements the policy checks between proposal and execution.
Deterministic control around nondeterministic reasoning
In Temporal, workflow logic must meet deterministic replay constraints. Model calls and other external operations therefore belong in Activities.
The agent loop can itself be represented as a workflow: call the model, inspect its response, execute permitted tools or hand off to another agent, update the context and repeat until a stopping condition is reached. Temporal's integration with the OpenAI Agents SDK for Python runs this orchestration inside a Workflow, with model calls executed as Activities and external tool operations explicitly given Activity boundaries. Similar integrations exist for Pydantic AI and Vercel's AI SDK. The workflow defines how the loop progresses and recovers; it does not have to predetermine which scientific action the agent will select. Agentic orchestration and workflow orchestration can therefore be responsibilities within the same running process, not necessarily separate layers or competing products.
The main agent is not limited to invoking individual Activities. In this architectural pattern, specialist agents can themselves be represented as Temporal Workflows. The parent workflow can delegate to a specialist agent or launch a child workflow that coordinates a complex scientific procedure. Each specialist agent can run its own reasoning loop, invoke tools and delegate further work, returning its result to the parent investigation. Activities perform external operations; workflows compose those operations into larger scientific capabilities. Agent-to-agent delegation becomes workflow-to-workflow orchestration. When started as child workflows, these executions have their own histories and configurable lifecycle relationships with the parent. Temporal child workflows
Loading architecture diagram…
Diagram source
flowchart TD
subgraph W["Parent Workflow coordinates the agent loop"]
M["Activity: call model"] --> D{"Finish or escalate?"}
D -->|"Yes"| F["Complete or escalate"]
D -->|"No: proposed action"| P["Check policy and approvals"]
P -->|"Permitted external operation"| T["Activity: call tool or submit job"]
P -->|"Revise proposal"| M
T --> U["Update agent context from result"]
U --> M
end
T -->|"Call, submit or monitor"| E["External system: service, Kubernetes or HPC"]
E -->|"Response, status or result"| T
P -->|"Permitted delegation"| C["Child Workflow: specialist agent or scientific procedure"]
C -->|"Result"| UThe parent workflow coordinates the loop and may execute an Activity or start a child workflow. An Activity can perform compute directly, call an external scientific service, or submit or monitor a job on Kubernetes or HPC. The external system performs the work; the workflow tracks progress and coordinates what happens next. A long-running job need not occupy one Activity for its entire lifetime: submission and status checks can be separate Activities, with the workflow waiting between them. A specialist agent can run its own loop within a child workflow, using its own Activities to call models and external systems. Platform code implements policy checks and any approval waits. During replay, recorded Activity results restore progress without requiring fresh model decisions.
A recorded Activity result can be used when reconstructing workflow state. That allows a workflow to recover without asking the model to make an already recorded decision again. An Activity whose completion was not recorded may still be retried. Temporal replay
Permissions, budgets and approval rules remain application responsibilities. The workflow provides durable points at which the platform can check them, record a decision and wait for intervention.
Before a proposed action executes, the platform can check:
-
which tools the agent may use;
-
which data it may access;
-
how much compute it may spend;
-
which actions require human approval;
-
how results must be registered;
-
which failures may be retried;
-
when the process must stop or escalate.
These checks allow reasoning to adapt while preserving the platform's operating rules. For a long agent loop, consequential tool calls also need checkpoints or durable execution boundaries. Wrapping the entire loop in one Activity does not automatically give each internal step independent recovery.
Recovery, reproducibility and scientific validity
Recovering a process and reproducing a scientific result require different evidence. Workflow replay reconstructs control state from recorded events; it does not establish that running a model or calculation afresh would produce the same answer.
A reproducible analysis needs identifiable input data, pipeline and software versions, model versions, parameters and, where relevant, random seeds and execution conditions. An agentic investigation also needs the context supplied to the model, the returned proposals and the actions taken. Even with those records, a fresh model call may produce a different plan.
The architecture should therefore support both resuming the investigation and inspecting or repeating its scientific methods. Durable execution helps with the first; versioned pipelines and scientific provenance support the second.
But traceability is not validity either. A calculation can be reproducible and still use an inappropriate model, an unsuitable dataset or a method that does not answer the scientific question. A successful agent run is not, by itself, a valid scientific result.
There are three separate responsibilities: reliable execution establishes what happened; scientific provenance connects outputs to their inputs and methods; scientific evaluation assesses whether the evidence supports the proposed conclusion. A workflow can coordinate all three, but its completion status cannot substitute for the latter two.
This also separates operational gates from scientific judgement. Permissions, budget limits and required approvals should be enforced by platform code, not left as instructions for the agent to remember. Evidence checks can assess completeness, consistency and method-specific acceptance criteria. Interpretation still requires appropriate scientific review; agreement between agents is not, by itself, independent validation.
A design make test learn workflow
Consider an illustrative agentic design-make-test-learn cycle.
A scientist begins with an objective rather than a fully specified computational protocol: identify candidate molecules that improve a measured property while meeting agreed constraints. An agent retrieves the current project context, identifies relevant models and proposes a plan involving property prediction, structural modelling and literature research. Before confirmatory experiments begin, the scientist agrees how the candidates will be evaluated and what evidence would support or refute the hypothesis. Exploratory work can remain more flexible, with that distinction recorded.
Platform checks confirm that the proposed tools and compute cost are permitted. Activities run calculations or submit versioned scientific pipelines to Kubernetes or HPC. Results are registered against their input data, parameters, model versions and candidate identities.
The agent interprets those results. It might run another calculation, ask a specialist agent to investigate an anomaly or conclude that the current route is unpromising. When it identifies candidates for experimental work, the workflow pauses for scientific approval.
Approved candidates are sent to a laboratory. The process may now wait days or weeks. When experimental results return, the platform checks sample identities, units, completeness and the agreed quality criteria, then registers the observations with their provenance. Suppose the laboratory request completed successfully, but an assay control failed. The workflow records that outcome and routes it for investigation; the agent should not treat successful delivery as evidence that the candidate met the objective.
The workflow resumes and the agent reasons over the accepted evidence. It may propose a further experiment or draft a conclusion for scientific review. Permission to run the experiment and acceptance of its conclusion are different decisions. The scientific record retains the observations, the proposed interpretation and the reviewed decision as distinct, linked objects.
This investigation crosses agents, models, compute infrastructure, people, laboratories and systems of record.
An agent can dynamically shape the route through the loop. A durable workflow preserves the identity and state of the loop as a whole.
Process state agent state and scientific state
Agentic platforms need to distinguish three forms of state.
Process state
This describes what is currently happening:
-
which work has started;
-
what is waiting;
-
what failed;
-
what should be retried;
-
which approval is outstanding;
-
whether the process has completed.
A workflow engine can own this state.
Agent state
This describes what the agent currently knows and is using:
-
the working context;
-
retrieved knowledge;
-
the current plan;
-
intermediate tool results;
-
summaries of earlier reasoning;
-
specialist-agent responses.
The agent runtime and context infrastructure own this state.
Scientific state
This describes what the organisation knows about the project:
-
compounds, sequences, materials and samples;
-
datasets and model versions;
-
experiments and measured results;
-
hypotheses, proposed interpretations and reviewed decisions;
-
provenance and accountable actors.
Scientific systems of record own this state. Their value accumulates across investigations: later agents should be able to reuse the evidence and methods even when models, prompts and workflow implementations change. The record must preserve the difference between an observation, a hypothesis and an accepted conclusion.
An agent trace is not automatically scientific provenance. A workflow history is not project memory. A context window is not a system of record.
These forms of state need clear ownership and links, even when they share storage infrastructure. Confusing them makes it harder to recover a process, rebuild agent context or establish which evidence supports a scientific decision.
What the workflow engine still provides
Agentic orchestration does not remove the hard distributed-systems problems beneath a scientific platform. In scientific work, retries can spend compute budgets or consume samples, and missing execution records can break the connection between a decision and its evidence.
A job accepted before a worker crashes
Consider an illustrative failure: an Activity submits a GPU job to Kubernetes. Kubernetes accepts it, but the worker crashes before Temporal records successful submission. The job continues running. When the Activity times out and retries, a second submission could duplicate expensive compute.
A robust integration assigns a stable job identifier to the requested operation. On retry, it checks whether the matching job already exists and reconnects to it, verifying that it belongs to the same request. The workflow then waits for its result through a callback or polling.
Temporal manages the retry; the integration must make repeated attempts safe. Its documentation recommends idempotent Activities so that retries do not duplicate side effects. The same concern applies when submitting a laboratory request, where repetition may consume physical samples as well as money. Temporal Activities and idempotency
Cancellation and compensation
Cancellation needs explicit integration logic: cancel external jobs, withdraw laboratory requests where possible and record irreversible work. Stopping workflow execution alone does not undo its effects.
Change over time
A scientific process may outlive several versions of its workflows, agents, prompts and tools. Changes need a versioning strategy that preserves both execution compatibility and the meaning of decisions already made.
Operational evidence
Execution history should identify requested actions, external job references, retries and completion. Linking those records to agent proposals and scientific outputs allows a team to investigate a failure or explain a result.
What has changed for the workflow engine
A scientific pipeline describes an analytical method. The wider workflow describes how that method participates in an investigation.
With agents, that wider workflow can define lifecycle stages, execution boundaries, approval points and failure policy, while the agent chooses detailed steps within them.
The workflow defines a durable contract for how the investigation may proceed, leaving the detailed scientific choices to the agent.
The central design task is to define where an agent's proposal becomes an authorised operation, how completion is recorded, and how the next decision receives the resulting evidence.
That boundary must remain explicit as the plan changes. Dynamic reasoning increases the value of clear execution contracts because the full path is not known in advance.
Different responsibilities within one process
This architectural pattern illustrates why agentic orchestration and workflow orchestration need not be competing approaches. As the agent-loop and child-workflow examples above show, adaptive reasoning and delegation can operate within durable workflows. The agent selects the next action; the workflow preserves progress and coordinates execution. These remain distinct responsibilities even when they form part of the same scientific process.
Temporal can make the agent's actions part of a recoverable process while the agent continues to choose what to do next.
A brief research task may need only an agent runtime. A scientific process that spans weeks, launches expensive compute or interacts with laboratories needs stronger recovery, approval and provenance mechanisms.
Choose the runtime against those requirements, rather than its category. A separate workflow engine is useful when it supplies execution guarantees the agent runtime does not; adding one without a clear responsibility only introduces more coordination.
Design for scientific execution guarantees
Scientific pipelines preserve methods that can be tested and reused. Agents select and adapt the next action. Durable execution carries those decisions across failures, long waits and external systems. These responsibilities remain even when one runtime provides all three.
The question is not whether an agent has replaced the workflow engine. It is whether the architecture still provides the guarantees the science requires.
Agents can reshape an investigation as new evidence arrives. The architecture must still preserve what was authorised, what actually happened, which evidence resulted and how a scientific conclusion was reached. Whether the execution guarantees come from a workflow engine, an agent runtime or both matters less than whether they exist. Reliable orchestration makes the investigation recoverable; scientific records and evaluation make its conclusions inspectable and defensible.
