Workflow or Agent: Deciding Which Steps Belong in Code and Which Belong to the Model
Contents 10 sections
- Two ways to build the same system
- Errors compound across steps
- A check that can fail is the cheapest reliability
- When a workflow wins, and when an agent is worth its cost
- What belongs in code and what belongs in the model
- Matching effort to the task
- A fixed retrieval pipeline or a model-directed search
- One agent or several
- When to take the model out of the critical path
- Conclusion
A system built on a language model can follow steps that its own code lays down, consulting the model only for particular judgments, or it can let the model choose each next step for itself. The first design is a workflow and the second an agent. Most useful systems sit somewhere between the two, and where the line falls decides how reliable the system is, what it costs to run and whether its behaviour can be tested at all. This article sets out how to place that line: which decisions are safer in code, which genuinely need a model, how much effort a task deserves, and when work should be divided among several agents.
Two ways to build the same system
A running example keeps the argument concrete. An online shop wants to automate its customer-support tickets, and a typical ticket reads “my order never arrived, I want my money back”. Handling it means finding the order, checking its shipping status, reading the refund policy, deciding whether a refund applies, and then either issuing it or writing a reply.
As a workflow, code fixes the steps and their order: look up the order, fetch the tracking status, call the model once to answer one narrow question (“does this ticket ask for a refund, and does the policy allow it?”), and act on the answer. As an agent, the model runs a loop: it reads the ticket, chooses a tool, reads the result, chooses again, and decides for itself when the task is finished. Code supplies only the tools and the limits. Anthropic’s engineering guide Building effective agents draws the same line: workflows are “systems where LLMs and tools are orchestrated through predefined code paths”, while agents “dynamically direct their own processes and tool usage”. Between the two poles lies a spectrum:
| Design | Who decides the next step | In the support example |
|---|---|---|
| Plain code | Code, with no model at all | a delivered order closes the ticket |
| One model call | Code, with a single judgment in the middle | classify the ticket as refund, tracking or other |
| Workflow | Code, calling the model at fixed points | classify, extract the order number, then draft the reply |
| Agent | The model, choosing tools in a loop | investigate an unusual complaint across orders, carriers and past tickets |
| Several agents | Several models, each with its own loop | one investigates, one drafts, one checks the policy |
Each row down the table buys flexibility and pays for it in predictability, money and testability. The same guide recommends “finding the simplest solution possible, and only increasing complexity when needed”, and the rest of this article is a way of applying that advice: choose the lowest row that still solves the task.
Errors compound across steps
Every step a model decides is a step that can go wrong. If each step succeeds independently with the same probability p, the chance that a task of n steps runs cleanly from start to finish is p multiplied by itself n times, pn.
Figure 1: Exact values of p to the power n, assuming every step succeeds or fails independently of the others.
A step that succeeds nineteen times in twenty sounds dependable, and a twenty-step task built from such steps runs cleanly only about one time in three. That is the central argument for workflows: every step moved out of the model and into correct code stops multiplying. Two caveats pull in opposite directions. Real agents can notice a mistake and recover, which makes the independent-steps picture too pessimistic; and mistakes can be correlated, so that one early misunderstanding spoils every later step, which makes it worse. The direction of the lesson survives both: fewer model-decided steps leave fewer chances to go wrong.
A check that can fail is the cheapest reliability
The first caveat, recovery, is worth engineering on purpose rather than hoping for. Suppose code checks the output of each step (that it parses, that the order number it cites exists, that the refund it proposes is within policy) and asks the model again when the check fails. The next figure computes what that buys under simple assumptions: each attempt succeeds with probability 0.95, the check catches a given fraction of failures and never rejects a correct answer, and a caught failure is retried.
Figure 2: Exact values. A step’s success is 0.95 times the sum of (0.05 × detection) to the powers 0 to r, where r is the number of retries; the task needs 20 such steps.
Two properties make this the best bargain in the design. The check is ordinary code, so it can be unit-tested like any other function. And the retries are cheap, because they are rare: at a 5 per cent failure rate, a perfect check with one retry adds one extra model call for every twenty steps on average. What the figure also shows is that the check’s coverage matters more than the number of retries. A check that sees half of the failures leaves most of the loss in place, however patiently the system retries.
When a workflow wins, and when an agent is worth its cost
The practical test is whether the flowchart could be drawn before the case is seen. “Where is my order?” tickets pass it: the steps never change, so they belong in code, with the model called only where the flowchart says “judge this”. An unusual complaint fails it. A customer who reports three damaged orders, where the tracking shows a change of carrier, needs an agent, because each finding decides the next question and no one can write the steps in advance.
| A workflow fits when | An agent is worth it when |
|---|---|
| the steps are known before the task starts | the next step depends on what the last one found |
| many similar tasks arrive | each case is an open-ended investigation |
| a wrong action is costly | a slow or imperfect answer is cheap |
| every decision must be tested, audited or explained | a person reviews the result anyway |
| the budget for cost and latency is tight | the task is rare and valuable enough to justify the cost |
The test is not hypothetical. On SWE-bench Lite, a benchmark of 300 real GitHub issues from popular open-source Python projects, the Agentless system replaced an autonomous agent with what its authors call “a simplistic three-phase process of localization, repair, and patch validation, without letting the LLM decide future actions or operate with complex tools”.
Figure 3: From Table 1 of the Agentless paper (version 2), keeping the systems that report an average cost. The systems use different models; the two that both use GPT-4o are labelled with it.
The lesson is not that structure always wins. The retrieval baselines in the same table, which fetch likely files once and ask for a patch once, resolve almost nothing: too little structure of the wrong kind. Agentless works because its fixed steps match how a repair actually proceeds (find the place, change it, test the change), and because its tests give it exactly the kind of check the previous section described. When the shape of the work is known, encoding that shape in code is worth more than granting a model the freedom to rediscover it on every run.
What belongs in code and what belongs in the model
The dividing rule is short: code takes anything with a right answer that can be computed or checked, and the model takes judgments about language and open situations.
| Put it in code | Give it to the model |
|---|---|
| the order of steps and the branches between them | reading messy text and deciding what it means |
| limits: maximum steps, token budget, timeouts | choosing among options described in words |
| permissions: which tools, which records, which amounts | writing replies, summaries and drafts |
| retries, made safe to repeat | judging relevance, intent and quality |
| checking output: its format, and that its facts appear in the source | planning when the path depends on what is found |
| anything exact: arithmetic, dates, lookups, counting | explaining a decision in plain words |
| a log of every step, for later audit |
Everything in the left column can be covered by ordinary unit tests and behaves the same way every time. The model is then used only where no test could capture the right answer, which is precisely where its judgment is worth paying for. A code-review assistant shows the split cleanly: linters, static analysis and the test suite run as code; the model judges only what they cannot, such as whether a design is sound or a name misleading; and code removes duplicate comments and enforces the budget. The companion article on tool calling makes the same argument one level down, for the limits a tool should enforce itself.
The rule also cuts the other way. Some tasks look like rule-following and are not. Deciding whether a support message is a complaint looks like a keyword rule until real messages arrive: “great, another late delivery” contains no complaint word and is a complaint, while “I’d complain if it weren’t so fast” contains one and is praise. When a hand-written rule keeps missing real cases, the task is a judgment, and it belongs to the model after all.
Matching effort to the task
Agents can also fail by doing too much: planning ten steps for a one-step task, re-checking what is already known, or reasoning at length about an easy question. The tendency is measurable. Asked “what is the answer of 2 plus 3?”, reasoning models of the o1 type consumed on average 1,953 per cent more tokens than conventional models to reach the same answer, in a study of their overthinking. Four controls keep effort proportional:
- Budgets. A hard cap on steps and tokens per task, enforced in code, with a plain message to the model as it approaches the cap.
- Plans sized to the task. Plan only as far ahead as the next finding can change, and replan when a result is surprising rather than at every step.
- Routing. Send easy, recognizable requests down a cheap path (a workflow or a small model) and keep the full agent for the rest. A router is itself a classifier, so its mistakes have to be measured like any other classifier’s.
- Stop conditions. Define “done” as a check that can fail, not as the model’s sense that it has finished.
A fixed retrieval pipeline or a model-directed search
Retrieval, finding the right documents before answering, is where the choice arises most often. A fixed pipeline searches once with the question, takes the top results and answers. A model-directed search lets the model decide what to look for, read, and search again when the first attempt misses. The FlashRAG toolkit ran both kinds under identical conditions, the same answering model, retriever and corpus, on six question-answering datasets.
Figure 4: From Table 3 of the FlashRAG paper (version 2). Exact match for NQ, TriviaQA and WebQuestions; F1 for PopQA, HotpotQA and 2WikiMultihopQA. Llama-3-8B-Instruct answers from five passages retrieved by E5-base-v2.
No method wins everywhere. IRCoT, which retrieves again at each step of its reasoning, gains most on the multi-hop datasets, where an answer requires combining facts from more than one document, and loses slightly on the two easiest. FLARE, which retrieves only when the model judges itself uncertain, loses heavily on three datasets; its scores track those of answering with no retrieval at all, which suggests it seldom chose to search. On 2WikiMultihopQA even that beats the fixed pipeline, so part of every gain in that column is the fixed pipeline’s own failure rather than the alternative’s success. The standard of proof follows directly. A model-directed search must win on the hard questions to justify its cost, and must not lose on the easy ones, and both pipelines have to be compared on the same questions with several runs each, as the companion article on evaluating agents describes.
One agent or several
Dividing a task among several agents, each with its own context and its own loop, has become a common next step after a single agent. It helps in some situations and does harm in others. Anthropic reports that its multi-agent research system, a lead agent directing parallel sub-agents, outperformed a single agent “by 90.2% on our internal research eval”, and also that “multi-agent systems use about 15× more tokens than chats”, against “about 4× more” for a single agent (How we built our multi-agent research system). The same article names the poor fits: “domains that require all agents to share the same context or involve many dependencies between agents”.
A controlled study makes the boundary sharper. Kim and colleagues compared a single agent with four multi-agent architectures across six agentic benchmarks and three model families, holding the total token budget fixed so that only the organization changed.
Figure 5: Every configuration that section 4.2 of Kim et al. (version 3) reports in its text, as the change in mean success against a single agent given the same total number of tokens. Independent agents work in parallel without communicating.
Financial analysis splits naturally into revenue, costs and market factors that can be studied separately, and there several agents did far better than one. Sequential planning, where each move depends on the state the previous move left behind, degraded under every architecture. The authors trace this to the fixed budget: tracking the state of the plan consumes most of the tokens before the agents can communicate, and the messages that follow degrade the reasoning rather than inform it. Across all six benchmarks and every architecture, the mean change was −0.3 per cent, so “use more agents” is not a default worth adopting.
Several agents help when the work divides into independent pieces that can run in parallel, when a piece needs a great deal of reading but only its conclusion matters, when roles need different permissions (a reviewer that cannot write), or when elapsed time matters more than total cost. They make things worse when the pieces depend on one another, when a handoff loses context the next agent needed, or when agents share state and can overwrite each other. A sub-agent is warranted for a task that is narrow, independent and returns a short result, such as “search these five files and report which one defines this function”. It is not warranted when it would need most of its parent’s context to do the job. Whatever the task, a handoff has to state the goal, the exact inputs, the rules that apply, what counts as done and what to return. A handoff is a prompt, and a vague one produces confident work on the wrong problem.
When to take the model out of the critical path
The decision also runs in reverse: a task once given to a model sometimes belongs back in code, or off the path a request has to wait for. The signals are recognizable. The model’s errors cost more than its judgment is worth, and the incidents keep coming. The task has become well defined, because enough cases have now been seen that correct rules can be written and tested. The recipe is fixed but fragile, so the model knows what to do and keeps getting one detail wrong. The latency or cost no longer fits the budget. Or every decision must now be explained, and the model’s decisions cannot be tested.
None of this requires the model to disappear. It can move from deciding to suggesting, with code making the final call, or from the live path to a background job whose results are checked before anyone relies on them.
Conclusion
The choice between a workflow and an agent is made step by step rather than once per system. Each step should be owned by code when its right answer can be computed or checked, and by the model when it calls for judgment about language or an open situation. Compounding errors make every model-decided step expensive, checks in code recover most of that expense cheaply, and the published comparisons point the same way: a fixed pipeline that matches the shape of the work resolved more issues than every agent that reported its cost, while spending less than most of them, a model-directed search earns its keep only on the questions that need it, and several agents help only when the work genuinely divides. The most useful design question is therefore not how capable an agent can be made, but how much of the task can be taken away from it without losing what only a model can do.