I ran the same contract-extraction task two ways. One was a fixed pipeline. The other was an agent that chooses its own steps, using a search tool and a calculator. The setup is deliberately small, so treat the numbers as an illustration, not a benchmark.
Accuracy was essentially tied. Across three runs, the pipeline missed 1 of 80 fields per run and the agent missed 2. The agent used about 5.6× as many tokens: 1,981 per contract on average, against 356. Neither approach changed a single answer between runs, but the agent’s cost did change. The largest run-to-run difference we observed on one contract was 1,848 tokens for the agent, against 7 for the pipeline.
So the question this post is about: if the steps are already known, what exactly are we buying by giving the model control of the flow?
Put control flow where the knowledge lives
The rule of thumb I use: decide who owns the control flow by asking who knows the next step, and when.
- If the sequence of operations is known when you design the system, encode it in code. Extracting four fixed fields is like this. Every document goes through the same steps, so asking the model what to do next has one right answer, and you already know it.
- If the next step depends on information you only discover at runtime, giving the model control can make sense. When step N+1 depends on what step N returned, code can’t express the path without enumerating every branch.
Anthropic’s Building effective agents draws the same line. Workflows orchestrate LLMs and tools through predefined code paths. Agents let the model direct its own process and tool use. Their advice is to add agentic complexity only when it demonstrably helps. This experiment is a small attempt to put a number on “demonstrably.”
The setup
Task: pull four fields out of vendor contracts: the two parties, the effective date, the annual value in whole dollars, and the termination notice in days.
Data: 20 short contracts written for the experiment, in five groups. Each group has a different kind of mess:
| Group | n | What makes it hard |
|---|---|---|
clean | 5 | nothing; the baseline |
date | 4 | dates in different formats (07/04/2024, the 1st of August, 2023, 2024.10.01) |
words | 3 | amounts written out in words |
math | 4 | annual value isn’t stated and has to be computed |
buried | 4 | the right number sits next to decoys: late fees, insurance limits, superseded prices |
Controls:
- Both approaches use the same model (
openai/gpt-oss-120bon Groq), the same field spec,temperature=0andmax_tokens=1024. - Both are graded by the same exact-match scorer, with light normalization for case,
$and commas. - Every contract runs three times through each approach.
Pipeline: one call with a fixed schema, then the JSON is parsed. No loop and no tools. In particular, no calculator: it has to do the math group’s arithmetic on its own.
resp = client.chat.completions.create(
model=MODEL, max_tokens=MAX_TOKENS, temperature=0,
messages=[{"role": "system", "content":
"Extract these fields from the contract and reply with ONLY a JSON object: " + FIELDS},
{"role": "user", "content": contract["text"]}])
Agent: a standard tool loop capped at 8 turns, with two tools:
find_line(keyword)returns up to three sentences of the contract that contain the keyword.calculate(expression)is an exact arithmetic evaluator that only accepts numbers and+ - * /.
The model decides which tool to call, in what order, and when to stop and answer.
Measured:
- Field-level accuracy.
- Total tokens per contract as reported by the API: prompt plus completion, which for gpt-oss includes its reasoning tokens.
- Whether answers or token counts changed across runs.
Results
Table 1: Overall
| Pipeline | Agent | |
|---|---|---|
| Field accuracy (mean of 3 runs) | 99% | 98% |
| Accuracy range across runs | 99–99% | 98–98% |
| Model calls per contract | 1.0 | 3.5 |
| Tokens per contract (mean) | 356 | 1,981 |
| Agent ÷ pipeline tokens | 5.6× | |
| Contracts whose answer changed between runs | 0 / 20 | 0 / 20 |
| Largest token swing on one contract across runs | 7 | 1,848 |
Table 2: By contract group
| Group | n | Pipeline acc. | Agent acc. | Pipeline tokens | Agent tokens | Ratio |
|---|---|---|---|---|---|---|
clean | 5 | 100% | 100% | 323 | 1,680 | 5.2× |
date | 4 | 94% | 94% | 356 | 1,764 | 5.0× |
words | 3 | 100% | 100% | 389 | 1,160 | 3.0× |
math | 4 | 100% | 100% | 346 | 3,039 | 8.8× |
buried | 4 | 100% | 94% | 382 | 2,130 | 5.6× |
Token columns are per-contract means. I computed the ratio column from them.
Every miss was in the parties field, and each one repeated in all three runs. Contract D1 was missed by both approaches. B3 was missed by the agent only. The scorer is exact-match, and the output I kept records which field was missed but not what the model returned. So I can’t tell you whether these are real misreads or formatting mismatches, such as a variant hyphen in “Weyland-Yutani”.
The math and buried groups were designed as the agent’s home turf, and that’s where it did no better. math was a tie at 100%, and buried is where the agent’s one extra miss came from.
The calculator result
In the math group the annual value is never written down. One contract says “a base fee of $40,000 per year plus support at $2,000 per month.” The answer is 64,000, and something has to compute it.
The agent had a calculator built for exactly this. The pipeline had to do the arithmetic on its own. Both scored 100%. The agent spent 3,039 tokens per contract on this group against the pipeline’s 346, about 8.8×. That is the widest gap of any group.
The calculator removed the arithmetic, but not the work around it:
- The model reasons about what to do and emits a tool call.
- The loop appends the call’s result.
- The next request re-sends the whole conversation (system prompt, contract, every earlier call and result) before the model can continue.
The pipeline sends the contract once. The agent sends an increasingly long conversation on every turn. Here is the loop, condensed slightly from the notebook:
for _ in range(max_turns):
resp = client.chat.completions.create(model=MODEL, messages=messages,
tools=TOOLS, max_tokens=MAX_TOKENS, temperature=0)
meter.add(resp) # each turn is billed for the full history
msg = resp.choices[0].message
if resp.choices[0].finish_reason != "tool_calls":
return parse_json(msg.content)
messages.append(msg)
for tc in msg.tool_calls:
messages.append({"role": "tool", "tool_call_id": tc.id,
"content": run_tool(tc)}) # history grows
return {} # out of turns counts as a miss
I didn’t save per-turn traces from these runs, but the loop shows why the cost is structural rather than incidental.
A tool only buys accuracy if the model would get the step wrong without it. For this model, multiplying 2,000 by 12 wasn’t the hard part, so the calculator added cost and nothing else. With a weaker model that might change. If it does, there’s a cheaper fix that keeps the known step in code: ask the model for {"amount": 2000, "per": "month"} and do the multiplication yourself.
Stable answers, unstable cost
To me, this is the more interesting result.
At temperature=0, neither approach changed any answer across the three runs: 0 of 20 contracts for both. The answers were stable. The work was not. The pipeline’s largest per-contract swing was 7 tokens. The agent’s was 1,848, nearly the size of its average cost per contract. One contract with the same model, settings and answer cost about 1,800 tokens more on one run than on another.
In this implementation, that comes from dynamic control flow. The number of turns is decided at runtime, and each extra turn re-sends a growing history, so a small difference in decisions becomes a large difference in tokens. Pipeline cost depends on input length. Agent cost also depends on the choices the model happens to make on that run.
Operationally, that shows up in a few places:
- Budgeting and forecasting. A mean cost per document hides the spread. Forecasting needs the distribution.
- Latency and timeouts. More turns mean more sequential round trips. A timeout or
max_turnscap tuned on typical runs turns the long ones into failures. - Rate limits and capacity. Token-per-minute limits are hit by peaks, not averages. The notebook’s run instructions warn about 429s specifically because the agent makes several calls per contract.
- Testing. An eval that checks answers says nothing about cost. Regression tests need cost assertions too.
None of this makes agents unreliable. It means that with dynamic control flow, resource consumption is a separate thing to budget, cap and monitor, alongside accuracy.
When an agent is worth it
This isn’t an argument that pipelines beat agents. It’s about where the boundary is. An agent becomes compelling when you can’t reliably write the path down ahead of time:
- Open-ended research, where what to look at next depends on what you just found.
- Iterative investigation, like debugging: read the error, decide what to inspect, repeat.
- Variable tool use, where the set and order of actions changes per input in ways you can’t enumerate.
- Too many paths for a workflow, where a fixed pipeline would need so many branches that the branching itself becomes the brittle part.
Even then, the right shape is often a pipeline with one bounded agent step for the part that genuinely branches, rather than an agent wrapped around everything.
Limitations
- The task is simple. Both approaches were close to the accuracy ceiling, which leaves the agent little room to buy anything.
- The contracts were written for the experiment. There are 20 short ones, with no appendices, scanned text or 40-page documents. Messier inputs are where an agent’s search tool might pay off.
- One model, three runs. It’s one model at
temperature=0, with no significance testing. A different model or provider could change both the accuracy and the ratio. - Tokens are a proxy. No dollar cost or wall-clock latency was measured.
- The scoring is strict. The three misses (all parties) weren’t separated into misreads versus formatting mismatches.
- The agent wasn’t optimized. A more tightly prompted agent might take fewer turns.
Most importantly: this experiment does not show that a pipeline is sufficient in general, or that agents are not useful. It measures a narrower question. When a task’s control flow is already well specified, how much additional value does runtime autonomy provide? Here, runtime autonomy bought us essentially the same answers, at roughly 5.6× the token cost and with a much wider cost distribution.
Reproduce the experiment
The full notebook has everything:
- the 20 contracts and gold labels
- both implementations and the tools
- the scorer and token meter
- the cells that produce every table above
It runs on Groq’s free tier: 02-contract-extraction-pipeline-vs-agent.ipynb. Absolute token counts will differ by model and provider. The ratio and the spread are what’s worth comparing.