Does agentic complexity actually beat a fixed workflow?
Grading three analytics systems on 38 tasks with known answers, including whether they claim causation the data can't support.
The question
Most agent evaluations measure whether the SQL executed. That tells you nothing about whether the answer was right. Because this warehouse has planted causal effects, agent conclusions can be scored against truth rather than against plausibility.
What I did
I generated a normalised SQLite warehouse with deliberate traps: a Simpson's paradox where one acquisition channel is best in aggregate and worst within every segment, returns that erode revenue, missing campaign records, and seasonality that gives "revenue declined" a boring correct explanation.
I built the 38-task benchmark before building any of the systems, so the questions weren't written to flatter them.
Then three system designs ran against the same model, the same database, and the same task set, differing deliberately in how much tool access and orchestration each one gets: a schema-only single call, a constrained SQL → validate → execute → summarise workflow, and an autonomous multi-tool agent. The single-call baseline holds no tools at all, which is the point of a baseline rather than a shortcoming in it.
Each run is graded on six dimensions independently (syntactic validity, data selection, numerical correctness, method selection, conclusion correctness, and causal language), so a system that produces a running query and a wrong conclusion scores like one.
Results
| System | Numerical acc. | Conclusion acc. | Unsupported causal claims | Tool calls/task | Tokens/task | p95 latency |
|---|---|---|---|---|---|---|
| A: single LLM call | 0.40 | 0.237 | 14.3% | 0.0 | 784 | 43.3s |
| B: fixed workflow | 0.60 | 0.408 | 0.0% | 1.13 | 1,285 | 67.2s |
| C: autonomous agent | 0.64 | 0.355 | 0.0% | 2.08 | 4,892 | 172.2s |
The autonomous agent did not earn its complexity. It spent 3.8× the tokens and 2.6× the p95 latency of the fixed workflow and scored lower on conclusion correctness, 0.355 against 0.408. Its one apparent win, numerical accuracy, is the weaker metric here: a number can be right while the conclusion drawn from it is wrong.
The more interesting result is how the two systems reached the same zero rate of unsupported causal claims. They did it by opposite mechanisms. The fixed workflow grounded 71.4% of its refusals in a query it actually ran. The agent declined 71.4% of them without examining the data at all. Equal scores, and only one of them is a system you would trust: an agent that refuses by not looking is not being careful, it is being absent. A refusal-rate metric alone cannot tell those two behaviours apart, which is an argument for grading the mechanism rather than the outcome.
Design constraints
- Safety lives at the tool layer, not in the prompt. Read-only SQL, query timeouts, row caps, sandboxed Python with no network. Prompt-level restrictions are suggestions; tool-level ones are guarantees.
- The systems provably cannot read the ground truth. If the agent can read the answer key, the benchmark is worthless, so that's a test, and the test is mutation-verified.
- Retries cap at two, and retry count is a metric rather than a log line.
What I'd flag in review
- The tool layer is in-process Python, not a real MCP server, despite what one filename suggests. Renaming it is on the list.
- The warehouse uses its own data-generating process rather than the causal study's, so cross-project comparison of specific effect sizes isn't apples to apples.