Does agentic complexity actually beat a fixed workflow?

Grading three analytics systems on 38 tasks with known answers, including whether they claim causation the data can't support.

The question

Most agent evaluations measure whether the SQL executed. That tells you nothing about whether the answer was right. Because this warehouse has planted causal effects, agent conclusions can be scored against truth rather than against plausibility.

What I did

I generated a normalised SQLite warehouse with deliberate traps: a Simpson's paradox where one acquisition channel is best in aggregate and worst within every segment, returns that erode revenue, missing campaign records, and seasonality that gives "revenue declined" a boring correct explanation.

I built the 38-task benchmark before building any of the systems, so the questions weren't written to flatter them.

Then three system designs ran against the same model, the same database, and the same task set, differing deliberately in how much tool access and orchestration each one gets: a schema-only single call, a constrained SQL → validate → execute → summarise workflow, and an autonomous multi-tool agent. The single-call baseline holds no tools at all, which is the point of a baseline rather than a shortcoming in it.

Each run is graded on six dimensions independently (syntactic validity, data selection, numerical correctness, method selection, conclusion correctness, and causal language), so a system that produces a running query and a wrong conclusion scores like one.

Results

Three systems across 38 tasks with known answers.
SystemNumerical acc.Conclusion acc.Unsupported causal claimsTool calls/taskTokens/taskp95 latency
A: single LLM call0.400.23714.3%0.078443.3s
B: fixed workflow0.600.4080.0%1.131,28567.2s
C: autonomous agent0.640.3550.0%2.084,892172.2s

The autonomous agent did not earn its complexity. It spent 3.8× the tokens and 2.6× the p95 latency of the fixed workflow and scored lower on conclusion correctness, 0.355 against 0.408. Its one apparent win, numerical accuracy, is the weaker metric here: a number can be right while the conclusion drawn from it is wrong.

The more interesting result is how the two systems reached the same zero rate of unsupported causal claims. They did it by opposite mechanisms. The fixed workflow grounded 71.4% of its refusals in a query it actually ran. The agent declined 71.4% of them without examining the data at all. Equal scores, and only one of them is a system you would trust: an agent that refuses by not looking is not being careful, it is being absent. A refusal-rate metric alone cannot tell those two behaviours apart, which is an argument for grading the mechanism rather than the outcome.

Design constraints

What I'd flag in review

← Back to home