02 · AI evaluation
Does agentic complexity actually beat a fixed workflow?
A benchmark that grades an analytics agent's conclusions against known causal truth: not whether the SQL ran, but whether the answer was right and whether it overclaimed.
Farida Sakr
Partner Data Scientist, Meta
Causal inference · AI evaluation · B2B decision science
I build things that make the difference measurable: synthetic data with planted ground truth, so estimates and agents can be scored rather than argued about.
Track record
Bell 2023–2025 · Meta 2025–present
$2M
cost savings from replacing third-party crowdsourced data with an internal solution
28%
reduction in congestion-related network impact
100+ hrs
of manual reporting eliminated per event cycle through the FIFA World Cup reporting pipeline
40+
users of Bell's Data-as-a-Service dashboard
Case studies
01 · Causal inference
A naive difference in means overstated a free-shipping promotion's effect by 3.6×. Cross-fitted AIPW recovered the planted truth to within 9%, and the economics flipped the recommendation.
Naive
$23
AIPW
$6
Error vs truth
9%
02 · AI evaluation
A benchmark that grades an analytics agent's conclusions against known causal truth: not whether the SQL ran, but whether the answer was right and whether it overclaimed.
The arc
Studies
What is likely to happen?
Bell
Why is it happening?
Meta
What should we do?
Now
Did what we did actually create value?
Scope
Canada · Bell, through 2025
United States · Meta, now
LATAM · regional scope