How the measurement playbook needs to evolve, and what companies should be doing to quantify AI impact
- Customer Experience Live

- 4 days ago
- 4 min read
I've been measuring things for nearly a decade. I started in market research, moved into business cases and value consulting. Along the way I’ve found myself consistently asking my customers: What do you want to do? How do you want to do it? How will you measure it?

In light of the pace of change in this time, what's striking to me is how much has stayed the same.
The fundamentals haven't moved. Everyone still cares about baseline plus uplift equals new number. A clear before, a measurable after, and a story you can tell a CFO in under two minutes. That will never change. Business cases live or die on simplicity.
But the landscape around that core has changed significantly. We can measure more things now, because we have access to more information. We care about different things, because we can impact them in ways we previously couldn't.
Now that AI has entered the chat, some of what used to be a given, how many marketing assets a team can produce in a quarter with a fixed set of resources, for example, is no longer fixed at all. Capacity is now a significant lever. Time spent on a task is a metric worth tracking because we can influence it significantly with technology.
The conundrum
FOMO took over. AI was new, the pressure to do something was real. So, organizations onboarded tools while throwing the traditional software assessment playbook out the window. Two or three years down the line, execs are asking: what was the impact? But the measurement framework was never placed, and teams are now backtracking to justify investments that were made on instinct rather than intent.
So, what's the solution?
Go back to basics
When things get complicated, the answer is almost always to simplify. The foundation of any good measurement approach is (and always has been) capturing the baseline before you start.
Before any agent goes live, you need a clear picture of how things work today. How long does a campaign cycle take? How many approval rounds does a piece of content go through? What is the current conversion rate? What does it cost to produce an asset? How many campaigns can this team deliver in a quarter?
Connect productivity to growth
The measurement playbook needs to track AI impact across two dimensions. Productivity captures how teams do more with less: faster campaign cycles, lower cost per asset, more campaigns produced with the same headcount. Growth captures what that efficiency makes possible: higher conversion rates, stronger revenue per visitor, better experiment win rates.

The sweet spot is where both move together, and the key is translating the numbers into language that lands with your execs.
If a team goes from producing 8 campaigns a month to 12, that's a 50% productivity gain. But if those four additional campaigns each drive incremental conversions, the story becomes: here is the revenue we generated that we couldn't have generated before, without adding any additional resources or headcount.
Build attribution discipline in from the start
One of the trickiest parts of measuring AI impact is isolating it from everything else happening simultaneously. A new campaign launches, the market shifts, a competitor pulls budget, and it becomes hard to pinpoint what specially moved the needle.
This is a problem we've had to solve ourselves. To measure the impact of Optimizely Opal (our Agentic Orchestration Platform) on our own customers, we built a methodology designed to control for confounding factors and ensure consistent attribution:
We compared users to non-users. Customers who adopted Opal were measured against a control group of customers who didn't, over identical time windows, giving us a clean like-for-like comparison rather than just before-and-after numbers in isolation.
We isolated Opal's contribution. We subtracted what the control group did over the same period to strip out the impact of seasonal patterns, product updates, broader market shifts.
We weighted results by baseline activity. Customers with more pre-adoption activity carry more weight in the final number, preventing outliers from distorting the headline result.
The output is a net uplift figure. For each metric, the number reflects real incremental change attributable to Opal, not just improvement over time.
It's not the only way to do it. But the principle, controlling for what you can't change and isolating what you can attribute, applies regardless of the tools you're measuring.
The measurement flywheel

The most successful programs have measurement within their DNA. Keeping things simple, measurable and explainable as a consistent habit goes a much longer way than doing an audit once every 2 years.
Define goals. Capture baseline. Execute. Measure uplift. Communicate results. Integrate learnings. Repeat. This allows translating learnings and compounding marginal gains over time.
That's not a new idea. I was doing a version of it in market research a decade ago. What's new is the scale at which it can now operate, and the range of things worth measuring. Don't get distracted by the complexity of agents and orchestration and agentic workflows. The fundamentals are the same as they've always been.
Baseline plus uplift equals new number. Know where you started. Know where you got to. Tell that story clearly. That’s all you need.
Source: Optimizely
.jpg)




Comments