article » Every AI ROI number you've seen is wrong

Every AI ROI number you've seen is wrong

April 20, 2026
9 min read

By Benjamin Falk

Guest essay. Originally published on LinkedIn on April 20, 2026; republished here with the author’s permission after he left the platform. Ben is an AI strategist, economist and founder based in London.

Billions are being spent on AI tools across financial services, professional services, and the broader knowledge economy. Boards are demanding returns, CEOs are mandating AI usage, CFOs are asking for business cases, and consultants are happily supplying answers.

The problem is that almost every methodology being used to measure those returns is fundamentally broken. There are 5 major problems with the way AI ROI is being measured:

Problem One: Productivity Is a Manufacturing Concept

The dominant framework for measuring AI ROI borrows from the manufacturing industry: outputs divided by inputs. The term itself originates in the early 19th century, and increased in popularity late in the same century as the Industrial Revolution gained steam. It’s a concept grounded in homogenous, tangible goods, where the volume of outputs are easily measured and where market prices are transparent and established.

But this concept, despite being widely applied, doesn’t work well for services. In this sector, there is no tangible “product” which renders the entire concept problematic. In law, consulting, finance, healthcare, and other services sectors, outputs are intangible, heterogeneous, quality dependent, frequently highly customized, and usually deeply relational.

This of course isn’t a new problem. Economists have been arguing about services sector productivity measurement for literally decades. Baumol’s cost disease, articulated in 1966, identified that services productivity is structurally harder to improve than manufacturing productivity, in part because quality and output are so difficult to disentangle. Forty years of national accounting methodology has grappled with exactly this without clear answers, only trade-offs to manage.

Yet when enterprises try to measure the ROI of an AI tool, they typically reach for the crudest possible version of the manufacturing framework: “We saved X hours. An hour costs Y. Therefore we saved X×Y.” This is not measuring what actually matters, as time freed is not equivalent to value created. An analyst who completes a task in two hours instead of four hasn’t necessarily produced twice the value because they may have produced work of inferior quality, or the time saved may have been absorbed into unproductive activity elsewhere. A simple hours-times-salary model ignores critical aspects of production entirely.

What you actually need to measure across your service workforce is a multidimensional outcome vector: task completion time matters, but so does output quality (which requires independent assessment), error rates, client or customer satisfaction, employee experience, and downstream business outcomes. Each of these has its own measurement challenges. Together, they give you something that actually resembles a productivity signal for knowledge work.

Problem Two: You Haven’t Run a Proper Experiment

Even if you get the outcome measurement right, most enterprise AI deployments make a second, equally fatal error; they don’t control for exogeneity.

Exogeneity is the economist’s word for the problem where variables or factors outside of the model remain unaccounted for. When you roll out an AI tool to your entire workforce and then measure productivity changes, you have no idea whether the changes you observe are caused by the AI tool, or by the ten other things that changed simultaneously such as macroeconomic conditions, team composition, process redesign, the fact that you just promoted three senior people, or the fact that your competitors are struggling and your deal flow improved.

As everyone knows (but doesn’t seem to correct for), correlation is not causation. The gold standard for establishing causation is the randomised controlled trial, the methodology that pharmaceutical research has used for decades to establish whether a drug actually works. The logic is simple; randomly assign your population to two groups, give one group the treatment and the other nothing, keep everything else constant, and measure the difference in outcomes.

Applied to enterprise AI, this means taking a sufficiently large cohort of employees with equivalent roles and responsibilities, randomly allocating them to an experiment group (with AI tool access) and a control group (without), assigning both groups equivalent tasks, and measuring performance across your full outcome vector. The randomisation is what does the causal work by ensuring that any systematic differences between the groups at baseline are eliminated by chance, so that observed differences in outcomes can be attributed to the tool.

Almost nobody in enterprise AI deployment is doing this, and therefore no one actually knows if AI is generating real gains.

Problem Three: You’re Only Measuring One Node in the Production Chain

There is a subtler measurement failure that sits between getting the outcome variables right and getting the experimental design right, and it is silently inflating results even when everything else is done correctly.

Most AI ROI studies measure productivity at the point of intervention. The worker using the AI tool completes tasks faster, that gets measures and its called a productivity gain. What isn’t being measured is what happens up and downstream of the task being monitored.

Knowledge work is not a series of isolated tasks. It is a production chain, and like any production chain, it has dependencies and a flow of value creation. The output of one node is the input to the next, meaning compressing one node can create pressure on adjacent nodes in the production pipeline.

You can think of it like squeezing a balloon. Compress one end and the other end expands. If an AI tool helps a junior analyst produce first drafts twice as fast, but those drafts now require more review time from a senior analyst because quality has declined in ways that aren’t immediately visible, the productivity gain at the junior level is partially or wholly offset by a productivity cost at the senior level. Your headline number looks impressive whilst your actual system-level output has barely changed or even deteriorated.

This happens in both directions. If using an AI tool requires more careful task specification, prompt engineering, or input preparation, the workers upstream feeding the process may be spending more time than before; time that never shows up in your measurement because you’re only watching the tool user. Likewise, if AI-generated outputs require more verification, correction, or client-facing explanation, the workers receiving those outputs absorb costs that are invisible to your ROI model.

The correct unit of measurement is not the individual worker using the tool. It is the full production chain within which that worker operates. This means mapping the workflow end-to-end before designing your measurement framework, identifying all the roles that touch the process, and measuring productivity across all of them not just the ones with direct AI access.

This is harder than it sounds, because production chains in professional services are rarely linear or formally documented. They’re emergent, relational, and context-dependent. But the difficulty of measuring them correctly is not a reason to measure them incorrectly. It’s a reason to invest in the mapping work before you run the experiment.

The organisations claiming the largest AI productivity gains are almost universally measuring at the node level. Until they measure at the chain level, those numbers should be treated as upper bounds, not realistic estimates.

Problem Four: RCTs Are Necessary But Hard to Execute

RCTs are the right methodology for measuring the ROI of AI. They are also genuinely difficult to implement well in knowledge work environments.

In pharmaceutical trials, the treatment group gets the drug and the control group gets the placebo. They don’t collaborate or discuss how they are feeling with one another. But in knowledge work, your experiment group and control group sit next to each other, share documents, discuss approaches in meetings, and ask each other questions. Treatment effects leak, which can contaminate the results. This is known as a SUTVA violation, or a breach of the Stable Unit Treatment Value Assumption and it can quietly destroy your experimental validity.

There are other difficulties as well. Hawthorne Effects, where both groups might know they’re being observed and measured, can cause workers to change their behaviors independently of the treatment. The experiment group may work harder because they feel privileged to have early access. The control group may work harder because they feel they’re being judged. Neither effect has anything to do with the AI tool.

Furthermore, assigning “equivalent tasks” to two groups of knowledge workers in a live enterprise environment is extremely hard without creating artificial conditions. If the tasks are too synthetic, the results won’t generalise; if they’re too real, equivalence is impossible to guarantee because outputs are so heterogeneous.

Finally, AI tools have learning curves as users gain experience leveraging them. A 30-day RCT will systematically understate long-run productivity gains because your treatment group is still learning how to best use the tool. Short trials bias toward null results, which may cause organisations to incorrectly abandon tools that would have delivered value at scale over time. But longer trials can be costly, requiring patience and consistent funding which are in short supply in the current hype-filled environment.

None of these problems are reasons to abandon RCTs. They are reasons to design them more carefully, with protocols that minimise contamination, blinding where possible, longer trial durations, and honest acknowledgement of residual uncertainty in the results.

Problem Five: The Portfolio Attribution Problem

There is a fifth problem that rarely gets discussed at all. Even a well-designed RCT on a single tool doesn’t tell you the ROI of your entire AI investment programme.

Most organisations are deploying multiple AI tools simultaneously whilst simultaneously redesigning processes, restructuring teams, and making other complementary capital investments. The productivity gains (or losses) that emerge from this environment are a function of the entire intervention portfolio, not any individual tool.

Attributing overall performance changes to a specific AI investment requires either heroic assumptions or a factorial experimental design that almost no enterprise is equipped to run. This doesn’t mean you shouldn’t measure at the tool level. It means you should be honest about what tool-level measurement can and cannot tell you about programme-level returns.

What This Means In Practice

If you’re a CFO, a CTO, or a board member being shown AI ROI numbers, ask the following 4 questions:

What exactly was measured? If the answer is hours saved multiplied by average salary, the number is not credible. Ask what happened to quality, error rates, and customer and employee satisfaction.

What is the counterfactual? If there was no control group, there is no causal claim. There is only a correlation, and correlations in dynamic business environments are essentially worthless as evidence of tool efficacy.

How long did you measure? Short-duration measurements in learning-curve environments will systematically understate returns. If the trial ran for less than a quarter, treat the results with extreme scepticism.

Did you measure the full production chain? If the answer is no and the measurement stopped at the desk of the worker using the tool, the resulting number is an upper bound, not an estimate. Ask what happened to the people upstream and downstream.

The AI industry has a strong commercial interest in producing impressive ROI numbers. The consulting firms advising on AI adoption have an equally strong interest in validating their clients’ decisions. Neither of these incentives is aligned with rigorous measurement.

The organisations that will compound genuine advantage from AI are those that invest in the infrastructure to measure it honestly and that have the patience to design experiments that actually tell them something true.

Everything else is post-hoc rationalisation dressed up as analysis.

Benjamin A. Falk is an AI strategist, economist and 2x founder. His background spans global macro hedge funds, an early-stage AI language modelling startup, a VC-backed data rights company, and 6 years leading emerging technology research and strategy for a Big 4 firm. He is based in London.

Share: