Let's Talk
The Four-Lens AI ROI System
ROI

The Four-Lens AI ROI System

Pablo Cruz Pou

· September 14, 2026 · 19 min read

Why One Measurement of AI Value Is Worse Than None

EXECUTIVE SUMMARY

This series has spent four articles on four lenses: operational, financial, experiential, custodial. This one argues they are not four partial views of the same picture. They are four instruments that disagree, and the disagreement carries the information.

A single lens does not under-report AI value. It produces a confident wrong answer. In a pre-registered field experiment with 758 BCG consultants, the same tool raised quality by more than 40 percent on one class of task and cut correct answers from 84.5 percent to between 60 and 70 percent on another. A firm that piloted only the first class would have rolled out enthusiastically. A firm that piloted only the second would have banned the tool. Both measured accurately. Both would have been wrong about the deployment.

The discipline that follows from this is not more measurement. It is Four-Lens Reconciliation: score the same deployment on all four lenses, then go directly to the pairs that disagree. Agreement across lenses tells you little you did not already suspect. Contradiction tells you where the value actually sits, or where it is leaking. The Monday morning action is one hour with one deployment and four numbers — and the only number that matters is the gap between the highest and the lowest.

THE PROBLEM

I. Two Surveys, Eight Months Apart, Almost Exactly Inverse

In July 2025, a report out of MIT's Media Lab asserted that 95 percent of organizations were getting zero return from generative AI. It became the most-quoted statistic in enterprise AI within a month, largely on the strength of a single trade-press write-up.

Look at how the authors built it. They ran 52 structured interviews and collected 153 survey responses from senior leaders at industry conferences. No probability sample, no sampling frame, and no published data behind the headline figure. Wharton's Kevin Werbach, quoted in a trade publication that examined the report, said there appeared to be no further support for the 95 percent claim and called on the authors to release the underlying data.

Now take March 2026. RSM surveyed 1,030 middle-market leaders — a disclosed sample, a stated margin of error of plus or minus 3.1 points, published field dates — and found that 97 percent reported moderate or high success from their AI pilots over two years, and 54 percent said AI investments had exceeded their ROI expectations.

Ninety-five percent getting nothing. Ninety-seven percent succeeding. Eight months apart, both surveying business leaders about enterprise AI.

Neither one measured a P&L.

The obvious response is to pick the more credible instrument, and RSM is plainly the more credible instrument. But that misses what is useful here. Both surveys asked people to grade their own AI deployments, and got answers that depended on what the question pointed at. One pointed at attributable financial return. The other pointed at whether a pilot had gone well. Different lenses, and they were never going to agree.

This is not a story about one bad survey. It is a story about what happens when somebody treats a single measurement as a verdict.

THE DIAGNOSIS

II. The Same Tool, Two Answers, Opposite Signs

A pre-registered field experiment, published in *Organization Science* and run inside a real professional services firm, shows the problem most clearly.

Researchers randomized 758 BCG consultants — roughly seven percent of the firm's consultant-level staff — across tasks with and without access to GPT-4. On eighteen realistic consulting tasks, the consultants using AI completed 12.2 percent more tasks, worked 25.1 percent faster, and produced work rated more than 40 percent higher in quality. The weakest performers improved most.

The researchers also included one task deliberately designed to sit outside what the model could do well. On that task, the control group reached the correct answer about 84.5 percent of the time. The AI-assisted groups scored 60 and 70 percent.

One tool. One firm. One population of consultants. Measure the first set of tasks and you have a transformational result. Measure the second and you have a tool that makes expensive professionals confidently wrong. Neither measurement is an error. The error is treating either one as the answer.

The researchers called the boundary a jagged frontier — capability that is uneven in ways that are not visible from outside. The practical consequence for anyone deploying AI is sharper than the metaphor: what you measure determines what you conclude, and the two conclusions can have opposite signs.

That is the difference between incomplete measurement and misleading measurement, and it is the whole argument of this article. An incomplete measurement gives you a number that is too small. A misleading measurement gives you a number pointing the wrong way, delivered with the confidence of arithmetic.

WHY THIS IS STRUCTURAL

III. The Divergence Is Not a Sampling Problem

It would be comfortable to treat the BCG result as a quirk of one experiment. It is not. METR — the group most careful about measuring AI productivity — has since shown the divergence is a formal property.

METR — an independent evaluation organization with no product to sell — published a result in 2025 that circulated widely: experienced open-source developers took 19 percent longer on tasks with AI assistance, while believing they had been about 20 percent faster. A roughly 39-point gap between what people experienced and what a stopwatch recorded.

The honest thing to say about that study is what METR themselves say about the work that followed it. In February 2026 they announced they were redesigning the experiment, and their reasons are the interesting part. In the later study, the developers who most wanted AI increasingly refused to work without it, and the design selected them out. Between 30 and 50 percent told METR they had chosen not to submit tasks they did not want to do without AI. Developers picked different kinds of tasks when AI was available, so the comparison stopped being like-for-like.

METR's position is narrower than a retraction, and worth stating precisely. Their newer raw results point the other way — toward speedup rather than slowdown, roughly 18 percent faster for returning participants and 4 percent for new ones, both with confidence intervals spanning zero. They judge those estimates a likely lower bound on true uplift, because of the selection effects above. And they think developers are probably more sped up in early 2026 than in early 2025 — while saying their own data is "only very weak evidence for the size of this increase." Meanwhile a peer-reviewed study in *Management Science*, covering 4,867 developers across three field experiments, found a 26 percent increase in completed tasks.

Read that sequence carefully, because it is the argument rather than a complication of it. A self-report lens said faster. A stopwatch lens said slower. Then the stopwatch, rebuilt, pointed weakly back toward faster — once the task mix stopped holding still underneath it. Three readings, three answers, one question — and the group that produced the original finding was the first to say so.

METR later formalized why this happens. When a tool changes what work people choose to do, the measured gain depends on which task mix you measure. Their result gives an ordering: the uplift measured on the old task mix is always less than or equal to the true uplift in value, which is always less than or equal to the uplift measured on the new task mix. In their worked example, the same deployment reads as plus 67 percent, plus 124 percent, or plus 200 percent depending on which of the three you pick.

Nobody in that example is measuring badly. They are measuring different things and calling all of them ROI.

THE PATTERN AT PORTFOLIO SCALE

IV. Where the Lenses Split, and What Each One Hides

Set the experiments aside and look at what happens at the scale a sponsor operates on. The same divergence appears, and it appears predictably.

Gallup surveyed 23,717 employed US adults in February 2026 — a probability-based panel with a margin of error of plus or minus 0.9 points. Sixty-five percent of employees at AI-adopting organizations said AI had improved their productivity. Eight percent strongly agreed that AI had transformed how work gets done in their organization.

Same people. Same deployments. The individual lens says the tool works. The organizational lens says almost nothing has changed. Both are honest answers to different questions, and the gap between them is where the value went — into individual time savings that never became organizational throughput.

The financial lens tells a third story. Humlum and Vestergaard linked large-scale adoption surveys — roughly 25,000 workers and 7,000 workplaces — to Danish matched employer-employee records, and found precise null effects on earnings and hours at both worker and workplace level, with confidence intervals ruling out effects larger than one percent. That null held for intensive users, for early adopters, and, most pointedly, for workers who themselves reported perceived time savings. A large firm survey published through NBER in early 2026 found nine in ten firms reporting no impact on productivity or employment, with an average reported productivity gain of 0.29 percent over three years. Those are executives estimating their own firms — a fourth instrument, with its own bias, pointing the same way as the second.

None of this proves AI does not work. Research on the same question also finds real, clean, causal returns. Seven randomized field experiments at a large e-commerce platform found sales effects ranging from zero to 16.3 percent depending on the deployment — a pre-sale chatbot delivered 16.3 percent, while AI-generated ad titles produced no detectable effect at all. Within one firm, seven deployments of the same technology, spanning nothing to substantial.

Which is the point. The evidence is not mixed because researchers are careless. It is mixed because the answer genuinely depends on the lens, the workflow, and the task mix — and any single reading of any single deployment will be confidently wrong about the others.

FIGURE 1 · WHAT EACH LENS MEASURES, AND WHAT IT HIDES

LensThe question it answersWhat it can miss entirelyThe false positive it produces
OperationalDid the workflow get faster or cleaner?Whether anyone decided where the released time went, and whether the task mix shifted underneath the measurementA genuine speed gain in a workflow that produces nothing more than it did before
FinancialCan we attribute the spend and account for the return?Value that lands as capacity, quality, or retention rather than as cost line movementA deployment written off as unproven because its return never had a line to land on
ExperientialDo the people running it believe it works?The difference between individual relief and organizational changeNinety percent satisfaction with a tool that has moved nothing measurable
CustodialCan we prove what we depend on, and would we catch it failing?Whether the tool creates value at all — custody is silent on upsideA fully documented, well-governed, named-owner tool that nobody needed

Read the last column across. Each lens, alone, is capable of producing a confident endorsement of a deployment that the other three would fail.

THE METHOD

V. Four-Lens Reconciliation, and the Contradiction Test

The instinct at this point is to measure more. That instinct produces a dashboard, and a dashboard averages away the exact signal worth having.

The discipline we run at AWSM LABS for this is Four-Lens Reconciliation. Score one deployment on all four lenses. Then ignore, at first, every lens that agrees with the others, and go straight to the pair that disagrees. The reconciliation is not the average. It is the argument between the readings.

Four pairs come up often enough in our own work to name, and each one has a different meaning and a different fix.

Operational high, financial flat. The workflow genuinely got faster and the money never showed up. This is the pair with the most boring diagnosis: the time released was real and nobody decided what it was for. Capacity nobody redeploys is not savings.

The fix is a decision, not a measurement. Somebody with the authority to move people says out loud where the released hours go. In practice that sentence sounds like: *the two analyst-days a week this gave us back go to diligence, starting this month, and I moved them.* Until someone says a sentence like that, the hours dissipate into slightly earlier evenings, and the financial lens keeps reading flat — correctly.

Operational high, experiential low. The dashboard says people use it, and that it is fast. The people using it describe tolerance rather than belief. The measured speed is real but narrow: a task got quicker while the workflow around it stayed exactly as it was, so the person doing the work absorbed the cost of the seam.

Picture the analyst whose draft now lands in forty minutes instead of three hours, and who then spends an hour reformatting it into the template the partner will actually open. The tool did its job. The seam on either side of it never moved. The fix is the seam, not the training.

Experiential high, financial flat. Everybody likes it and nothing moved. This is the pair most likely to produce a bad decision, because enthusiasm reads as evidence. Sometimes the value is real and landing somewhere the financial lens does not look — retention, quality, error rates avoided. Sometimes the tool is pleasant and marginal. The two look identical on a satisfaction score, and the only way to separate them is to name in advance where the value should have landed and go look.

Everything positive, custodial blank. The deployment scores well on three lenses and nobody can say who owns it or what would happen if it were wrong. This is the pair that has no urgency attached to it and the longest tail. A tool that works, that people like, that pays for itself, and that no one is accountable for is not a success with a caveat. It is an unbounded position. Somewhere there is a model that has been quietly mispricing something for eleven weeks with nobody's name on it — and if this tool is load-bearing, the other three scores are what make it dangerous rather than what make it safe.

FIGURE 2 · THE CONTRADICTION TEST

If these two disagreeThe likely readingWhat to do first
Operational high · Financial flatReleased capacity was never redeployedName where the hours go, and who decides
Operational high · Experiential lowA task got faster inside a workflow that did not changeRedesign the seam, not the training
Experiential high · Financial flatValue may be landing where finance does not look — or may not existName the expected landing place, then check it
All three positive · Custodial blankAn unowned dependency that is now load-bearingAssign a named owner before the next renewal
All four agree, positiveRare, and worth studyingFind what made this one work and try to repeat it
All four agree, negativeAlso rareStop paying for it

The spread between the highest and lowest lens score is more informative than any single score. A deployment reading 4/4/4/4 tells you less than one reading 5/3/4/2, because the second one tells you where to look.

THE HONEST GAP

VI. What This Framework Cannot Yet Prove

Governance vendors make one claim in this territory constantly, nobody has established it, and it would be convenient for us to repeat it.

The claim is that governance maturity predicts AI value — that the well-governed deployment delivers more. Every quantified version of that claim we have been able to find is cross-sectional, self-classified on both the governance side and the outcome side, and produced by an organization selling governance services. Nobody has tested the obvious alternative: large, well-resourced, AI-mature firms both govern more and extract more value, and no published work separates the two.

We do not know whether custodial rigor is a leading indicator of value. Nobody does. It is a testable claim. Nobody has tested it.

What the custodial lens does reliably is different and narrower, and Article 4 in this series made the case: it is the only one of the four where somebody else sets the standard — a buyer, from outside your firm, on their schedule. That is an argument about exposure, not about return. Keeping the two apart is the difference between a framework and a sales pitch.

The same discipline applies to the rest of this. The four lenses do not produce a number. They produce a disagreement you can act on, and the honest version of the claim is that acting on the disagreement beats acting on any single reading.

MONDAY MORNING

VII. One Deployment, Four Numbers, One Hour

You do not need to instrument the portfolio to find out whether you have a single-lens problem. One deployment will tell you, and the exercise fits inside an hour — yours, or the hour you ask a portfolio CEO to spend before your next board meeting. Pick the deployment together. Let them score it. Read the spread yourself — you are the one quoting these numbers upward.

1. Pick the deployment everyone agrees is working. Not the troubled one. The success story — the tool somebody names in the quarterly review. The contradiction test is most useful where confidence is highest, because that is where a single lens has been doing all the work.

2. Score it four times, one to five, in this order. Operational: did the workflow get better — not whether a task got faster — and did the released time go somewhere? Financial: can you attribute the spend and name where the return landed? Experiential: would the people using it keep using it if it were optional? Custodial: who owns it by name, and how many days would a wrong output run before anyone noticed?

3. Write down the spread. Highest score minus lowest. That single number is the finding.

4. Go to the disagreement, not the average. A spread of zero or one means your lenses agree — trust the story, and note which lens you never actually scored. A spread of two means one reading is doing more work than it has earned; open that pair. A spread of three or more means at least one of the four numbers you have been quoting in board updates is wrong, and Figure 2 tells you which pair to open first.

If you cannot score one of the four lenses at all, that is not a gap in the exercise. That is the finding, and it arrived in under an hour. Custodial is the one we see come back blank most often, and the survey data says to expect that. A blank custodial score is not a prediction that the deployment will underperform — Section VI said plainly that nobody has established that link. It is a statement about what you are carrying without knowing you are carrying it. Forty-four percent of enterprises have an incident response procedure written for AI systems. Accountability for AI risk is scattered rather than absent: the most common single answer, at 37 percent, points to the CIO or head of IT — which means roughly two in three organizations locate it somewhere else, and locating it is exactly what the fourth lens asks you to do.

CONCLUSION

VIII. Measure the Argument, Not the Average

Four articles, four lenses, and the same conclusion arrived at four different ways: the number you trust most is usually the one you checked least.

Run one lens and you get a number. Run all four and you get an argument — and the argument is the product.

That is a harder discipline than it sounds, because it means the reward for doing this well is not a cleaner story. It is a messier one, with a contradiction in it that somebody now has to resolve. Firms that measure on one lens get a clean answer every quarter and are wrong at an unknown rate. Firms that measure on four get an uncomfortable answer and know exactly where to spend the next month.


AI ROI ASSESSMENT · ANALYSIS TOOL

The AI Accountability Gap

Find out where your AI program really stands — and where the need is greatest.

Start the assessment→

9 QUESTIONS · ~2 MINUTES · INSTANT TAILORED REPORT


WORKS CITED

Cui, Kevin Zheyuan, Mert Demirer, Sonia Jaffe, Leon Musolff, Sida Peng, and Tobias Salz. The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers. Management Science, 2026. Peer-reviewed; 4,867 developers across three field experiments.

Dell'Acqua, Fabrizio, et al. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality. Organization Science, 2025. Pre-registered field experiment; 758 BCG consultants.

Fang, Yuan, Zhang, Donati, and Sarvary. Generative AI and Firm Productivity: Field Experiments in Online Retail. Working paper (arXiv:2510.12049 / MSI Report 26-112), October 2025. Seven randomized field experiments.

Gallup. Rising AI Adoption Spurs Workforce Changes. Gallup, April 2026. Probability-based panel; 23,717 employed U.S. adults, fielded 4–19 February 2026; margin of error ±0.9 points.

Humlum, Anders, and Emilie Vestergaard. Large Language Models, Small Labor Market Effects. NBER Working Paper 33777, April 2025 (revised). Two adoption surveys covering roughly 25,000 workers and 7,000 workplaces, linked to Danish matched employer-employee records; outcomes tracked through mid-2024. Working paper, not peer-reviewed.

METR. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. arXiv:2507.09089, July 2025. Preprint; randomized controlled trial, 16 developers, 246 tasks.

METR. We Are Changing Our Developer Productivity Experiment Design. METR, February 2026. Authors' own account of selection and task-substitution bias in the 2025 result.

METR. Task Substitution and Uplift. Cunningham and Whitfill, METR, May 2026. Formal ordering of three uplift measures.

MIT Project NANDA. The GenAI Divide: State of AI in Business 2025. July 2025. Not peer-reviewed; 52 interviews and 153 conference-intercept survey responses. Cited here as an object of analysis, not as evidence.

RSM US. RSM Middle Market AI Survey 2026. RSM, 2026. 1,030 respondents (827 US, 203 Canada), fielded 5–16 March 2026; margin of error ±3.1 points.

Schellman. *2026 State of AI Governance Report.* Schellman, July 2026. 525 U.S. professionals at organizations with 500+ employees and $100M+ revenue; non-probability sample. Accountability and maturity figures are respondent self-reports.

Yotzov, Ivan, Jose Maria Barrero, Nicholas Bloom, Philip Bunn, Steven J. Davis, et al. Firm Data on AI. NBER Working Paper 34836, February 2026.

\

AWSM DSPTCH

Get the dispatch.

Perspectives on AI activation — ROI, frameworks, and lessons from the frontier — sent when we publish. No noise.

How often?

We respect your inbox. Unsubscribe anytime.