The Gap Between Perception and EBIT in Enterprise AI
80% of professionals perceive higher productivity with AI, but only 37% of companies see impact on their EBIT; the article explains why and how to measure it.
The 80% and the 37%: the gap that decides whether your AI reaches EBIT.
Eight out of ten professionals say AI makes them more productive. Only 37% of companies see it in their profit and loss statement. The difference isn’t the model they chose: it’s that nobody has measured.
Your company has two AI figures. One comes from your employees. The other comes from your accounting. They don’t match.
According to the global McKinsey survey published this August, 80% of professionals state that AI has improved their personal productivity. 37% of organizations say AI has contributed to their EBIT.
A forty‑three‑point difference. And the second number has been stuck at the same level for two years while investment and deployment keep rising.
This article is about why that gap exists, and how much, in euros, it costs to keep it open.
The data, and whose they are
The source is "The state of AI in 2026: On the road to ROI", the annual survey by McKinsey & Company authored by Dan Tinkoff, Lieven Van der Veken and Michael Chui, with fieldwork conducted between May 4 and June 8, 2026 and 1,719 participants in 97 countries. EBIT is earnings before interest and taxes: the line a board looks at when deciding whether something has worked.
Adoption is moving forward without debate. 44% of organizations are scaling AI across the enterprise, up from 38% the previous year. In companies with more than a billion dollars in revenue, the use of agents has jumped from 27% to 40% in twelve months.
The return, not so much. The 37% that attribute any effect to EBIT is unchanged from 2025. And the high‑performance group, those who credit AI with 5% or more of their EBIT, remains at 6% of the sample, also flat.
More deployment and more spending, same result on the bottom line. That’s the fact to explain.
80% is not a measurement
The comfortable reading is that productivity is already there, it just needs to surface. It’s worth resisting, because that 80% is declared perception, and there’s an experiment that shows how far perception can drift.
It was done by METR (Model Evaluation & Threat Research), an independent nonprofit that evaluates model capabilities. In July 2025, Joel Becker, Nate Rush, Beth Barnes and David Rein published a randomized controlled trial —the standard design of clinical research, extremely rare in productivity studies—: 16 seasoned developers, 246 real tasks from their own repositories, each task randomly assigned to “with AI” or “without AI”.
Three numbers from the same group of people:
- Before starting, they estimated AI would make them 24% faster.
- After doing the work, they thought they had been 20% faster.
- The stopwatch showed they were actually 19% slower.
They didn’t miss the magnitude. They missed the sign. And they kept missing it after doing the work with their own hands.
A nuance that is almost never mentioned when this figure is cited: METR flagged it as historical. In February 2026 the team announced that the measurement corresponded to early‑2025 tools, that acceleration is now likely, and that their follow‑up experiment became unusable due to selection bias—more and more developers refuse to take part if it means working without AI, so they are redesigning the study.
That correction does not weaken the lesson. It sharpens it. What holds is not “AI makes you slower.” It is that perceived productivity does not replace measured productivity, even when the perceiver is an expert with five years on that codebase.
Applied to the survey: 80 % is not measuring productivity, it’s measuring enthusiasm. The 37 % figure is the only one of the two that has gone through an accounting.
The accounting that needs to be put on the table
Let’s translate it into dollars, because that’s where it stops being a methodological debate.
Take a company with 200 qualified employees and an AI tool priced at €30 per user per month. That amounts to €72,000 a year in licensing fees alone, not counting deployment, training, or the time of the people who integrate it.
The usual justification is that it saves “around 20 % of the time.” With a total labor cost of €45,000 per person, that 20 % would be worth €1.8 million. Against €72,000 in licenses, the case gets approved in two minutes.
Now suppose the real saving is 3 %. It’s still positive: €270,000. The project works, but the committee approved a plan sized for a return sixty times larger, and built the rest of its roadmap on that figure.
And if the actual effect were what METR measured in its negative context, the company would be paying €72,000 a year to go slower, while 80 % of the staff happily confirm they’re moving faster.
The three scenarios are indistinguishable from the inside. No internal survey separates them. The only thing that does is a control group.
The method is already in your company, in another department
If you have a marketing team with some analytical maturity, you already have the methodology your AI program is missing.
No serious marketer launches a campaign, looks at the month’s sales, and claims the lift. A control group is set aside (a holdout: an equivalent population that does not receive the impact), you compare, and the difference is the result. It’s exactly the METR trial design under a different name, with twenty years of practice in direct marketing.
An AI pilot is measured the same way:
- Choose a process that already has a metric. Ticket resolution time, cost per published piece, reconciliation hours at closure. If that data doesn’t exist, the first phase of the pilot is to measure, not to roll out.
- Leave out a comparable part. It doesn’t have to be half the team: it just needs to resemble the one that does participate.
- Measure outcome, not activity. Consumed tokens, generated messages, and adoption percentage are consumption. The outcome is the metric from point 1.
- Set the read‑out date before starting. Without a prior date, the pilot is evaluated on the day the numbers look good.
- Collect perception, but in a separate field. It says a lot about adoption and friction, which matter. It’s not evidence of impact.
Cost: a few design hours and the discipline of not contaminating the control. In return, the next time someone asks “Is this working?”, there’s an answer that doesn’t depend on who shouts the loudest in the meeting.
What makes the 6% different
McKinsey characterizes the group that does see impact, and the pattern aligns with everything above:
- Redesign the workflow instead of plugging AI into the existing one. Almost three quarters have done so, versus a quarter of the rest. It's the biggest gap between the two groups in the whole report.
- Measure. They are twice as likely to have defined processes for quantifying the impact of their initiatives.
- Don't chase efficiency alone. They add growth and innovation goals to savings.
- Define where a person fits. Deciding when a model output needs human validation is another practice where they gain an advantage.
- Treat cost as a design constraint. One in five respondents already acknowledges that operational costs, tokens included, are limiting their AI use.
Five operational decisions. None is a better model.
And a note on where to start: in the report, revenue increases are concentrated in marketing and sales; cost reductions are in supply chain, service operations and manufacturing. These are two distinct conversations and they are measured differently. A revenue case needs a control group and an attribution window; a cost case needs a baseline of hours or units. Mixing them up is the quickest path to a pilot that “works” while EBIT stays flat.
Takeaways
- Declared productivity and measured productivity can point in opposite directions, and the declarer may be a convinced expert. 80% is a sign of adoption, not a return.
- Without a control group, a 20% saving, a 3% saving, and a negative saving look exactly the same from the inside.
- What separates the 6% that reaches EBIT is not the vendor: it’s redesigning the process and having a method to measure it.
- First the metric and the reading date. Then the use case. In that order fewer projects get approved and many fewer fail.
Sources
- McKinsey & Company, "The state of AI in 2026: On the road to ROI", August 2026. Dan Tinkoff, Lieven Van der Veken and Michael Chui, with Tara Balakrishnan. Online survey of 1,719 participants from 97 countries, fielded from May 4 to June 8, 2026 — https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai
- Joel Becker, Nate Rush, Beth Barnes and David Rein (METR — Model Evaluation & Threat Research), "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity", July 2025 — https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
- METR, "We are Changing our Developer Productivity Experiment Design", February 2026. The team itself labels the previous result as historic and explains the selection bias of its follow‑up experiment — https://metr.org/blog/2026-02-24-uplift-update/
- The license calculation and the three savings scenarios are self‑generated and illustrative: they serve to show the order of magnitude of the risk of not measuring, not as an estimate for any specific case. Survey data are self‑reported and global, and should be read as direction, not as audited measurement.