AI coding tools have moved from pilot to budget line, and the question has shifted from engineering to finance. Most engineering leaders can show developers use the tools and feel faster. Far fewer can show where that speed went.
In McKinsey’s 2026 State of AI survey, 80% of respondents say AI improved their own productivity, while only 37% attribute any EBIT impact to it.
In my view, the gap comes from a decision most companies skip. Time saved by engineers is capacity, and it becomes return only when leadership decides before rollout what that capacity should produce, records a baseline to measure it against, pairs the target with a stability guardrail, and names an owner for the result.
Hours saved are capacity, and capacity has to be assigned
The figure most teams report first is hours saved, and it is the least reliable. Gartner’s survey of 12,004 employees and managers found 19% reported no time saved with AI, and warned that leaders are mistaking access and adoption metrics for transformation, an “enablement illusion.”
Even where the hours are real, the figure says nothing about where they went, and that is the part finance is asking about.
Where the speed is real, part of it gets spent downstream. The 2025 DORA report found AI adoption now improves delivery throughput while still increasing delivery instability, so some of the time saved writing code comes back as time spent fixing what shipped.
Gross hours saved therefore overstate the capacity gained, and only your own before-and-after data can show the net figure. That number has to come from your own before-and-after data.
Stay updated with Simform’s weekly insights.
Decide what the capacity becomes before you measure it.
Match the target to your constraint
The conversion target should follow whatever constrains the business today. When the backlog is the constraint, target throughput and measure release frequency and lead time.
When cost is the constraint, target avoided spend and measure contractor hours and planned hires against budget.
When production incidents are the constraint, target stability and measure change failure rate and rework.
When speed to market is the constraint, target cycle time on the few features that carry revenue. Pair each speed metric with a stability metric, since DORA’s data shows the two can move in opposite directions.
Case in point
Salesforce made its conversion decision in public. On the February 2025 earnings call, Marc Benioff said, “we’re not going to hire any new engineers this year,” citing a 30% productivity increase in engineering, according to the call transcript.
The target was flat engineering hiring. What Salesforce has not published is how it measured the 30%. Whatever target you choose, you should be able to state it before rollout and defend it with a baseline when the CFO asks.
When headcount is the wrong target
Evidence from automation more broadly points the same way. A Gartner survey of 350 executives at billion-dollar companies deploying autonomous technologies found about 80% had cut headcount, at nearly equal rates among those reporting higher and lower ROI.
In Gartner’s words, “workforce reductions may create budget room, but they do not create return.” For engineering specifically, DORA’s ROI guidance advises reinvesting reclaimed capacity.
Measure at the level where the question lives
Three levels, three questions
Adoption data shows whether people are using AI. Delivery data, meaning lead time, deployment frequency, and change failure rate, shows whether engineering performance changed. Business outcome, the conversion you chose, is where finance can recognize that change as return.
McKinsey’s study of nearly 300 companies found top performers track quality improvements (79%) and speed gains (57%), while bottom performers track adoption alone, which showed little correlation with performance.
AI ROI has a learning curve
DORA’s ROI guidance describes a J-curve, a productivity dip while teams learn the tools and adapt review and testing, and advises budgeting for it before rollout.
A Gartner survey of infrastructure and operations leaders found 57% of those reporting AI failures had expected too much, too fast. The sample is outside software development, but the risk of judging returns too early is the same. Record the baseline before you scale, and agree on a learning window before the first formal ROI review.
The same AI productivity gain can create two different returns
Take an illustrative 200-engineer product company whose roadmap backlog keeps growing. Its target is throughput, tracked monthly as release frequency and lead time against the six months before rollout, with change failure rate as the guardrail.
A services-heavy company of similar size, squeezed by contractor spend, targets avoided contractor hours against the approved plan, and reports separately which hours became realized savings and which were redeployed to delivery.
Track the costs and stability issues that can erase the gain
Gartner projects that AI coding costs will exceed the average developer’s salary by 2028 as pricing shifts toward consumption, so tool and usage costs belong inside the calculation.
Re-baseline when tools or team structure change, and give one named owner a quarterly review with finance. If lead time improves while rework or change failure rises enough to absorb the reclaimed capacity, treat the ROI case as unproven until stability recovers.
Make the four ROI decisions before rollout
A CFO can accept an ROI figure for AI-assisted development when those four decisions are made before rollout. Most teams make them after rollout and then try to reconstruct what the capacity became.
A team that starts without them holds a productivity hypothesis, and the next budget cycle is the moment to turn it into an ROI model.
If you are setting up that model now, Simform’s enterprise AI adoption and change management team builds the strategy, governance and change foundation for it before tools scale.