Why CFOs Could Reject the Business Case for AI in Software Development
Will your business case for AI deployments stand up when stakeholders and Finance ask where the claimed value actually came from? If not, here's how to respond.
.webp)
One enterprise went into a leadership conversation with a $500K AI savings figure.
It sounded impressive.
But when someone asked where the number came from, the story disintegrated.
We’re hearing versions of that conversation across large engineering organizations. AI coding tools are already deployed – embedded, even. Developers are using them widely. Budgets have been committed (and spent).
Now stakeholders want to know what changed. It’s a fair question – and with the 2027 budget round underway, usage or adoption metrics alone can’t answer it.
Across multiple conversations we’ve had recently, the language varies but the pressure is noticeably consistent: build the business case, show the return, translate it into something leadership can act on.
For engineering leaders, it’s creating a new challenge. They need to connect AI spend to what’s actually happening in software development; and make that connection credible enough to survive the next question.
AI adoption was easier to measure than AI value
The first phase of enterprise AI rollout produced plenty of data.
- How many licences were assigned?
- How many developers activated them?
- How often are the tools being used?
That information matters: if you’re spending millions on AI tooling, you should know whether anyone is using it. But adoption only gets you so far. A board or finance team looking at next year’s budget is now more likely to ask what that usage produced.
- Did engineering output improve?
- Did we free up meaningful capacity?
- Did the software get easier or harder to maintain?
- Did one tool work better than another?
- Did the investment create enough value to justify the next chunk of spend?
That’s where many AI business cases start coming unstuck.
We see the same pattern across sales conversations, industry narratives and practitioner forums: organizations have moved faster on AI adoption than on proving what the investment delivered.
A headline productivity number rarely ends the conversation
Suppose you tell Finance that developer productivity increased by 15%. It sounds like good news – and it is on one front.
Then the trickier questions start: how was productivity measured, how much of the increase came from AI, and what happened to code quality?
The maintainability question (how easy it is to understand and change code) is especially important when looking at the return on investment of AI in software development.
Our research into the GenAI period found productivity increasing while codebase maintainability declined. The same maintainability research found a stark difference in what happened when incidents occurred in different parts of the codebase.
Incidents involving the least maintainable quarter of code had a median recovery PR duration of 65.2 hours. In the most maintainable quarter, it was 1.7 hours.
That’s a 38-fold observed difference.
A productivity gain can still be valuable, but the business case changes if faster output is accompanied by software that gets more expensive or difficult to change later. Which is why a single number rarely settles the AI ROI question.
AI performance also varies depending on the work
There’s another reason broad claims about AI productivity (rightly) face scrutiny from CFOs and other stakeholders: AI coding performance isn’t uniform.
Our BARE benchmark research tested 57 LLMs on maintainability-oriented refactoring using real production source code.
We found even the best-performing models remained below 23% overall success on this class of task. Performance also varied greatly depending on the programming language and the type of refactoring required.
JavaScript averaged 31.9% refactoring success across the models tested; C was around 3.7%. More localised changes were considerably easier than architectural restructuring.
That makes a company-wide statement like “AI gives us a 20% productivity uplift” hard to use as an AI business case on its own. The result may be very different across teams, technologies and types of engineering work.
And those differences matter when someone has to decide – and account for – where the next pound, euro or dollar of AI budget should go.
What a business case for AI in software development needs
A useful AI business case should connect the investment to an outcome someone outside engineering can understand.
That demands more than a perfect ROI formula. It requires evidence that can answer a few basic questions.
- What changed after AI was introduced?
You need a credible before-and-after view of engineering outcomes.
- How much of that change can reasonably be linked to AI?
Hiring, project mix, organizational changes and dozens of other factors can move engineering numbers.
- What happened beyond output?
Productivity needs context from quality, maintainability and downstream operational impact.
- What decision does the evidence support?
Increase spend? Change tools? Expand adoption in some areas? Pull back elsewhere?
Measurement is valuable when it changes a decision like more spend, a different tool, a narrower rollout, or a rethink.
The 2027 budget conversation has already started
For many engineering organizations, this isn’t a theoretical problem for next year. Some have already raised 2027 budget timing in conversations with our team.
The second wave of AI investment will face a different level of scrutiny from the first.
Leadership now has real spend to look at. They have months of usage data. And they have enough engineering output to start asking whether the original assumptions held up.
A business case built mainly on licenses deployed, developer sentiment or a single productivity figure is likely to unravel into more questions.
You need evidence that connects AI investment to what changed in the software itself, and then connects those engineering outcomes to the decisions Finance and the board care about.
And that evidence has to come from your own engineering data.
Some metrics can show that the broader problem exists, where AI performs well and where it struggles. They can’t tell you whether your specific AI investment is working. That answer sits in your codebase itself.
Preparing for your next AI budget conversation?
See how BlueOptima measures AI’s impact on engineering productivity, quality and maintainability using evidence from your own software estate.

.png)
.webp)
.webp)

.webp)
.webp)
.webp)
.webp)
.webp)
.webp)
.webp)
.webp)

