Businesses are reporting productivity gains from artificial intelligence, but many are struggling to show what those improvements are actually worth.
An EY Parthenon CEO Outlook Survey released in October found that half of 1,200 CEOs surveyed across 21 countries saw AI as the biggest contributor to productivity gains over the past year. Yet 23 percent said those gains were being absorbed before reaching the bottom line, while only 16 percent said they had clear, real-time visibility into AI costs and returns.
For companies, the challenge is moving from demonstrating that AI can improve a process to proving that the improvement creates enough value to justify the investment.
Productivity Doesn’t Mean Profit
An employee might use AI to complete a task in one hour instead of three. That creates additional capacity, but it does not automatically mean the company has saved two hours of salary costs. The financial benefit depends on what happens with that extra time.
Speaking to Newsweek, Tahir Nisar, professor of strategy and economic organization at the University of Southampton in England, said businesses should focus on what changes as a result of AI rather than simply measuring how widely it is used.
“Companies should measure outcomes, not adoption. The important question is not how many employees are using AI, but whether it is reducing costs, increasing revenue, improving quality or enabling people to do more valuable work.”
That distinction can make the difference between a productivity claim and a financial result.
“The biggest mistake is to assume that time saved equals money saved,” he said. “If AI saves an employee five hours a week but those hours are not converted into greater output, better decisions or lower costs, the financial return may be negligible.”
Proving What AI Actually Changed
Even when a company’s performance improves after AI is introduced, it can be difficult to establish how much of that improvement came from the technology.
Nisar said businesses should establish a credible comparison before making a claim about AI’s impact.
“Companies need to ask a simple counterfactual question: what would have happened without the AI? Comparing AI-enabled teams or processes with credible baselines or similar non-AI groups can help distinguish the impact of AI from changes in demand, staffing, management or wider market conditions.”
Taylor Treese, founder of Timbuc, an advisory platform that uses AI to analyze company data and generate growth plans, said businesses often struggle to prove the value of advisory services and technology investments because they fail to establish a baseline before implementation.
“Coaching and consulting ROI [return on investment] was never a coaching problem. It was a baseline and measurement problem,” he told Newsweek. “Historically, business coaching has suffered from a measurement problem: ROI was tied to subjective recollection rather than hard financial data. We built TIMBUC to turn business diagnosis and growth execution into an objective science.”
Iavor Bojinov, associate professor of business administration at Harvard Business School, told Newsweek companies must build measurement into AI deployments from the outset rather than trying to calculate ROI after implementation.
“You might have high adoption, high individual productivity gains, but lower financial performance in the short term,” Bojinov said. “There is typically a productivity J-curve where performance gets worse before it gets better.”
For Nisar, genuine ROI requires more than adoption figures or demonstrations.
“The strongest evidence is sustained improvement in business outcomes that can reasonably be attributed to AI and exceeds the full cost of deploying it,” he said.
How GM Is Measuring AI
General Motors has said it is trying to build that discipline into its AI projects by starting with specific business and engineering problems and setting measurable outcomes before expanding a deployment.
Jason Fischer, GM’s executive director of virtual integration engineering, said the company looks at factors including speed to market, productivity, cost, quality and safety.
“At GM, our AI strategy is practical by design,” Fischer told Newsweek. “We start with a specific business or engineering challenge, apply the right tool, and scale only what measurably improves the work.”
GM says it begins by establishing a baseline and defining what success should look like. Its approach includes pilots, measurement against corporate performance indicators and human oversight before wider deployment.
The company says AI can now evaluate airflow across a vehicle surface from an initial two-dimensional sketch in under a minute, compared with a traditional process that could take days, weeks or months.
Software testing provides another example.
“Our testing process is now catching 10x as many defects during development and catching them earlier.”
Fischer said the value is not simply completing individual tasks faster. Earlier feedback allows engineers to explore more designs and make better informed decisions.
“We start with a 10x improvement mentality.”
GM describes this as a “virtual first, physically proven” approach, using virtual tools to accelerate development before validating results physically.
The company is also setting longer-term targets, including reducing the time needed to take a new vehicle concept to production from around four or five years to about two. It is working toward a 65 percent reduction in material and tooling spending. Both are targets, not achieved results.
GM also acknowledges the difficulty of assigning every improvement to AI.
“That helps us avoid attributing every improvement to AI when other changes or technologies may also be involved,” Fischer said.
The EY findings suggest companies are entering a more demanding phase of AI investment.
For Nisar, the standard is clear.
“Genuine return means measurable improvements in costs, revenues, quality or capacity that would probably not have occurred otherwise,” he said.