Ask a finance leader to prove an AI return and you usually get one of three answers: a usage number, an hours-saved number, or a promise that the value is real but hard to isolate. None of those survives an audit committee. In our September finance data, proving ROI as the main blocker on AI moved from 53 percent of respondents in the cohort through July to 65 percent in the August cohort. That is not a measurement problem arriving. That is a measurement problem that has already arrived and is now the loudest thing in the room.
I have written before about why most companies still cannot prove AI ROI. This piece is the other half: what actually goes on the page.
The clock is not the problem
Six months is a reasonable window. It is short enough to keep a vendor honest and long enough for a workflow to change. The trouble is that most six-month AI cases are built out of measures that cannot become money in six months, or in sixty.
Seat licences used, prompts run, documents processed: these are adoption measures. They tell you the tool is being touched. They tell you nothing about whether the company is better off. Hours saved is the more seductive one, because it looks financial. It is not, unless the hours are removed from a budget or redirected to work that produces revenue. An hour saved that stays in the same salaried week is an hour the company still paid for.
That is the conversion problem. Every line on an AI scorecard has to answer one question: what changes on a financial statement, and when?
The four lines
Each line below states the measure, the conversion, and the evidence finance should ask for. Any AI case that cannot fill in at least two of these lines is not ready for a purchase order.
1. Headcount avoided. The measure is roles that were in the plan and are no longer in the plan. The conversion is direct: fully loaded cost times months. The evidence is the previous hiring plan and the revised one, both dated, both signed by the function head.
This is the line finance is most likely to actually get. In the September data, 21 percent of finance-lane respondents fund AI with money that would have gone to headcount, rising to 34 percent at the CEO seat (base 290). The money is already moving this way. What is usually missing is the paper trail that turns it into a provable number.
2. Software retired. The measure is licences cancelled or not renewed. The conversion is contract value. The evidence is the termination notice.
This is the cleanest line on the scorecard and the most frequently skipped, because retiring a tool requires someone to own the migration. If an AI purchase is meant to replace something, name the something and put its renewal date on the scorecard.
3. Cycle time on a revenue-linked process. The measure is elapsed days on a process that sits in front of money: quote to cash, claim to settlement, application to decision. The conversion is working capital or throughput, not effort. The evidence is the process timestamp before and after, on the same volume mix.
This line is where most of the genuine value sits and where most cases go wrong, because teams measure the step the AI touched rather than the end-to-end elapsed time. If the AI shortened one step and the queue simply moved downstream, the cycle time did not change.
4. Error and rework cost. The measure is the rate of a defined failure and the cost of correcting it. The conversion is corrections avoided times unit cost. The evidence is the existing quality log.
Only use this line if the failure was already being counted before the AI arrived. A baseline invented after the purchase is not a baseline.
The two gates
The scorecard is worthless if it is filled in after the fact by the team that wanted the tool. Two things have to be settled before the money is committed.
Gate one: name the prover, and it cannot be the signer. Our data keeps showing the same structural problem. The CEO is the most-named signer of AI purchases, at 47 percent of finance-lane sign-off mentions on a base of 290 and 51 percent of investor responses on a base of 245. But the seat that will be asked to produce the return is finance. When the signer and the prover are different people and neither has agreed the measure in advance, the argument happens six months later with no shared definition.
Write the prover's name on the purchase paper. Give them the right to define the baseline before go-live.
Gate two: date the number. Not "we expect a return in the first year". A specific figure, on a specific line, by a specific month. The Optimism Gap exists precisely because this does not happen: 70 percent of CEO-seat respondents expect measurable payback inside six months against 42 percent of the finance seat, a 28-point difference on seat bases of 80 and 52. Two people can leave the same meeting with different expectations because nobody wrote the date down.
What finance should do with the extension
The most interesting movement in the September data is that finance is stretching the clock while the demand for proof rises. Under-six-month expectation fell from 55 to 48 percent, and six-to-twelve months rose from 22 to 28. Read together with the blocker moving from 53 to 65 percent, that looks like finance leaders concluding that the honest payback window is longer than the one they were given.
If that is your position, say it out loud at purchase rather than at review. A twelve-month case with four convertible lines is more defensible than a six-month case built on hours saved. The risk of the shorter promise is not that you miss it. It is that missing it teaches the organisation that AI numbers are soft.
Limitations
Everything above rests on one instrument, answered by finance leaders as they applied to Open Future Forum finance events. The full base is 290 and the August cohort within it is 54, which is a single month and should be read as direction rather than a trend line. These are people who chose to attend an AI-focused finance session, so they sit ahead of the average finance function, and the comparison you should draw is with your peers rather than with the market. Where a question allowed more than one answer, the shares add to more than 100 percent.
The scorecard itself is not from the data. It is drawn from what finance leaders in these rooms describe as the cases that survived review, and it should be adapted to how your business actually recognises value.
Compare your scorecard with your peers
The CFO Executive Forum convenes finance leaders working on exactly this question, off the record. If you want to see where your payback assumptions sit against the rooms, the September CFO AI Leverage Report carries the full instrument with response bases stated on every figure.
Last updated: September 9, 2026
Frequently Asked Questions
Join the CFO Executive Forum
The CFO Executive Forum convenes finance leaders working on AI measurement, off the record and by application.