Agentic AI expense audit and AP automation | AppZen

AI agents in finance and the numbers that actually matter

Written by AppZen | Sep 14, 2026, 3:53:34 AM

AI agents in finance are measured on four things, namely autonomous rate, auto-approval rate, exception quality, and cost per transaction forecast in advance. Hours saved is an input measure that flatters every product equally. Two of those four measures are usually missing from the proposal on your desk.

Key takeaways

  • Hours saved assumes you take the hour off the payroll, and in most finance functions you do not. Work gets reallocated, which is a fine outcome and a poor business case.
  • An agent with a 90 percent autonomous rate that sends noise into the remaining 10 percent has moved the work rather than reduced it. Ask for the overturn rate on escalations.
  • Four questions make vendor benchmarks comparable, covering the denominator, whether the vendor counts escalations in the numerator, the period and volume, and the baseline.
  • Build the case around a population rather than a headcount, and include recovered leakage as a separate line so it survives scrutiny on its own.

Vendors lead with hours saved, which is the wrong metric. The measures that predict whether an agent pays back are different, and there are four of them.

What AI agents in finance are measured on

Four measures describe agent performance in a way an internal auditor and a CFO both accept.

Autonomous rate. This measure is the share of transactions the agent completed with no human touch. Qualcomm raised autonomous invoice processing from 14 percent to 61 percent after deploying 21 agents across six categories. Applied Industrial Technologies reached 87 percent autonomous AP on more than 500,000 invoices a year. Those are population numbers, and they are the ones to compare.

Auto-approval rate. In audit workflows, this measure is the share of items cleared without review. Spectrum Brands reached 72 percent auto-approvals on expense reports and cut processing time from three weeks to three hours. The rate matters more than the speed, because the speed follows from it.

Exception quality. This measure describes what arrives in the human queue. Ask for the false positive rate on escalations, and ask what share of escalations a reviewer overturned.

Cost per transaction, forecast in advance. This measure is not cost per seat, and it is not a total contract value. Deloitte's Q2 2026 CFO Signals survey, published July 2026 from 200 North American CFOs at billion-dollar companies, found 46 percent naming cost uncertainty as their single biggest internal AI concern. Agents priced per token make this number unknowable before execution.

What enterprises are reporting in 2026

Two things are true at once, and reading only one of them produces a bad decision.

Scaling is real at the top end. McKinsey's State of AI survey, published August 2026 from 1,719 respondents in 97 countries, found 40 percent of respondents at organizations above $1 billion in revenue scaling AI agents, up from 27 percent the year before. The share at organizations under that threshold was 22 percent, unchanged from the year before.

Financial impact is concentrated. In the same survey, 37 percent attributed at least some earnings impact to AI, about the same share as the year before, while only 6 percent reported AI contributing 5 percent or more of earnings before interest and taxes. The separator was workflow redesign, present in nearly three-quarters of high performers against about one-quarter of everyone else.

That finding is the practical one. Agents applied on top of an unchanged process return the value of a faster clerk. Agents applied to a process rebuilt around them return structural cost reduction. Keep the approval chain, the review step, and the exception routing exactly as they were, and you land in the 37 percent rather than the 6 percent.

The four questions that separate a real number from a demonstration

Vendor benchmarks are usually accurate and rarely comparable. Four questions make them comparable.

  1. What is the denominator? An 80 percent autonomous rate on purchase-order-backed invoices from known suppliers is a different claim from 80 percent of all invoices. Ask what was excluded.
  2. Before or after the exception queue? Some rates count an escalated item as handled. Ask whether the vendor counts escalations in the numerator.
  3. Over what period, and at what volume? A rate measured in month one on a clean subset regresses. Ask for the rate at month six on full volume.
  4. Measured against what baseline? The comparison that matters is the agent against your current team on the same transactions, not the agent against a vendor's estimate of manual effort.

Building the business case for AI agents in finance

Structure the case around a population rather than a headcount.

Start with the transaction volume in one workflow and the share currently touched by a person. Multiply by your fully loaded cost per touch, which most finance teams already hold from a shared-services or outsourcing analysis. That is the addressable pool. Apply a conservative autonomous rate, meaning the low end of what the vendor showed on a comparable denominator, and hold the exception volume flat rather than assuming it shrinks.

Then add the two lines finance leaders leave out. The first is recovered leakage, meaning duplicate payments prevented, out-of-policy spend caught before reimbursement, and fraudulent invoices stopped. That is a hard cash number where sampling previously guaranteed misses. The second is cycle-time value, meaning early payment discounts captured and close days removed.

Subtract the cost of the exception queue that remains, the integration work, and the ongoing governance effort. Finance teams routinely model the exception queue and forget the governance effort, and governance is not free.

Deloitte's Q4 2025 CFO Signals survey, published January 2026, found 54 percent of CFOs planning to integrate AI agents as a transformation priority and half naming digital transformation of finance their top priority for the year. Budget exists. What blocks approval is a case built on hours rather than population.

Where the market falls short

Two gaps show up in nearly every evaluation.

The first is unmeasurable pricing. Where agent cost varies with context length, retries, and how many sub-agents chain together, cost per invoice cannot be forecast, and a business case with an unknown denominator does not survive procurement.

The second is that horizontal agent platforms report generic metrics. Task completion rate and tool-call success are engineering measures. They do not map to touchless rate, days payable outstanding, or audit coverage, so the finance sponsor ends up translating them, and translation is where business cases go to die.

How we approach measurement

We report the measures finance teams already run their functions on. Our platform benchmarks each Agent against historical and live transactions before deployment, so the comparison is the Agent against your own team's decisions on the same population rather than against a generic manual baseline. Every Agent action is logged and auditable, which makes exception quality inspectable instead of asserted.

We price in fixed Agent credits rather than variable tokens, so cost per execution is known before approval. Across our customer base, CFOs report reductions in finance operating costs of up to 50 percent and automation rates above 80 percent, with results across AP, expense, and compliance workflows tracked on those same four measures.

The bottom line

The business case for AI agents in finance rests on autonomous rate, auto-approval rate, exception quality, and a forecastable cost per transaction. Any proposal missing exception quality and a forecastable cost per transaction is asking you to sign for an unknown. Pull the transaction volume for one workflow, ask the vendor for its autonomous rate on a denominator matching yours at month six, and build from there.

Frequently asked questions

How do you measure the return on AI agents in finance?

Measure the share of transactions completed with no human touch, the auto-approval rate in audit workflows, the quality of what still escalates, and cost per transaction forecast before execution. Compare each against your current team's performance on the same population rather than against a vendor estimate of manual effort.

What autonomous rate is realistic for accounts payable?

Published enterprise results range widely by invoice mix. Qualcomm reported moving from 14 percent to 61 percent, and Applied Industrial Technologies reported 87 percent on more than 500,000 invoices a year. Rates depend heavily on the share of purchase-order-backed invoices and the number of suppliers, so compare denominators before comparing rates.

Why do most AI agent deployments fail to show earnings impact?

McKinsey found in August 2026 that 37 percent of respondents attributed at least some earnings impact to AI while only 6 percent reported 5 percent or more. The separator was workflow redesign, present in nearly three-quarters of high performers and about one-quarter of the rest. Agents layered on an unchanged process return incremental speed rather than structural cost reduction.

Should the business case include fraud and leakage recovery?

Yes, and it is usually the strongest line. Where a team previously sampled 10 to 20 percent of expense reports, moving to full pre-payment coverage converts previously guaranteed misses into prevented cash outflow. Treat it as a separate line from labor savings so it survives scrutiny on its own.