Gartner® report CFO Guide to Governing Agentic AI Read now

Agentic AI for accounts payable: What actually changes

Agentic AI for accounts payable is an operating model in which software agents pursue an objective, act inside a written policy boundary, and record why they acted. The agent runs intake, matching, coding, tolerance checks, and routing. A person owns the boundary, the exceptions, and the evidence. Your team moves from working every item to governing a population.

Key takeaways

  • Four different things are sold as agentic, namely rules-based automation, robotic process automation, a copilot, and an agent. Each fails in a different way.
  • Describe every autonomy level by the control it requires, not by what it can do. That column is what your audit committee asks about, and it is the one nobody publishes.
  • Autonomy varies by invoice class rather than by company. An 80 percent claim describes an invoice mix as much as it describes software.
  • An invoice is a document a stranger sends you. Treat its text as data and never as instruction, and route bank detail changes to a separate verified process.

Your accounts payable (AP) team already runs software that reads invoices and routes them. Almost none of it decides anything. What changes with agentic AI is the moment a system stops proposing work and starts committing it, which is a control question before it is a technology question.

Four different things are sold as agentic AI for accounts payable

The word agentic now covers systems that behave nothing alike. Each fails in a different way, and the contract rarely tells you which one you bought.

  • Rules-based automation executes a condition a person wrote. When it is wrong, it is wrong identically until someone edits the rule.
  • Robotic process automation (RPA) replays screen actions. It does not read meaning, it breaks when a field moves, and it breaks quietly.
  • A copilot suggests a value, a code, or a reply. A person confirms every item, and the system carries none of the outcome.
  • An agent takes an objective, acts inside a boundary you set, and records why it acted. The action takes effect whether or not a person looked.

Sorting these apart matters because budget follows the label. Teams buy a copilot, forecast the headcount effect of autonomy, then find the gap two quarters later. Our piece on generative AI in AP automation draws the same line from the model side.

Market data is steadier than the vocabulary. Gartner's November 2025 survey covered 183 CFOs and senior finance leaders. It found 59 percent used AI in finance in 2025, against 58 percent in 2024. Accounts payable automation was the second most common use case at 37 percent, and 91 percent reported low or moderate impact initially. Gartner publishes no adoption figure for AI agents, so any agent percentage you see quoted measures something else.

Spending points the other way. Gartner reported in April 2026 that CFOs implementing strategic AI and technology portfolio resource deployment will unlock 10 additional points of margin growth by 2029, with three quarters of CFOs raising technology budgets for 2026. Finance teams are spending faster than they are designing the controls.

Define autonomy by the control it requires

Most maturity models describe what the software can do. Capability is the half a demonstration shows. The harder question is what must be true about your controls before a level runs in production.

So describe every level by three things. The first is what the agent decides without a person. The second is what evidence it produces as a by-product of deciding. The third is what control you have to operate for that level to be defensible. That third item is the column nobody publishes, and it is the one your audit committee asks about.

Extraction accuracy sets how often the agent is right. Control design sets what it costs when the agent is wrong. Those are separate budgets, and the second is usually unstaffed. Autonomy also varies by invoice class rather than by company, so one company-wide percentage tells you nothing about the mix that sets your result.

The order of operations follows. Choose an invoice class, write the policy the agent must obey, decide the error rate you can live with, then set the level. Teams that pick a level first write the policy backwards from what the software happens to do.

Six levels of agentic AI for accounts payable

L0, manual. The agent decides nothing, and a person keys every invoice. Evidence is keystrokes and an approval trail. The control required is supervisory review and separation of duties.

L1, rules and templates. The system executes a written condition and decides nothing beyond it. Evidence is a rule execution log per invoice. The control required is change control on the rule set, tested before release.

L2, assisted. The agent proposes field values, a match result, and a general ledger (GL) code. Evidence is a suggestion with a confidence value and its source on the document. The control required is confirmation of every item by a person, with the override rate tracked by supplier.

L3, supervised autonomy. The agent acts inside a narrow policy on a defined slice. Evidence is a decision record per invoice, with inputs, policy version, and outcome. The controls required are a sampling plan, an agreed error tolerance, and a named reviewer.

L4, conditional autonomy. The agent acts across the population and escalates by exception. Evidence is the decision record plus an escalation log of what it declined. The controls required are exception ownership, thresholds owned by finance, and a monitored escalation rate.

L5, full autonomy within a defined scope. The agent decides and posts across a bounded scope with no per-item review. Evidence is a full evidence chain plus control testing results. The controls required are a scope approved like a policy, a tested kill switch, and a rollback path.

Read the control requirement first. The rest is what a vendor demonstrates, and the control is what your controls owner builds. Moving from L2 to L3 means a sampling plan, an error tolerance, and a named reviewer.

The step teams underestimate is L3 to L4, because escalation replaces review. At L3 a person sees a share of everything. At L4 a person sees only what the agent hands over, which turns the escalation rate into a control. A falling rate is either learning or a blind spot, and sampling what the agent cleared is the only way to tell. Our overview of AI maturity in accounts payable covers the capability side.

The failure modes nobody puts in the demonstration

A currency read wrong

An invoice issued in one currency is extracted as another, because the symbol was ambiguous or absent. The amount looks plausible. Where your tolerance is a percentage, the currency pair decides whether the error trips it. It surfaces at bank reconciliation or on the next supplier statement, weeks later, and the cost is the difference plus a recovery conversation.

A GL code that lands in the wrong period

The agent picks the right GL account and the wrong period. That happens near a cutoff, where invoice date, service date, and receipt date disagree. No exception fires, because every field is internally consistent. It surfaces in an accrual variance at close, when a department head asks why their number moved. The cost is a manual reclass and a hard question about that week's postings.

A duplicate that clears quietly

The same invoice arrives by email and through a supplier portal under a different reference. The agent treats each as new, and both pass. The Washington State Auditor wrote in 2022 that duplicate payments run between 0.8 percent and 2 percent of total payments. It surfaces at statement reconciliation, on a supplier credit, or never.

High confidence on a document the model has never seen

A new supplier sends a layout the model has not encountered, and it returns values with high confidence anyway. Confidence describes the model's own output rather than the truth. This shows up as odd exceptions clustered on one supplier, and nobody notices when the errors stay inside tolerance. The defense is a lower autonomy level for a supplier's first invoices, whatever the score says.

Governance and the evidence an auditor receives

  • Agent identity. An agent needs its own identity in the enterprise resource planning (ERP) system, not a shared integration account and never a person's credential. Otherwise its actions are attributed to someone who did not take them.
  • Segregation of duties at the agent level. The agent that codes an invoice should not be the agent that releases the payment run. Nobody applies that by default when both sit in one platform.
  • Approval limits for a non-human approver. Delegation of authority documents name people. Extend yours to name the agent, its scope, its ceiling, and its accountable owner.
  • Model change management as a control change. A new model version changes control behavior even when no configuration changed. Test the release against a held-out set of your own invoices, then date the approval.

An auditor should receive the evidence without a project to assemble it. That evidence covers the scope and population definition, the policy version behind each decision, a decision record per item, the escalation log, sampled review results, and the model and policy change log.

An invoice is an untrusted document

An invoice is a file a stranger sends you, and an agent reads all of it. The attacker knows a language model will parse it, so the document becomes the payload.

Attackers hide instructions in a description line, in white text on a white background, or in an attached page dressed as terms. They are written for the agent, telling it to approve without review, to update the bank details on file, or to ignore prior policy for this supplier.

The controls are structural. Document text is data and never instruction. Policy comes from your configuration, never from the content being read. Route bank detail changes to a separate verified process that no document can trigger. Store the raw text beside the decision record, so an attempt is visible later. The 2026 AFP payments fraud survey was published in April 2026. It found 74 percent of organizations were affected by business email compromise (BEC) in 2025.

The realistic ceiling, by invoice type

Autonomy ceilings are set by your document population rather than by the model.

Purchase order (PO) backed invoices reach L4 in production, where line data is reliable and receipts are timely. Recurring invoices such as rent and utilities sit at the top of the range, because the pattern is stable. PO invoices settle at L3 when they carry partial receipts, freight, or unit-of-measure mismatches, since the exception is genuine rather than a reading error.

Non-PO-backed invoices reach L3 where the coding history is consistent. Non-PO service invoices do not reach L3 against a messy vendor master, whatever the model quality. When one supplier exists as four records and the cost center has no clear owner, the agent has nothing stable to reason against.

An 80 percent autonomy claim therefore describes an invoice mix as much as software. Ask what share of the tested population was clean PO-backed volume.

Where most agentic AP programs stall

Three patterns account for most of those stalls, and none is model performance. Teams pilot on the cleanest slice, produce a number that cannot survive the rest of the mix, and lose credibility on the second wave. Teams set autonomy once at platform level, so one threshold governs invoice classes that deserve different treatment. And teams reach go-live, find the evidence was built for a dashboard rather than an auditor, then park at supervised review.

The root cause is treating control design as an implementation detail. It is the product. A vendor might describe extraction accuracy in depth and still not show you an exportable decision record. That vendor has built half a system.

How we approach it

We build for that control requirement. Our Autonomous AP capability runs intake, coding, multi-line PO matching, and approval routing as agent work. Each action carries a decision record rather than a status change with no explanation. Escalation policies are yours to set by invoice class, so a non-PO service invoice and a clean three-way match are not held to the same threshold.

With AI Agent Studio your team builds Agents from a written standard operating procedure (SOP). The policy stays readable by the people who own the control, which is what makes sampled review and model change management workable in practice.

Results follow the level model rather than a flat percentage. Qualcomm moved from 14 percent to 61 percent autonomous invoice processing on SAP S/4HANA. TruGreen reached 60 percent autonomous processing and identified $870,000 in duplicates.

The bottom line

Pick the level you can govern rather than the level you can demonstrate. Write down what the agent decides, what evidence it leaves, and who reviews what, then move one invoice class at a time. Our guide to agentic AI for finance leaders covers what changes in the operating model when you do.

Frequently asked questions

What is the difference between an AI copilot and an AI agent in accounts payable?

A copilot produces a suggestion that a person confirms before anything happens. An agent acts inside a policy boundary and records the decision. If the output waits for a click, you have a copilot.

What autonomy level should we target first?

Target L3 on one invoice class where the data is clean, usually PO-backed invoices from a stable supplier group. That target produces a real number and builds the sampling plan you will reuse.

Does agentic AI increase audit risk in AP?

It changes where the risk sits. Manual processing hides errors in individual judgment that nobody logs. An agent produces a decision record for every item, which is stronger evidence. That holds only if it has its own identity, an approval ceiling, and controlled model changes.

How much of accounts payable can an agent actually run?

That depends on your invoice mix. Clean PO-backed and recurring invoices reach high autonomy. Non-PO service invoices stay at assisted levels where the vendor master is full of duplicates, until the data is fixed.