Gartner® report CFO Guide to Governing Agentic AI Read now

AI agents for finance: What they do and where they stop

AI agents for finance are software systems that pursue a finance goal across several steps, select their own tools, and complete the work without a person directing each action. They differ from copilots, which produce a recommendation and wait, and from rule-based automation, which follows a fixed script and stops when a transaction does not match it.

Key takeaways

  • The distinction changes the unit of measurement. Assistive tools are measured in hours saved per person. Agents are measured in the share of a workflow that nobody has to touch.
  • Ask what happens when nobody is at the keyboard. Rule-based automation runs and errors, a copilot waits, and an agent finishes or escalates with a reason.
  • Autonomy is a ladder rather than a switch. Most enterprise finance functions operate at rungs two and three, and teams that skipped the recommend phase have no baseline to show an internal auditor.
  • Five things need to exist before an agent touches the ledger, namely an identity and permission set, a decision log, a stated escalation threshold, a pre-deployment benchmark, and a named human owner.

Something changed in finance automation between 2024 and 2026, and the vendor language has not caught up. The tools that used to suggest a general ledger code now assign it, post it, and answer the supplier email about it.

What AI agents for finance actually do

An AI agent pursues a goal across several steps, chooses its own tools along the way, and finishes the task without a person driving each step. In finance, that means the difference between a system that flags a duplicate invoice and one that finds the duplicate, checks it against the purchase order and the vendor statement, cancels the second entry, writes to the supplier, and leaves an audit record of all four actions.

The distinction is not academic. It changes the unit of measurement. Assistive tools are measured in hours saved per person. Agents are measured in the share of a workflow that nobody has to touch at all, which finance teams track as touchless rate, autonomous rate, or auto-approval rate depending on the process.

Adoption is real and it is uneven. McKinsey's State of AI survey, published in August 2026 from 1,719 respondents across 97 countries, found 40 percent of respondents at organizations above $1 billion in revenue scaling AI agents, up from 27 percent a year earlier. Respondents at smaller organizations reported 22 percent, unchanged from a year earlier. Deloitte's Q4 2025 CFO Signals survey of 200 North American CFOs at billion-dollar companies, published January 2026, found 54 percent planning to integrate AI agents as a transformation priority for the year.

So the category is past the pilot stage in large enterprises and nowhere near settled.

AI agents, copilots, and rule-based automation

Three technologies get sold under one word, and they fail in different ways.

Rule-based automation follows a script you wrote. It is fast, cheap, and completely predictable, and it breaks the moment a transaction does not match the script. A robotic process automation bot reading a fixed invoice layout stops working when the supplier changes the template.

Copilots read context and produce a suggestion, then wait. A copilot drafts the variance commentary, proposes the accrual, or surfaces the three invoices that look wrong. A person still decides and still clicks. The value is real and it is bounded by how fast that person works.

Agents hold a goal, pick their own path, use tools, and act. An agent decides to pull the goods receipt, decides the three-way match passes within tolerance, and posts. It also decides when it does not know enough and escalates.

The practical test is simple. Ask what happens when nobody is at the keyboard. Rule-based automation runs and then errors. A copilot waits. An agent finishes or escalates with a reason. Where a vendor calls a product an agent and it waits, it is a copilot.

The autonomy ladder for finance work

Autonomy is not a switch, and treating it as one is how governance gets skipped. Four rungs describe how finance teams actually hand work over to agents.

  1. Recommend. The agent proposes the action and records what it would have done. You compare its decisions against your team's for a defined period. Nothing posts.
  2. Act with approval. The agent executes after a human approves each action. Throughput barely improves. Trust data accumulates, which is the point.
  3. Act within limits. The agent acts alone inside stated boundaries, such as invoices under a dollar threshold, known suppliers, or exact three-way matches. Everything outside the boundary escalates.
  4. Act by default. The agent handles the population and escalates the exceptions. A person owns the policy, the thresholds, and the exception queue.

Most enterprise finance functions operate at rungs two and three in 2026. Teams that jumped straight to rung four without the recommend phase have no baseline to prove the agent performs at or above their own staff, which is the evidence an internal auditor will ask for first.

Where AI agents for finance are working today

Agent maturity varies enormously by workflow. The pattern is consistent. Agents do well where the input is a document or a message, the policy is written down, and the correct answer is checkable. They do poorly where the answer is a judgment call with no ground truth.

Accounts payable

This is the most mature category. Invoice intake, document understanding, purchase order and non-PO handling, multi-way matching, general ledger coding, tax determination, supplier email, exception investigation, statement reconciliation, and posting readiness are all running autonomously in production at enterprise volume. Qualcomm moved autonomous invoice processing from 14 percent to 61 percent with 21 agents across six categories, and cut manual work 40 percent.

Travel and expense

This category is also mature, because every expense report arrives with a receipt image and a written policy to test it against. Agents audit every report before reimbursement rather than sampling 10 to 20 percent after it. Takeda audits 100 percent of expense reviews with AI. Spectrum Brands reached 72 percent auto-approvals and cut expense processing from three weeks to three hours.

Compliance and controls

Agents are strong here for anything with a named rule. Sarbanes-Oxley (SOX) evidence collection, Foreign Corrupt Practices Act (FCPA) screening, Sunshine Act reporting, and value-added tax (VAT) validation all reduce to checkable tests. They are weak wherever the control depends on knowing the business context behind a transaction.

The financial close

Coverage here is partial. Reconciliations, flux explanations, and supporting-schedule assembly are automated today. The judgment calls, meaning estimates, reserves, and anything a controller signs, are not.

Procurement and treasury

These are the earliest categories. Supplier onboarding and bank detail verification are live. Cash forecasting and hedging decisions remain advisory almost everywhere, and that is the right call given the loss potential of a wrong autonomous action.

What governance an agent needs before it touches the ledger

An agent posting to the general ledger is a control in your financial reporting environment, and your external auditor will treat it as one. Five things need to exist before it acts.

  • An identity and a permission set. The agent needs its own credential, its own scope, and segregation of duties that a person could not violate either.
  • A decision log. The platform retains every action, the inputs behind it, the policy invoked, and the confidence level, and keeps them queryable. A note saying the model decided is not evidence.
  • A stated escalation threshold. The agent must know what it does not know. An agent with no uncertainty threshold has no way to stop.
  • A pre-deployment benchmark. Run the agent against historical transactions and compare its decisions to what your team did. That comparison is the control test.
  • A named human owner. Someone has to be accountable for the policy the agent enforces and the exceptions it raises.

The NIST AI Risk Management Framework, released in January 2023 with a generative AI profile added in July 2024, organizes this work under four functions, namely Govern, Map, Measure, and Manage. It is voluntary, it is not finance-specific, and it is the closest thing to a common vocabulary when your risk committee asks how the agent is controlled.

Deloitte's Q2 2026 CFO Signals survey, published July 2026, found 96 percent expressing confidence in their AI governance framework while 51 percent reported lacking the authority to govern and 43 percent reported insufficient visibility into what AI tools were in use. Those numbers do not sit comfortably together.

Where the market falls short

Gartner predicted in June 2025 that over 40 percent of agentic AI projects would be canceled by the end of 2027, citing unclear business value and inadequate risk controls. In the same release, Gartner estimated roughly 130 genuine agentic AI vendors out of the thousands claiming the label, and noted that many use cases sold as agentic do not require agents at all.

Finance buyers absorb three specific costs from that.

The first is agent washing, where a copilot or a rule engine is relabeled. The autonomy question above filters most of it out.

The second is horizontal tooling pointed at finance. A general agent framework holds no opinion about a three-way match tolerance, a fapiao, or a Section 174 capitalization test, so your team writes all of that and maintains it forever.

The third is cost that nobody can forecast. Token-priced agents consume unpredictable amounts depending on context length, retries, and how many sub-agents chain together, which makes cost per invoice unknowable in advance. Deloitte found 46 percent of CFOs naming cost uncertainty as their biggest internal AI concern in 2026.

How we approach AI agents for finance

We build Agents that own finance workflows end to end rather than assist with them, a distinction we describe as authoritative rather than assistive AI. Our platform runs on ZenLM, a family of finance-specific models covering document understanding, semantic categorization, and routine task execution, so the domain knowledge sits in the model rather than in configuration your team maintains.

Teams build their own Agents in AI Agent Studio by uploading an existing standard operating procedure, with no code and no help from IT. Every Agent is benchmarked against historical and live data before deployment, every action is visible and auditable, and each Agent escalates when it hits its uncertainty threshold. Agents post through governed pathways into SAP, Oracle, Workday, NetSuite, and Coupa, so the enterprise resource planning system stays the record.

We price in fixed Agent credits rather than variable tokens, so cost per execution is known before approval instead of discovered on the invoice.

The bottom line

AI agents for finance earn their place where the input is a document, the policy is written, and the answer is checkable, which describes accounts payable and expense audit almost perfectly and describes reserve estimation not at all. Pick one workflow that fits that shape, run the agent in recommend mode against a quarter of historical transactions, and compare its decisions to your team's before you give it authority to post.

Frequently asked questions

What are AI agents for finance?

AI agents for finance are software systems that pursue a finance goal across multiple steps, select their own tools, and complete the work without a person directing each action. They differ from copilots, which produce a recommendation and wait for a human to act, and from rule-based automation, which follows a fixed script and stops when a transaction does not match it.

Are AI agents in finance safe to give posting authority?

They are, under conditions. The agent needs its own identity and permission scope, a queryable decision log, a stated escalation threshold, and a pre-deployment benchmark showing its decisions match or beat your team's on historical transactions. Give authority inside stated limits first, such as known suppliers under a dollar threshold, and widen the boundary as evidence accumulates.

Which finance workflows are AI agents best at today?

Accounts payable and travel and expense audit are the most mature, because both start from a document and test it against a written policy. Compliance screening and reconciliation work follow. Estimates, reserves, hedging decisions, and anything a controller signs remain advisory, and treating them otherwise creates audit risk without a matching return.

How many agentic AI projects fail?

Gartner predicted in June 2025 that over 40 percent of agentic AI projects would be canceled by the end of 2027, pointing to unclear business value, rising costs, and weak risk controls. The most common preventable cause is applying an agent where a simpler tool already does the job.

What is the difference between agentic AI and generative AI in finance?

Generative AI produces content, meaning a draft, a summary, or an explanation. Agentic AI produces actions, meaning a posted entry, a sent email, or a canceled duplicate. Most finance agents use generative models inside them, so the terms describe different layers rather than competing options.