Auditing expense reports means testing every claim against the document submitted to support it and against written policy. An auditor runs nine tests in a fixed order, from receipt reconciliation and arithmetic through document provenance, threshold proximity, cross-period history, and attendee validation. Each test produces a distinct finding, and the order keeps two reviewers consistent.
Key takeaways
- AI-generated documents rose from zero percent of flagged fraudulent receipts in March 2025 to 70.8 percent by mid-May 2026 in our own platform data. A crisp, well-formed receipt is now weak evidence.
- Read the shape of the report before opening a receipt. The expected trip cost, items near a policy limit, the category mix, and the submitter's last six reports take about 30 seconds to review.
- Document inspection and cross-period duplicate testing are the two tests that stop working by hand. Both need the whole population in view rather than a sample.
- A finding survives challenge when it names the line item and amount, quotes the policy clause by number, describes the evidence as observed, states the disposition, and gives a response window.
Two auditors reading the same expense report often reach different conclusions. The gap is usually sequence rather than skill. Nine tests do the work, and each one narrows what the next test has to examine. Run in a fixed order, they produce a more consistent result than an unstructured read. Every line item raises four questions. An auditor asks whether the claim has supporting evidence, whether it is correctly valued, whether it falls inside policy, and whether it is documented as policy requires.
Why auditing expense reports became harder in 2026
Expense documents have stopped working as reliable evidence. Most audit technique still assumes that they do.
PYMNTS reported our platform data in 2026, in coverage of AI-generated fake receipts. Generative artificial intelligence (AI) means software that creates realistic images on request. Documents produced that way rose from zero percent of flagged fraudulent receipts in March 2025 to 70.8 percent by mid-May 2026. That analysis covered 1,471 fake receipts from 745 employees at 174 companies, claiming a combined $148,143 in fabricated reimbursements. The average AI-generated fake was worth roughly $100, against $182 for older fakes built from reused templates.
The second figure matters more in practice. Employees submit cheaper fakes in greater numbers, and they fall below the amounts that trigger manual review. A program that scrutinizes only the largest line items therefore misses most of the volume.
HR Executive reported a 2026 employee AI receipt survey covering 2,000 workers in the United States and the United Kingdom. Four in ten US employees said they had used AI to create a fake receipt. Nearly 20 percent had fabricated a purchase outright, about 15 percent had inflated a real one, and 6 percent had replaced a receipt they had lost. That last group matters for technique. A regenerated receipt for a real purchase shows the same technical traces as a fabrication. A test identifies the document rather than the intention behind it, so the finding has to describe the document.
Auditors learned to treat smudged photographs of crumpled paper as the norm. They can no longer rely on that expectation.
Read the shape of a report before opening a receipt
Experienced reviewers start at the report level. The shape of a report shows where its risk is, and reading it takes about 30 seconds. Four checks do most of the work.
- Estimate what the trip should cost before opening any receipt. A three-day domestic conference has a rough expected total, and a report at double that figure has a reason worth locating.
- Identify line items within a few dollars of a policy limit. Proximity to a limit is a stronger signal than size.
- Compare the mix of spending categories against the stated business purpose. Client entertainment with no associated travel warrants a second look.
- Review the submitter's previous six reports. An auditor who has read those six sees the current one differently.
The standard of proof follows from the four questions above. An audit does not determine whether an employee is honest. Holding that line keeps findings easier to write and harder to dispute.
Decide in advance what evidence would change your assessment of a suspicious item. That habit stops one odd receipt from turning a review into a search for confirmation.
The nine tests for auditing expense reports, in order
The sequence begins with the simplest comparisons and ends with the tests that depend on wider context. Each one produces a distinct signal and a distinct type of finding.
1. Reconcile the receipt to the line item
Compare the claimed amount, date, currency, and merchant against the attached document, field by field. Most mismatches are innocent, and they still have to be cleared. A claim dated a day after the receipt often reflects a time zone difference or a posting date. A claim above the receipt total is a valuation finding regardless of intent. Record the two values and the difference.
2. Check the arithmetic on the document
Add the line items and compare the sum to the printed subtotal. Then check tax and tip against that subtotal. Fabricated documents fail this test often, because the software that produces them generates plausible figures rather than consistent ones. Tax that matches no rate in the relevant jurisdiction is the strongest arithmetic signal. A tip that computes to 42.7 percent is close behind. State the arithmetic in the finding rather than characterizing it.
3. Test the itemization against the total
Ask whether the itemized detail supports the total claimed. Group meals and hotel folios matter most here. A folio is the running statement of charges a hotel issues at checkout, and it bundles room rate, tax, resort fees, parking, and in-room charges. Personal items routinely appear on it. Where policy requires an itemized receipt and only a card slip is attached, the result is a documentation finding, the type most often waived.
4. Test merchant plausibility
Assess whether this merchant, in this category, at this location, at this time, fits the stated business purpose. Four signals do most of the work. The first two are a merchant that does not exist at the address shown, and a category that conflicts with the trip type. The others are a charge in a city the traveler was not in that day, and a merchant name formatted differently from the same merchant elsewhere in the data. Fabricated merchants look correct in isolation and wrong across a population, so compare each one against history.
5. Inspect the document rather than only its contents
Examine the file itself. Image provenance, the record of where a file came from, and metadata, the technical details stored inside the file, come first. Consider next resolution that is unusually uniform, fonts that do not match the merchant's template, and alignment that is too precise. Repeated visual traces across receipts from different merchants point to a shared generator. Most manual programs never run this test, because it means looking at the file rather than reading it. Record what the document shows, not what the employee is presumed to have done.
6. Measure threshold proximity
Sort the line items by their distance from the nearest policy limit. A $74 meal against a $75 cap is a signal, and three such items on one report form a pattern. Sustained clustering just below a threshold across several periods is the most reliable behavioral indicator available to a manual auditor. It is also the most likely to be innocent on a single report, so hold it in a trend finding until the pattern persists.
7. Run cross-period history and duplicate testing
Compare the report against the submitter's earlier submissions and against colleagues on the same trip. Look across periods rather than within one. The same receipt filed twice in different periods, the same expense claimed by two attendees, and a personal card claim covering something already on the corporate card all surface here. Duplicate expense detection covers the mechanics. Duplicates are the clearest case in which a single report contains no evidence at all.
8. Validate attendees and business purpose
For meals and entertainment, confirm three things. The attendee count matches the receipt, the named attendees are real, and the amount per person falls inside policy. Government officials and healthcare professionals among the attendees move the item into a different compliance regime, which is a finding even when the spending is within limits. A business purpose recorded as client meeting with no client named is incomplete documentation.
9. Write the finding so that it holds up
A finding survives challenge when it contains five elements.
- Name the specific line item and its amount.
- Quote the policy clause by number.
- Describe the evidence as observed rather than as characterized.
- State the proposed disposition.
- Give the response window.
Two versions of the same conclusion show the difference. "Receipt looks fake, denying the claim" gives the employee nothing to answer and the controller nothing to stand on. The alternative reads as follows. "Line 4, $128.40, dated March 12. Section 6.2 requires an itemized receipt for meals above $75. The attached document shows a subtotal of $118.00 with tax of $10.40, which is 8.8 percent against a jurisdiction rate of 6.5 percent. Requesting the itemized receipt within five business days." The second version is checkable, and it states the same finding. Characterizations of intent weaken what an investigator has to work with later, so they belong outside the file.
Where auditing expense reports by hand reaches its limit
Every test above works when it is run. A team cannot run several of them by hand once volume rises.
Document inspection is the clearest example. Checking metadata and template consistency on one receipt takes about a minute. Doing that on every receipt on every report is not a staffing problem an organization can hire its way out of. Cross-period duplicate testing has the same shape, because it needs an entire history in view rather than a single report.
Sampling makes both problems worse. Reviewing part of the population rather than all of it means threshold clustering never becomes visible, since the pattern is in the reports nobody opened. The expense report audit process sets how much of this sequence a team runs at all. A vague expense policy leaves findings with no clause to cite. Technique alone corrects neither condition.
How we approach auditing expense reports
We built Expense Audit to run this sequence on every report rather than a sample. Our AI reads every line of every receipt before reimbursement and applies the tests a practitioner would apply. We cover 100 percent of expense reports, and we review card transactions as purchases post.
The detection layers map onto the sequence above. Image provenance and metadata cover test 5. Pattern recognition covers threshold clustering and cross-period history. Merchant authentication covers plausibility, mathematical validation covers arithmetic and tax, and completeness verification covers documentation and attendees. More than 40 pre-built travel and expense audit models apply these tests across 42 languages and 97 countries. That includes validation of fapiao, the official invoices issued under Chinese tax rules, and of value-added tax, which manual programs apply inconsistently.
For an auditor the change is one of attention. Routine findings are already written against a line item and a clause, which leaves judgment calls and pattern work to people.
The bottom line
Auditing expense reports is a sequence rather than a single judgment. Run tests 1 through 9 in order across a set of recent reports and note which ones produced results by hand. The tests that produced nothing make the strongest case for automation, because nobody was really running them. Our overview of expense report auditing describes how a team runs the full sequence before payment.
Frequently asked questions
What does auditing expense reports involve?
It means testing each claim against its supporting document and against policy. The work covers reconciliation, arithmetic and tax validation, merchant plausibility, document provenance, threshold proximity, cross-period history, and attendee validation. Each test produces its own type of finding.
What are the signs of a fake receipt?
Tax or tip figures that do not compute against the subtotal are one sign, as is a merchant that does not exist at the address shown. The file itself may show uniform image quality, unusually precise alignment, fonts that differ from the merchant's template, and repeated visual traces across receipts from different merchants.
How should an expense audit finding be written?
A durable finding names the line item and amount, quotes the policy clause by number, describes the evidence as observed, states the proposed disposition, and gives a response window. Judgments about intent are left out, because they weaken the record.
Can manual review still catch AI-generated receipts?
It catches some of them. Arithmetic errors and merchant inconsistencies stay visible to a careful reviewer. Metadata checks on every receipt and duplicate testing across periods are the parts that stop working once volume rises.