Today we are introducing ZenLM Plus, a new family of finance-specialized large language models built for the decisions at the center of finance operations. The models read financial documents, interpret transactions, apply company policy, identify risk, and explain the evidence behind each decision. ZenLM Plus extends the existing ZenLM intelligence layer. It gives finance teams better-than-frontier quality on the tasks where specialization matters, without forcing every expense line through a premium frontier model. As more ZenLM Plus capabilities enter the agent harness, customers gain a path to more capable agents, lower operating costs, and models that improve specifically for finance work.
Key takeaways
- ZenLM Plus is available today to select customers in AppZen Expense Audit, with general availability expected in December 2026.
- In the current benchmark table, ZenLM Plus led six of seven comparisons against seven frontier models.
- AppZen’s Mastermind Platform routes each finance task to the model or method that performs it best, including a frontier model when one leads.
- Under the benchmark’s modeled assumptions, the least expensive frontier model cost about twice as much as ZenLM Plus, while Opus 5 cost about 50 times as much.
- Task-level evaluation and governed feedback loops allow ZenLM Plus to improve for the finance decisions customers need it to make.
- ZenLM Plus is available today to select customers using AppZen Expense Audit and will be generally available in December 2026.
Why autonomous finance operations need specialized language models
Finance transformation is entering a new phase. Teams are moving from AI that assists them to authoritative AI agents that perform work autonomously. That shift raises the bar on every decision an agent makes. An agent that approves an expense needs to apply the company’s policy correctly and show the evidence behind its decision.
Frontier models bring broad general intelligence to that work. A customer’s policy is a different matter. On a general benchmark, each question usually has one correct answer. An audit control has no single correct answer across companies. One company approves a hotel bill that another flags because each sets its own thresholds, categories, documentation rules, and risk tolerance.
An audit model can make two kinds of errors, each with a different consequence. A false positive sends a compliant report to an auditor and asks an employee to defend a legitimate purchase. A false negative reimburses a policy violation. The first consumes auditor time and employee goodwill. The second leaks spend. A model built for audit work must keep both errors low at a cost that allows the company to evaluate every report.
“Autonomous finance requires more than a powerful model. It requires the complete stack, with finance-specialized models, proprietary data, intelligent workflows, and enterprise governance working together to make decisions and take action with accuracy, transparency, and control. ZenLM Plus now outperforms the frontier models we tested on expense audit tasks while operating at a lower cost. Combined with AppZen’s agent-first platform, it delivers on the promise of finance transformation, with our AI Agents safely performing more work, autonomously, at enterprise scale.”
Anant Kale
CEO and co-founder, AppZen
How ZenLM Plus works
ZenLM Plus is AppZen’s family of proprietary, finance-specialized large language models. AppZen builds the models on open-weight LLM foundations, then adds finance-specific training, document understanding, customer policy configuration, transaction search, business rules, confidence controls, and evidence an auditor can follow.
Continuous learning for finance
Frontier models continue to improve broadly. ZenLM Plus can improve specifically for the finance decisions our customers need to make. Because we control the task-specific training, evaluation, validation, and deployment process, we turn operational learning into measurable improvements instead of waiting for the next general model release.
As our Mastermind Platform processes more financial documents and decisions, it can identify new patterns, add difficult cases to evaluation sets, learn from auditor outcomes, and validate changes before deployment, subject to enterprise privacy, security, and governance requirements. This is not uncontrolled learning in production. It is a governed feedback loop designed around finance work.
The same loop enables task-based learning in our platform. We measure which model or method performs best for a defined task, route work accordingly, and use task-level outcomes to improve specialized models. As additional ZenLM Plus models become available through the agent harness, customers will benefit from agents that become more capable and less expensive to operate.
What ZenLM Plus covers in Expense Audit
ZenLM Plus supports controls used by travel and expense teams in AppZen Expense Audit. The current capabilities span four practical areas.
- Policy and category decisions. Apply customer-defined rules to premium travel and upgrades, electronic devices and gifts, personal expenses, and merchant categories.
- Receipt validity and transaction verification. Determine whether evidence qualifies as a merchant-produced receipt and whether the merchant, date, amount, expense type, and currency support the submitted expense.
- Documentation and itemization requirements. Verify that receipts contain the required line-item or daily detail.
- Cross-report duplicate detection. Search the submitting employee’s reports and other employees’ reports for the same receipt or underlying purchase.
Depending on the control, ZenLM Plus can read employee-entered fields, receipt extractions, payment type, customer-defined categories and thresholds, merchant information, card data, and related expense reports. Its judgment reflects how the customer defines risk.
ZenLM Plus led six of seven expense audit control comparisons
Expense Audit is the first proof point of ZenLM Plus’s capabilities. We compared ZenLM Plus on expense audit tasks against Opus 5, Sonnet 5, Gemini 3.1 Pro, Gemini 3.5 Flash, GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.6 Luna. For each task, every system received the same available expense data, receipt images, supporting documents, and customer configuration.
We also measured the models’ precision and recall for each task. Precision is the share of flagged expenses that were true violations. Recall is the share of true violations the system caught. The F1 score combines the two on a scale of 0 to 100, where 100 indicates perfect precision and recall.
ZenLM Plus’s largest advantage appeared in premium travel and upgrades. It scored 97.4 F1, 10.8 points above GPT-5.6 Sol, the strongest frontier result for that task. It also led non-conforming receipt detection by 7.2 points and electronic devices and gifts by 4.2 points.

Figure 1. F1 score of ZenLM Plus and seven frontier models on each evaluated audit task. Higher is better.
Where specialization makes the difference
Non-conforming receipt detection
This control tests judgment about the role a document plays in a transaction. A card payment slip proves the employee paid, yet it does not qualify as a merchant receipt. An emailed order confirmation may qualify even though it looks nothing like a traditional receipt. On this task, ZenLM Plus scored 92.4 F1, 7.2 points above Opus 5 at 85.2.
Policy and category decisions
Premium travel can involve cabin class, paid upgrades, or luxury and special vehicle rentals. Electronic-device and gift policies require the system to identify what was purchased and apply the customer’s category definition. Merchant category matching can use web search, merchant databases, merchant category codes, and social media signals to determine what a merchant actually does. ZenLM Plus led each comparison, though the merchant-category margin was nearly tied at 0.2 points.
Duplicates and itemization
ZenLM Plus scored 97.9 F1 on duplicates across reports, 2.4 points above Gemini 3.5 Flash. It scored 95.8 on receipt itemization verification, 2.1 points above Gemini 3.5 Flash. Both controls require the system to connect evidence across fields or documents rather than classify a single line in isolation.
Where a frontier model led
Receipt verification was the one comparison where ZenLM Plus did not lead. It scored 93.3 F1, while Gemini 3.1 Pro and Sonnet 5 each scored 94.1. Model leadership is task-specific. When a frontier model performs best on a task, the Mastermind Platform can assign that task to it.

Figure 2. F1 score of every evaluated system on each audit task.
“With ZenLM Plus, we are building the intelligence layer finance teams actually need. Our models are trained on real financial documents, policy decisions, and audit outcomes, then evaluated task by task before they reach production. That approach delivered stronger results on most controls we tested at a fraction of the modeled inference cost. And when a frontier model is better for a specific task, AppZen can use it. Our goal is the best outcome for the customer.”
Kunal Verma
CTO and co-founder, AppZen
Cost changes what can be automated
ZenLM Plus delivered the lowest modeled inference cost among the systems evaluated. GPT-5.6 Luna, the least expensive frontier model tested, cost about twice as much. Opus 5 cost about 50 times as much.
That gap decides coverage. A lower cost per expense line makes it practical to run specialized AI across every report instead of reserving it for a sample or a small set of investigations. It also lets us use a more expensive frontier model selectively, when the improvement on a particular task justifies the premium.
These are relative estimates, not total-cost-of-ownership figures. Frontier-model costs use API pricing, while ZenLM Plus reflects GPU-provisioning costs.
How we ran the evaluation
We drew our test cases from anonymized, production-derived expense patterns and targeted scenarios. Each system classified the expense in a structured format. For duplicate detection, we first retrieved the strongest candidate transactions, then gave each model those transactions and their images. We verified the reference labels for the headline results against the receipt images.
This is a focused benchmark of expense audit tasks, not a general ranking of AI models. ZenLM Plus was trained and tuned for this domain; the frontier models received the benchmark instructions and supplied evidence without our specialized stack. The results should be read as evidence about these tasks under the tested configuration. Companies should validate performance against their own policies, languages, document mix, and risk tolerance.
Availability of ZenLM Plus
ZenLM Plus is available today to select customers using AppZen Expense Audit and will be generally available in December 2026. Expense Audit is the first proof point, not the endpoint. AppZen will continue expanding task-specific models across finance workflows and making them available through Mastermind’s agent harness.
The future of autonomous finance will not be defined by a single model. It will be defined by a governed system that can choose the right intelligence for each task, learn from finance outcomes, and improve the balance of quality, control, speed, and cost.
Using ZenLM Plus in AI Agent Studio
ZenLM Plus also runs inside the AppZen AI Agent Studio, the intelligence and orchestration layer behind AppZen’s AI Agents. AppZen’s agent harness does not assume that one model should handle every decision. It can route a task to ZenLM Plus, deterministic logic, retrieval and validation services, a frontier model, or a person, depending on the task, the available evidence, and the customer’s risk tolerance.

Figure 3. AppZen Mastermind routes each finance task to the model or method that performs it best.
Frequently asked questions
What is ZenLM Plus?
ZenLM Plus is AppZen’s family of proprietary, finance-specialized large language models. The models run inside the AppZen Mastermind Platform and power decisions in AppZen Expense Audit.
How does ZenLM Plus compare with frontier models?
In the current expense audit benchmark table, ZenLM Plus led six of seven comparisons against seven frontier models. Receipt verification was the exception, where Gemini 3.1 Pro and Sonnet 5 each scored 0.8 points higher.
Does AppZen still use frontier models?
Yes. The AppZen Mastermind Platform routes each finance task to the model or method that performs it best, including a frontier model when one leads.
How does ZenLM Plus improve over time?
AppZen uses task-level evaluation and governed feedback loops to identify difficult cases, learn from auditor outcomes, validate improvements, and deploy better finance-specific models.
When will ZenLM Plus be generally available?
ZenLM Plus is available today to select customers using AppZen Expense Audit and is expected to become generally available in December 2026.