I spend a lot of time helping CEOs make better decisions with data. So when a new benchmark drops that actually measures AI performance on real professional work, rather than standardized tests, I pay attention.
Mercor and Ramp just published APEX-Accounting, a benchmark built to test whether AI agents can complete genuine accounting work at professional standards. The results are worth understanding before your next conversation with a vendor pitching AI for your finance function.
Let me break down what the data actually says.
What this benchmark actually measures
Most AI benchmarks test whether a model can produce the right answer once. That is fine for an exam. It is not fine for accounting.
Closing the books requires something harder: an agent must reconcile conflicting files, apply company-specific context, carry conclusions across multiple steps, and produce the correct result consistently. A model that drops a correct intermediate finding can still generate a bad journal entry.
APEX-Accounting was built to test exactly that. The benchmark includes 160 tasks across 10 simulated companies. Each company is fictional but internally coherent, complete with its own accounts, records, business history, and documents including spreadsheets and PDFs, all frozen at month-end close.
The experts who created the tasks had a median of 11 years of experience. More than half had worked at a Big Four firm. Each task was graded against a rubric with an average of 13.7 criteria.
This is about as close to real accounting work as a benchmark can get without using your actual books.
The headline number you should sit with
The benchmark ran every model on every task eight times. That repetition matters. Accounting work must be repeatedly correct, not correct once.
Even the best model in the test solved just 2.6% of tasks correctly across all eight runs.
Let that settle for a moment. The most consistent AI agent available today, tested against professional accounting work, was fully reliable on fewer than three tasks out of a hundred.
The top performer on the primary leaderboard scored 56.4%. That means even the best model fails to complete roughly four out of ten tasks that a human professional would handle. And 58% of tasks were never fully solved by any model on any run.
Where models are actually failing
This is the finding I think matters most for a CEO making real decisions.
The benchmark team worked with accounting and bookkeeping experts to categorize the failures. Roughly seven in ten failures came from flawed reasoning, not from an inability to find the right information.
A model might correctly identify a discrepancy early in a workflow, then omit or contradict that finding in its final journal entry. The model found the answer. It just failed to carry it through.
This is a judgment problem, not a retrieval problem. Better data access will not fix it. What is missing is the discipline to carry conclusions consistently across a complex, multi-step workflow. That is harder to patch than a knowledge gap.
The cost picture is more complicated than vendors will tell you
The benchmark also tested performance at different spending budgets: $1, $5, $10, and $50 per task. More budget generally allows more token usage, which can improve results.
But the relationship is not linear or predictable. One top model scored 11.8% on a $1 budget and improved to 55.2% at $50. Another top model was already strong at $1 and barely improved with more money. At the $50 cap, one model actually spent around $32 per run. Another spent around $5. Their scores were within 4 percentage points of each other.
If you are evaluating AI tools for your finance team, the cost-to-performance ratio is genuinely unpredictable across different models. You need to test, not trust the pitch deck.
What this means if you are a CEO making AI decisions
Here is how I translate this for the CEOs I work with.
First, AI can handle a meaningful share of accounting work today. Scoring above 50% on complex, professional-grade accounting tasks is not nothing. These models are useful tools in the right hands.
Second, “useful in the right hands” is the operative phrase. The failure mode is not that AI gets nothing right. It is that AI is inconsistent and fails in ways that are hard to detect from the output alone. A plausible-looking journal entry with an error buried in the reasoning is more dangerous than an obvious failure.
Third, the 58% of tasks never fully solved by any model defines the current boundary. If your finance team is planning to deploy AI on month-end close, they need a clear picture of where that boundary sits for your specific workflows, with your specific data, not a vendor’s benchmark on simulated companies.
That distinction matters. APEX-Accounting itself acknowledges it does not evaluate tax, audit, consolidation, multi-entity or multi-currency accounting, external reporting, or how agents handle requests for clarification. The scope is month-end close and bookkeeping. Your scope may be different.
The decision you should actually be making
The AI readiness question for your finance function is not “which model has the highest benchmark score.” It is “do we have the data infrastructure and oversight processes to deploy these tools safely at our scale.”
Inconsistent AI on top of inconsistent data does not produce better results. It produces confident-looking errors that are harder to catch.
Before adding AI to your accounting workflows, I would want to know your current data quality on the inputs those models would consume, your reconciliation error rate today as a baseline, and how your team would detect a reasoning failure buried in a plausible output.
Those are not AI questions. They are data strategy questions. And they belong in front before any tooling decision. It is the discipline I keep coming back to with clients: diagnosis before dollars. Diagnose before you spend. Pour the foundation before you frame the walls.
Find out how to make better decisions with your data through the professional data minds at JLytics.
Original source: Introducing APEX-Accounting (Mercor)
