The APEX Accounting Benchmark puts a number on what most teams have been experiencing qualitatively: AI models now outperform licensed professionals on structured accounting tasks. On a standardized 160-task suite evaluated by Mercor, Claude Opus 5.5 reached 61.8% accuracy — compared to roughly 37% for CPAs with 5.5 years average experience when they were benchmarked 18 months ago.
The benchmark gap is real. So is the caveat buried in the comparison.
The Numbers
The Mercor study evaluated 12 licensed CPAs against the full APEX Accounting Benchmark across simplified accounting tasks.
| Model | APEX Benchmark Score |
|---|---|
| Claude Opus 5.5 | 61.8% |
| Claude Fable 5.1 | 61.0% |
| GPT-6 Astra | 57.9% |
| CPAs (18 months ago) | ~37% |
The progression from 37% to 61.8% represents what 18 months of model capability improvement looks like in a domain that requires precision on numerical data, rule application, and instruction following — exactly the category of reasoning tasks where large language models have improved most substantially.
What the Benchmark Tests
The APEX benchmark is designed around the structured, rule-intensive parts of accounting work: following procedures, interpreting ledgers, applying tax rules, identifying correct account classifications. These are tasks where precision matters more than contextual judgment, and where AI models’ strength in “hunting down details and following instructions precisely” creates a genuine advantage over human performance on the same isolated tasks.
The Caveat: 40% Remains Incomplete
No model fully solved 60% of the complete benchmark. More precisely, Mercor’s researchers note that “AI models can’t close the books without oversight yet” — a specific claim that the capabilities demonstrated on benchmark tasks don’t translate to fully autonomous accounting operations.
The failure modes are predictable from the task composition:
- Client communication: Accounting requires explaining findings to non-accountants, negotiating on ambiguous interpretations, managing client expectations under deadline pressure. Benchmark tasks don’t test this.
- Colleague coordination: Audit teams, tax partners, management: the organizational dimension of accounting work doesn’t appear in isolated task performance.
- Contextual experience: Recognizing which transactions are unusual for a specific client, understanding the history that explains current anomalies, knowing which partners have which risk tolerances — this is the tacit knowledge that 5.5 years of experience builds.
The 61.8% accuracy and the 40% remaining incomplete aren’t in tension; they’re measuring different things. The benchmark captures the rule-following precision of structured accounting tasks. The 40% gap reflects the parts of accounting that require judgment, communication, and contextual experience.
The Implication for Accounting-Adjacent AI Systems
The relevant implication isn’t “AI will replace accountants” — that’s the framing the benchmark doesn’t support. The relevant implication is this: the structured, repeatable, rule-intensive parts of accounting work are now reliably solvable by models at rates above what licensed professionals achieve. In an enterprise context, that means:
- Document processing and classification: Automated
- Rule application and compliance checking: Automated, with review
- Exception identification: Automated, with escalation
- Client communication, judgment calls, relationship work: Human
The benchmark score gap isn’t argument for replacing CPAs. It’s evidence that the time allocation inside accounting work can change: professionals spend less time on high-volume mechanical tasks and more on the judgment-intensive work that isn’t captured in the benchmark at all.
The So What
Claude Opus 5.5 at 61.8% on accounting tasks where CPAs scored 37% is a leading indicator for how AI tools will be integrated into professional services workflows over the next 18–24 months. Not as replacements, but as productivity amplifiers on the rule-intensive work that consumes the most professional time for the least professional value.
For teams building or evaluating AI tools for financial services workflows: the APEX benchmark is a useful evaluation framework. For finance and accounting leaders: the benchmark gap makes the case for AI-assisted workflows on structured processing tasks — and the remaining 40% is precisely where human expertise retains its highest value.
Content created with AI assistance and reviewed for accuracy.
Join the conversation
Stack Insiders is our free community for readers who want to go deeper — share resources, ask questions, and connect with others across every vertical we cover.
Join Stack Insiders →