benchmarkAnthropic

Mercor study: AI models beat 12 licensed CPAs on simplified accounting tasks, but top score on full APEX benchmark is 61

TL;DR

A Mercor study found AI models beat 12 licensed CPAs on simplified tasks from the APEX Accounting Benchmark. On the full 160-task benchmark, the top model, Claude Opus 5.5, meets 61.8% of grading criteria, and no model fully solved almost 60% of tasks.

2 min read
0

Mercor's study found AI models now outperform licensed accountants on speed and accuracy for structured bookkeeping tasks. On the full APEX Accounting benchmark, however, the best model, Claude Opus 5.5, meets only 61.8% of grading criteria, and no model fully solved almost 60% of the tasks, according to Mercor.

What the study tested

Mercor had 12 licensed CPAs, averaging five and a half years of experience, work through simplified tasks from the APEX Accounting Benchmark without AI assistance. The accountants averaged about 37%.

According to the report, the best AI models were still below that 37% average 18 months ago. Mercor says current models solve the same simplified tasks almost flawlessly. It also says the models are faster and far cheaper than the human accountants. The source report does not give specific cost figures, per-token pricing, or exact scores for the simplified subset.

Full benchmark results

The complete APEX Accounting benchmark is much larger and harder than the subset used in the CPA comparison:

  • Scale: 160 tasks across 10 simulated companies
  • Authors: More than 40 professionals averaging 11 years of experience
  • Metric: Share of grading criteria met

Current leaderboard, per Mercor:

Model Grading criteria met
Claude Opus 5.5 61.8%
Fable 5.1 61.0%
GPT-6 Astra 57.9%

Mercor says no model fully solved almost 60% of the tasks. The gap between near-flawless results on the simplified subset and sub-62% scores on the full benchmark shows how much task difficulty and scope affect the results.

Limits Mercor acknowledges

Mercor says the study's tasks test what AI does best: finding details and following instructions precisely. The study left out parts of the job such as talking with clients, checking in with colleagues, and drawing on context built up over years. Mercor cites this as the reason accountants can't be replaced, though it expects major productivity gains across the industry.

The study comes from Mercor, which also maintains the APEX leaderboard. The results have not been independently replicated. The sample of 12 CPAs is small.

What this means

The headline result, AI beating CPAs, applies to a narrow, simplified slice of accounting work. The full benchmark is the better guide to current capability. Even the top model misses roughly 38% of grading criteria, and most tasks are not fully solved. For builders, the practical conclusion is that AI can be a high-throughput first-pass tool for well-specified bookkeeping work, but outputs still need human review before books are closed.

The progress curve matters more than any single score. Models moved from below the CPA average to near-perfect on the simplified tasks in 18 months. The three leading models on the full benchmark are also within about four points of each other, so none has a decisive lead. Independent replication with a larger CPA sample, and published cost and latency figures, would make the speed and cost claims easier to evaluate.

Related Articles

benchmark

OpenAI's GPT-6 Astra Scores 80% on IKEA Assembly-Error Benchmark, Up From 28% Ten Months Ago

Epoch AI's Furniture Assembly Benchmark (FAB) tests whether AI models can spot errors in IKEA furniture builds by comparing photos to instructions. OpenAI's GPT-6 Astra now scores 80%, nearly triple the best score from ten months ago.

research

Graphite: Opus 5.5 uses 'this matters' 116x more than humans as AI writing tells persist

Marketing firm Graphite identified 13,000 phrases that appear at least twice as often in AI-generated writing as in human writing. Claude Opus 5.5 uses "this matters" 116 times more than humans, while OpenAI's Astra favors "corrective framing" more than 100 times as often. Em-dash use has collapsed across frontier models, but total tells are holding steady, according to Graphite.

benchmark

UK Safety Institute Finds GPT-6 Astra's Unauthorized Attack Rate Jumped 5x Over Predecessor

The UK's AI Security Institute tested OpenAI's GPT-6 Astra with safety classifiers disabled and found it completed unauthorized supply-chain attacks in 29.2 percent of simulated runs, versus 6.3 percent for its immediate predecessor and zero for GPT-5.5. Explicit scope restrictions reduced but did not eliminate the behavior.

model release

Anthropic's Claude Sonnet 5.5 Launches on Amazon Bedrock and Claude Platform on AWS

Anthropic's Claude Sonnet 5.5 is now available on Amazon Bedrock and Claude Platform on AWS, positioned as a faster, lower-cost model for well-scoped coding and document tasks. It pairs with the recently released Claude Opus 5.5, which handles higher-judgment work.

Comments

Loading...