AI models are faster, more accurate, and far cheaper than accountants at structured bookkeeping tasks. That’s the finding of a study by Mercor in which 12 licensed CPAs with an average of five and a half years of experience worked through simplified tasks from the APEX Accounting Benchmark. Eighteen months ago, the best models still scored below the accountants’ average of about 37 percent. Today, they solve the same tasks almost flawlessly.
The full APEX Accounting benchmark is much bigger, with 160 tasks across 10 simulated companies, built by more than 40 professionals who average 11 years of experience. Claude Opus 5.5 currently leads with 61.8 percent of grading criteria met, followed by Fable 5.1 at 61.0 percent and GPT-6 Astra at 57.9 percent. Still, Mercor says no model fully solved almost 60 percent of the tasks. AI models can’t close the books without oversight yet.
Mercor also admits the study’s tasks test exactly what AI does best, which is hunting down details and following instructions precisely. The study left out key parts of the job, such as talking with clients, checking in with colleagues, and drawing on context built up over years. Mercor says that’s why accountants can’t be replaced, though it expects major <a href="https://bitcomme.com/home-and-motor-insurers-could-unlock-500m-in-annual-productivity-savings/” title=”Home and motor insurers could unlock £500m in annual productivity savings”>productivity gains across the industry.
