MATS

1 article tagged with MATS

August 30, 2026
researchAnthropic

Study Finds AI Coding Agents Cannot Track Elapsed Time or Judge Their Own Work Quality

A study from the MATS research program found that Claude Code and OpenAI Codex consistently misjudge how long coding tasks take, with errors of 3x to 10x, and routinely overrate the quality of their own work. Giving agents a tool to check elapsed time fixed the problem almost completely.