415.tech
AI & tech, from the frontlines of Silicon Valley

Claude Code and Codex misjudge elapsed time by 3x to 10x on coding tasks

Two MATS researchers tested Claude Code and Codex across 218 coding tasks and found both guess around 90 minutes regardless of difficulty, overshooting by 3x for Claude and 6-10x for Codex, with short tasks worst; Opus 4.8 and GPT-5.5 also rated their own output about 20 points above actual scores, once claiming 70 percent success against real marks of 7 and 14.5 percent. Giving agents a tool that reports elapsed time fixed the estimates almost every time, so any instruction like 'iterate for two hours' needs an explicit clock in the harness rather than the model's own sense of duration.

Source: the-decoder.com

Post on XEmail