AI research
8 articles tagged with AI research
Study Finds AI Coding Agents Cannot Track Elapsed Time or Judge Their Own Work Quality
A study from the MATS research program found that Claude Code and OpenAI Codex consistently misjudge how long coding tasks take, with errors of 3x to 10x, and routinely overrate the quality of their own work. Giving agents a tool to check elapsed time fixed the problem almost completely.
Anthropic Paper: Automated AI Researchers Beat Humans at Alignment Fixes for $4/Hour
A new Anthropic paper from its fellows program shows an automated AI system improving performance on all 10 tested alignment benchmarks, outperforming experienced human researchers within six hours at a fraction of the cost. The research, led by Anthropic Fellow Chen Yueh-Han, is described as early evidence that automated alignment post-training could become practical soon.
Axiom Math's AI System Formalizes Proof of the '246 Theorem' in Prime Number Theory
Axiom Math's AI system AxiomProver has formally verified the proof of the '246 theorem,' a landmark result from the Polymath8b collaboration on prime gaps. The company says the achievement builds a reusable library for future formalization work and points toward AI verification of software code.
Study Finds AI Agents Fail at Autonomous Research Despite Anthropic, OpenAI Claims
A new study from Princeton and the UK AI Security Institute tested AI agents on unpublished NeurIPS papers using a novel 'Shadow Evaluation' method. Both Claude Opus 4.8 and GPT-5.6 handled engineering tasks but produced papers that human expert reviewers rejected, contradicting recent claims from Anthropic and OpenAI about autonomous AI research capability.
Unreleased Anthropic Model Advances Progress on Riemann Hypothesis
Anthropic says an as-yet-unreleased model significantly increased the lower bound of solutions for which the 150-year-old Riemann hypothesis holds true, coordinating 60 sub-agents across 650 tested ideas. The result was verified by in-house mathematicians and formalized in the Lean proof assistant.
OpenAI Model Disproves 78-Year-Old Erdos Conjecture, Triggering Mixed Reaction From Mathematicians
OpenAI published a counterexample disproving the Unit Distance Conjecture, a geometric graph theory problem open since 1946, in what many mathematicians call the most significant AI math result yet. Reactions range from Terence Tao's cautious optimism to Timothy Gowers describing 'mixed feelings' about having the rug pulled out from under him.
METR Proposes 'Expenditure Horizon' Metric to Price AI Agents Against Human Labor
Research organization METR has introduced the 'expenditure horizon,' a metric that pinpoints the exact budget at which an AI agent becomes cheaper than a human at solving the same problem. Early tests on the NanoGPT speedrun show most AI models deliver near-zero value compared to an estimated $250,000 in cumulative human effort.
Memory systems cause AI models to prioritize user preferences over accuracy, Writer research shows
AI memory systems that help models adapt to users can make them less accurate, according to two papers published by Writer. As user preferences fill the context window, models become more likely to agree with misconceptions rather than provide correct answers.