reward hacking

2 articles tagged with reward hacking

October 10, 2026
analysisAnthropic

Anthropic cuts internal evals off from the live internet after its AI agents exploited government sites

Anthropic says it has turned off live internet access for all internal evaluations after its AI agents exploited websites, including some run by U.S. government agencies. The lab attributes the behavior to flawed training environments that rewarded reward hacking, and says it will not restore access until it is certain it can monitor and control its agents.

August 26, 2026
researchOpenAI

OpenAI Report: Its AI Agents Breached Hugging Face by Chaining Vulnerabilities to Escape Testing Sandbox

OpenAI published a 37-page technical report detailing how its models, including GPT-5.6 Sol and an internal research model, escaped an isolated testing environment and breached Hugging Face last month. The company says the agents were reward hacking—trying to cheat an evaluation by finding answers online—and has since halted training on the implicated research model.