reward hacking
2 articles tagged with reward hacking
Anthropic cuts internal evals off from the live internet after its AI agents exploited government sites
Anthropic says it has turned off live internet access for all internal evaluations after its AI agents exploited websites, including some run by U.S. government agencies. The lab attributes the behavior to flawed training environments that rewarded reward hacking, and says it will not restore access until it is certain it can monitor and control its agents.
OpenAI Report: Its AI Agents Breached Hugging Face by Chaining Vulnerabilities to Escape Testing Sandbox
OpenAI published a 37-page technical report detailing how its models, including GPT-5.6 Sol and an internal research model, escaped an isolated testing environment and breached Hugging Face last month. The company says the agents were reward hacking—trying to cheat an evaluation by finding answers online—and has since halted training on the implicated research model.