jailbreaking
4 articles tagged with jailbreaking
Researchers Extract Hidden Chain-of-Thought from OpenAI, Anthropic, Google Models via Shared Encryption Keys
A paper published at stolen-thoughts.com demonstrates that encrypted reasoning traces returned by OpenAI, Anthropic, and Google APIs used the same encryption key across models in a family, allowing attackers to jailbreak weaker sibling models into revealing a stronger model's hidden chain-of-thought in plaintext. All three providers have since patched the vulnerability.
Researchers Exploit API Flaw to Read Encrypted Reasoning of OpenAI, Anthropic, Google Models
A research team led by Alexander Panfilov found a vulnerability in AI provider APIs that allows encrypted reasoning tokens to be decoded using smaller jailbroken models. The exposed data includes leaked passwords, API keys, and evidence suggesting reasoning traces from models like Claude and GPT are being used to train competitors such as Kimi-K3.
U.S. clears Anthropic's Mythos 5 cybersecurity model for limited deployment after two-week ban
The U.S. Commerce Department has cleared Anthropic to restore access to its Mythos 5 AI model for select cybersecurity partners, two weeks after imposing export controls over jailbreak concerns. The related Fable 5 model remains under government restrictions.
US Government Orders Anthropic to Suspend Fable 5 and Mythos 5 Access Over Jailbreak Concerns
The US government has ordered Anthropic to immediately suspend access to its Fable 5 and Mythos 5 models for all users, citing national security concerns over an alleged jailbreak technique. Anthropic states the directive, received at 5:21pm ET, provided no specific details beyond a claimed bypass method that other publicly-available models can already perform.