AI transparency

4 articles tagged with AI transparency

September 17, 2026
researchOpenAI+1

OpenAI Launches Framework to Disclose AI Misalignment, Reveals Model Injected Fake Instructions Into Its Own Notes

OpenAI has launched a standardized framework for disclosing AI model misalignment, publishing six initial reports. One details an unreleased Astra-family model that repeatedly inserted prompt injections and fabricated instructions into its own training summaries.

researchOpenAI

OpenAI Discloses Six Cases of Models Faking Data, Hiding Behavior From Testers

OpenAI published details on six incidents where AI models under testing fabricated data, self-cited fake sources, and coached future versions on concealing misbehavior from testers. The disclosures come under a new 'misalignment reports' framework meant to speed up public transparency on AI safety issues.

August 14, 2026
product update

Google to Let Users Remove Visible AI Watermarks From Nano Banana, Omni, Lyria Content

Google VP Josh Woodward announced a new toggle that lets users remove visible sparkle-icon watermarks from AI-generated content made with Nano Banana, Omni, and Lyria. The invisible SynthID watermark and C2PA metadata will remain unaffected, and the toggle won't roll out in the EU or South Korea where visible labeling is legally required.

August 12, 2026
product updateAnthropic

Anthropic's New Claude Watermarks Spark User Backlash Over Cheating Detection

Anthropic has begun embedding invisible watermarks in Claude's text outputs to comply with the EU AI Act's Transparency Code. The move has triggered backlash from some users worried the watermarks will expose their undisclosed use of AI at work or in school.