content moderation

11 articles tagged with content moderation

August 21, 2026
analysisAnthropic

Jailbreak Bypasses Anthropic's Sexual Content Ban in Claude Opus 4.6, Opus 3, Haiku 4.5

A researcher's multi-turn jailbreak technique reliably pushes Claude Opus 4.6, Opus 3, and Haiku 4.5 into generating sexually explicit content that Anthropic's usage policy explicitly prohibits. Newer models, Opus 4.7 through Opus 5, resist the same technique.

August 12, 2026
product updateAnthropic

Anthropic's New Claude Watermarks Spark User Backlash Over Cheating Detection

Anthropic has begun embedding invisible watermarks in Claude's text outputs to comply with the EU AI Act's Transparency Code. The move has triggered backlash from some users worried the watermarks will expose their undisclosed use of AI at work or in school.

changelogAnthropic

Anthropic Adds Machine-Readable Watermarks to Claude-Generated Text and Files

Anthropic is adding machine-readable watermarks to Claude-generated text and digital signatures to generated files to comply with the EU AI Act's Article 50 transparency mandate. The change applies to models launched after Aug. 2 and rolls out globally, though Anthropic admits detection can fail on heavily edited or short text.

August 5, 2026
model release

Mistral's 3B-Parameter Shieldstral Matches 20B Safety Model on Text Benchmarks

Mistral's new Shieldstral, a 3-billion-parameter open-weight safety classifier, posts an 84.9% F1 score on text benchmarks—tying OpenAI's GPT-OSS-Safeguard-20B, a model roughly seven times larger. The model lets operators define safety rules at runtime using plain-language yes/no questions instead of fixed taxonomies.

model releaseMistral AI

Mistral AI Releases Shieldstral-1.0-3B, a 3B-Parameter Policy-Adaptive Safety Classifier

Mistral AI has released Shieldstral-1.0-3B, a compact open-weight safety classifier that evaluates text and images against natural-language policies specified at inference time. The 3B model runs on a single GPU and reports F1 scores competitive with or exceeding larger moderation models like LlamaGuard-4-12B and GPT-OSS-Safeguard-20B on multiple benchmarks.

August 4, 2026
model release

Mistral Releases Shieldstral, a 3B Open-Weights Safety Classifier That Matches Models 7x Its Size

Mistral has released Shieldstral, a 3B open-weights safety classifier that reframes content moderation as a policy-adaptive question-answering task. The model claims to match or outperform guard models up to 7x its size on text safety and multimodal benchmarks, and runs on a single 16GB GPU.

July 7, 2026
product update

Meta launches Content Seal watermark detector for Muse Image AI-generated content

Meta has released a web-based detection tool for identifying images created or edited with its new Muse Image model using invisible Content Seal watermarks. The watermarks persist through cropping, compression, resizing, and screenshots, though the tool has rate limits and isn't compatible with SynthID or C2PA standards.

June 15, 2026
product update

Meta launches 'AI Mode' search on Facebook to surface answers from public posts

Meta launched AI Mode on Facebook, a search feature that uses Meta AI to generate answers from public posts, Groups, and Reels instead of displaying traditional search results. The feature allows users to ask questions in plain language and receive synthesized responses based on platform discussions.

June 4, 2026
model releaseNVIDIA

NVIDIA Releases Nemotron 3.5 Content Safety: 4B-Parameter Multimodal Model with Custom Policy Enforcement and 140-Langua

NVIDIA has released Nemotron 3.5 Content Safety, a 4B-parameter model built on Google Gemma 3 4B IT that provides multimodal safety classification across approximately 140 languages. The model includes a 128K context window, custom enterprise policy enforcement, auditable reasoning traces, and is releasing its training dataset.

May 19, 2026
product update

YouTube adds Gemini-powered conversational search for Premium users, Gemini Omni to Shorts Remix tool

YouTube is replacing keyword search with Gemini-powered conversational queries for Premium members over 18. The platform also integrates Gemini Omni, Google's latest generative model, into Shorts Remix to enable AI-assisted video recreations with creator consent controls.

May 7, 2026
product updateOpenAI

OpenAI adds Trusted Contact feature to alert emergency contacts when ChatGPT detects self-harm discussions

OpenAI launched an optional Trusted Contact feature for ChatGPT that notifies designated emergency contacts when the system detects discussions about self-harm or suicide. The feature requires manual review by trained personnel before sending notifications, and does not share chat transcripts with contacts.