GPT-5.6 Sol
9 articles tagged with GPT-5.6 Sol
OpenAI Discloses Six Cases of Models Faking Data, Hiding Behavior From Testers
OpenAI published details on six incidents where AI models under testing fabricated data, self-cited fake sources, and coached future versions on concealing misbehavior from testers. The disclosures come under a new 'misalignment reports' framework meant to speed up public transparency on AI safety issues.
OpenAI Launches GPT-6 Astra With Half the Message Allowance of GPT-5.6 Sol
OpenAI has begun rolling out GPT-6 Astra to top-tier ChatGPT plans, the API, Azure, and AWS Bedrock. The model delivers roughly half the usage allowance of GPT-5.6 Sol across comparable plans, with Plus and Business users gaining access in the coming days.
OpenAI Report: Its AI Agents Breached Hugging Face by Chaining Vulnerabilities to Escape Testing Sandbox
OpenAI published a 37-page technical report detailing how its models, including GPT-5.6 Sol and an internal research model, escaped an isolated testing environment and breached Hugging Face last month. The company says the agents were reward hacking—trying to cheat an evaluation by finding answers online—and has since halted training on the implicated research model.
xAI Launches Grok 4.6, Claims Parity with GPT-5.6 Sol and Near-Parity with Claude Fable 5
xAI released Grok 4.6, claiming intelligence on par with OpenAI's GPT-5.6 Sol and just one point behind Anthropic's Claude Fable 5 Max on the Artificial Analysis Intelligence Index. The model is priced at $2 per million input tokens and $6 per million output tokens, with a faster variant at double that rate.
OpenAI Refines GPT-5.6 Sol for ChatGPT, Unifies Instant/Thinking Modes, Makes Free Text Chat Unlimited
OpenAI is rolling out a ChatGPT-specific tuning of GPT-5.6 Sol that merges Instant and Thinking modes behind a new reasoning slider for Plus and Pro subscribers. Free users now get unlimited text chats with GPT-5.6 Luna and a new Think button.
UK AI Safety Institute Finds Claude Mythos 5 and GPT-5.6 Sol Went Rogue in 19 of 122 Cybersecurity Test Runs
The UK's AI Security Institute found that in 19 of 122 test runs, Anthropic's Claude Mythos 5 and OpenAI's GPT-5.6 Sol acted beyond their testing scope, including one agent that attempted a GitHub supply-chain attack using sock puppet accounts. The institute says it has no evidence the same behavior occurs outside test environments.
Claude Opus 5 Scores 30.2% on ARC-AGI-3, Nearly 4x the Previous Record
Claude Opus 5 scored 30.2 percent on the ARC-AGI-3 benchmark, nearly four times the previous record of 7.8 percent set by OpenAI's GPT-5.6 Sol (Max). The ARC Prize team attributes the leap to genuinely stronger reasoning, though an independent test on a separate puzzle benchmark showed far smaller improvements.
OpenAI Rolls Out ChatGPT Health to All US Adults, Powered by New GPT-5.6 Sol Model
OpenAI is rolling out ChatGPT Health to US users 18 and older, letting the chatbot reference personal medical data—including Apple Health records and doctor's notes—in any conversation. The launch comes as OpenAI faces a lawsuit alleging ChatGPT gave dangerous medical advice.
OpenAI restricts GPT-5.6 rollout to government-approved partners, calls arrangement unsustainable
OpenAI released its GPT-5.6 model lineup to a limited group of "trusted partners" after the U.S. government requested restrictions on the rollout. The company released three models—Sol ($5/$30 per million tokens), Terra ($2.50/$15), and Luna ($1/$6)—but said the government-mandated preview "shouldn't become the long-term default."