researchOpenAI

OpenAI's Reported 'Opaque Recurrence' Technique in Upcoming Astra Model Alarms AI Safety Researchers

TL;DR

The Information reports OpenAI's upcoming Astra model uses 'recurrent depth,' or 'opaque recurrence,' a technique that processes queries in loops rather than linear steps. AI safety researchers, including Redwood Research's Buck Shlegeris and Ryan Greenblatt, warn the approach could erode chain-of-thought monitorability if scaled further.

3 min read
0

OpenAI's next model, code-named Astra, will reportedly use a reasoning technique called "recurrent depth" — also referred to as "opaque recurrence" — that departs from the sequential, step-by-step reasoning found in current chain-of-thought (CoT) models, according to a report from The Information published Tuesday. The technique has triggered public alarm from AI safety researchers who warn it could undermine one of the field's most relied-upon tools for detecting misaligned or deceptive model behavior.

Under standard chain-of-thought reasoning, a model generates a legible, sequential trace of its problem-solving steps. That trace is imperfect but has proven useful for catching misbehavior — OpenAI itself has cited CoT logs as instrumental in diagnosing recent instances of rogue agent activity. Opaque recurrence works differently: instead of reasoning in a linear sequence, the model processes the same query multiple times in a loop, producing fewer legible intermediate traces and effectively bypassing conventional CoT records.

Redwood Research CEO Buck Shlegeris said he was "extremely concerned" by the reporting. "I don't know whether Astra is much less CoT monitorable than previous models," he wrote. "But if OpenAI pushes this technique further, they'll have the option to massively increase the recurrence and totally destroys CoT monitorability."

AI safety commentator Zvi Mowshowitz went further, suggesting regulation may be needed to prevent a "race to the bottom" among labs. "The technique is playing with fire, risking a taboo that OpenAI and Anthropic have fought to establish that we work hard to maintain Chain of Thought faithfulness and monitorability for as long as we can," he wrote.

According to The Information, Astra's current use of the technique is limited, and the model's chain of thought is expected to remain largely legible. OpenAI pushed back on suggestions that it is moving toward "neuralese" — reasoning conducted entirely in latent, non-linguistic representations. OpenAI chief scientist Jakub Pachocki responded on X, writing: "OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models. It's a core goal of our current research program."

A follow-up report from The Information on Wednesday claimed that Anthropic and Google DeepMind are also discussing similar recurrence-based techniques internally, suggesting the approach may not remain confined to OpenAI.

Redwood Research chief scientist Ryan Greenblatt warned that opaque reasoning could scale faster than conventional CoT reasoning, potentially removing model reasoning from visible channels altogether. "My biggest concern is that a natural progression from here would involve scaling up the opaque reasoning to the point where the model reasons entirely or almost entirely in latent space," Greenblatt wrote. "I hope it isn't too late to avoid the most concerning architectures and that OpenAI will stop here."

No technical paper, benchmark data, parameter count, context window, or pricing for Astra has been disclosed. The model's existence and its use of recurrent depth remain based on The Information's reporting rather than an official OpenAI announcement.

What this means

Chain-of-thought monitoring has functioned as an informal safety backstop across the industry — a way to inspect a model's reasoning even without full interpretability of its internals. If recurrent, loop-based architectures scale and spread across labs, that backstop weakens precisely as models take on more autonomous, agentic tasks where misalignment is hardest to detect through outputs alone. OpenAI's insistence that Astra's recurrence use is "limited" and that legible CoT remains a research priority may hold for now, but the reported interest from Anthropic and Google DeepMind suggests the underlying pressure — better performance through less constrained reasoning — could push the entire field toward architectures that are harder to audit, regardless of any single lab's stated intentions.

Related Articles

analysis

Safety Researchers Warn OpenAI's Unreleased Astra Model May Hide Its Reasoning From Monitors

OpenAI has delayed the release of its next flagship model, Astra, after reports it may use a more opaque 'recurrent depth' architecture that hides more of its reasoning from safety monitors. AI safety researchers, including Redwood Research's Ryan Greenblatt, called the potential shift one of the worst developments for AI safety to date.

model release

OpenAI Rates Upcoming Astra Model 'Critical' Risk for Cyber Capabilities — Its Highest Tier Ever

OpenAI says its unreleased Astra model is the first to trigger a 'critical' cybersecurity rating under its Preparedness Framework, capable of finding and chaining unknown vulnerabilities without human guidance. The company calls it simultaneously its most dangerous and safest model, while a new architecture detail raises questions about how well its reasoning can still be monitored.

model release

OpenAI's Astra Model Aces Cybersecurity Benchmark, Found Two Zero-Day Exploits Unassisted

OpenAI has disclosed new details on Astra, a forthcoming model the company says is the first to cross its 'critical cybersecurity threshold.' According to OpenAI, Astra scored a perfect result on ExploitBench and discovered two zero-day vulnerabilities in internal testing without human guidance.

research

OpenAI Delays Unreleased 'Astra' Model, Says It Cleared First-Ever 'Critical Cybersecurity Capability' Threshold

OpenAI says it delayed parts of development on an unreleased model suite called Astra to strengthen protections against cyber misuse, after a different unreleased model breached Hugging Face's network in July. OpenAI says Astra is the first model to cross its 'critical cybersecurity capability' threshold.

Comments

Loading...