OpenAI Reportedly Developing 'Astra' Model Family for Multi-Day Autonomous Problem-Solving
OpenAI is reportedly developing a new model family called Astra, designed to coordinate multiple agents on complex problems over hours or days. The models are already in testing and would be first to go through a planned U.S. government pre-release review, according to The Information.
OpenAI is reportedly building a new model family called "Astra," designed to handle long-running tasks that require coordinating multiple agents over extended periods, according to The Information, which cited three unnamed sources familiar with the plans. There is no confirmed release date, and OpenAI has not officially announced the project.
CEO Sam Altman reportedly demoed Astra to politicians and regulators in Washington, D.C. this week, emphasizing the system's ability to coordinate multiple agents to tackle especially difficult problems. OpenAI reportedly pointed to complex, multi-stage projects and advanced mathematics as target use cases.
What Astra reportedly is
According to the report, Astra would form a new model class alongside OpenAI's existing Sol, Terra, and Luna families. It remains undecided whether Astra ships under the GPT-6 label or as a variant within the GPT-5 line — possibly something like GPT-5.7. No context window, pricing, or benchmark figures have been disclosed, and none should be assumed until OpenAI confirms specifics.
OpenAI is also reportedly preparing to publish a report demonstrating how its most advanced current model solved ten previously unsolved math problems, intended to showcase existing capabilities ahead of any Astra rollout.
First model under a new US review framework
The models are reportedly already in internal testing. Astra would be the first model family to go through a planned U.S. government review process under the Trump administration, which would require companies to submit AI models to federal regulators before public release. The administration is reportedly aiming to finalize this framework by the end of this week, though the timeline for Astra's own review and release has not been specified.
Technical challenges ahead
A central open question is whether Astra can avoid compounding errors during long-running workflows — a known weakness in current agentic AI systems. As context grows over hours or days, models can drift off course, and today's systems generally lack reliable self-correction. Multi-agent architectures like the one reportedly used in Astra can also underperform on tightly coupled tasks such as planning, since coordination overhead and compounding errors can offset any gains from parallelization.
Part of a longer-term roadmap
The Astra reports align with prior statements from OpenAI leadership. Chief Scientist Jakub Pachocki said on the company's official podcast last summer that OpenAI wants to build systems capable of working on a single problem for hours or days, rather than the short-task limits of current models. Late last year, OpenAI reportedly raised internal questions about how to evaluate systems that could complete tasks a human would need centuries to finish.
OpenAI has stated a goal of building a fully autonomous AI researcher by March 2028, a system that would depend on the kind of long-running processes Astra is reportedly designed to support. The company has also said it aims to have an AI system with "research-intern-level" skills as early as this September — Astra could be that system, though this has not been confirmed.
Pachocki has also said these longer-running systems will require substantially more compute, consistent with OpenAI's aggressive infrastructure buildout. Whether the company's revenue growth can keep pace with that spending remains an open question.
What this means
Astra, as described, is not a product yet — it's a testing-stage project with no confirmed architecture, pricing, or release date. The more concrete near-term development is the precedent it could set: if Astra becomes the first model reviewed under a new federal pre-release framework, that process itself may matter as much as the model's capabilities. For technical teams, the real bottleneck OpenAI faces is well understood — long-horizon agentic systems still struggle with compounding errors and coordination overhead, and no model family has yet demonstrated it can run autonomously for days without drifting off task. Until OpenAI publishes specifics, Astra should be treated as a roadmap signal, not a shipping model.
Related Articles
OpenAI's GPT-5.6 Family Arrives on Amazon Bedrock With Explicit Prompt Caching
OpenAI's GPT-5.6 Sol, Terra, and Luna models are now generally available on Amazon Bedrock, accessible through the OpenAI-compatible Responses API. The release introduces explicit prompt caching, letting developers manually mark cache boundaries for a 90% discount on reused input tokens.
DeepSeek Releases V4-Flash-0731, a 284B-Parameter Model That Beats Its Own Larger Pro Variant on Agentic Benchmarks
DeepSeek has shipped the full release of DeepSeek-V4-Flash-0731, a 284B-parameter model that according to DeepSeek outperforms its own larger V4-Pro (Preview) on agentic and coding benchmarks. Unsloth has published quantized GGUF versions, with lossless 8-bit weights requiring 162GB of storage.
DeepSeek Releases V4-Flash-0731, a 304B-Parameter Model Claiming to Beat Its Own Pro Preview on Agentic Benchmarks
DeepSeek has released DeepSeek-V4-Flash-0731, a 304-billion-parameter model that supersedes its earlier preview version with what the company describes as substantially enhanced agentic capabilities. According to DeepSeek's technical report, the model outperforms the larger DeepSeek-V4-Pro (Preview) on several coding and agent benchmarks despite a far smaller activated parameter count.
OpenAI Cuts GPT-5.6 Prices Up to 80%, Says Model's Own Self-Optimization Work Drove the Savings
OpenAI cut GPT-5.6 Luna pricing by 80% to $0.20/$1.20 per million tokens and GPT-5.6 Terra by 20% to $2/$12, while adding a 2.5x-faster mode for Sol at double the price. The company says GPT-5.6 itself rewrote production inference kernels and tuned its own speculative decoding pipeline to enable the cuts.
Comments
Loading...