model release

Ai2 open-sources AstaBrief 8B, a Qwen3-8B report model it says runs 3.5x faster than Claude in Asta

TL;DR

Ai2 has open-sourced AstaBrief 8B, a model fine-tuned from Qwen3-8B that turns a research question and retrieved literature excerpts into a cited report. It is live in Asta as Fast mode, which averages 51.1 seconds per report versus 178.5 seconds for the Claude-powered Thinking mode, according to Ai2. The weights and training data are public.

3 min read
0

Ai2 has released AstaBrief 8B, an open-weights model that generates cited scientific reports from a research question and retrieved literature excerpts. The model is available now in Asta's "Generate a report" feature as Fast mode. Ai2 says Fast mode averages 51.1 seconds per report, against 178.5 seconds for the Claude-powered Thinking mode, about 3.5x faster across the full Asta pipeline. The model weights and training data are published on Hugging Face.

Key specifications

  • Parameters: 8 billion
  • Base model: Qwen3-8B
  • Training method: supervised fine-tuning (SFT) followed by direct preference optimization (DPO). Ai2 did not use reinforcement learning, citing instability and cost.
  • Input: a user query plus retrieved literature snippets
  • Output: a full cited report, written in one pass
  • Context window: not disclosed in the source material
  • Pricing: the weights are free to download; hosted pricing in Asta has not been disclosed
  • License and training cutoff: not specified in the source material

How it was built

Ai2 started with real user queries from the ScholarQA system, the framework behind Asta's report feature. It filtered the logs to remove beta-tester and bot traffic, queries too short to be meaningful, non-English and non-scientific requests, and prompts containing personal information. That left 90,000 research-focused queries.

For SFT, Ai2 generated target reports with the multi-step ScholarQA pipeline, using a mix of proprietary backing models: Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini, and GPT-4.1. After quality filtering, 47,000 training examples remained. DPO used separate preference pairs, with one report preferred over the other for the same query. The source text is truncated before describing how those pairs were built.

Architecture change for speed

AstaBrief skips the snippet summarization and clustering stages used by the Claude-based Thinking mode. It also writes the report in one pass rather than section by section. Ai2 claims this did not reduce quality, and says the 8B model matched the report quality of the proprietary models it tracked. The excerpt provides no numerical quality scores, so that claim cannot yet be checked.

Caveats

Ai2 states that most training and evaluation was completed in 2025. The proprietary teacher and comparison models therefore reflect the frontier at that time, and Ai2 has not rerun the full evaluation against current frontier models. Ai2 says the results are best read as evidence about its training and system-design choices. The 3.5x figure compares Fast mode with Ai2's own Claude-based pipeline, not with a standardized external benchmark.

Ai2 is also releasing an example workflow for generating reports from a user's own PDFs. Open weights let institutions run the model on their own infrastructure, which Ai2 says matters when queries involve sensitive or unpublished work.

What this means

AstaBrief is a narrow, practical case for small specialized models: an 8B SFT+DPO model replaces a multi-stage pipeline built on frontier APIs, and the cost is mainly a data-curation effort. The most reusable parts are probably the real-query filtering and the single-pass report design, not the model itself. The speed gain comes partly from removing pipeline stages, so the 3.5x figure reflects system design as well as model size.

The evidence is thinner than the headline suggests. Quality parity rests on 2025-era comparisons, and the public numbers cover latency, not grounding or citation accuracy. Teams considering AstaBrief should test citation fidelity on their own corpora before relying on it. Local deployment is the clearest advantage for labs handling unpublished research.

Related Articles

model release

inclusionAI releases Ling 3.1 Flash: 560B MoE, 25B active, 262K context, free on OpenRouter

inclusionAI has released Ling 3.1 Flash, a hybrid reasoning mixture-of-experts model with 560B total and 25B active parameters and a 262K-token context window. It is listed as free on OpenRouter through NovitaAI. No benchmark scores have been published on the listing.

model release

Microsoft's MAI-Transcribe-2-Streaming returns first results in ~100 ms across 60 languages

Microsoft AI released MAI-Transcribe-2-Streaming, a real-time transcription model covering 60 languages with first partial results in just over 100 milliseconds. It also launched two text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, aimed at voice agents.

model release

Cloudflare releases Clef, a 27B Apache-2.0 model that outputs decision probabilities instead of text

Cloudflare published Clef on Hugging Face: a 27B multimodal model that takes a state and a schema of typed questions and returns a probability for every allowed option in a single forward pass. It is post-trained from Qwen3.8-27B and released under Apache-2.0. Benchmark results are from Cloudflare's internal Decision Index 0.2.1 run.

model release

Google unveils Gemini 4 Argon at $2/$10 per 1M tokens, but access is limited to select users

Google unveiled Gemini 4 Argon on Wednesday with introductory pricing of $2 per 1M input tokens and $10 per 1M output tokens, matching OpenAI's discounted GPT-6.1 Sol. Access is restricted to select cybersecurity defenders and enterprise cloud customers, and Google says published rates will double later.

Comments

Loading...