Ai2 open-sources AstaBrief 8B, a faster report model for Asta

Ai2 has announced AstaBrief 8B, a small model trained to turn a research question and retrieved literature excerpts into a cited report. It is live in Asta's Generate a report feature as Fast mode, next to the existing Claude-powered Thinking mode. Ai2 says it is open-sourcing both the model and its training data so others can study, reproduce and build on the approach.
The motivation comes from how scientists use Asta, Ai2's agentic platform for scientific work. Users tend to bring substantial context and many constraints rather than keyword searches, and many return to generated reports later as working research artifacts. Ai2 wanted to test whether a small, open model trained specifically for scientific report generation could match the report quality of the proprietary models it was using, while cutting generation time and serving costs.
On speed, Ai2 reports nearly an order-of-magnitude reduction in report generation time compared with the proprietary models it tracked. Across the full Asta pipeline, Fast mode averages 51.1 seconds per report versus 178.5 seconds for Thinking mode, about 3.5 times faster. Much of the gain comes from a redesigned pipeline: AstaBrief writes the full report in one pass from the user query and retrieved snippets, skipping the snippet summarization and clustering stages that Thinking mode uses and not writing section by section. Ai2 says it found this possible without sacrificing performance.
AstaBrief starts from Qwen3-8B. Ai2 considered reinforcement learning, as in its earlier DR Tulu work, but chose supervised fine-tuning (SFT) plus direct preference optimization (DPO), calling RL unstable and expensive and preferring a setup that is cheaper and easier to debug. That put the weight on data quality.
The data began with real queries from the system behind Ai2's paper on synthesizing scientific literature with retrieval-augmented LMs and ScholarQA, the framework that now underpins Generate a report. Ai2 stripped out beta-tester and bot traffic, dropped queries that were too short, and used an LLM pass to catch non-English queries, non-scientific requests and prompts with personal information. That left 90K research-focused queries. For SFT, Ai2 generated full-report targets with the multi-step ScholarQA pipeline, drawing on Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini and GPT-4.1. After quality filtering, 47K usable examples remained.
For DPO, Ai2 used a separate subset of queries. One report per query came from the existing ScholarQA pipeline, typically backed by Claude 3.5 or 3.7 Sonnet; the competing report came from o3, o4-mini, DeepSeek-V3 or DeepSeek-R1 working from the same retrieved excerpts. Two judges, GPT-4.1 and DeepSeek-R1, picked a winner for each pair, and only pairs where both agreed were kept. Ai2 says the judges were aligned with human preferences (95% agreement). The final DPO set came to about 6K examples.
The main evaluation target was SQABench-CS2, 200 user-written computer science research questions, tracked with four metrics: rubric score (how much necessary content is covered), answer precision (whether each paragraph is relevant), citation precision (whether each citation supports its claim) and citation recall (whether the report's claims are fully supported by the citations). Secondary evaluations for the final model were DeepScholarBench, a 63-query benchmark built from recent ArXiv papers, plus two pairwise comparisons against the Claude-powered pipeline: an LLM-judged one on SQABench-CS2 and a small human study.
First SFT runs improved overall content quality but still lagged the Claude-powered pipeline on answer precision and citation quality. That led Ai2 to test four statistics-based filters for weaker synthetic training examples, starting with the output-to-input token ratio (very high ratios often meant noisy answers written from too little evidence), average citation relevance (low averages suggested reliance on lower-ranked evidence) and citation density. Ai2 also flags a gap in its own metrics: they focused on relevance, coverage and citation grounding, not on whether a report preserves the scope and strength of its sources' claims, such as turning a finding about one sample into a claim about a whole population.
Alongside the weights, Ai2 is releasing an example workflow researchers can adapt to build reports from their own PDFs, as a starting point for local report generation. It adds that open weights let institutions run AstaBrief on their own infrastructure, which it calls necessary when research questions reveal sensitive or unpublished work.
Key facts
- AstaBrief 8B, built from Qwen3-8B, turns a research question plus retrieved literature excerpts into a cited report; Ai2 is open-sourcing the model and its training data.
- Inside Asta, Fast mode (AstaBrief) averages 51.1 seconds per report against 178.5 seconds for the Claude-powered Thinking mode, about 3.5 times faster.
- Ai2 claims nearly an order-of-magnitude cut in report time versus the proprietary models it tracked, by writing the whole report in one pass, which it says cost no performance.
- Training used SFT on 47K examples (from 90K filtered real queries) and DPO on about 6K pairs where both judges, GPT-4.1 and DeepSeek-R1, agreed.
- Evaluation centred on SQABench-CS2 (200 questions); Ai2 notes most work was done in 2025 and has not rerun the full evaluation against today's frontier models.
Why it matters
Report generation for scientists is a long-form task where answers must stay tied to evidence, and Ai2 is arguing that a small open model can do it faster than proprietary systems. The 8B model is already in production in Asta as Fast mode, so this is a shipped feature, not only a research artifact. Ai2 also frames AstaBrief as a test case for building open language models adapted to the specific demands of science, a theme tied to its NSF OMAI initiative.
Who it affects
Researchers who use Asta get a quicker preliminary report they can then iterate on in later turns. Institutions that handle sensitive or unpublished work can, according to Ai2, run the open weights on their own infrastructure. Teams building scientific report writers can study the training data and the one-pass design, and researchers with their own PDFs get an example workflow to adapt for local report generation.
How to use it
In Asta, choose Fast mode in the Generate a report feature; Thinking mode, powered by Claude, remains the alternative. Ai2 positions Fast mode for quick first drafts that users refine in follow-up turns. To run it yourself, Ai2 is releasing the model weights, the training data and an example workflow for generating reports from your own PDFs.
How solid is it
The speed figures are specific: 51.1 versus 178.5 seconds per report across the full Asta pipeline, about 3.5 times (178.5 divided by 51.1 is 3.49). The post says 'averages' but does not state the number of reports measured. The 'nearly an order of magnitude' claim compares against proprietary models Ai2 tracked, a different comparison from the 3.5 times figure. The visible text gives no benchmark scores for the final model, and it does not include the results of the pairwise LLM-judged comparison or the human study, so the claim of matching Claude-powered quality cannot be checked here. The one claim Ai2 makes directly is that the one-pass design did not sacrifice performance. This is Ai2's own account of its own product.
Risks and caveats
Ai2 itself lists limits. Most training and evaluation was done in 2025, so the proprietary comparison models reflect the frontier of that time, and the full evaluation has not been rerun against current frontier models; the results are best read as evidence about the design choices tested. Early SFT runs lagged the Claude-powered pipeline on answer precision and citation quality, which is why Ai2 added data filtering. Its development metrics also focused on relevance, coverage and citation grounding, not on whether claims keep the scope and strength of their sources, so a cited sentence could still overstate what a study showed. No license for the weights or data is named in the visible text.
“Interestingly, we found it was possible to do so without sacrificing performance.”
— Ai2, AstaBrief announcement post