Ai2 on Oct. 2 open-sourced AstaBrief 8B, the model behind Fast mode in Asta, its agentic platform for scientific work. Fast mode finishes cited reports about 3.5 times faster than Asta's Claude-powered Thinking mode. AstaBrief 8B is an open-weights language model that turns a research question and retrieved paper excerpts into a cited report. Ai2 released the weights under the Apache 2.0 license and also published the training data.
Speed is the main gain: 51.1 seconds per report on average, against 178.5 seconds for Thinking mode. The benchmark picture is closer and more mixed than that.
Key facts
- Model: 8 billion parameters, fine-tuned from Qwen3-8B. Weights on Hugging Face, Apache 2.0.
- Speed: Fast mode is about 3.5× faster than Thinking mode across the full Asta pipeline.
- Quality: within two points of the systems it was compared with on Ai2's main benchmark, and last of three on DeepScholarBench.
- Users: Fast mode's positive-feedback rate is 84.2%, against 85.2% for Thinking mode. Ai2 calls the feedback too sparse for strong conclusions.
- Use: research and education, under Ai2's Responsible Use Guidelines.
Where the speed comes from

Thinking mode builds a report in stages. It summarizes and clusters the retrieved snippets, then writes section by section. AstaBrief skips those stages and writes the whole report in one pass from the question and the snippets. Ai2 says this did not cost it performance on its own tests.
Ai2's post also mentions "nearly an order-of-magnitude reduction in report generation time" compared with the proprietary models it tracked. Ai2 does not say what that figure measures; its end-to-end number for Asta is 3.5×.
How it scores against other report systems

Ai2's model card compares three systems: AstaBrief-8B; DR-Tulu-8B, an Ai2 research model whose work showed that reinforcement learning can improve long reports; and Asta ScholarQA, the multi-step pipeline behind Asta's report feature, according to Ai2. The win rate below is measured against that pipeline.
| Measure | AstaBrief-8B | DR-Tulu-8B | Asta ScholarQA |
|---|---|---|---|
| SQA-CS2 test score | 87.0 | 88.8 | 86.2 |
| DeepScholarBench score | 53.50 | 56.26 | 60.25 |
| LLM-judged win rate vs Asta ScholarQA (SQA-CS2 test) | 72% | 54% | reference |
On the SQA-CS2 test split, a set of computer science research questions written by real users, the three scores are 2.6 points apart, by ZIPMEX's calculation. On DeepScholarBench, a 63-query benchmark built from recent arXiv papers, AstaBrief scores lowest. Ai2 warns that the two benchmarks use different metrics and should not be compared with each other. On the SQA-CS2 dev split, AstaBrief scores 86.3, last of the three, and its LLM-judged win rate against Asta ScholarQA is 55%, against 36% for DR-Tulu.
In a 14-question human study, three researchers ranked reports from the three systems. DR-Tulu won on overall preference. Two of the three researchers preferred AstaBrief on citation accuracy.
How Ai2 trained it

Ai2 chose a simple recipe: supervised fine-tuning (SFT), then direct preference optimization (DPO), a method that trains a model on pairs of answers where one is marked better. It skipped reinforcement learning, which the post says "can be unstable and expensive."
- Queries: real user queries from Ai2's ScholarQA systems, filtered for quality, relevance and privacy, left 90,000 research questions.
- SFT data: reports generated by Asta's pipeline with Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini and GPT-4.1. After filtering, 47,000 examples remained.
- DPO data: a separate set of queries, two competing reports each. GPT-4.1 and DeepSeek-R1 judged every pair, and Ai2 kept only pairs where both judges agreed. About 6,000 pairs remained.
Ai2 calls its filtering result one of the clearest lessons of the project. Ai2 tested four signals for weak training reports. Citation density is the share of statements that had at least one citation. Dropping reports where it was low gave the strongest gains. Combining filters or filtering harder did not add much.
Across both training stages combined, SFT and DPO, citation precision on the model card's test set rose from 76.2 for base Qwen3-8B to 90.5 for the final model. Citation recall rose from 64.6 to 78.2. Answer precision fell from 90.6 to 89.0.
What early users did
Ai2 reports that 374 Asta users have tried Fast mode. Of those, 29.1% used it on two or more days, and users generated 3.67 report threads with it on average. Nearly a quarter, 23%, kept using it and never switched back to Thinking mode. Another 18% switched between modes and used Fast mode for about 40% of their threads.
Fast mode drew positive feedback at a rate of 84.2%, against 85.2% for Thinking mode. Ai2 says the feedback is "too sparse to draw strong conclusions."
How to run it yourself
- Try it: Fast mode in Asta, under Generate a report.
- Download: AstaBrief 8B weights and the training data collection.
- Build on it: Ai2's example workflow generates reports from your own PDFs. The model does no retrieval itself: you supply the paper excerpts, and the model card warns that a prompt format other than Ai2's recommended one "may lead to degraded or inconsistent behavior." Its example code loads the model with Transformers and vLLM. By ZIPMEX's estimate, 8 billion parameters in BF16 precision take about 16 GB for the weights alone, before memory for long inputs.
Institutions can run it on their own hardware, behind their own firewall, Ai2 says. That matters, Ai2 argues, because research questions can reveal sensitive or unpublished work.
Analysis: what the numbers show
ZIPMEX's reading: AstaBrief is mainly a speed and control release. It is close to the Claude-powered pipeline on SQA-CS2, Ai2's main development benchmark, and trails both rivals on DeepScholarBench. The 72% win rate also needs caution. Ai2 itself notes that AstaBrief, unlike DR-Tulu, was optimized for this kind of pairwise ranking during DPO.
Ai2 expects its data and filtering lessons to generalize beyond this model. In Ai2's tests, one simple filter, whether a report cites its claims, beat the more complicated combinations it tried.
Limits Ai2 states
Most training and evaluation was done in 2025, and Ai2 has not rerun the comparison against today's frontier models. Its metrics check whether a citation supports a claim, not whether the claim stretches the study's scope, for example by turning a finding about one sample into a general rule. Ai2 lists that as a gap for future evaluations.
Sources: Ai2, "Open-sourcing AstaBrief, the fast report-generation model in Asta" (Oct. 2, 2026) · AstaBrief 8B model card · AstaBrief collection · ai2-scholarqa-lib · Ai2 on DR Tulu.
ZIPMEX promotes trading on trade.zipmex.com. Promotions are labelled and kept separate from news coverage, which is selected and checked without regard to them.