Ai2's AstaBrief 8B Writes Cited Research Reports 3.5x Faster in One Pass
Ai2 open-sources an 8B Qwen3-based model that generates cited scientific reports 3.5x faster than Claude-powered pipelines, with full training data.
- Ai2 released AstaBrief 8B, an open-weights Qwen3-based model for cited scientific report generation.
- Generates reports in one pass, averaging 51.1s vs 178.5s for Claude-powered Thinking mode (3.5x faster).
- Trained with SFT on 47K examples plus DPO on 6K preference pairs from real research queries.
- Citation-density filtering on training data was the single biggest driver of grounding quality.
- Scores 87.0 on SQA-CS2 test, competitive with Asta ScholarQA (86.2) and DR-Tulu-8B (88.8).
- Example repo ships for running local report generation over your own PDFs.
AstaBrief 8B writes cited research reports in one pass
Ai2, the Allen Institute for AI, has released AstaBrief 8B, an open-weight Qwen3-8B model fine-tuned to turn a research question and retrieved paper excerpts into a structured, cited report. The Apache 2.0 release includes the full training dataset, allowing developers to inspect, reproduce, and extend the post-training process.
Ai2 reports average end-to-end latency of 51.1 seconds for its Asta Fast pipeline, compared with 178.5 seconds for Thinking mode, a 3.5-fold speedup. The comparison covers complete pipelines and includes the time saved by removing intermediate summarization, clustering, and section-drafting stages.
One generation replaces four stages
Agentic literature-synthesis systems such as Ai2’s Claude-backed Asta service typically retrieve passages, summarize them, group them into sections, and draft each section separately. AstaBrief receives the research question and retrieved snippets together, then generates the complete report in one pass. Ai2 reports that this simplification preserved comparable quality on its main evaluation.
The model still depends on an external retrieval system to find papers and select relevant passages. Applications must prepare those excerpts, supply source identifiers for citations, and fit the evidence within the model’s context window. Ai2 provides a local PDF example that demonstrates the workflow.
Curated preferences shape the model
Ai2 trained AstaBrief with supervised fine-tuning, or SFT, followed by direct preference optimization, or DPO. SFT teaches the model from complete example reports. DPO then uses fixed pairs of preferred and rejected reports to reinforce desirable output patterns. Ai2 chose this lower-complexity process because judge-driven reinforcement learning can be expensive, unstable, and harder to debug.
- Query collection: Ai2 started with real prompts submitted to ScholarQA and earlier Asta systems. Filtering removed bot traffic, short prompts, non-English requests, non-scientific questions, and prompts containing personally identifiable information. About 90,000 research queries remained.
- SFT targets: The multi-stage ScholarQA pipeline generated reports using Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini, and GPT-4.1. Quality filtering produced about 47,000 training examples.
- DPO pairs: For a separate query subset, one report came from ScholarQA and another came from o3, o4-mini, DeepSeek-V3, or DeepSeek-R1. GPT-4.1 and DeepSeek-R1 judged each pair. Ai2 reports 95% agreement between the automated judgments and human preferences. Keeping only pairs on which both judges agreed left about 6,000 preference examples.
One filter improves the citations
Early SFT checkpoints produced fluent reports but sometimes invented citations or omitted them. Ai2 evaluated four filters for the synthetic reports: output-to-input token ratio, citation relevance, citation density, and citation diversity. Citation density measures how consistently a report attaches sources to claims that require support.
Removing examples with low citation density produced the largest improvement. Stricter thresholds, combinations of multiple filters, and learning-rate sweeps added little. In Ai2’s experiments, selecting targets with strong citation coverage improved grounding more reliably than additional filtering complexity.
DPO raises citation recall
ScholarQA-CS2 contains user-written computer science research questions. Its reported metrics separately measure answer quality, whether citations support their associated claims, and whether claims that need evidence receive citations. Higher scores are better.
| Model | Average | Answer precision | Citation precision | Citation recall |
|---|---|---|---|---|
| Qwen3-8B | 77.3 | 90.6 | 76.2 | 64.6 |
| AstaBrief-8B-SFT | 83.7 | 90.4 | 87.7 | 71.3 |
| AstaBrief-8B | 87.0 | 89.0 | 90.5 | 78.2 |
Relative to the SFT checkpoint, DPO raises the average score from 83.7 to 87.0. Citation precision gains 2.8 points and citation recall gains 6.9 points, accompanied by a 1.4-point reduction in answer precision.
On the same test set, AstaBrief scores 87.0, compared with 86.2 for Asta ScholarQA and 88.8 for DR-Tulu-8B. It also wins 72% of head-to-head comparisons with Asta ScholarQA. Performance is weaker on DeepScholarBench: AstaBrief scores 53.5 against Asta ScholarQA’s 60.25, and Ai2 reports that DR-Tulu also scores higher there.
Fast synthesis has firm boundaries
- Best fit: AstaBrief suits applications that already retrieve relevant scientific passages and need fast, structured synthesis with dense citations.
- Outside its scope: The model does not perform open-web research, query decomposition, iterative search, or multi-turn tool use. Those functions require a surrounding application or agent.
- Human preference: DR-Tulu received the higher overall preference in a small study with three reviewers. Two of the three reviewers rated AstaBrief higher for citation accuracy.
- Fidelity risk: A citation can point to a relevant paper even when the generated claim overstates its findings. For example, the model may generalize a result beyond the study’s sample. AstaBrief was not explicitly trained to prevent this failure mode.
Open weights bring reports on premises
An 8B checkpoint can run on a single high-memory consumer GPU with suitable precision and context settings. Local deployment can keep unpublished papers, sensitive queries, and generated reports inside an institution’s network, removing the need to send research material to a hosted frontier-model provider.
Ai2’s release post also highlights the effect of post-training data composition. AstaBrief’s results came from examples that demonstrated grounded synthesis and consistent attribution, followed by preference training that strengthened those behaviors. For teams building domain-specific writing systems on open models, SFT, DPO, and citation-density filtering provide a reproducible baseline before adopting more complex reinforcement-learning pipelines.