Back Home

最新模型

Ai2 Releases AstaBrief 8B, Simplifying Citation-Backed Scientific Reporting in a Single Generation

Built on Qwen3-8B, AstaBrief releases model weights and training data, combining research questions and retrieved passages into reports. Ai2 reports that the end-to-end workflow is about 3.5× faster, though the comparison mainly uses models and evaluations from 2025.

投稿者が撮影 · Public domain · Image source
zh-Hant

On October 2, Ai2 released AstaBrief 8B, enabling research institutions to deploy their own model for generating scientific reports with citations. Built on Qwen3-8B, it is already used in Asta’s Fast mode. Ai2 also released the training data and an example workflow for producing PDF reports locally. The model weights are licensed under Apache 2.0. Official announcement, model card

The main change is in the generation workflow: the model takes a research question and retrieved literature passages, then writes a complete report in one pass. This replaces the existing Claude workflow of summarizing passages, clustering them, and drafting the report section by section. Ai2 reports that the average time for the full Asta workflow falls from 178.5 seconds in Thinking mode to 51.1 seconds in Fast mode, about 3.5× faster. This comparison covers the end-to-end workflow and should not be taken as a direct measure of improved model inference speed. Workflow and testing details

Training used supervised fine-tuning and offline direct preference optimization. The team created about 47,000 fine-tuning examples from real research queries and found that removing synthetic reports with low citation density contributed most to quality improvements. For the preference stage, they retained report pairs on which two judge models agreed. The model card specifies that DPO used eight H100 GPUs, BF16 precision, and a maximum training sequence length of 16,000 tokens, providing reproducible settings. Data construction details, training configuration

On the computer science research question test set reported in the model card, citation precision rose from 76.2 for the base model to 90.5, while citation recall increased from 64.6 to 78.2. Answer precision, however, fell from 90.6 to 89, showing that the metrics did not all improve in tandem. For deployment, Ai2 recommends using the training prompt format; replacing it with an interactive format may lead to inconsistent behavior. Evaluation and usage

For engineering teams, AstaBrief offers a concrete way to use a small model for literature synthesis, but retrieval quality still needs to be validated separately. Ai2 explicitly notes that most training and evaluation were completed in 2025 and that it has not rerun a full comparison against current frontier models. Before adoption, teams should test Traditional Chinese reports, cross-disciplinary data, and whether citations preserve the scope of the original studies; citations do not prevent evidence from being overgeneralized. Official limitations

Sources

  1. Open-sourcing AstaBrief, the fast report-generation model in Asta
  2. allenai/AstaBrief_8B 模型卡