Open-sourcing AstaBrief, the fast report-generation model in Asta
Ai2 has open-sourced AstaBrief 8B, the report-generation model behind the Fast mode in its Asta scientific research platform. It takes a research question plus retrieved literature excerpts as input and writes a full cited report in a single pass — no section-by-section generation. Weights and the training data are both being released.
It's built on Qwen3-8B, with most of the engineering effort going into post-training data, evaluation, and the report-generation scaffolding rather than the base model.
Performance and setup
- Across the full Asta pipeline, Fast mode averages 51.1 seconds per report, compared with 178.5 seconds for the Claude-powered Thinking mode — about 3.5× faster, described as nearly an order-of-magnitude reduction in generation time vs the proprietary models they tracked.
- AstaBrief is live today in Asta's Generate a report feature as Fast mode, running alongside Thinking mode.
- Ai2 also ships an example workflow that researchers can adapt to generate reports from their own PDFs, aimed at local report generation.
- Open weights let institutions run it on their own infrastructure, which matters when research questions involve sensitive or unpublished work.
Training recipe: SFT + DPO, not RL
Ai2 explicitly chose a simpler recipe: supervised fine-tuning plus direct preference optimization. They point to their earlier DR Tulu work showing RL can improve long-form report generation for open-weights models, but say RL-based training is unstable and expensive. They wanted a setup that's cheaper to run, easier to debug, and easier to iterate on — which pushed the emphasis onto training data quality.
Building AstaBrief required tens of thousands of real research queries, citation-focused filtering, preference data, and a redesigned pipeline that writes the report in one pass. The model's stated goals are answer quality, relevance, structure, and citation grounding — keeping answers tied to what the evidence actually supports instead of silently broadening a study's conclusions.
Caveats
Most training and evaluation was completed in 2025, so the proprietary models used to generate training data and as comparison points reflect the frontier at that time. Ai2 has not rerun the full evaluation against current frontier models, and frames the results as evidence about the specific training and system design choices they tested. The work also ties into NSF OMAI, the U.S. initiative led by Ai2 to build open AI infrastructure and models for scientific discovery.
If you run local models and want cited long-form synthesis over your own PDFs, the released example workflow is the place to start.
📖 Read the full source: HN AI Agents
👀 See Also

OpenRouter's Healer Alpha stealth model appears to be unreleased Qwen 3.5-Omni variant
OpenRouter has deployed a free anonymous omni-modal model called Healer Alpha with 262,144 context window and multimodal capabilities. Forensic analysis suggests it's an unreleased Qwen 3.5-Omni variant from Alibaba.
YouTube Terminates 20 'Ghost Creator' Channels for AI Spam
YouTube removed 20 channels linked to a network that used AI-generated scripts and human actors to pose as political commentators, violating spam policies.

Claude Code 2.1.76 adds MCP elicitation, worktree improvements, and fixes for context limits
Claude Code version 2.1.76 introduces MCP elicitation support for structured input during tasks, adds worktree.sparsePaths for large monorepos, and fixes 'Context limit reached' errors on 1M-context sessions. Version 2.1.75 made 1M context windows default for Opus 4.6 on Max, Team, and Enterprise plans.

Hy3 LLM Tops OpenRouter Rankings: Cheapest Model or Something Else?
Hy3 preview, a Tencent open-source LLM, surged to the top of OpenRouter's model rankings by token usage, surpassing Claude and DeepSeek V4 Flash. Priced at $0.066/1M input tokens, it's the cheapest major model, but benchmarks show quality far below leaders.