Pimp My IDE / garage dispatch
Back to garage
October 2, 2026 | research agents / citation scope

A citation is not a torque wrench.

AstaBrief shows how a small open model can draft cited reports quickly. It also exposes the harder job: keeping every claim inside the evidence that supports it.

Check retrieval, citation support, claim scope, and independent review as separate parts. Fluent prose cannot combine them into proof.
Claim apertureScope mismatch
Evidence widthClaim exceeds source

Fast reports move the bottleneck.

Ai2 released AstaBrief-8B to turn a research question and retrieved literature excerpts into a cited report. The project post reports a 51.1-second average for Fast mode and 178.5 seconds for its Claude-backed Thinking mode across the full Asta pipeline. Ai2 also says most training and evaluation work was done in 2025 and was not rerun against current frontier models.[1]

That speed can make a preliminary report cheap enough to revise. It does not make retrieval complete or claims accurate. Faster drafting puts more weight on the inspection step.

A citation can point to the right paper while the sentence still claims too much.

Support and scope are different checks.

The release tracks ingredient recall, answer precision, citation precision, and citation recall as separate measures. That split matters. A report may cover the requested material but attach a weak source. It may cite a relevant source but broaden a result from one sample into a general rule.[1]

Ai2 names several scope failures directly: changing a past finding into a universal present-tense claim, turning a descriptive result into a recommendation, or removing the population and setting that bounded the result. Citation matching alone does not catch these changes.

The model is one part of the machine.

The model card says AstaBrief-8B expects a research question plus retrieved excerpts in a specific prompt format. The card reports results on a 100-question computer science test set and a second 63-query benchmark. These rows measure named datasets and judges. They do not establish quality for another field, retrieval collection, or question type.[2]

The current inference example on the AstaBrief-8B card names allenai/AstaBrief_8B_SFT in its code. The page itself documents the later DPO checkpoint. Pin the exact model ID, revision, prompt, retriever, corpus, and parser before comparing results.

Open weights do not bundle the evidence.

The Hugging Face API lists the Apache-2.0 model artifact at revision 3a4e553a40c87d0f426c255267a367c7f721a153. It includes four safetensor shards, configuration files, tokenizer files, and a model card. The listed storage is about 32.8 GB because the repository also carries PyTorch binary shards.[2]

Ai2 points to a ScholarQA Lite path for local report generation. We verified that the public repository path contains prompt construction, response parsing, and the one-pass report runner. We did not download the weights or run a report, so this is an artifact inspection, not an inference result.[3]

Keep the retrieval receipt.

The earlier OpenScholar paper describes a larger retrieval system over 45 million open-access papers and evaluates long-form literature synthesis across several fields. Its claims belong to that system and benchmark, not to every report made with AstaBrief.[4]

  1. Save the question and every constraint.
  2. Save the retrieved excerpt, paper identifier, and locator for each claim.
  3. Compare the sentence with the source population, method, time, and uncertainty.
  4. Mark missing coverage instead of filling it with fluent guesses.
  5. Run domain review before a report guides research or practice.

The useful output is not a smooth report by itself. It is a report paired with enough evidence to reject a sentence.

Interactive makeover / report grounding dyno

Put the claim on the rack.

Traditional purpose replaced: a citation checkmark. Better version: select the claim type, route four distinct inspections, and copy a review packet that keeps required evidence blank.

Claim setup

The controls define what the packet should request. They do not inspect a report or verify a citation.

Claim type
Review packet sections
Generated review packet

Separate the evidence paths

The route lights discrete sections. A downstream-only selection stays downstream. It does not close earlier gaps.

Retrieved evidence
Citation support
Scope comparison
Independent review
Observation packet incomplete.No review-packet section is selected.
All four sections selected means the packet structure is ready. It does not prove retrieval coverage, citation support, preserved scope, or reviewer approval.

Sources read

Source log and evidence boundary
  1. Ai2, "Open-sourcing AstaBrief, the fast report-generation model in Asta", published and read October 2, 2026. This is the first-party source for the release, training account, reported pipeline timings, evaluation setup, and stated limits on claim scope.
  2. AstaBrief-8B model card and Hugging Face model API, read October 2, 2026. These pages supply the model ID, revision, license, files, expected prompt shape, code sample, benchmark rows, and intended-use notes.
  3. ScholarQA Lite implementation path, read through the GitHub page and API on October 2, 2026. We confirmed the public path and its four Python files. We did not execute the workflow.
  4. OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs, submitted November 21, 2024 and read October 2, 2026. This is the earlier system paper for the retrieval corpus, ScholarQABench, and the broader retrieval-plus-synthesis design.
  5. Hacker News item 49938783, resolved through the official Hacker News API on October 2, 2026. It was the discovery signal and supports none of the technical claims.

Evidence boundary. The performance and quality statements are Ai2 results on named pipelines, benchmarks, models, and judges. We verified the public model metadata, file inventory, revision, and example workflow path. We did not download the 32.8 GB repository, run inference, reproduce a benchmark, inspect private training queries, or judge report quality.