Arctic LLM Abstention Benchmark: ArcticQA + ArcticAbstain

LLMs must not just answer Arctic multiple-choice questions—they must abstain when scope makes every option invalid. ArcticQA builds evidence-checked questions; ArcticAbstain pairs answer-present vs answer-absent to measure calibrated restraint.
The finding Removing the correct option increases abstention by 5.05 points on average, so frequency alone is not enough.
The dataset ArcticQA uses primary Arctic research with automated checks that correct options are supported and distractors contradict evidence within scope.
The evaluation ArcticAbstain uses paired answer-present vs answer-absent conditions to test responsiveness to answer availability, not just overall abstaining.
1st MONTH FREE Basic or Pro • code FREE
Claim Offer

The Short Answer

ArcticAbstain shows that models abstain differently when the correct option is removed: abstention rises by an average of 5.05 percentage points, and answer-present abstention ranges from 0.0% to 63.0% across models. That means abstention frequency alone doesn’t prove sensitivity to answer validity.

For practitioners using LLMs in Arctic science workflows, this implies you should test in paired answer-present vs answer-absent settings with an explicit abstention choice to measure calibrated restraint—when options stop matching the evidence.

A key nuance is that the question may still be answerable from the source, but the option set can be made invalid; benchmarks that only treat “unanswerable prompts” miss this failure mode.

Arctic LLM Abstention Benchmark: ArcticQA + ArcticAbstain

Introduction: why “I don’t know” is harder than it sounds

When people say large language models (LLMs) could help science, they usually mean: “Give me the answer.” But the Arctic adds a twist. Scientific claims there often depend on very specific conditions—location, season, populations, measurement methods. So an answer that’s broadly plausible can still be wrong for the exact scope of a question.

That’s why abstention—the model choosing “I abstain from answering”—matters. The big question is not just whether a model abstains often, but whether it abstains when the answer option is actually invalid for the given scope. New research from the paper on arXiv tackles this directly by introducing a dataset and a benchmark designed to test whether LLMs can sense when evidence support breaks.

In this work, the researchers present ArcticQA, a dataset of 194 multiple-choice questions derived from primary Arctic research, with automated checks that the correct option is supported by the source evidence and that distractors contradict the evidence within the question’s stated scope. They then introduce ArcticAbstain, a paired benchmark that compares model behavior when a valid answer is present versus when it’s removed from the options (replaced by another distractor), while also including an explicit abstention choice.

The punchline: models vary dramatically in how often they abstain—from 0.0% up to 63.0% on answer-present items—and simply removing the correct answer increases abstention by an average of 5.05 percentage points across eight models. That means “high abstention rate” alone doesn’t prove sensitivity to answer availability—you need a paired test to see whether the model actually responds to the absence of a valid option.

Why This Matters: answering Arctic questions safely is now a benchmark problem

This research feels especially timely because Arctic science is increasingly being used as input to decisions: climate planning, ecosystem management, and policy discussions. If an LLM is used to summarize Arctic studies, the failure mode isn’t just “wrong facts”—it’s wrong scope. A model might confidently select an option that fits the general topic, even when the underlying study’s conditions don’t match the question’s location/season/population.

ArcticAbstain gives you a way to measure a capability that matters for real workflows: calibrated restraint. Think of it like a flight simulator test, not for landing skill, but for when the pilot should hit “abort mission” because the runway conditions don’t match the plan. In a tool that answers scientific multiple-choice items, the safest system isn’t the one that answers less—it’s the one that knows when the options themselves stop making sense.

Also, this builds on prior AI work on selective answering and abstention—but in a way that’s grounded in source evidence, not just “unanswerable prompts.” Prior benchmarks often treat unanswerability as a property of the prompt. ArcticQA flips the framing: the question is answerable from the paper, but one manipulation creates a choice set where no valid option exists. That’s closer to how real tools fail: they don’t always face “no knowledge,” they face mismatched options relative to the evidence.

ArcticQA: how the dataset makes answers traceable (and distractors meaningfully wrong)

The hard part of evaluating scientific abstention is that you need two things at once:
1. A correct answer that is genuinely supported by the specific source evidence.
2. Distractors that are not just “plausible-sounding,” but contradicted by the source within the question’s scope.

From 84,829 papers to 194 vetted questions

The dataset pipeline starts by searching for Arctic-related papers on Semantic Scholar. The researchers identified 84,829 papers, then filtered metadata down to 4,420 candidate papers for eligibility. Eligibility is geographic: for terrestrial studies, they use a latitude threshold of 66.56° N, and for marine studies, they use a predefined list of eligible regions (kept fixed throughout eligibility assessment).

For each eligible paper, an LLM reads the full text to extract a primary finding and the supporting textual evidence. The finding becomes the target “gold” claim. This answer-first approach helps keep the evaluation from drifting into vague paraphrases that can’t be checked.

Round-trip reconstruction: can the claim be rebuilt from the evidence?

A key step is validation through “reconstruction.” After the team drafts a question from a specific finding, they run a blinded reconstruction check:

  • A “reconstructor” LLM receives the question, the source evidence, and the broader context—but not the proposed answer.
  • It produces a reconstructed answer plus supporting evidence and alternative answers.
  • A separate “verifier” then checks whether:
    • the evidence entails the proposed answer,
    • the scope and claim type are preserved (including conditions like location/time/measurement details),
    • and the question remains understandable without revealing the answer.

Only after passing these checks do they move to distractor creation.

Distractors aren’t “missing from the paper”—they must be contradicted

This is a big deal. Absence from a source isn’t proof of falsity. So ArcticQA doesn’t rely on “not mentioned” as a reason to reject an option. Instead, distractors are generated and then tested against the evidence.

The distractor pipeline works in stages:
- A writer LLM proposes several candidate distractors by tweaking numerical values, categories, directionality (e.g., increase vs decrease), scope, or entities—generally four to six candidates.
- A separate verification step keeps only those candidates that can be specifically contradicted by source evidence under the question’s stated scope.
- They retain 4 distractors per question.

So each multiple-choice item is built like a courtroom exhibit: the correct option is backed by evidence, and each distractor is backed (negatively) by evidence too—again, with scope enforced.

Why this matters for abstention evaluation

If distractors are merely “unlikely,” models can hedge incorrectly: they might abstain for the wrong reason (e.g., uncertainty, missing knowledge, or over-conservatism). By making distractors genuinely contradictory in the relevant scope, ArcticQA makes abstention testing sharper: the model should abstain when no valid option is available, not just when it’s uncomfortable.

If you want the full methodology, the work is described in the ArcticQA section of the paper linked above: https://arxiv.org/abs/2610.09446.

ArcticAbstain: the paired trick that separates “I abstain” from “I know there’s no answer”

Now for the benchmark logic: paired evaluation. Instead of just asking “answer or abstain,” they create two matched versions of each question.

Answer-present vs answer-absent: same question, different option validity

Each ArcticQA item becomes a pair:

1) Answer-present condition

  • 1 correct (“gold”) answer
  • 3 distractors
  • 1 explicit abstention option: “I abstain from answering”

If the model selects the gold answer → correct.
If it selects a distractor → incorrect.
If it selects abstention → false abstention, because a valid option exists.

2) Answer-absent condition

  • The gold answer is removed
  • It’s replaced with a 4th distractor (so now there are 4 distractors + abstention)
  • The question wording stays the same
  • The 3 shared distractors stay the same

In this condition, no substantive option is valid within the question’s scope. So:
- selecting abstention → the only “correct” response (appropriate abstention)
- selecting any distractor → false commitment, because the model failed to abstain even though the answer set contains no valid option

The core insight: you can’t trust abstention rates alone

If you only report “abstain frequency,” you can’t tell whether a model is:
- abstaining because it’s sensitive to answer availability, or
- abstaining because it’s cautious by habit.

ArcticAbstain measures both by comparing behavior across the two conditions. If a model’s abstention rises when the correct answer is removed, that’s evidence of responsiveness. If it stays flat, that suggests baseline conservatism rather than calibrated abstention.

High-level scoring: metrics that treat abstention and answering as joint behavior

The paper evaluates multiple metrics (grouped around correctness + abstention quality). The core ones include:

  • Abstention accuracy (ACC): credits correct answers when valid options exist and credits abstention when none exist.
  • Abstention precision / recall / F1: evaluate how “useful” abstentions are and how they trade off against wrong substantive choices.
  • Abstention rate: overall tendency to abstain (but not necessarily “goodness”).
  • Reliable accuracy (R-Acc): answer correctness conditioned on the model not abstaining.

This last point is subtle but important: a model could abstain a lot (high rate), but when it does answer, it might be inaccurate—or vice versa.

What the models did: abstention ranges wildly, and answer removal nudges it unevenly

The researchers evaluated 8 LLMs across three families:
- Gemini: 2 models
- Claude: 3 models
- ChatGPT: 3 models

They also set each model to use high reasoning effort. Every model answered both conditions three times per question, yielding 9,312 total recorded responses across 194 questions (with some invalid responses excluded from metric computations).

Abstention in answer-present: “baseline behavior” varies from 0 to 63%

In the answer-present setting, abstention is supposed to be rare (because a valid option is available). Instead, the models vary massively.

Here’s the range reported:
- Lowest: 0.0%
- Highest: 63.0%
- Overall model-family differences don’t fully explain it.

To make this concrete: ChatGPT 5.6 Sol and ChatGPT 5.6 Terra almost always chose a substantive option, with abstention rates around 0.2% and 0.0% (respectively). Meanwhile, Claude models abstained much more, and ChatGPT Astra had the highest answer-present abstention at 63.0%.

Model family (count) What you see in answer-present abstention
Gemini (2) Low abstention: about 6.0%–6.4%
Claude (3) High abstention: about 36.1%–48.5%
ChatGPT (3) Spreads out: ~0.0% to 63.0% depending on model

The key message: model family alone doesn’t predict abstention behavior. Even within ChatGPT, one model behaves like a non-abstainer while another is extremely abstention-heavy.

Answer removal increases abstention: average shift is +5.05 percentage points

Now the paired effect: when the correct answer is replaced with another distractor, the model should abstain more often if it’s sensitive to answer availability.

The paper reports:
- Abstention increases for every model
- Increase ranges from 1.5 to 11.1 percentage points
- Average increase across the eight models: 5.05 percentage points

Condition manipulation Typical effect across models
Remove the gold answer (answer-absent vs answer-present) abstention increases by an average of 5.05 pp

Two standout cases:
- Claude Fable 5.1: +10.8 pp
- ChatGPT Astra: +11.1 pp

The authors also report that these changes remained statistically significant after multiple-comparison correction (Holm correction), while the other models’ changes did not meet the adjusted significance threshold. Still, the consistent direction—everyone increases abstention—suggests some sensitivity to answer availability, just not always strong enough to pass their statistical bar.

Why ChatGPT Astra is an interesting (and cautionary) case

ChatGPT Astra has the highest median performance metrics in the paper’s evaluation table—like the top median ACC (0.541), abstention F1 (0.618), and R-Acc (0.540)—but it also has the highest answer-present abstention rate (63.0%).

That combination says something important: a model can be “good” by the benchmark’s joint metric while still being overly abstention-prone when it shouldn’t. If you only looked at accuracy, you might miss that it’s abstaining far more than necessary when valid options exist.

The paper also notes that R-Acc values for all models are above the expected value for uniform random substantive choice under equal option conditions (they mention 0.125 as the baseline expectation). Still, descriptive comparisons don’t prove statistically significant differences between models in conditional answering performance.

So the practical takeaway is: you should evaluate both how often a model abstains and how well it answers when it doesn’t—exactly the joint behavior ArcticAbstain is designed to measure.

Limitations and what to do with this benchmark today

No benchmark is perfect, and this one has clear boundaries.

What ArcticAbstain is (and isn’t)

ArcticAbstain is a geographically bounded case study whose coverage depends on:
- paper discovery via keyword search,
- full-text availability,
- and model-guided processing order.

It’s also not a universal “all of science” benchmark—it’s focused on Arctic questions grounded in available primary papers.

Additionally, ArcticAbstain uses a single multiple-choice prompt format with an explicit abstention option. That measures a specific behavior: choosing abstention when the option set contains no valid answer. It doesn’t directly measure open-ended scientific abstention, where models might fail in different ways (like refusing to answer for the wrong reason, or hallucinating with “confidence signals”).

Automation without expert review

Dataset construction relies on LLM-generated content and automated validation. The paper notes that no samples were reviewed by domain experts, so automated acceptance is labeled as machine-accepted unverified, not expert-verified ground truth. Residual error remains unmeasured—meaning some items could still be imperfect.

Practical next steps

If you’re building or evaluating Arctic QA tools, ArcticQA + ArcticAbstain give you a blueprint:
- Use paired conditions so abstention sensitivity can be measured.
- Don’t trust abstention rate alone—monitor false abstention and false commitment.
- Anchor evaluation in specific evidence and scope, not generic knowledge.

If you want a system that’s safe for scientific use, this kind of benchmark is a strong starting point: it tests the exact failure mode where a model picks an option that doesn’t match what the evidence actually supports.

Key Takeaways

  • ArcticQA provides 194 multiple-choice questions grounded in primary Arctic research, with automated checks that:
    • the correct option is supported by source evidence, and
    • distractors are contradicted within the question’s stated scope.
  • ArcticAbstain measures abstention properly using paired answer-present vs answer-absent conditions, including an explicit abstention choice in both.
  • Models show huge baseline differences: answer-present abstention rates range from 0.0% to 63.0% across eight evaluated models.
  • Removing the correct answer increases abstention for every model, with an average jump of +5.05 percentage points (range 1.5 to 11.1 pp).
  • The study shows why you must evaluate jointly:
    • false abstention (abstaining when a valid option exists),
    • false commitment (answering when no valid option exists),
    • and accuracy when the model does answer.
  • For the future: this benchmark approach—evidence-grounded questions plus paired option-availability tests—sets a template for more trustworthy scientific QA evaluation beyond the Arctic.

The dataset and benchmark are available at: https://github.com/BenWilcox8/arctic-qa.

Sources Used

This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:

Frequently Asked Questions

Limited Time Offer

Unlock the full power of AI.

Ship better work in less time. No limits, no ads, no roadblocks.

1ST MONTH FREE Basic or Pro Plan
Code: FREE
Full AI Labs access
Unlimited Prompt Builder*
500+ Writing Assistant uses
Unlimited Humanizer
Unlimited private folders
Priority support & early releases
Cancel anytime • 10,000+ members
*Fair usage applies on unlimited features to prevent abuse.