AI-Powered Access to Scientific Figures: What BLV and Sighted Scientists Really Need

Scientific figures and tables are where results live—but BLV scientists have long lacked usable access. New interview-based research tests ChatGPT and Gemini for multimodal QA, showing what helps, what breaks trust, and design changes that make figure access truly usable.
The finding Scientists can use AI to query figures interactively, but vague or incorrect answers cause them to drop the workflow.
The workflow Researchers often read by zooming from text into figures or tables, and AI usefulness depends on how well it supports that verification step.
The design need Multimodal scientific QA must improve interpretability and correctness so users can confirm axis ranges, thresholds, and sub-panels—not just get generic summaries.
1st MONTH FREE Basic or Pro • code FREE
Claim Offer

The Short Answer

AI question-answering can make scientific figures usable for both BLV and sighted scientists, but vague descriptions and incorrect outputs make researchers abandon the workflow.

Practitioners should frame queries with specific, checkable details (numbers, axes, thresholds, sub-panels) and treat AI as an access bridge only when responses are sufficiently verifiable for their domain needs.

The key limitation is trust: when the AI output is incomplete or wrong, BLV scientists may lack an easy way to correct it, and even sighted scientists stop using the AI mid-task.

AI-Powered Access to Scientific Figures: What BLV and Sighted Scientists Really Need

Scientific papers are packed with diagrams, graphs, and tables—not just prose. And if you’re blind or low-vision (BLV), getting real value from those visuals has historically been hard, often depending on static alt-text written by authors. Now, AI systems like ChatGPT and Gemini make something new possible: interactive question-answering (QA) where you can ask about a figure, a chart trend, or what a table implies. New research from the paper (based on interviews with 10 scientists across STEM) digs into what actually happens when real scientists try this workflow—how they ask questions, what they trust, and where AI responses break down.

This post walks through the study’s key insights: the difference between using AI as an “access bridge” versus an “efficiency tool,” why vague or incorrect figure descriptions can make people abandon AI mid-workflow, and what design changes could make multimodal scientific QA genuinely usable for people with different visual access needs.

Why This Research Matters for Scientific Work (Not Just Accessibility Research)

This matters right now because AI-powered assistants are increasingly woven into research workflows: summarizing papers, extracting methods, and helping with literature review. But scientific understanding isn’t delivered through text alone—figures encode the relationships, parameters, and spatial logic that papers rely on. If AI can’t interpret those reliably, it doesn’t just reduce convenience; it can reduce scientific rigor.

Here’s a real-world scenario where this lands immediately: imagine you’re doing a clinical or materials science literature search and you spot a promising figure or table. You want to quickly extract key quantitative details (thresholds, axis ranges, experimental setup parameters) without manually parsing an inaccessible PDF. Today, an AI assistant might give you something “close enough,” but if it’s vague (“high values” without baselines) or wrong (wrong number, wrong scale, missing sub-panels), a researcher may not have a safe way to verify. For BLV scientists, the paper itself may not be directly checkable, so an error isn’t just inconvenient—it can feel impossible to confidently correct.

This study also builds directly on earlier AI research that focuses on model benchmarks or synthetic datasets. Instead of asking only, “How well does the AI answer questions about images?”, it asks, “How do real scientists use these systems in practice, and what makes them stop?” That’s a big shift from performance metrics toward workflow reality—exactly the kind of evidence you need before deploying AI tools as “assistants” for scientific reasoning.

How Scientists Actually Read Papers: The Multimodal Reality Before AI

The paper starts from a blunt truth captured by a participant: science is inherently non-accessible when figures matter as much as text. That’s not rhetoric—it’s reflected in how scientists read.

What both BLV and sighted scientists agreed on: figures and tables are central

Both groups described a common pattern: they often start broad, then zoom in.

  • Many sighted scientists said they look at figures early to “walk through” a method or grasp a hypothesis fast, with text acting as confirmation.
  • Some domains are especially visual. One neuroscience researcher noted that in their field, the visuals are often treated as the main signal, with text secondary.
  • Other domains lean toward tables. Medicine-related participants described prioritizing tables over figures.

What differs: BLV scientists need visuals to be interpretable, not just describable

The study highlights a key gap in standard accessibility: alt-text is static and may not support the level of verification and specificity that scientific work demands. In the interviews, BLV researchers described situations where captions or alt-text don’t provide the precision needed to interpret axes, trends, or spatial relationships.

A participant described practical needs like understanding exactly where a data point relates to a geographic region—the kind of spatial nuance that’s hard to capture with generic descriptions.

The workaround landscape: scripts, screenshots, tactile rendering, and prompting

Because PDFs are frequently inaccessible, BLV scientists have built custom strategies. Examples from the interviews include:

  • Writing a Python script to extract text from PDFs.
  • Asking AI systems to reconstruct visual information—sometimes by requesting mathematical or code-based representations.
  • Creating tactile renditions of complex images using an embosser. One participant reported that tactile images could take 2–3 hours per picture, which is a major commitment when you’re doing real research quickly.
  • Using prompting tricks to reduce hallucinations—like explicitly asking the AI to cross-check with external sources.
  • Turning to colleagues for expert validation when the description needs scientific judgment (not just “what’s in the picture”).

Those workarounds matter because they set up what AI is expected to replace: not convenience alone, but access to the scientific meaning embedded in figures.

The “Funnel” Pattern: How Scientists Query AI (and Why It Breaks)

To understand usage, the researchers interviewed 5 BLV and 5 sighted scientists across STEM domains and then had participants use two free AI tools—ChatGPT and Gemini—on uploaded PDFs. They collected 115 queries and responses from these interactions.

All participants could ask whatever they wanted, but they were asked to include at least one query about a figure or table. They also worked through familiar vs. unfamiliar papers and used their normal workflows to navigate.

A shared behavior: start with overview, then dive deep

A consistent “funneling” strategy appeared across groups:

  1. General/overview queries (e.g., summarize the paper, key findings)
  2. Then granular deep-dives into figures, tables, methods, or specific claims

This is an important finding because it suggests the workflow expectation is not “AI should answer everything perfectly.” Instead, people want scaffolding: first orient, then zoom.

The critical divergence: AI as access bridge vs AI as efficiency tool

Where the groups diverged was how they used AI once they got oriented.

From the paper’s categorization of query types, BLV participants spent a much larger share of questions on figures/tables, while sighted participants spent more on methods/findings synthesis.

Group Share of queries focused on Figures/Tables Share of queries focused on Methods/Findings What that usually means in practice
BLV scientists (n=5) 49% (remainder across other categories) AI is a primary bridge to visual content—axes, trends, relationships
Sighted scientists (n=5) 29% 56% AI is used more as an efficiency/synthesis layer over content they can already access

This is one of the study’s most actionable insights: designing for “universal usefulness” might fail if you ignore different roles. BLV users need AI to provide interpretive access; sighted users often need AI to accelerate understanding and cross-checking.

Domain shape: spatial fields push harder toward figure interpretation

The paper also shows that query intent is shaped by discipline:

  • One BLV researcher in marine geology had 67% of queries about figures—because the subject matter is inherently spatial.
  • Materials science and medicine participants leaned toward extracting experimental parameters and outcomes.
  • Computer science and neuroscience participants asked for conceptual clarification—sometimes wanting “timeline-style” or intuition-based explanations.

So the AI isn’t being asked generic questions; it’s being asked to match the cognitive shape of the domain.

When AI Responses Aren’t Precise Enough, Trust Collapses Fast

The study finds a common failure mode: AI answers about multimodal content can be vague, incomplete, inconsistent, or subtly wrong. That’s not a small issue; it can change whether researchers keep using AI at all.

“Vague” isn’t good enough for science work

Multiple participants said that current AI descriptions often lack the nuance needed for scientific use. Examples of gaps include:

  • Missing chart axes scales or ranges, making it impossible to evaluate trends.
  • Using relative language like “high” without a baseline.
  • Omitting essential quantitative details that determine meaning (e.g., exact percentages or where a curve peaks).

A BLV participant described an interaction where AI gave a measurement (e.g., peak location “about 10 degrees”) and then conflicted with itself—first “you’re wrong, it’s actually 15 degrees,” then later “I believe it’s 10 degrees.” Even if the AI corrects itself within the conversation, the researcher is left doubting which value is reliable.

Another participant criticized descriptions that sounded like literal visual labeling without scientific interpretation—for instance, stating a line color represented a subduction zone while failing to say where those zones are located in the figure.

Incomplete figure coverage is especially dangerous

In science, “missing part of the figure” often means missing the conclusion.

The paper reports a failure case where AI summarized only a subset of sub-panels in a multi-part figure. For a BLV scientist, that changed the perceived structure of the figure (fewer panels than exist), producing confusion that couldn’t be easily resolved because the missing content wasn’t accessible through the AI answer.

Why BLV participants lose agency more severely

Here’s the part that makes the findings feel urgent: both BLV and sighted scientists can get misled by wrong AI outputs, but sighted participants often have a recovery path.

  • Sighted researchers can often validate by looking at the source PDF directly during the workflow.
  • BLV researchers may not have a comparable way to verify correctness when the document’s visuals are inaccessible.

So while everyone experiences mistrust, BLV scientists experience it as a harder-to-fix problem—because the “source of truth” may not be directly checkable.

Validation, Transparency, and Privacy: The Three Barriers to Safe Use

The study digs into what happens around verification and how that affects willingness to keep using AI systems.

First usage vs second usage: validation changes behavior

In the study design, participants initially queried an uploaded paper via AI without opening the PDF separately. Later, they were allowed to open the paper to validate.

This created a natural observation:

  • Some participants (including BLV4 and S1) said they would stop using the tool if they couldn’t validate outputs.
  • For others, AI acted as a directional aid: identify relevant figures or mechanisms, then check the manuscript afterward.

In the interviews, cross-checking often happened after participants suspected an error—like S1 manually inspecting figures after receiving information he believed was incorrect.

Trust is “fragile” and errors are intolerable

Participants described a “one strike” mentality in high-precision domains. In materials science, for example, even very small numeric differences can imply different chemical conclusions—so trust can’t be probabilistic. If the AI is wrong once, the user may abandon the workflow.

A BLV participant explained that sighted people can often visually detect hallucinations, but as a blind researcher they might need an accessible version of the document to determine whether the AI is hallucinating at all.

Users also struggle to tell synthesis vs quoting

Another issue: AI sometimes produces responses that are hard to classify as either (a) directly derived from the paper or (b) synthesized by the model.

One participant suggested the difficulty was partly due to tone—AI doesn’t signal uncertainty or “this is my interpretation,” making it harder to know whether you’re seeing the authors’ claim or the model’s addition.

Privacy: not uniform, but real for some BLV scientists

Most participants didn’t raise privacy concerns, but one BLV participant did. The concern was that AI systems may appear to claim they won’t use information, while still doing so in practice.

Interestingly, another BLV participant had a different stance: they explicitly asked for “blind-friendly” outputs and did not fear the system inferring blindness. That contrast suggests privacy and disclosure preferences are personal, not universal—so “one setting for everyone” isn’t the answer.

What Inclusive AI Scientific QA Should Do Next (Practical Design Guidelines)

The paper’s recommendations are grounded in the observed failures and workflow differences. Here are the most important design implications, translated into plain language.

1) Figure descriptions must be domain-aware, not just visually accurate

Right now, AI may describe colors and lines but omit scientific meaning. Users want:

  • axis scales, ranges, and where data lives
  • what spatial regions correspond to scientific entities
  • how trends support a claim

If AI can’t deliver that, it becomes decoration rather than analysis.

2) Go beyond “one paragraph alt-text”

Users asked for alternative representations that support deeper reasoning: math expressions, code, tables, and more structured breakdowns. One participant loved math-and-code focused reconstructions of how a figure was generated.

The theme here is that scientific understanding often works in multiple representational modes. Accessibility shouldn’t be limited to “describe the picture.” It should enable analysis of the picture.

3) Support progressive scaffolding from overview to detail

Since both groups followed a funnel workflow, QA tools should mirror that structure:

  • overview first (orienting context)
  • then staged answers that keep the broader context intact while zooming into details

This reduces cognitive burden—especially for BLV scientists who may be building their own “figure map” as they go.

4) Make it clear what’s quoted vs inferred vs synthesized

Trust depends on transparency. Systems should label:

  • what is directly supported by the paper
  • what is inferred or generalized
  • what might be uncertain

This is particularly important for BLV users who can’t easily verify against inaccessible visuals.

5) Provide accessibility options without forcing users to reveal identity

The study highlights a tension: some users directly request “blind-friendly” outputs; others prefer to phrase queries in a way that doesn’t require disclosing disability status.

QA systems should support accessible formatting on demand without requiring permanent profiling of the user’s disability.

If you want a compass for system design goals, the guidance in this paper is a helpful one—again, see the original work for the full context and dataset discussion.

Key Takeaways

Key Takeaways

  • Scientists query AI in a “funnel” workflow: start with overview questions, then drill down into figures/tables/methods.
  • BLV scientists use AI primarily for visual access: 49% of their queries focused on figures/tables, while sighted scientists had 29% there and 56% on methods/findings.
  • AI responses often fail where science needs precision: missing axis scales/ranges, vague trend language (“high”), incorrect numbers, and omitted sub-panels.
  • Trust collapses quickly when errors are uncheckable, especially for BLV participants who can’t easily validate against inaccessible PDF visuals.
  • Users want more than alt-text: they need interpretive, domain-aware descriptions and “beyond text” representations like math, code, or structured breakdowns.
  • Accessible QA tools must support transparency about whether answers are quoted from the paper vs synthesized/inferred.
  • Privacy preferences differ across individuals, so systems should offer accessibility without requiring explicit disclosure.

If you’re building or evaluating AI tools for scientific QA, the biggest lesson is simple: in scientific work, “almost correct” is often not usable. And for BLV researchers, usability also depends on whether they can recover from mistakes. This research shows where current multimodal QA systems fall short—and what inclusive design would look like when the stakes are real.

Sources Used

This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:

Where To Go Next

Structured DB Search vs ChatGPT: How Close Are We Really?

Spotting AI-Powered Stack Overflow Answers: SOGPTSpotter’s BigBird-Siamese Detector

Grounded AI That Knows Its Ground: A New OCT-Powered Coach Elevates PCI Planning Beyond General Models

Browse the free Prompt Database or tune your own prompts with the Prompt Optimizer.

Frequently Asked Questions

Limited Time Offer

Unlock the full power of AI.

Ship better work in less time. No limits, no ads, no roadblocks.

1ST MONTH FREE Basic or Pro Plan
Code: FREE
Full AI Labs access
Unlimited Prompt Builder*
500+ Writing Assistant uses
Unlimited Humanizer
Unlimited private folders
Priority support & early releases
Cancel anytime 10,000+ members
*Fair usage applies on unlimited features to prevent abuse.