Grice Rules for VQA: How VLMs Handle Misleading Questions

When a vision-language model answers “not what I meant,” you’re seeing a communication problem. New research evaluates VQA under Grice’s cooperative principles and finds misleading, human-like violations hurt VLM accuracy—plus humans and VLMs recover differently.
The finding When VQA questions violate cooperative conversational norms, VLM performance consistently diminishes.
The method Researchers use VLM-based “question modifier” engines to inject non-essential, ambiguous, or false details into otherwise valid questions.
The takeaway Humans and VLMs recover differently from violations, so real-world VQA needs robustness-focused prompt handling.
1st MONTH FREE Basic or Pro • code FREE
Claim Offer

The Short Answer

VLM accuracy drops when VQA questions violate Grice’s cooperative principles by adding non-essential, ambiguous, or false information. The study also finds humans and VLMs differ in how they recover from these “communication glitches.”

So what: if your VQA prompts may include irrelevant details, unclear phrasing, or injected assumptions, you should expect worse performance and different behavior than humans—meaning you need prompt validation and normalization for reliability.

Caveat: the impact depends on the type of violation, and the study’s evaluation is based on question modifiers and VQA v2.0 sampling (including yes/no, numeric, and open-ended subsets), so results may vary with different datasets and prompt-generation methods.

Grice Rules for VQA: How VLMs Handle Misleading Questions

Introduction

If you’ve ever asked a vision-language model a question and thought, “Wait… that’s not what I meant,” you’ve bumped into a real problem: these systems don’t always reason the way humans do in conversation. New research from Evaluating VQA in Vision Language Models using Cooperative Principles digs into exactly that. The paper tests how Vision Language Models (VLMs) answer Visual Question Answering (VQA) questions when the questions violate human communication norms—specifically the Cooperative principles behind Grice’s maxims.

The core idea is simple but powerful: humans don’t just read text literally. We also infer intent. We know when something is irrelevant, overly detailed, ambiguous, or even untrustworthy. Humans often “repair” meaning using context. The study asks whether VLMs can do the same, and whether they behave differently when the misleadingness is human-like versus AI-made.

To test this, the researchers use VLMs in two roles: (1) as regular VQA answerers (e.g., ChatGPT4o, Claude-3.5-Sonnet, Gemini-flash-1.5, LLaVA-7B) and (2) as “question modifier” engines that add extra details to otherwise valid human-written questions—details designed to break Gricean norms. Across experiments with 995 VQA questions sampled from VQA v2.0 (including 438 yes/no, 42 numeric, 493 open-ended), they find a consistent pattern: modifiers hurt accuracy, but the kind of violation matters a lot. Even more interesting, humans and VLMs differ sharply in how they recover from these communication glitches.

Why This Matters

This research is significant right now because VQA is moving beyond lab benchmarks and into real workflows—think content moderation, robotics (“Is that object safe to grasp?”), accessibility (“What does the sign say?”), and even customer support with screenshots. In those settings, you don’t just get “clean” questions. You get noisy instructions: someone adds irrelevant context, a teammate writes a slightly ambiguous prompt, or a tool accidentally injects a wrong assumption. The question becomes: will the system stay robust when language becomes less cooperative?

A concrete scenario you can imagine today: say you’re helping a device understand a photo of a street crossing. A user asks, “Is the light on?” but the UI later appends extra text like “If the pedestrian button is pressed, the light must be off”—even when that condition isn’t true in the image. Humans would likely ignore the irrelevant or suspicious bit and still answer based on what’s visible. This paper tests whether VLMs can do that kind of pragmatic “filtering” rather than just trusting the text.

How does this build on previous AI research? Earlier work explored hallucinations and false premises in VQA, often by crafting adversarial questions or synthetic image-question pairs. This paper takes a different angle: instead of only asking “Is the model factually right?”, it asks “Does the model reason like a cooperative conversational partner?” It frames VQA robustness through a communication lens (Grice maxims), connecting multimodal reasoning to pragmatic inference—exactly the gap between how humans interpret messy dialogue and how machines often interpret instructions.

What the Researchers Actually Measured

The experiments are built around one measurement theme: what happens when you intentionally make the question less cooperative—then see whether VLMs can still answer correctly.

How they define “cooperative” versus “non-cooperative” questions

Grice’s maxims (quality, quantity, relation, manner) are rules humans follow to make conversation effective: be truthful, don’t overload with unnecessary details, stay relevant, and be clear. In this VQA setting, the researchers treat a question as “Grice-optimal” if it can be answered from the image alone—without needing extra assumptions.

Then they create non-Grice-optimal questions by adding “modifiers” that are designed to be:
- Non-essential (violating quantity),
- Ambiguous (violating manner),
- Irrelevant or misleading (violating relation and/or quality),
- Or some mix of those.

Two types of modifier sources: model-generated vs human-generated

They run comparisons in two broad worlds:

  1. Model-generated violations

    • They use VLMs to produce modified questions from VQA v2.0’s original questions.
    • These modifiers are supposed to preserve the original answer when the modifier is added with the image available, but may still break pragmatic norms (like being overly specific or ambiguous).
  2. Human-generated violations (more realistic)

    • For a separate dataset, they use human-created modifiers from prior work (Britton et al.).
    • The key point: humans tend to violate maxims in more natural ways. AI-modified questions might be weird in ways humans don’t do.

This lets the study ask a question that matters for real deployment: do VLMs handle human-like messiness better than AI-like messiness?

How Question Modifiers Change VQA Accuracy (and Why the Source Matters)

The main VQA experiment tests VLM accuracy when the question gets altered. The researchers sample 995 questions from VQA v2.0’s test split and group them into response types:
- Binary: yes/no (438 questions; unique ground truth)
- Quantitative: numbers (42 questions; unique ground truth)
- Open-ended: free-form text (493 questions; ground truth collected via AMT)

For closed-weight VLMs, they test:
- ChatGPT4o (GPT)
- Gemini-flash-1.5 (Gemini)
- Claude-3.5-Sonnet (Claude)
and also open-source:
- LLaVA-7B (LLaVA)

They consider the modifier-model (which created the modified question) and the response-model (which answered it). A consistent headline emerges: modifiers reduce accuracy almost across the board.

Accuracy drops: modifiers make VLMs worse at answering

In their main closed-form evaluation (binary + quantitative where ground truth is unique), the accuracy decreases by different amounts depending on the VLMs involved.

At an overall level, they report:
- Claude shows the smallest accuracy reduction: 4.82%
- LLaVA shows the largest accuracy reduction: 13.28%
- With Gemini as the modification-model, the average reduction can exceed 14% (they highlight this as particularly damaging)

They also observe an interesting “source symmetry” pattern:
- Gemini introduces strong violations that hurt others, but its own accuracy doesn’t tank as much.
- ChatGPT4o behaves like the opposite: it can do relatively well on its own modified questions, but much worse on modifiers created by others.

Here’s a compact way to read the pattern described in the paper: the difference between a VLM’s effect when modifying vs when responding is not the same for all models.

The results summarized by model pair behavior

The paper reports that the “row vs column averages” reflect how much change a model induces as a modifier compared to how much its own responses change when faced with others’ violations. Ranked by that difference, they state:

Model Reported effect direction in paper
Gemini Strong inducing of difficult violations (others’ accuracy drops more)
Claude Smaller inducing difference
LLaVA Minimal inducing difference
ChatGPT4o Negative inducing difference (its own modifications may be easier for it; others’ modifications hurt it more)

(This is a “behavioral pattern” summary based on their narrative around row/column averages and standard deviation.)

Were these drops statistically meaningful?

For binary + quantitative responses, they use McNemar’s test (on correct vs incorrect) to check whether the modifier-induced changes are significant. They find:
- In most cases: null hypothesis rejected with p < 0.05
- Notable exception: Claude, where violations don’t significantly change responses for 3 of 4 modification-models

So the story isn’t just “accuracy dropped.” It’s “accuracy dropped in ways that look real and repeatable,” with Claude being comparatively robust.

Open-Ended Answers: When “Being Wrong” Gets Trickier to Measure

Open-ended VQA isn’t as clean as yes/no because the “right” response might be phrased many ways. The researchers handle this by using an AMT evaluation design where workers compare two answer pairs:
1) original question + model’s answer
2) modified question + model’s answer
and choose which answer is better (or if both are acceptable).

The key finding for open-ended questions matches the closed-form ones: modifiers still hurt.

They report (in their AMT-based comparison):
- Gemini introduces violations that make reasoning harder across all response-models
- LLaVA generates worst open-ended outcomes under violations (its responses get worse more reliably)

They also do a sanity check: AMT workers directly answer modified open-ended questions, and they compute whether VLM answers exactly match one of the human answers. They then correlate those “exact match” accuracies with the AMT comparison results. Correlation values are modest overall (overall correlation = 0.21), but positive, supporting that their evaluation isn’t totally disconnected from reality.

Do VLMs Ignore False Modifiers the Way Humans Would?

So far, the modifiers were designed to preserve the original answer (and were validated by humans when image context was available to the modifier-generator). But humans also need to handle a bigger failure mode: what if the question contains false or irrelevant information?

To test that, the paper introduces a second experiment:
- The modifier-model generates changes without being given the image
- That forces the modifier-model to hallucinate visual/relational modifiers
- Then the response-model sees the image and the (now unreliable) modified text

Can VLMs “pragmatically repair” the meaning?

They compute the percentage of times a response-model’s answer stayed unchanged when the false modifier was introduced—because if the model ignores the misleading text and uses the image context, it should often answer the same as before.

Their results indicate:
- ChatGPT4o had a smaller “unchanged response” rate (around 43%)
- Other models had > 60% unchanged response rates

One interpretation offered by the authors: ChatGPT4o may be more literal about the question text, while other models more often rely on visual context to filter out the noise.

Comparing humans vs VLMs under hallucinated modifiers

They also ask humans to answer the same hallucinated-modifier questions on AMT and compare:
- For binary/numerical: VLM vs human agreement (majority human label)
- For open-ended: semantic similarity via BERTScore F1 (average across three human responses)

They find humans and VLMs differ. ChatGPT4o, Gemini, and Claude match humans better than LLaVA in these settings.

A particularly revealing breakdown (binary-only):
- When humans say “no”, VLMs often agree closely.
- When humans say “yes”, VLMs only say “yes” roughly about half the time.

The paper’s explanation is pragmatic: humans can treat partially true statements with uncertainty (and decide “yes” depending on relevance to the image), while VLMs behave more “logically constrained” (they become negative more readily when any part of a condition seems wrong).

Human vs AI Violations: Are VLMs Better at “Human-Like Mess”?

This is one of the most important comparisons in the paper: do VLMs handle violations that come from humans differently than violations generated by AI?

They use a dataset from Britton et al., constructed by AMT workers who modified VQA questions in natural ways. They sample:
- 500 binary questions
- 130 open-ended inference questions (who/which/where/why-style)

Then they compare each VLM’s accuracy on:
1) original human questions
2) human-modified violations
3) AI-modified violations (generated by VLMs using their modifier mechanism)

What happens to accuracy?

They report (binary):
- All 4 models have lower accuracy on AI-modified questions than on human-modified ones
- Claude performs best for both human and AI violations
- LLaVA performs worst for both

For open-ended responses, they use AMT preference judgments (which VLM answer is better when comparing responses to human-modified vs AI-modified questions). The pattern holds:
- LLaVA is significantly worse when violations are AI-generated
- Claude is best in this comparison setting

They also report a summary preference stat (for Claude): only about 26% of the time humans judged the AI-modified response worse than the human-modified response (lower is better).

A grounding/realism angle using CLIPScore

To understand why “AI violations” are harder, they analyze alignment with visual content using CLIPScore.

They find:
- For both binary and open-ended settings, CLIPScore is different for human-modified questions than for each VLM’s AI modifications.
- Claude has the most similar alignment to human-modified questions in binary (0.639 vs 0.612).
- For open-ended, human alignment is closest to Gemini.

One striking point: LLaVA’s CLIPScore is highest partly because its modifiers are longer (average modifier lengths):
- human: 2.1 words
- GPT: 2.5
- Gemini: 2.6
- Claude: 2.7
- LLaVA: 3.8

So LLaVA adds more text, but it may be less realistic—leading to worse pragmatic behavior even if the alignment metric looks superficially high.

Bottom line: VLMs seem better when the violation looks like something a human would plausibly say, not just something that is “textually adversarial.”

Human Cognitive Effort vs VLM Pragmatic Reasoning: What the Time-on-Task Says

The paper also asks a psychological-style question: if humans work harder to interpret a violation, do VLMs also struggle?

They measure human cognitive effort using time-on-task in an AMT experiment where workers answer modified binary questions and then simplify them by removing parts they didn’t need. They track elapsed time (with outliers filtered beyond 500 seconds).

They compare:
- humans resolving human-modified questions vs
- humans resolving VLM-modified questions

On 400 binary modified questions (100 per VLM, and the same set modified by humans), they find:
- mean time for humans on human modifications: 73.4
- mean time on VLM modifications: 74.2

So humans spend slightly more time on VLM-induced violations overall (the distributions look similar), but the authors emphasize the qualitative relationship: humans seem to recognize and simplify unnatural violations more easily, requiring less effort, while VLMs don’t always become more accurate in those cases.

They connect this to model training habits: VLMs are trained on natural language, so they may handle human-like violations better. Humans, meanwhile, can quickly filter or rewrite what doesn’t sound right.

Key Takeaways

  • Gricean violations hurt VQA performance. Across 995 VQA v2.0 questions, adding modifiers generally decreases accuracy for ChatGPT4o, Claude, Gemini, and LLaVA.
  • Not all violations are equal. VLMs handle human-like violations better than AI-generated violations—especially noticeable when comparing human-modified vs AI-modified question sets.
  • Model identity matters (modifier-model vs response-model). Some models (like Gemini) can generate particularly difficult violations for others, while ChatGPT4o shows strong sensitivity to whose modifiers it’s answering.
  • Humans repair meaning differently from VLMs. When false/irrelevant modifiers appear, humans often treat conditions with uncertainty and use the image to decide what to trust. VLMs behave more “textually constrained,” especially in binary cases where humans show more nuanced “maybe/uncertainty” behavior.
  • Evaluation design matters for open-ended VQA. Open-ended answers were assessed via AMT comparisons, and the trends largely align with binary/quantitative results.
  • Practical implication: If you’re building systems that ask VLMs to answer from images, you should assume users/tools will produce messy prompts. Designing guardrails that detect “non-cooperative” or irrelevant additions could improve reliability—especially for models like LLaVA.
  • Future-facing insight: The work suggests that making VLMs handle pragmatic reasoning isn’t just about accuracy—it’s about producing and tolerating human-realistic linguistic behavior. The more the “mess” looks like real human dialogue, the better the models seem to cope.

If you want, I can also turn this into a quick “checklist” for prompt designers (how to avoid triggering Grice violations) or a “model risk profile” summary for GPT, Claude, Gemini, and LLaVA based on the paper’s patterns.

Sources Used

This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:

Where To Go Next

On-Demand UI That Follows HCI Rules: Skills + a Space to Think

Browse the free Prompt Database or tune your own prompts with the Prompt Optimizer.

Frequently Asked Questions

Limited Time Offer

Unlock the full power of AI.

Ship better work in less time. No limits, no ads, no roadblocks.

1ST MONTH FREE Basic or Pro Plan
Code: FREE
Full AI Labs access
Unlimited Prompt Builder*
500+ Writing Assistant uses
Unlimited Humanizer
Unlimited private folders
Priority support & early releases
Cancel anytime • 10,000+ members
*Fair usage applies on unlimited features to prevent abuse.