The Short Answer
Most image-upload tasks for multimodal LLMs are multi-step capability chains (perception → cognition → generation), not single-step “image → answer” requests.
Practically, benchmark design and product evaluation should target chained outcomes (e.g., interpreting and then producing a deliverable like code, text, or decisions), not only perception and fixed-answer reasoning.
A key caveat is that the findings reflect observed usage patterns in specific datasets and depend on how image-upload conversations are categorized into capability demands.
On this page
- Why This Research Matters for Multimodal AI Right Now
- The Capability Chain: How Image Uploads Turn Into “Jobs”
- What Users Try to Do: A Demand Map That Goes Beyond “Answering”
- Why Benchmark Coverage Still Misses Real Image-to-Artifact Workflows
- The Privacy-Aware Method: How They Did This Without Raw Images
- What This Means for Building Better Multimodal Assistants (and Better Benchmarks)
- Key Takeaways
Image-to-Task Mapping for Multimodal LLMs in the Wild
Multimodal LLMs are everywhere now—people can chat and also drop in screenshots, charts, or photos. But here’s the big question that most teams (and benchmarks) haven’t really nailed: when users upload images, what are they actually trying to accomplish? New research from the original paper digs into this using an enormous real-world dataset of image-upload conversations, and the results are surprisingly clear: most image-based workflows aren’t single-step “look and answer” tasks—they’re multi-step chains that mix perception, reasoning, and generation.
In the study, the authors analyzed over 40,000 de-identified image-upload conversations from Microsoft Copilot—specifically 42,617 multimodal sessions—and then validated the same analysis on an independent ChatGPT dataset (23,413 multimodal sessions). Their goal wasn’t to guess. It was to systematically characterize the capabilities people rely on when images enter the conversation, and then ask whether existing benchmarks measure what’s common in real life.
The paper’s core idea: understand multimodal use not as a novelty (“the model can see”), but as a demand map (“the model is being used to do certain jobs, in certain combinations, with certain end products”). Once you do that, it becomes obvious where benchmark design is biased—and where multimodal assistants still feel frustratingly incomplete.
Why This Research Matters for Multimodal AI Right Now
This matters right now because multimodal assistants are shifting from “demo mode” to “everyday tool.” When users start relying on image uploads for work—debugging, summarizing docs, converting screenshots into editable formats—the bar changes. People don’t just want OCR. They want an outcome: a draft, a piece of code, a decision-ready summary, a workflow that keeps going after the model reads the image.
What makes this research especially timely is the mismatch it reveals between benchmark incentives and user reality. Many multimodal evaluations still behave like: image → answer (or image → caption/explanation). But the paper shows users more often follow image → interpret → reason → produce something useful. That changes what “good performance” should mean.
A concrete scenario you can picture: you’re a product manager or marketer reviewing analytics. You upload a screenshot of a chart, ask Copilot to interpret it, and then—without wanting to copy everything into the chat—you request “turn this into a strategy memo for leadership” or “draft the next email campaign based on these trends.” This isn’t just perception. It’s perception feeding into generation, often with knowledge grounded in your context.
This work builds on earlier AI research that studied how people use LLMs in real settings, but it pushes deeper into multimodality. Previous analysis often focused on text-only conversations. Here, the authors ask what changes when visual input enters the loop—and they quantify not only what capabilities appear, but how frequently they combine. That combination part is the real game changer.
The Capability Chain: How Image Uploads Turn Into “Jobs”
The paper frames multimodal tasks as a chain of capabilities across three functional layers:
- Perception: pull structured signals from the image (e.g., identify what’s in it, read text from it)
- Cognition: use knowledge to interpret and reason about those signals
- Generation: write or output what the user requested (text, code, documents, etc.)
If you like analogies: think of image upload as dropping a “rough sketch” onto the desk. Perception is the step where someone deciphers the sketch into usable details; cognition is the step where they figure out what it implies; generation is where they produce the final deliverable.
The paper’s taxonomy: 10 multimodal capability categories
The authors derive a hierarchical taxonomy with ten capability categories, then label each image-upload session with a primary and (when relevant) secondary capability. They don’t do this from raw images—because of privacy restrictions. Instead, they use privacy-preserving summaries (more on that later).
Across the Copilot dataset, the most frequently invoked capabilities are:
- Recognition
- Knowledge
- OCR & Text Extraction
In other words: users most often upload images to help the system understand what’s there, extract readable content, and then apply knowledge/meaning.
Most sessions are compositional (not single-capability)
One of the strongest findings is how often image uploads trigger chains. The paper reports that:
- 74.17% of image-upload sessions invoke multiple capabilities
So, rather than “the model reads the chart,” it’s often “the model reads the chart and then reasons and drafts/explains/code-generates something.”
The top recurring combinations include examples like:
- Knowledge + Language & Text Generation (the single most frequent combination)
- Recognition + Reasoning
- Knowledge + Recognition
The practical implication: if your multimodal assistant performs well on isolated “vision questions,” it still might fail the user’s real workflow—because the workflow is multi-step by default.
What Users Try to Do: A Demand Map That Goes Beyond “Answering”
The paper doesn’t stop at listing capabilities. It also characterizes what those capabilities look like in practice by extracting keywords from user task summaries using TF–IDF. Even without raw user text, you can infer the kinds of intents people have.
For instance:
- Under Recognition, the verb–noun patterns include things like “provide description” and “identify elements.”
- Under Language & Text Generation, common patterns include “generate resume” and “create caption.”
So far, that may sound intuitive. But here’s the bigger surprise: image-upload usage occupies a broader and more diverse task space than text-only interactions.
Multimodal tasks extend beyond text-only tasks (asymmetrically)
To compare multimodal vs text-only task spaces, the authors build task summaries for both:
- multimodal sessions (images included)
- matched-size text-only sessions (no uploaded images)
Then they evaluate the overlap using three methods:
1. cross-modality perplexity
2. out-of-distribution (OOD) detection in embedding space
3. t-SNE visualization of task embeddings
Their perplexity result shows a clear asymmetry:
| Comparison | Result |
|---|---|
| Model trained on multimodal evaluated on text-only | lower perplexity: 65.55 |
| Model trained on text-only evaluated on multimodal | higher perplexity: 76.61 |
Interpretation: text-only training doesn’t cover multimodal task diversity as well as multimodal training covers text-only tasks.
OOD detection shows multimodal includes “extra territory”
They also run OOD detection using KNN distance and Isolation Forest on embedding space. The asymmetry continues:
| OOD detector | Fraction marked OOD (trained on text-only → evaluated on multimodal) | Fraction marked OOD (trained on multimodal → evaluated on text-only) |
|---|---|---|
| KNN | 76% | 43% |
| Isolation Forest | up to ~30% | <<1% |
That’s strong evidence the multimodal task space isn’t just “text tasks with extra visual flavor.” It includes categories that are harder to reach from text-only patterns.
Human audit: some tasks are uniquely multimodal
To test whether images are actually required (not just decorative), the authors perform a human audit of 200 multimodal sessions. Two annotators judge whether each task could reasonably be completed from text alone.
Result:
- 30% of sessions (60/200) are judged uniquely multimodal
That’s an important nuance: even if a workflow could be approximated from text, a substantial portion depends on actual visual grounding.
Why Benchmark Coverage Still Misses Real Image-to-Artifact Workflows
Once you know what users do, you can ask the uncomfortable question: do current multimodal benchmarks test those real-world capability chains?
The authors do exactly that by mapping their taxonomy onto 253 existing multimodal benchmark papers, extracting 1,173 evaluation tasks, and then filtering out non-image tasks and “umbrella” benchmarks. After exclusions, they end up with 866 tasks to analyze.
How they measure coverage: supply-to-demand (S/D) ratio
For each capability pair combination, they compute a supply-to-demand ratio:
- S/D > 1 means benchmarks cover that combination more than users demand it
- S/D < 1 means benchmarks under-cover it
- They classify coverage as:
- Strong:
S/D ≥ 1 - Limited:
0.3 ≤ S/D < 1(the paper’s bands use a thresholding scheme; key point is “not fully covered”) - Incidental:
S/D < 0.3
- Strong:
Where benchmarks do well: perception → reasoning → fixed answers
Coverage is strongest when the image maps to a clear, constrained output like:
- labels
- values
- short explanations
- reasoning toward a fixed answer
Some of the strongest matches (real-world frequent combos that benchmarks also test well) include:
- Recognition + Reasoning: S/D = 3.12 (Strong)
- Knowledge + Reasoning: S/D = 1.92 (Strong)
- Recognition + Language Generation: S/D = 1.91 (Strong)
- OCR + Recognition: S/D = 2.00 (Strong)
So if your benchmark suite is about “what’s in the image and what does it mean,” it’s probably doing okay.
Where benchmarks underperform: reading the image isn’t the end
Here’s the gap the paper calls out sharply: real users don’t stop at short answers. They use image inputs as a starting point for bigger outputs—often documents, code, and decision-ready artifacts.
The mismatch shows up most clearly for the most common real-world combination:
Knowledge + Language & Text Generation
- Real-world frequency: 3,563 sessions (8.4%)
- Benchmark task share: 13 tasks
- S/D = 0.22 → Incidental
Why? The benchmark tasks tend to end at shorter outputs (like an explanation), whereas real usage often continues into “write the full structured thing” (e.g., a memo, a report, a policy-style summary).
The same pattern holds for:
- OCR + Language & Text Generation: S/D = 0.43 (Limited)
Code and data workflows are the rarest in benchmarks
The paper emphasizes that code and data generation show the weakest benchmark alignment. Examples:
- Code Generation + Recognition: S/D = 0.09 (Incidental)
- Code Generation + Reasoning: S/D = 0.11 (Incidental)
- Data Analysis + Recognition: S/D = 0.39 (Limited)
And the reason isn’t subtle: benchmarks often test code-related tasks that look like “reconstruct the plot from the image” or “edit CSS to match the screenshot.” But real workflows look more like “take this UI mockup + instructions and produce the working implementation” or “extract chart data, then compute and recommend.”
If perception and reasoning are the “front-end” of multimodal AI, then benchmarks are mostly benchmarking the front-end while missing the “build and ship” part.
The Privacy-Aware Method: How They Did This Without Raw Images
A practical detail that matters: the study is constrained by privacy.
The authors analyze de-identified conversations and use an eyes-off approach. That means:
- They don’t access raw images.
- They transform each session into privacy-preserving summaries:
1. a task summary (≤ 50 words)
2. an image summary (≤ 50 words), describing the uploaded image and requested operation
This is then fed into an automated taxonomy induction pipeline based on gpt-4o-mini within a framework adapted from TnT-LLM (as described in the paper).
To validate that summaries still preserve the multimodal signals needed for capability labeling, they report multiple checks, including:
- robustness across prompt configurations and datasets
- agreement between LLM and human annotations (N=200)
- an eyes-on audit (N=200) confirming the summaries preserve the capability signal with limited loss
So while we don’t see the raw image-level details, the system is still capable of identifying the capability structure—because the summaries include the user’s intent and the requested operation tied to the image.
What This Means for Building Better Multimodal Assistants (and Better Benchmarks)
If you’re designing a multimodal feature—say, “upload a screenshot to get help”—this paper suggests you should treat uploads like structured work inputs, not just “extra context.”
1) Support “extract → interpret → produce,” not “extract → answer”
Because the most common real-world sessions are compositional (74.17% multi-capability), assistants should explicitly support multi-step workflows:
- read/recognize
- reason with user constraints
- generate complete artifacts
2) Reward end-to-end usefulness, not just correctness of a single label
Benchmarks often grade on the final fixed answer. The paper argues for scoring outputs on:
- fidelity
- executability (especially for code)
- decision relevance (for data analysis)
3) Treat formatted content as first-class input
The paper notes that users frequently upload structured textual content (code snippets, HTML/CSS, Excel/table) rather than retyping. That aligns with the idea of cognitive offloading: screenshots reduce the friction of reproducing layout and context in chat.
Practically: systems should “recover” underlying structure from images (tables, code blocks, layouts) instead of treating everything as a flat picture.
4) Expect multimodal errors to propagate
One open direction the paper flags: upstream perception mistakes (e.g., OCR errors) can cascade into downstream reasoning/generation. So robust multimodal assistants need to manage uncertainty and recover gracefully—not just “answer confidently.”
And if you want to see this research in its original form, the full details live in the paper.
Key Takeaways
- Image-upload tasks are usually multi-step. In the Copilot dataset, 74.17% of sessions invoke multiple multimodal capabilities.
- Top demanded capabilities are Recognition, Knowledge, and OCR & Text Extraction. These appear most often in real user workflows.
- Multimodal tasks cover a broader space than text-only tasks. The comparison shows an asymmetry: models trained on multimodal task summaries cover text-only better than the reverse (perplexity 65.55 vs 76.61).
- A meaningful chunk of real tasks require vision for real. In a human audit of 200 sessions, 30% were judged uniquely multimodal (not reasonably doable from text alone).
- Benchmarks still skew toward fixed-answer evaluation. Capability combos involving perception+reasoning toward a clear answer are well covered (e.g., Recognition + Reasoning S/D = 3.12).
- Most common real workflows involving generation are under-tested.
- Knowledge + Language & Text Generation is 8.4% of real sessions but only incidental in benchmarks (S/D = 0.22).
- Code and data workflows are especially scarce in benchmark coverage (e.g., Code Generation + Recognition S/D = 0.09).
- Going forward, benchmarks should mirror “image-to-artifact” workflows, and scoring should focus on output usefulness (fidelity, executability, decision relevance) rather than only a single correct label or short answer.
If you’ve ever used a multimodal assistant and thought “it understood the chart… but didn’t actually finish the job,” this paper explains why—and points to exactly what needs to change next.
Sources Used
This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:
- From Images to Tasks: Characterizing Multimodal LLM Interactions in the Wild — arXiv
- Authors: Authors: Jinyi Ye, Scott Counts, Gaurav Verma, Kate Lytvynets, Weiwei Yang