Design AI workflows that “fit” professional tasks, not just prompts

People think AI can “help with anything,” but you still must structure the work. New research compares two scaffolded workflows for negotiation prep—and shows task-fit design improves analytic coverage beyond chat.
The finding Scaffolded workflows beat open-ended chat for negotiation preparation coverage by embedding task structure.
The method Two interfaces used the same negotiation scaffold: one prefills a completed analysis, the other grows incrementally as users and AI co-develop it.
The takeaway To get better analytic results, redesign the workflow—not just your prompts—so the UI structures what to request, verify, and evaluate.
1st MONTH FREE Basic or Pro • code FREE
Claim Offer

The Short Answer

Scaffolded AI workflows improved preparation coverage more than open-ended AI chat, and user-directed incremental scaffolds made the process feel less mentally effortful while eliciting broader analytic requests. The key is embedding professional task structure into how users and AI build the analysis.

So what: instead of asking for a finished answer, design or choose interfaces that start from your professional steps and let AI fill and refine each component alongside you. This helps you catch what’s missing and aligns AI output with your workflow.

Caveat: the study reports similar overall coverage between scaffold implementations, so the main differentiator is often how the workflow shapes engagement (request variety and perceived effort) rather than dramatically different coverage.

Design AI workflows that “fit” professional tasks, not just prompts

People love the idea that ChatGPT can “help with anything.” But new research from Ma and colleagues in this paper points to a problem nobody can brute-force away with better prompting: general-purpose AI doesn’t tell you how to structure a professional task. Every time you ask for help, you still have to figure out what matters, what’s missing, how to decompose the work, and how to evaluate what you get back.

The core question in the study is deceptively simple: should AI give you a completed analysis, or should it help you build that analysis step-by-step inside a professional workflow? The researchers compared two AI “scaffold” designs for frontline humanitarian negotiation preparation—using the same professional negotiation framework—but implemented differently: one interface filled the scaffold with an AI-produced answer up front, and the other started empty and grew as the user and AI co-developed it.

What they found is even more interesting than “the AI helped.” Scaffolded workflows improved preparation quality more than open-ended chat, and the way the scaffold is populated changed how people work: the user-directed, incremental approach elicited a broader range of analytic requests and felt less mentally effortful—even when overall quality (coverage) ended up similar.

Why This Matters (and why it’s a big deal right now)

This research is significant because we’re at an awkward phase of AI adoption: organizations are buying “AI chat” tools, but their workflows haven’t been redesigned around what real work actually requires. In practice, professionals don’t just need a summary—they need structure for deciding what to do next, what to verify, what to trade off, and what to defend later. The moment you put a chat box in front of someone, you’ve basically outsourced the workflow design to their intuition.

A concrete scenario: imagine you’re a policy analyst preparing for a negotiation with partners (or a vendor, or a regulator) this week. You open a ChatGPT-like tool and ask: “Help me prep.” The model can generate plausible content—but it won’t reliably guide you through the specific checklist your organization uses (what you can concede, what you can’t, where interests overlap, what risks break your plan). That’s exactly the metacognitive burden the paper highlights: people often don’t know how to translate their professional process into “good prompts,” and they may not notice what the interface is leaving out.

This study builds on earlier AI research about automation and decision support, but it reframes the main design lever. Instead of asking only “how good is the model output?” it asks how the workflow between user and AI should be engineered. Previous work often compared task-specific scaffolds versus open-ended interfaces; this paper goes further by comparing two implementations of the same scaffold—one that externalizes a finished product and one that externalizes the process of building it.

In other words: even when the “content framework” is the same, workflow design changes what people do with AI, how hard they feel they’re working, and what kind of ownership they experience over the result.

What the researchers compared: four interfaces, same professional scaffold

The study ran a four-condition, between-subjects experiment with N = 1,712 participants originally, and N = 800 completed and met inclusion criteria:
- Reader: 181
- Chatbot: 260
- AI-Prefilled: 206
- Co-Evolving: 153

Everyone did the same timed preparation task for frontline humanitarian negotiation. Participants reviewed a realistic case using 15 documents (6,977 words total) and had 15 minutes to prepare under time pressure—then answered written questions in an assessment phase where AI access was removed.

The professional scaffold came from the Centre of Competence on Humanitarian Negotiation (CCHN) Field Manual. It included three recurring preparation modules:
- Iceberg (stated positions vs. underlying interests)
- Island of Agreements (where parties agree vs. core tensions)
- Paths to Agreement (a workable package, with boundaries/redlines and constraints)

The key experimental twist: the same scaffold could be used in two different ways.

Condition What the scaffold looks like When AI fills it What the user does most of the time
Reader Notes + document reader No AI Read + take notes
Chatbot Notes + open-ended assistant N/A Ask anything; build preparation “through prompts”
AI-Prefilled Completed scaffold artifact + transcript Before user starts Inspect, verify via citations, and optionally revise modules
Co-Evolving Empty scaffold artifact that updates Incrementally during interaction Direct what to pursue; AI populates modules step-by-step

The paper also includes a practical parallel: after preparation, participants kept whatever artifacts the interface produced, so the comparison wasn’t just “what did AI generate,” but what work participants ended up performing and retaining.

If you want to see the design logic described in the paper, it’s laid out in their experimental overview and interface section—see the original paper.

The headline results: AI improved coverage, scaffolds beat chat, but risk coverage stayed stubborn

The researchers scored participants’ written responses using a rubric developed by frontline negotiators and applied consistently (with an LLM judge). The main dependent measure was overall preparation coverage, essentially: how many key elements of a well-prepared negotiation response the participant included.

Here’s what happened:

1) AI support improved coverage vs. no-AI.
When pooling the AI-supported conditions (Chatbot, AI-Prefilled, Co-Evolving) and comparing to Reader, preparation coverage increased by 3.65 points (95% CI [1.25, 6.05], p = .003, effect size g = 0.25).

2) Scaffolded workflows beat open-ended chat.
Pooling AI-Prefilled and Co-Evolving improved coverage by an additional 2.36 points over Chatbot (95% CI [0.11, 4.61], p = .040, g = 0.16).

3) Incremental vs. completed scaffold produced similar overall coverage (with some uncertainty).
The Co-Evolving advantage over AI-Prefilled was estimated at 3.19 points, but was less precise: 95% CI [−0.14, 6.52], p = .060, g = 0.21.

So yes: scaffolds help. And the system didn’t need to “train users to prompt better” to get a measurable quality bump—structure in the interface did the heavy lifting.

But there’s an important nuance:

Package-risk coverage didn’t improve detectably

The paper reports that risk coverage and the match between identified risks and the proposed package didn’t differ detectably in the nested contrasts (all p_BH ≥ .45). That matters because it suggests scaffolds can make certain parts of professional preparation more complete while leaving other, harder-to-infer components untouched.

And the interface usage logs reinforce this: participants rarely asked AI to verify, challenge, or defend claims. So even if a scaffold names “risk” explicitly, it doesn’t automatically make users treat the risk analysis as something to scrutinize.

How the workflow changed interaction: chat gets summaries, co-evolving gets real negotiation thinking

Quality improvements are one thing. But this paper’s most practically useful insight is about how people used the AI.

They analyzed every user assistant “turn” across AI conditions (total turns: 2,992) and classified each interaction by:
- request form (passage-only vs passage + request, or authored message)
- purpose (reading support, fact retrieval, own-side analysis, counterpart analysis, cross-party synthesis, strategy/package development, verification/challenge, etc.)

A few key patterns:

Scaffolded workflows shifted what participants asked for

Relative to Chatbot, the scaffolded conditions reduced reading support by 30.1 percentage points (adjusted q < .001) and increased:
- counterpart analysis by 16.6 points (q < .001)
- cross-party synthesis by 13.0 points (q = .006)
- strategy/package development by 11.8 points (q = .020)

That’s a subtle but huge design lesson: Chat makes many tools available, but it doesn’t tell users what to use them for. When the scaffold is visible, users tend to ask the kind of questions the scaffold implies.

Co-evolving elicited a broader analytic repertoire

Compared to AI-Prefilled, Co-Evolving increased the incidence of each analytic purpose by 30.9 to 40.2 percentage points (all adjusted q < .001). It also increased coverage breadth:
- share of participants covering at least two analytic purposes: +52.3 points
- share covering all four analytic purposes: +23.5 points

They even checked this under interaction volume control. Among participants with at least three turns, the expected number of distinct analytic purposes in a three-turn sample was 1.68 for Co-Evolving vs. 0.91 for Chatbot and 0.88 for AI-Prefilled (with Co-Evolving comparisons p_Holm < .001).

Verification stayed rare

In scaffolded workflows, participants did author more explicit requests (participant-authored requests were +25.2 points vs. Chatbot, while passage-only was −20.2 points). Still, explicit verification/checking/challenging remained uncommon:
- only 72 of 2,992 turns asked for verification-like behavior (2.4%)
- across 52 of 481 assistant users (10.8%)

So the scaffold helped users move from “summarize what I read” to “analyze negotiation elements,” but it didn’t reliably make them demand justification or adversarial checking of the model’s outputs.

Engagement and psychology: agency held steady, but ownership and effort split in surprising ways

If you care about human factors (you should), the paper digs into three subjective metrics:
- Decision agency (did users feel in control of process?)
- Psychological ownership (did users feel the output was “theirs”?)
- Subjective effort (mental effort required)

Effort: incremental development felt less demanding

The paper found a specific difference between the two scaffolded designs:
- Co-Evolving reported less preparation effort than AI-Prefilled
(6.17 vs. 6.55, g = −0.35, p_Holm = .005)

But effort didn’t drop below the Reader baseline; rather, AI redistributed the effort. The completed AI analysis up front can be cognitively expensive to digest, while incremental development lets users process in smaller chunks.

Agency: surprisingly similar across conditions

Decision agency did not differ detectably between conditions (values hovered around 5.04 to 5.12), including across the nested comparisons. In other words, users felt they controlled the process about equally often—regardless of workflow.

Ownership: AI reduced it across the board

Psychological ownership was lower in every AI condition vs. Reader (overall g = −0.42, p_Holm < .001). And it also differed within AI conditions: ownership was lower in the scaffolded workflows than in chat (g = −0.23, p_Holm = .015), with no clear detectable difference between Co-Evolving and AI-Prefilled (p_Holm = .49).

This is a key design insight: preserving “control” isn’t the same thing as preserving “ownership.” Even if users can choose what to work on, the resulting artifact may still feel less authored—especially when AI generates most of the content.

Practical implications: how to design “professional AI” workflows that users actually benefit from

This paper’s conclusion is basically: workflow design is not a packaging detail—it’s the product. If you deploy AI as a generic assistant, you’re forcing users to reconstruct the professional scaffold themselves.

Here are the most actionable lessons, translated from their findings into design principles:

1) Don’t assume chat removes workflow design—scaffold it

Chat interfaces let users ask anything, but the study shows many users defaulted to “read and summarize” style requests. Professional scaffolding reduced that behavior by making negotiation-relevant components visible.

Practical move: build interfaces where the professional framework is persistent and inspectable (like the Iceberg / Island / Paths modules), not something users have to remember or prompt for.

2) Choose whether AI fills a scaffold or co-develops it—those are different experiences

AI-Prefilled and Co-Evolving produced similar overall coverage, but differed in:
- analytic breadth (co-evolving better)
- subjective effort (co-evolving less effortful)

Practical move: if your users need to explore and reason (not just consume), prefer co-evolving workflows that let the user steer what gets generated next.

3) Treat verification as a first-class interaction goal, not an optional “nice-to-have”

Even with a scaffold, explicit verification requests were rare. If your domain requires justification (medical decisions, legal analysis, compliance), your UI probably needs to operationalize verification:
- prompts that ask users to confirm key claims
- required “evidence links” for risk statements
- explicit slots for “what would change my mind?”

This directly addresses the paper’s observation that risk coverage didn’t improve detectably.

4) Ownership is fragile—support lightweight authorship

Agency stayed steady, effort didn’t magically disappear, and ownership dropped when AI contributed a lot of content. If you care about user trust and long-term accountability, you need authorship mechanics that create meaningful user contribution.

Practical move: generate adaptable scaffolds while requiring users to:
- select modules
- refine claims
- write concise justification notes
- approve the final package/rationale

Ownership may be less about “who clicked the button” and more about “who shaped the content that matters.”

5) Don’t rely on domain expertise or “AI literacy training” alone

The paper reports that prior negotiation experience improved performance in the no-AI condition, but that advantage didn’t reliably translate into better AI prompting behavior. In other words: knowing the domain doesn’t automatically teach you how to delegate well to AI.

So your organization shouldn’t just train employees to prompt better. It should redesign the interface so the right workflow is easy to follow.

Key Takeaways

  • General-purpose AI isn’t enough. Users must still structure professional work; open-ended chat leaves that burden on them.
  • Scaffolded workflows improved preparation coverage over no-AI (overall +3.65 points) and over open-ended chat (+2.36 points).
  • Two ways of using the same scaffold matter. Co-Evolving (incremental scaffold development) elicited a broader analytic repertoire and felt less mentally effortful than AI-Prefilled, even though overall coverage was similar.
  • Risk analysis didn’t improve detectably. Package-risk coverage stayed stubborn, suggesting scaffolds alone don’t force verification-grade thinking.
  • Users felt in control (agency) but less ownership. Decision agency stayed similar across conditions, while psychological ownership dropped in every AI condition.
  • Design implication: treat “professional AI” as work design. Structure not only what AI outputs, but also how users direct, inspect, verify, and take responsibility for the analysis.

If you’re building or choosing AI tools for knowledge work, this paper is a strong argument that the winner won’t just be the “best model.” It’ll be the system that engineers the right workflow around the way professionals actually think and defend decisions.

Sources Used

This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:

Where To Go Next

Rethinking AI Mental Well-Being Design: Supplements, Drugs, or Primary Care?

Endings for AI Companions: Designing Safe Closures for Human–AI Bonds

AgentIF-OneDay: Benchmarking Everyday AI Agents in Daily Tasks

Browse the free Prompt Database or tune your own prompts with the Prompt Optimizer.

Frequently Asked Questions

Limited Time Offer

Unlock the full power of AI.

Ship better work in less time. No limits, no ads, no roadblocks.

1ST MONTH FREE Basic or Pro Plan
Code: FREE
Full AI Labs access
Unlimited Prompt Builder*
500+ Writing Assistant uses
Unlimited Humanizer
Unlimited private folders
Priority support & early releases
Cancel anytime 10,000+ members
*Fair usage applies on unlimited features to prevent abuse.