On-Demand UI That Follows HCI Rules: Skills + a Space to Think

Generative UI can produce interfaces, but they often miss the usability details. This research argues for procedural HCI: a “Space to Think” workspace where task decomposition generates an interface and machine-enforces HCI principles via runtime skills.
The finding The research targets a “generation gap” where generative UIs can be functional yet miss usability details.
The method It proposes a shared Space to Think plus runtime-loaded HCI skills to make interface craft executable.
The caveat Usable outcomes depend on the quality and coverage of the machine-readable HCI skills for the interface context.
1st MONTH FREE Basic or Pro • code FREE
Claim Offer

The Short Answer

On-demand UI follows HCI rules by generating a working interface from task text while enforcing interaction design knowledge via machine-readable “skills” loaded at runtime. The research argues the key shift is from producing UI code to reliably applying HCI craft.

Practically, your task decomposition becomes structured work in a “Space to Think,” and the system materializes UI that supports that thinking while accessibility, learnability, and cognitive-load constraints are treated as invariants rather than after-the-fact checklists.

A major caveat is that usability depends on how well the HCI skills capture the relevant principles for the target interaction context—skills must be accurate and maintainable, otherwise enforced constraints can’t guarantee the intended user experience.

On-Demand UI That Follows HCI Rules: Skills + a Space to Think

HCI is about to get a major upgrade. Instead of designers anticipating users’ needs ahead of time and shipping a fixed app, new research (based on the paper) argues we can generate working, usable user interfaces on demand from plain-language task descriptions—and generate them well, not just “working.” The big shift: move from “the model can output UI code” to “the system can reliably apply HCI craft.”

This work describes a framework built around a Space to Think—a structured shared workspace where a human and an AI collaborate. As the user decomposes a task in dialogue, the system mirrors that thinking by materializing an interface that supports the work. And crucially, it proposes turning classic HCI guidance (accessibility, learnability, cognitive load, mixed-initiative behavior) into something machine-executable: skills—small, readable, version-controlled instruction files that the generator loads at runtime.

Why This Matters More Than It Sounds: The “Generation Gap” Is Now the Problem

Right now, “generative UI” is impressive—but it’s also uneven. Many systems can produce interfaces that look plausible and run, yet still miss the subtle stuff that makes interfaces usable: consistent affordances, keyboard reachability, error recovery, or keeping cognitive load low. The research calls out a missing ingredient: we need a conceptual frame and a mechanism for procedural HCI—making good interaction design a property of the generation process itself, not a hoped-for outcome of post-hoc human review.

Here’s a scenario where this becomes real today: imagine you’re a small business owner trying to set up an internal workflow tool. You don’t know what buttons, forms, and validations you’ll need—you just describe the goal: “I need to track invoices, send reminders after 14 days, and flag overdue items for follow-up.” Today’s apps force you into a menu-driven world designed for the average user. Even with AI assistance, the UI often ends up requiring workarounds: “Why is this button here?” “Why can’t I tab to that field?” “Why does this form hide the important info until later?”

The research’s angle is that this doesn’t have to be a guessing game. The proposed approach builds a shared workspace where your decomposition of the task becomes UI structure, and the system enforces HCI rules automatically via skills like accessibility-wcag.skill.md and cognitive-load-progressive-disclosure.skill.md. Instead of treating accessibility and usability as release-time audits, it treats them like invariants.

And it builds on earlier AI work in a specific way. Previous efforts in LLM agents and mixed-initiative interaction showed how AI could cooperate, plan, and act through tools. But generative UI made the collaboration visible as screens. This paper’s contribution is to connect the dots: cooperation becomes a “Space to Think,” and the design knowledge becomes enforceable constraints—shifting HCI from checklist culture to executable craft. You can think of it as the missing bridge between “AI can produce artifacts” and “the artifacts meet interaction design standards by construction.”

From “Chat + Output” to a Space to Think Where UI Becomes Work

The core conceptual move in the paper is that the dialogue should be more than a query-answer exchange. It becomes a Space to Think: a shared, persistent, structured workspace where both human and AI externalize and refine task understanding. The paper emphasizes three properties:

  1. Shared: both participants can read and write the workspace content.
  2. Structured: the workspace isn’t just text—it has scaffolding like sub-goals, hypotheses, evidence, decisions, and tools-under-construction.
  3. Operational: it supports the user’s work directly, which is the part that distinguishes it from a chat log.

If that sounds abstract, here’s the analogy: a Space to Think is like the difference between (a) talking to someone while they sketch in a notebook off to the side, versus (b) having them draw directly on a whiteboard you both can edit in real time. In the second case, the whiteboard becomes the working surface. Similarly, an on-demand UI isn’t “the output of the space”—it’s an extension of it. The interface that appears is tied to the structure of the task decomposition happening in conversation.

Task decomposition becomes UI structure, not a hidden step

The paper also leans on mixed-initiative interaction: either the human or AI can propose sub-goals, revise the plan, and contest uncertainty. The big idea is that uncertainty becomes a signal for when the AI should ask rather than act.

Then it gets practical: each sub-goal identified in the space produces corresponding affordances in the UI. So the UI doesn’t just represent the final form—it represents the current state of thinking. That’s why this could change how people learn and navigate complex tasks: you’re not learning a static app designed for someone else, you’re learning the tools as your needs become explicit.

You can find this framing directly in the paper’s overview of the system architecture and its dialogue/UI coupling, described in their structure for https://arxiv.org/abs/2610.02369.

Generative UI Isn’t Enough: Making HCI a Procedural Constraint

A lot of “generative UI” research (and industry tooling) focuses on producing functional interfaces: turn a task description and some data model into code. This paper argues that the missing step is making UI generation incorporate HCI principles every time, not as occasional audits.

Skills are the mechanism: skill.md files the generator can load at runtime

The proposed mechanism is surprisingly simple in spirit: encode HCI design knowledge as skills—machine-readable directories containing a skill.md file plus supporting assets (schemas, scripts, templates, example components).

A skill declares:

  • When it applies (via a triggering description)
  • What the AI must do (procedural steps and constraints)
  • What counts as success (verification criteria)

The key point is that this knowledge becomes declarative, inspectable, version-controlled, and editable—meaning HCI researchers and practitioners can update and audit it without retraining the underlying UI model. The HCI community owns the “how to design well” layer, not just the model weights.

Think of it like linting + compiler checks for interaction design. Today, we often accept that UI code is “generated” but then we validate it with human review. Skills aim to validate it automatically during generation, turning good HCI into an invariant: good UI becomes a property of the process.

Turning Classic HCI Guidance Into Skills That Actually Run

This is where the paper gets concrete: it lists a set of skills that map directly to well-known HCI guidance. The idea isn’t to reinvent the field—it’s to operationalize it.

Nielsen heuristics become generation requirements

heuristics-nielsen.skill.md encodes Nielsen’s 10 usability heuristics as constraints. The paper frames examples like:

  • every action produces visible feedback
  • vocabulary matches the user’s domain (extracted from dialogue)
  • destructive actions are reversible
  • primary actions are recognizable without prior exposure
  • layouts obey a consistent component library

So instead of “run a heuristic evaluation later,” the generator is instructed to follow these rules while composing the UI.

Norman affordances become signifier compliance at the component-library level

affordances-norman.skill.md enforces Norman’s distinction: controls need to visually signal their real action. The paper describes this as a component-level check—because the generator emits structured component code, not raw pixels, the system can reject controls that fail the expected “signifier” pattern before they reach the user.

Example intuition: if you generate a control that doesn’t look pressable as a button, the skill blocks it. That’s a big deal for consistency and perceived reliability.

WCAG becomes a hard constraint, not a best-effort promise

accessibility-wcag.skill.md treats WCAG success criteria as hard constraints. The paper lists automated checks like contrast, focus order, ARIA roles, alt text, keyboard reachability, and screen reader support expectations.

The practical consequence is huge: accessibility isn’t “one more thing” that might be fixed later; it becomes a built-in requirement for any emitted component.

Cognitive load and progressive disclosure become interactive policies

The paper argues that cognitive load should guide a “minimum sufficient surface” UI policy—show only the affordances justified by what the dialogue has established so far.

So progressive disclosure isn’t just a design-time principle. In this system, it’s an artifact of interaction: as soon as your articulated need requires new capabilities, the UI is expanded.

Human-AI Mixed Initiative With Guardrails: How the UI Responds to Uncertainty

Another major theme is mixed-initiative design: when the AI acts, it should do so in a way that preserves user control and reduces risk.

Skills for direct manipulation, reversibility, and confirmation

The paper lists a mixed-initiative.skill.md skill inspired by Horvitz’s principles. While the specific text in the summary focuses on the idea, the described behavior includes instrumentation like:

  • confirmation for actions affecting the world
  • reversibility
  • audit trails
  • visible confidence signals

Meanwhile, direct-manipulation.skill.md encodes Shneiderman’s criteria: visible objects of interest, rapid reversible incremental actions, and “replacement of command syntax with object manipulation.”

Accessibility and personalization aren’t the same thing

The paper also introduces ability-based-personalization.skill.md, which goes beyond WCAG conformance. It references the idea of ability-based accommodations (including adjustments like larger targets, simplified gestures, type scaling, longer time budgets, reduced surface area, and clearer labels).

This is where the on-demand paradigm matters: because each UI is generated for a specific user (based on profile info they articulate), personalization can happen at generation time—not as a one-size-fits-all setting added after the fact.

Skills vs. model-only approaches: what changes in practice?

One way to see the build is that the model is now responsible for interface composition, while skills are responsible for interaction constraints and guarantees.

Aspect Earlier “LLM-generated UI” approach Skills + Space to Think approach
HCI rules Often post-hoc (human review, best-effort prompt guidance) Procedural constraints during generation
Accessibility A checklist or “try to comply” behavior Hard constraints via accessibility-wcag.skill.md
Affordances Can be inconsistent across components Enforced via component-library signifier checks (affordances-norman)
Cognitive load Usually design-time and manual Procedural policy tied to dialogue progress (cognitive-load-progressive-disclosure)
Update speed Slow (retrain or retool model) Faster (edit/replace skills; no retraining required)
Auditability Hard to know “why” a control exists Skills give explainable provenance back to dialogue justification

That shift is the paper’s big bet: the field’s heuristics become something you can inspect, version, and run.

What This Could Look Like at Scale (And What Still Needs Work)

The paper is optimistic about an HCI future where designers, researchers, accessibility advocates, and organizations contribute “runnable craft” artifacts. Their horizon is basically: skills become an ecosystem comparable to component libraries.

Why the skills format could be a turning point

The authors highlight properties that hand design doesn’t have:

  • Uniform: every component gets checked, not only the parts a designer remembered.
  • Inspectable: users (and researchers) can ask why something behaves a certain way, and the system can cite the skill used.
  • Editable: skills are readable and auditable, so updating standards doesn’t require retraining.
  • Verifiable: because skills are text-based procedural specs, they can be compiled into automated tests.

That addresses a long-standing pain point in HCI: proving “good design” is expensive and uncertain even for human-created interfaces.

The hard part: validation and prediction

The paper also doesn’t pretend this is solved. It calls out a real concern: LLM-generated code and UI are hard to verify, and user interfaces resist formal proof of usability the way code can be formally tested.

So future work needs:

  • formal validation criteria for interaction quality
  • machine-interpretable heuristic and accessibility checks
  • predictive models for cognitive load and learnability
  • “verification skills” that act alongside generation skills

It even suggests that over time, the validation toolkit may sit as a parallel knowledge base to generation skills.

A realistic next step: open-sourcing an initial skills library

The research team plans to open-source an initial skills library and apply it within their Space to Think research, aiming to enable on-demand UIs that externalize reasoning in human-AI interaction. In other words: they’re not just proposing an idea—they’re planning to ship the artifact that makes it testable.

If you want a quick mental model: generation skills are “how to build,” while verification skills are “how to check.” The long-term vision is for HCI publications to ship with runnable skill artifacts, not just PDF persuasion.

Key Takeaways

  • HCI is moving toward on-demand UI: LLM systems can generate executable interfaces, but the next step is generating them well.
  • The paper introduces a Space to Think: a shared, structured cognitive workspace where dialogue decomposition directly shapes the UI that appears.
  • The core mechanism is skills (skill.md files): machine-readable, inspectable, version-controlled HCI knowledge loaded at runtime.
  • Classic HCI guidance becomes procedural constraints, including:
    • heuristics-nielsen.skill.md (usability heuristics)
    • affordances-norman.skill.md (signifier/affordance compliance)
    • accessibility-wcag.skill.md (WCAG checks as hard constraints)
    • cognitive-load-progressive-disclosure.skill.md (minimum sufficient surface area)
    • mixed-initiative.skill.md and direct-manipulation.skill.md (control, reversibility, visibility)
  • This reframes “good UX” from a post-launch audit into a property of the generation process—more consistent, faster to update, and easier to verify.
  • The remaining challenge is validation: we still need better automated and predictive methods for learnability, cognitive load, and interaction quality—likely via future “verification skills.”

Sources Used

This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:

Where To Go Next

SKILL.md vs AGENTS.md: Agent Skills vs Custom Instructions 2026

Claude Code Skills vs OpenAI Codex Skills: SKILL.md Comparison 2026

How to Use AI Agent Skills in 2026: The Complete Guide

Browse the free Prompt Database or tune your own prompts with the Prompt Optimizer.

Frequently Asked Questions

Limited Time Offer

Unlock the full power of AI.

Ship better work in less time. No limits, no ads, no roadblocks.

1ST MONTH FREE Basic or Pro Plan
Code: FREE
Full AI Labs access
Unlimited Prompt Builder*
500+ Writing Assistant uses
Unlimited Humanizer
Unlimited private folders
Priority support & early releases
Cancel anytime • 10,000+ members
*Fair usage applies on unlimited features to prevent abuse.