CLEAR-Med: Dual-Agent SQL Validation for Clinical Tables

AI can sound confident with the wrong clinical number. CLEAR-Med fixes that by running generated SQL and having a second agent validate cohort, denominators, and calculations—fail-closed instead of guessing. See how this dual-agent approach targets auditable clinical table analysis.
The finding CLEAR-Med ensures clinical SQL outputs are validated against cohort and denominator logic, abstaining when checks fail.
The method A dual-agent pipeline separates SQL generation/execution from independent validation using deterministic checks plus model-based review.
The use case It supports faster, eligibility-aware subgroup summaries from clinical registries while maintaining auditability of the resulting numbers.
1st MONTH FREE Basic or Pro • code FREE
Claim Offer

The Short Answer

CLEAR-Med validates AI clinical outputs by generating and running SQL in one agent, then having a second agent verify cohort fidelity and denominator consistency before accepting the result. If validation fails, it abstains instead of guessing.

For clinical data workflows, this turns “AI numbers” into an auditable evidence chain tied to executable operations over governed clinical tables. It can speed up draft subgroup or cohort summaries while keeping the numeric outcome verifiable.

The caveat is that the system is intentionally fail-closed: results may be withheld when checks can’t confirm cohort logic or calculation validity, so you may need a repair/iteration path.

CLEAR-Med: Dual-Agent SQL Validation for Clinical Tables

Introduction

If you’ve ever seen an AI system answer a clinical question with a confident-looking number, you’ve probably also wondered: where did that number really come from? New research from Dehkalani et al. tackles exactly this problem for structured clinical data—where the “right” answer isn’t just plausible language, but the output of a specific cohort definition, denominator, and calculation.

The paper introduces CLEAR-Med (Clinical Enhanced Analysis and Retrieval System for Medicine), a dual-agent framework that treats the database as the numerical source of record. One agent generates and runs SQL; a second agent independently validates the draft response using both deterministic checks (like “does the denominator exist and match the cohort?”) and an additional model-based review. The big idea is simple: no execution, no trusted number—and if validation fails, the system abstains instead of guessing.

They demonstrate the approach using neonatal hypoxic-ischemic encephalopathy (HIE), built from two multicenter Neonatal Research Network trials across 21 sites, totaling 532 de-identified infants and ~1,300 variables per record. That’s a data setting where cohort fidelity is everything: mess up the selection rule, and even correct arithmetic becomes clinically misleading.

Why This Matters

This work is significant right now because healthcare is actively moving toward AI-assisted analysis and trial-like research workflows—but a lot of today’s “AI answers” still behave like a highly fluent storyteller rather than a verifiable calculator. CLEAR-Med pushes back on that by forcing the answer to be grounded in executable operations over governed clinical tables. In other words, it doesn’t just ask the model to sound right; it asks it to prove that the number follows the table.

A realistic scenario where this could be applied TODAY: imagine a research coordinator at a hospital trying to rapidly generate an eligibility-aware subgroup summary from an internal registry (e.g., “percentage of infants meeting condition X with outcome Y, stratified by treatment”). Today, that often means manual SQL or spreadsheets—slow, error-prone, and hard to audit. CLEAR-Med’s approach could accelerate the drafting (“here’s the SQL and result”), while the validation stage checks cohort and denominator fidelity, arithmetic sanity, and missingness—then either accepts, repairs once, or abstains. That’s the kind of pipeline that fits real operational needs: speed without sacrificing traceability.

How does it compare to previous AI research? Prior work in clinical NLP often focuses on generating answers from text (summarization, dialogue, retrieval-augmented generation). There’s also a big line of “text-to-SQL” systems that convert questions into executable queries. CLEAR-Med builds on that by treating SQL generation as only the first stage, not the endpoint. The novelty isn’t “can it write SQL?”—it’s “can it ensure that the emitted number is the outcome of validated execution over the intended cohort,” with an explicit fail-closed pathway. If you want the practical difference in one sentence: CLEAR-Med turns an AI analysis into an auditable evidence chain with controlled decision-making.

How CLEAR-Med Keeps Clinical Numbers Grounded (Not Just Worded Well)

CLEAR-Med is designed around a problem that’s easy to underestimate: in clinical tables, the value is determined by more than the dataset. It depends on:

  • the exact columns used,
  • the joins (how tables connect),
  • the filters (how cohorts are selected),
  • the aggregation operator (count, mean, odds ratio, etc.),
  • and the denominator (who is actually eligible to be counted).

A model can produce a “plausible” number while still being wrong—because fluent language doesn’t reveal whether the value came from the intended population. CLEAR-Med’s approach is basically: don’t let the model’s plausibility be the authority. The authority is the executed SQL against a fixed database.

The HIE Dataset: Why Cohort Fidelity Is Hard Here

The paper’s testbed is neonatal hypoxic-ischemic encephalopathy (HIE). HIE involves impaired oxygen and blood supply around birth and is associated with brain injury visible on MRI. The research data combine:

  • treatment assignments,
  • physiological measurements,
  • imaging variables,
  • and longitudinal outcomes.

They use a harmonized dataset drawn from two multicenter trials spanning 21 sites. The analysis table contains:

  • 532 participant rows
  • ~1,300 variables per record
  • de-identified identifiers that follow Neonatal Research Network (NRN) definitions

That combination matters because it makes the denominator and cohort-selection logic non-trivial—precisely the areas where ungrounded AI answers are most likely to drift.

The Dual-Agent Pipeline: SQL First, Then Independent Validation (With Abstention)

CLEAR-Med’s design is a dual-agent system with clear separation of responsibilities. Think of it like two people double-checking a math problem:

  1. Invocation Agent: does the first pass—writes and runs SQL, then drafts an answer.
  2. Clinical Validation Agent: does the second pass—checks what happened, and decides whether to accept, repair, or abstain.

Stage 1: The Invocation Agent Generates, Executes, and Drafts

The Invocation Agent takes a natural-language question (e.g., “association between cord blood pH and neurodevelopmental outcome”) and turns it into an executable pipeline with four steps:

  1. Parsing and planning: extract relevant tables and fields, read the schema.
  2. SQL generation: produce schema-guided SQL.
  3. SQL execution: run it on an SQLite database.
  4. Draft response: convert the raw result into natural language.

A key detail: they convert the source Excel data into an SQLite database so the system can rely on a real relational interface. SQLite is lightweight and schema-enforcing, which helps prevent “floating variables” problems where an AI guesses column names or types.

To connect the LLM tool-use mechanics, the paper describes use of LangChain’s SQLDatabaseToolkit for generating and executing SQL. The executed database result is kept as evidence, not discarded after drafting.

Stage 2: The Clinical Validation Agent Checks Evidence Like a Gatekeeper

The Clinical Validation Agent doesn’t just re-summarize the draft. It evaluates a complete evidence record consisting of:

  • the natural language question,
  • the executed SQL,
  • the raw executed database result,
  • and the Invocation Agent draft, plus deterministic-check outputs.

Validation has two layers:

(1) Deterministic checks (rule-based, deterministic)

These include checks for:

  • read-only SQL,
  • complete provenance (the system must keep what was executed),
  • cohort and column fidelity,
  • non-missing denominators,
  • arithmetic sanity,
  • declared ranges.

This is where most “quiet failures” can be caught: the draft might sound right, but the cohort or denominator might be wrong, or a required value might be missing.

(2) Independent model-based review (operational separation)

Then a separately invoked cross-provider language model reviews the question, SQL/result/draft, deterministic-check reports, and a prospectively authored clinical constraint card. Importantly, it must return one structured decision:

  • accept
  • repair
  • abstain

The independent aspect is operational: the validator is not the model that generated and executed the SQL.

Repair Is Limited: One Shot, Then Fail-Closed Abstention

A particularly important safety behavior is that CLEAR-Med allows at most one repair. If the validation agent asks for repair:

  • it issues a structured instruction (not a replacement number),
  • the Invocation Agent generates new SQL and reruns it,
  • and the whole validation flow repeats.

If repair fails again, or evidence is missing/malformed/unavailable, the system abstains (denoted as output ⊥). That’s a fail-closed design: it prefers no answer to a wrong-but-confident answer.

What CLEAR-Med Actually Guarantees (And What It Doesn’t)

CLEAR-Med’s paper includes a formal argument that the system’s acceptance is tied to encoded, executable properties. That means if the deterministic checker set and validation flow establish property satisfaction, then acceptance implies the emitted answer satisfies those properties.

But the guarantee is conditional: it doesn’t claim omniscient clinical correctness for everything. Instead:

  • If a clinical claim is encoded as an executable check (like “denominator must be non-missing,” “SQL must be faithful,” “arithmetic must match constraints”), then the system can guarantee that property for accepted outputs.
  • If clinical reasoning is outside encoded checks (e.g., nuanced individualized prognosis), the guarantee doesn’t cover it.

This distinction is crucial in healthcare—because otherwise people start treating AI systems like magic oracles.

How the System Decides: Accept vs Repair vs Abstain

At a high level, CLEAR-Med maps each attempted evidence record into:

  • accept (emit draft as final, after passing all checks)
  • repair (trigger one more SQL generation + execution + validation)
  • abstain (emit nothing / explicit refusal)

That design creates a clear boundary between what the system can prove and what it only suspects.

Performance: SQL Quality, Scalability Limits, and Oracle-Scored Accuracy

CLEAR-Med evaluates both the SQL-generation capability and the end-to-end grounded analysis performance, using an oracle-based scoring strategy.

Choosing the SQL-Generation Configuration: Speed vs Success

The paper compares two OpenAI configurations for the Invocation Agent on the SQL-generation task:

Configuration tested Outcome Key result
ChatGPT Thinking (OpenAI) Successful SQL generation No errors; usable SQL produced
ChatGPT Instant (OpenAI) SQL errors and no valid SQL 3 errors; iteration limit hit; no valid SQL

Runtime also differs: SQL-generation dominated time, with ChatGPT Instant taking 29.23s vs 167.44s for ChatGPT Thinking. But since ChatGPT Instant failed to produce valid SQL in the reported trials, they prioritize success over speed and keep ChatGPT Thinking (OpenAI) for the development configuration.

Can CLEAR-Med Scale to Large Structured Tables?

This is where CLEAR-Med distinguishes itself from prompt-only approaches. For the “big table” test, they compare interfaces under controlled constraints (timeouts, GPU memory limits, success criteria).

CLEAR-Med keeps the relational database as-is in SQLite and queries it. Other models in the scalability screen serialize tables into context (which is much harder to scale).

They test six nominal table configurations (row × column sizes), including:

  • 50×50, 50×100, 100×50, 100×100, 50×1000, and 500×1300
  • The largest configuration corresponds to the nominal HIE table size used in their “scalability” setup.

CLEAR-Med completed 6 of 6 fixed configurations, including the largest 500×1300 case, under the tested SQL-tool interface. For the comparison systems, success was lower (some reached GPU memory limits, others timed out), with fewer than 6 successes in their grid.

The paper also reports GPU memory behavior: CLEAR-Med used about 19.71GB of 39.56GB available GPU memory for the largest test, whereas several comparison models hit CUDA out-of-memory beyond smaller sizes.

Oracle-Scored Accuracy on Real HIE Questions

For numeric correctness, they use 25 natural-language clinical query templates covering:

  • counts,
  • percentages,
  • group comparisons,
  • odds ratios,
  • multi-condition filters.

Each question was posed five times (so 125 responses per condition and 250 total responses across two methods). The comparator baseline was an ungrounded ChatGPT approach answering directly from an in-context table (no database tools).

They score responses against deterministic numerical oracles, with tolerances such as:

  • counts/percentages within 5%
  • percentage-point differences within 2 points
  • odds ratios within 10%

Results (Invocation Agent accuracy vs ungrounded baseline):

  • Invocation Agent: 83/125 = 66.4% correct
  • Baseline: 15/125 = 12.0% correct
  • Reported paired improvement: +54.4 percentage points (95% CI 36.8–72.0)

They also report that for 12 percentage-valued queries (60 responses per condition), the Invocation Agent was within tolerance on 39/60 vs 4/60 for the baseline.

This is strong evidence that the database-grounded workflow (SQL generation + execution + validated evidence) is doing real work—rather than the LLM simply “guessing until it sounds right.”

Practical Implications: What This Means for People Building Clinical AI Workflows

CLEAR-Med isn’t just a new model; it’s a workflow pattern that’s especially relevant for structured clinical analysis.

If You Build Clinical Analytics, You Need an Evidence Chain

A lot of systems blur the line between “answering” and “computing.” CLEAR-Med’s separation makes it harder to hide mistakes.

Instead of returning a number generated from language alone, it returns a number that can be traced back to:

  • the executed SQL,
  • the raw database result,
  • the deterministic checks,
  • and the validation decision.

That’s the kind of provenance that can matter for auditing, governance, and reproducibility—especially in retrospective research settings.

This Is Not Just RAG, and That’s the Point

The paper contrasts CLEAR-Med’s approach with passage-based retrieval augmentation (RAG). RAG retrieves text passages and then synthesizes. Evidence is still “text,” and the system must interpret it.

But for structured aggregates, the result is determined by the executed relational operations. So CLEAR-Med treats execution as evidence, not something to narrate around.

Transfer Requires Dataset-Specific Validation

The paper is clear that adapting to other clinical datasets likely requires:

  • dataset-specific constraint cards,
  • schema adaptation,
  • and validation that encoded checks match local definitions.

So this isn’t a plug-and-play universal clinical calculator. It’s a repeatable architecture for systems where you can encode key properties and rely on faithful database execution.

Key Takeaways

  • CLEAR-Med provides grounded clinical table analysis by splitting the system into an Invocation Agent (generate + execute SQL) and a Clinical Validation Agent (deterministic checks + independent review).
  • The database is the numerical source of record: the executed SQL and raw result are preserved as evidence rather than being replaced by fluent language.
  • Fail-closed behavior matters: the system allows at most one repair; if validation can’t be satisfied, it abstains instead of emitting a potentially wrong number.
  • On a harmonized HIE dataset (532 infants, ~1,300 variables), the Invocation Agent achieved 66.4% correct answers vs 12.0% for an ungrounded baseline across 25 oracle-scored query templates (each asked 5 times).
  • Scalability is operational: CLEAR-Med completed 6/6 nominal table configurations under the tested SQL-tool interface, including the largest 500×1300 setup, while several other approaches hit memory limits or timeouts.
  • For readers building real clinical AI workflows: this is a strong blueprint for how to turn “AI answers” into auditable, verifiable computations—especially when cohort definitions, denominators, and aggregation logic are non-negotiable.

If you want, tell me who your target audience is (clinician? data scientist? trial operations?) and I can rewrite this post with a more tailored angle and examples.

Sources Used

This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:

Where To Go Next

Agentic Web Search: How Chatbots Decide, Query, and Attribute Facts

AI chatbots vs experts: can they retrieve clinical studies?

Embodied AI Agents That Can Be Trusted After Every Move

Browse the free Prompt Database or tune your own prompts with the Prompt Optimizer.

Frequently Asked Questions

Limited Time Offer

Unlock the full power of AI.

Ship better work in less time. No limits, no ads, no roadblocks.

1ST MONTH FREE Basic or Pro Plan
Code: FREE
Full AI Labs access
Unlimited Prompt Builder*
500+ Writing Assistant uses
Unlimited Humanizer
Unlimited private folders
Priority support & early releases
Cancel anytime • 10,000+ members
*Fair usage applies on unlimited features to prevent abuse.