Prompt-Content Alignment Beats “AI SEO” Fixes for LLM Citations

“AI SEO” advice for LLM citations often ignores what the data says. New production research across 2.1M citations finds prompt-content alignment is the dominant predictor—while common AEO checklists shrink once domain authority is controlled. Here’s what to do instead.
The finding Prompt-content alignment is the strongest page-level driver of LLM citation frequency across production engines.
The method The study uses real citation logs from four engines over six months and tests 60+ page features with domain fixed effects and multiple robustness checks.
The caveat Common AEO checklist items can shrink or vanish after controlling for domain authority, so treat schema/speed wins as conditional.
1st MONTH FREE Basic or Pro • code FREE
Claim Offer

The Short Answer

Prompt-content alignment is the dominant predictor of how often LLM engines cite a web page in production, outpacing common “AI SEO” checklist factors. The study tests 2.1M citations across four commercial engines and finds page-to-prompt language/meaning overlap matters most.

So what: instead of only adding FAQs, schema, and performance tweaks, prioritize updating pages so their wording matches the prompt language your buyers actually use—then measure citation lift.

Caveat: FAQ blocks, structured data, and Core Web Vitals show positive pooled effects but can reverse or collapse to zero once domain-level authority is controlled, so don’t assume checklist items will reliably drive citations.

Prompt-Content Alignment Beats “AI SEO” Fixes for LLM Citations

Introduction

If you’ve spent time trying to “optimize for AI” (AEO, GEO, LLM SEO—whatever your team calls it), you’ve probably seen the same checklist everywhere: add FAQs, sprinkle schema markup, improve page speed, make sure Core Web Vitals look good. But here’s the uncomfortable part: the evidence behind those ideas has mostly come from small case studies and pattern matching, not large-scale, confound-controlled measurements of what actually drives citation frequency in production systems.

New research from Moore & Dunne takes a big step toward answering the real question practitioners care about: what page-level factors actually predict how often LLM engines cite a web page alongside their answers? Instead of toy examples or synthetic prompts, they observe production behavior over six months across four commercial engines—ChatGPT, Claude, Google AI Overviews, and Gemini—using a corpus of roughly 2.1 million citations tied to the prompts that triggered them. Then they crawl 10,042 distinct cited pages and measure 60+ page features (alignment, structure, recency, and performance signals).

The headline result is striking: the dominant predictor isn’t schema, isn’t speed, and isn’t “AI-ready formatting” in the usual sense. It’s prompt-content alignment—basically, how well the page’s wording and meaning matches the language buyers actually use in prompts. And when the authors account for domain-level authority properly, several popular AEO checklist items shrink, flip, or vanish. Let’s unpack what they did and what it means for anyone trying to earn citations from modern LLM search-and-answer workflows.

Why This Matters

This is significant right now because “LLM citations” are quickly becoming part of how B2B buyers do research. When an engine cites a page, it’s not just providing an answer—it’s routing attention. If you’re a SaaS team trying to win deals, losing citation visibility can quietly mean losing the top-of-funnel conversation before your sales team ever gets a chance.

The practical scenario is simple: imagine you manage a documentation + resources site for a SaaS product. Your growth team reads that “adding FAQs and structured data improves citations,” so you roll out schema markup and FAQ blocks across high-intent pages. Months later, citations might look up… but the key issue is whether that increase was caused by your changes—or whether the pages that were already “winning” also happened to be on domains with stronger existing authority. This paper shows why that confusion is almost guaranteed if you look at pooled correlations without proper controls.

Also, the work builds on earlier AI research in a way that matters for decision-making. Prior studies often examined attribution quality (does the citation support the claim?) or verifiability audits, frequently using synthetic prompts or small samples. This paper flips the direction: it asks which properties predict citation frequency across a large set of real-world cited pages, while using a multi-method consensus approach to reduce the chances of “finding effects” that are actually artifacts of confounding. If you’ve ever felt stuck between “everyone says do schema” and “we don’t see consistent wins,” this is the kind of evidence you’ve been waiting for.

What the Researchers Actually Measured: Citations, Pages, and Confounding Controls

The study is observational, but it’s not casual observational work. It uses two joined datasets:

  1. A production citation corpus
    Collected from four production engines over ~six months:

    • ChatGPT
    • Claude
    • Google AI Overviews
    • Gemini

    The dataset includes about 2.1×10^6 citation events, represented as tuples like ⟨prompt, engine, cited URL, position⟩. It also records the prompt’s funnel stage label and the citation position within the engine response (position matters for “share of voice” analysis).

  2. A page corpus built from the cited URLs
    They crawl 10,042 distinct cited URLs, but only after filtering and sampling so the dataset doesn’t get dominated by a few brands or categories. They also require each URL be cited at least twice, because single-citation pages are basically noise for prediction.

Then they extract 60+ page features across five families:
- Alignment: lexical overlap (Jaccard-style token overlap) and semantic similarity (computed at title/intro/best-paragraph levels)
- Structural: things like word count, outbound links, FAQ/TLDR blocks, author bio presence, and schema flags
- Recency: page age and publication-date flags
- Infrastructure: real-user Core Web Vitals (LCP/INP/CLS) and Lighthouse-style synthetic scores
- Page type: labels inferred from URL patterns and content cues (article, comparison, how-to, listicle, pricing, etc.)

The confounding problem they explicitly tackled

If you’ve worked in SEO/marketing analytics, you’ve seen this trap: big brands tend to rank/cite more, and big brands also tend to have more resources—so speed, formatting, structure, and content quality are correlated with “brand authority.” If you just pool everything, you get misleading “effects” that are actually brand-level differences.

Their key move is to use domain fixed effects in the main analysis (and test lots of robustness). That means they ask: within the same domain, do pages with certain features get cited more often than other pages on that same domain? That’s a much closer approximation to “what you can influence” than pooled correlations.

Who had “agency” in the test

They also distinguish between:
- ownbrand pages (the domain matching the workspace’s brand)
- competitor
brand pages
- truethirdparty pages

For the “within-domain” page-level analysis, they focus on a brand-controlled subset (own_brand + competitor_brand) with N=4,015 pages across D=297 domains—the population where page-level changes would be meaningful.

The 1 Big Driver: Prompt-Content Alignment Predicts Citation Frequency

Let’s get to the best part: the dominant predictor that survives every check.

Prompt-content alignment is the strongest page-level predictor

The paper defines prompt-content alignment using a Jaccard overlap between page tokens and the full workspace prompt corpus—importantly, including both citing and non-citing prompts to avoid circularity.

In their headline mixed-effects model, alignment has:
- β = +0.37
- 95% CI: [+0.33, +0.41]
- q ≈ 10^-73 (FDR-adjusted)

What does that mean practically? Their interpretation is: a one-standard-deviation increase in workspace-level alignment corresponds to about a +0.37 log2 unit increase in citation count (roughly ~30% higher citation count at the geometric mean).

And it’s not just a “model artifact” (robustness checks all agree)

The authors didn’t stop at one regression. They required their headline effect to pass a nine-method consensus protocol. For alignment, it passes:
- Stability-selection: selected in 200/200 bootstrap samples (i.e., selection probability Π* = 1.0)
- Still significant after double machine learning (DML): their orthogonalized partial effect estimate is θ_align = +0.35 (with very small SE)
- Non-linear checks: GAM smooth is monotone increasing across the full data range (no weird saturation)
- Sensitivity splits: sign and magnitude preserved across five subsets
- Domain leave-one-out: still holds when dropping high-volume domains
- Temporal hold-out replication: the effect remains positive when the model is trained on earlier data and refit using a later partition

Interestingly, the magnitude shifts between the training/holdout partitions (they report attenuation in the hold period), but the direction stays positive and the effect remains significantly above zero.

Where on the page engines seem to “pull” citations from

Alignment isn’t just about what the page says—it also shows up in where citations come from. They compute a “best matching paragraph depth” (top is 0, bottom is 1) for cited pages and find:
- Median depth is 0.36
- The distribution rejects a uniform null at p < 10^-20

So engines preferentially cite paragraphs in the top third of pages (at least when there is at least one citation).

A practical takeaway: “Content positioning” beats “content formatting”

If you’re building an AEO plan, this finding reframes the work. The alignment lever points to something closer to voice-of-customer content design than classic “page optimization”:

  • Harvest language from first-party transcripts, sales calls, discovery docs, and support tickets
  • Write pages whose lexical and semantic substance mirrors the prompts buyers actually use
  • Put the most relevant “answerable” material earlier on the page so the engine lands there when it retrieves

In other words: engines reward pages that sound like the question they’re trying to answer—not pages that merely have the right metadata.

This is where the paper delivers a big reality check.

The AEO checklist effects mostly collapse after adding domain fixed effects

The standard AEO playbook includes things like:
- FAQ blocks
- TLDR/BLUF blocks
- author bios
- schema markup
- Core Web Vitals / speed

In pooled analysis (ignoring domain), some of these look helpful. But when the authors introduce domain fixed effects and focus on within-domain variation, those effects:
- shrink toward zero
- reverse in sign
- or become non-significant

They find (on true third-party pages) small positive effects for:
- FAQ: β = +0.07
- TLDR/BLUF: β = +0.05
- author bio: β = +0.02 (and it collapses after controlling for content-depth)

For real-user Core Web Vitals factors, they find no significant effects after domain control. Synthetic Lighthouse metrics similarly look null.

Simpson’s paradox: the “speed” story is mostly confounded

The clearest example: speed.

Pooled results can suggest that faster pages are cited less (or other odd relationships), but once domain fixed effects are applied, these associations flip/collapse across every speed feature. The paper explicitly calls this Simpson’s paradox with practical consequences for the AEO literature.

Interpretation: large incumbents often have more citations and also have systematic differences in page speed, structure, and formatting. Without within-domain controls, pooled regression confuses “incumbent advantage” with “your speed improvements cause citation gains.”

Comparison: what changes when you control for domain?

Here’s the conceptual pattern the paper reports (not a new measurement, but the relationship they describe):

AEO signal family Pooled association (ignoring domain) After domain fixed effects What that means
Prompt-content alignment Strong positive Still strong positive Page-level work can matter
Structured/format signals (FAQ/TLDR/author bio) Often positive Small or collapses to ~0 Effects may be driven by domain authority
Speed / Core Web Vitals Can look meaningful Collapses to null Likely confounded by who already wins
Synthetic Lighthouse metrics Can show patterns Null Less evidence once confounding is handled

Domain-Level “AI Authority” Is Real—But It’s Not Purely Off-Page Magic

So are we saying “off-page doesn’t matter”? Not exactly. The paper finds that domain-level AI authority dominates within their variance decomposition.

Domain authority explains more than the strongest non-alignment page feature

In their SHAP analysis of a gradient-boosted model, domain authority is among the top contributors. They quantify it with mean absolute SHAP values:
- domain authority: 0.381
- strongest non-alignment page-level feature: 0.060

The authors add an important nuance: this is not a perfectly “like-for-like” comparison because domain-level features are shared across pages, meaning they naturally absorb more between-domain variance. So the comparison is an upper bound rather than a precise causal estimate of effect size.

Most citation variance lives between domains, not within domains

In their fixed-effects regression framework, once they account for domain effects (α_d), within-domain residual variance attributable to page-level differences is about 14% of total variance (they report this approximately).

This can sound like: “page-level optimization won’t matter.” But the authors argue that framing can be misleading because:
- domain authority is not exogenous (it’s built from years of content strategy, links, brand activity, and AI visibility)
- improving alignment-first content is a pathway that can increase domain authority over time
- and the page-level increment after accounting for domain effects is still meaningful commercially

They also report an R^2 comparison:
- domain-only model: R² = 0.320
- full feature set: R² = 0.450
- page-level increment: ΔR² = 0.130

That’s not “nothing,” even if domain is the biggest piece in a snapshot model.

Important implication: you can’t just “optimize pages”—you’re also shaping the domain’s citation profile

If you’re a team deciding where to spend engineering time, this suggests a strategy:
- Use alignment-first content work to improve citation likelihood at the page level
- Recognize that domain authority will still dominate outcomes, so align your measurement and incentives at both levels (page assets and domain-level “visibility accumulation”)

Cross-Engine Reality: Citation Behavior Isn’t the Same Everywhere

One reason “AI SEO” advice is so messy is that engines behave differently. The paper shows substantial cross-engine heterogeneity in citation dynamics.

They report (examples):
- Citation age distributions differ by engine
Median citation age ranges from about 5–8 months depending on engine (Claude reaches 60% cumulative share within six months, while ChatGPT takes about twelve).
- Brand share-of-voice varies widely
They report brand-controlled URL share differences from 13.8% to 39.3% across engines, and even wider differences under position-weighting (e.g., ChatGPT gives 53% position-weighted share to brand-controlled pages vs 24% on Gemini).
- Some sources are effectively engine-specific
In a third-party analysis involving LinkedIn citations: out of 23,908 third-party LinkedIn citations, 23,097 (96.6%) come from Google AI Overviews alone, and Gemini contributes zero.

This matters because it means there may not be a single “universal AEO checklist.” Engines may retrieve and rank content differently—even when your domain and pages are identical.

Key Takeaways

  • Prompt-content alignment is the dominant within-domain page-level predictor of LLM citation frequency. In the headline model, β ≈ +0.37 with a tight CI and extremely small FDR-adjusted q ≈ 10^-73.
  • Many popular AEO checklist ideas don’t hold up once domain authority is controlled. Effects for FAQ/TLDR/author bio and especially speed/Core Web Vitals largely collapse to zero or become unstable after domain fixed effects (Simpson’s paradox).
  • Domain-level “AI authority” matters a lot, with SHAP indicating domain features dominate over the strongest page-level non-alignment feature—but domain authority is itself shaped by long-term content and visibility efforts.
  • Citations come from earlier page sections. The median “best matching paragraph” depth is 0.36, indicating citations concentrate in the top third of pages.
  • Don’t expect one-size-fits-all advice across engines. Citation recency patterns and brand share-of-voice differ by up to an order of magnitude across engines.
  • What you should do today: prioritize voice-of-customer alignment (matching buyer prompt language and meaning), then use structural/speed improvements as hygiene—without assuming they’ll produce citation lifts independent of brand authority.
  • What this means for the future: observational studies about LLM citation behavior should use within-domain or fixed-effect-style controls; pooled correlations will likely keep repeating confounded “wins.”

If you want, tell me what kind of site you’re optimizing (docs? pricing? comparisons? thought leadership?) and which engine(s) matter most for your business. I can translate the alignment-first insight into a concrete content and measurement plan that your team can run this month.

Sources Used

This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:

Where To Go Next

Future-Proof Your Prompts: How Language Turns Into Action

Design AI workflows that “fit” professional tasks, not just prompts

Citations That Gatekeep Mental Health Answers (and How to Audit Them)

Browse the free Prompt Database or tune your own prompts with the Prompt Optimizer.

Frequently Asked Questions

Limited Time Offer

Unlock the full power of AI.

Ship better work in less time. No limits, no ads, no roadblocks.

1ST MONTH FREE Basic or Pro Plan
Code: FREE
Full AI Labs access
Unlimited Prompt Builder*
500+ Writing Assistant uses
Unlimited Humanizer
Unlimited private folders
Priority support & early releases
Cancel anytime • 10,000+ members
*Fair usage applies on unlimited features to prevent abuse.