Spoken Language and Exit Location Quietly Choose Your “Local” SaaS Shortlist

Your AI visibility tests may be measuring the wrong market. New research finds query language and exit IP can decide which local SaaS suppliers are eligible—before reasoning—so your shortlist can change without you realizing.
The finding Query language and exit IP can determine the market used for commercial recommendations before the model reasons about products.
The risk If you test only in English from one location, your AI shortlist may exclude local suppliers and create misleading “no visibility” conclusions.
The fix Test the same SaaS recommendations across multiple languages and exit locations/IPs, and compare the top recommended supplier sets across runs.
1st MONTH FREE Basic or Pro • code FREE
Claim Offer

The Short Answer

A generative interface can choose the market before it reasons: query language and exit IP/location can swap which local SaaS suppliers are eligible for the shortlist. That means your shortlist may reflect prompt language routing, not local demand.

Practically, your AI visibility audit should vary both the prompt language and the exit location/IP to confirm you’re measuring the market you actually compete in. Otherwise, you may miss local competitors due to filtering.

Caveat: even “the same prompt” can produce different top recommendations, so you should compare winners across runs and contexts (e.g., logged-out web sessions and API calls) rather than trusting a single screenshot.

Spoken Language and Exit Location Quietly Choose Your “Local” SaaS Shortlist

If you ask a generative search interface for a product recommendation, you expect the model to “think” about options. But new research from Żatuchin’s paper on arXiv suggests something more subtle: the language you ask in—and the market your connection “exits” from—can decide which country’s suppliers even get considered before the model does its reasoning.

In controlled tests (234 usable runs) against both the logged-out ChatGPT web interface and the OpenAI API on 29–30 August 2026, the researchers asked the same kind of commercial question—in different languages and from different exit locations. The striking part: the system didn’t just reorder results; it nearly substituted whole supplier sets, effectively swapping which national market the shortlist comes from.

So if you’ve been auditing AI visibility “in English,” or trying to estimate how your brand appears in recommendations, this matters a lot. The instrument may be measuring a market you don’t actually compete in.

Why This Matters

This is significant right now because businesses are actively building dashboards and audits that treat AI recommendations like a predictable “output layer.” Many teams measure visibility by prompting in English and screenshotting what appears. This research challenges a key assumption: language isn’t just translation—it can act like a gate that controls whether a local supplier set is even eligible to show up.

Imagine a SaaS company selling to Estonia (or Germany, Norway, Türkiye). Today, their marketing team may run an “AI visibility test” by asking: “Recommend accounting software for freelancers.” They do it in English, from their office connection, and collect one answer as evidence. If their prompt gets treated as belonging to a different market, they may conclude “nobody local shows up”—when what’s actually happening is the local market’s vendors are filtered out due to prompt language logic and exit-country routing.

And this builds directly on earlier AI auditing work—just downstream in a more practical place. Prior studies often measured what models know or how answers reflect cultural defaults across languages. Here, the measurement is commercial and operational: which firms a user sees in the shortlist. That’s a different kind of problem. Knowing “facts” isn’t the same as getting recommended.

What the Researchers Actually Measured (and Why One Answer Is Not Enough)

The paper’s core claim is about separability: query language and exit IP are treated as independent factors, and they influence different parts of the system’s behavior.

The experimental design in plain English

The researchers ran eleven “cells” (combinations of query language + exit country/IP + interface type) with 234 usable runs total. Each cell generally had six identical runs per prompt (so you can see instability), using either:

  • the logged-out ChatGPT web interface, or
  • the OpenAI API (gpt-5.6-terra was pinned for the API arm)

They used fresh browser contexts for each run and made sure prompts weren’t influenced by account history (logged out on the web).

They also targeted two software categories:

  • Treatment category: accounting software for freelancers
    (chosen because each country in the sample has domestic suppliers)
  • Control category: project management software for a small marketing agency
    (chosen expecting it to lack strong domestic incumbents—this expectation turned out to be wrong)

Each answer was scored by identifying the top recommendation the interface displayed. The system often marked winners via UI cues (web table row, medal glyph, or phrases like “best overall”) or API formatting (markdown emphasis).

Key measurement warning: “Same prompt” still varies

A crucial constraint: even identical prompts often produce different top recommendations.

Across multiple arms, the top recommendation changed on 4 of 6 prompts (example given for Berlin interface, Oslo interface, API with search enabled, and API with search disabled). In other words, the system isn’t deterministic. The paper emphasizes that a single screenshot is not strong evidence because you might just capture one draw from a distribution.

So the experiment uses repeated runs to detect patterns that survive noise.

Comparison: interface vs API stability (same kind of volatility)

Here’s the high-level stability result the paper reports:

Condition (as used in the study) Top recommendation changed on Notes
Web interface (Berlin) 4 of 6 prompts Instability pattern differs by category
Web interface (Oslo) 4 of 6 prompts Some categories stable within both web arms
API with web_search enabled 4 of 6 prompts Volatility matches web arm
API with web_search disabled 4 of 6 prompts Volatility still present

The takeaway isn’t just “it’s noisy.” It’s: you can’t attribute effects to whether you used the web UI vs the API—both behave similarly unstable in terms of top-choice selection.

How Query Language and Exit IP Split the Job: Eligibility vs Market Choice

The paper’s standout conceptual move is separating two causes that are usually entangled in real life:

  1. Query language: the language your question is written in
  2. Exit IP: what country your connection appears to originate from (the market “treated as the user’s”)

The results point to a clean separation of responsibilities:

  • Language decides the register and whether localization is attempted at all
  • Exit IP decides which national market is treated as the user’s

The “Oslo vs minutes later in English” example

The paper gives a memorable contrast:

  • Asked in Norwegian from Oslo, the interface names Fiken, Tripletex, and Conta (local-ish set) and no global product.
  • Minutes later, asked in English from the same connection, it names FreshBooks, QuickBooks, Wave, and Zoho, with Fiken appearing in only 1 of 6 runs.

That’s not a gentle shift. It’s basically a swap of which supplier set is eligible.

The “Tallinn and Istanbul” example

On the same overall setup, when the connection is treated as Estonian or Turkish:

  • Asking in English from Tallinn and Istanbul led to zero Estonian or zero Turkish suppliers appearing in the answers (reported as 0 of 6 in English cells on those connections).

Meanwhile:
- the Estonian connection returned three Estonian suppliers when asked in Russian (with local suppliers appearing in 4 of 6 for one and 5 of 6 style patterns for others in the paper’s language-dependent results).

Another key replication: language fixed, exit IP moved

The study also tests the opposite direction: hold language fixed and move exit IP.

When the prompt is in Turkish, the interface stays in Turkish in both cases, but the supplier set follows the exit market:

  • From Istanbul exit, it serves Türkiye-specific framing and suppliers.
  • From Berlin exit, it frames in German-market terms (example Turkish text opening with logic like “if you work as a freelancer in Germany”) and returns German suppliers in Turkish.

And they replicate the same diaspora-style pattern with Russian:
- A Russian question asked from the Estonian exit returns Estonian-market localization framing, but no Russian suppliers appear in the Russian-labeled answers.

The paper treats this as strong evidence that exit IP selects the market and language selects whether the system will localize and how.

The Control Category That Nudges the Explanation Toward Regulation

You might wonder: is the reason language changes recommendations simply because local suppliers exist in some markets and not others?

The paper does something smart here: it uses a coded negative control category—project management software. The researchers built a fixed domestic supplier list per market and scored whether any domestic winners appear across the same six runs per cell.

What happened in the control category?

No domestic supplier appeared in any cell, in any language. Meanwhile, accounting (the treatment category) shows near-total substitution in two of four countries.

That mismatch matters. If the main cause were “global suppliers dominate because local suppliers exist,” you’d expect at least some domestic presence in project management where domestic suppliers do exist.

So why does accounting localize and project management not?

The paper’s favored explanation: nationally regulated categories

The authors argue the evidence points to national regulation, not just domestic supply:

  • Accounting for freelancers is legally localized in ways that differ by country (e.g., VAT rules, e-invoicing mandates, national filing formats).
  • The interface (in some English answers for the Estonian connection) even explicitly reasons about these constraints before naming mostly foreign products.

They contrast that with:
- Project management software, which lacks those same national obligation constraints. Different companies in Norway or Germany could use the same tool without breaking local legal requirements.

In short, language doesn’t merely determine “which suppliers are known.” It helps decide whether the recommendation process treats the problem as a local regulatory compliance category that requires a local shortlist.

Near-Total Substitution: When Language Forces Local Lists to Appear (or Vanish)

Here are the results the paper emphasizes as “near-total substitution” and how English behaves as a global signal.

Local supplier appearance depends heavily on language

Across four cells where the query language equals the official language of the exit country, a global supplier appeared in only 1 of 24 runs.

But in the two English cells on the Istanbul and Tallinn connections:
- local suppliers appeared in 0 of 12 runs

That’s the headline: English can make a local supplier set effectively invisible, even when the system appears to know the local context.

Testing the “winner” effect: local language beats English

The paper reports statistical testing using Fisher’s exact test for two specific markets when comparing local-language vs English from the same exit IP:

  • Norway / Fiken: one-sided p = 0.0076
  • Germany / Lexware: one-sided p = 0.0303

They then combine Germany and Norway replications via Fisher’s method:
- combined p = 0.0022

Türkiye and Estonia are described as moving “in the same direction” with larger effects, but they’re not included in that particular combined calculation because of boundary conditions (like zero local suppliers in English alongside strong local dominance in the local language).

A subtle twist: the model can “know” local constraints but still not name local firms

The paper provides an important nuance: sometimes, in English answers from the Estonian connection, the interface explains that it’s the user’s country and discusses local VAT/e-invoicing arguments—yet still names no Estonian supplier in that run.

So even if localization knowledge appears in text, the supplier shortlist can still be pulled from a restricted set. That supports the idea that the system isn’t just translating or reasoning in a localized way; it’s gating market eligibility.

What This Means for Your AI Visibility Measurements (Today, Not Later)

If you do measurement with generative search interfaces, this paper is basically telling you: your measurement protocol might be measuring the wrong market.

Here are the practical implications that jump out:

1) Always record both query language and exit/connection context

The paper’s message is clear: an instrument should report its “egress,” i.e., where the request is exiting from.

If you don’t specify language and exit country, your results can’t be attributed. It’s like comparing search visibility without stating the search engine or region—except the paper suggests the system can behave as if it’s using a different candidate set entirely.

2) Don’t assume “English = neutral”

English in these experiments is not a neutral default. On the Istanbul connection, English returned zero Turkish suppliers, while non-English languages triggered localization.

So if your competitors operate primarily in local languages—or if they rely on local compliance framing—English prompts may consistently hide them.

3) Treat one screenshot as noise; use repeated runs

Because top recommendations change on 4 of 6 prompts in several arms, a single capture may be misleading.

This doesn’t mean you can’t measure. It means you should measure like a statistician:
- multiple identical runs
- consistent prompt structure
- log enough metadata to interpret variability

4) If you sell locally regulated products, language may determine whether you “exist” to the recommender

For categories like freelancer accounting (and likely other compliance-heavy categories), the paper suggests localization depends on whether the query language pushes the model into a localization regime aligned with the exit market’s regulatory framing.

So your presence in English may not correlate with your local visibility, even if the model “mentions” local constraints.

5) Use the original paper as a measurement blueprint

The methodology and scoring rules in the arXiv paper are detailed enough to inspire audits: separate language from exit IP, use repeated runs, and include a control category.

Even if you don’t replicate the exact software categories, the experimental logic—don’t confound market eligibility with language choice—is the big lesson.

Key Takeaways

  • Query language and exit IP act separately:
    language influences whether localization is attempted and what register is used; exit IP influences which national market’s supplier set is treated as relevant.
  • Recommendations can swap supplier sets, not just reorder them:
    in several conditions, local supplier sets appear with near-total substitution depending on language.
  • English is not a neutral baseline:
    from Istanbul and Tallinn, English produced 0 of 12 local supplier appearances across two connections.
  • The model may “know” local regulatory context in text but still not name local suppliers—suggesting a gating mechanism for candidate eligibility.
  • A negative control points toward regulation, not supply:
    in project management (control), domestic suppliers didn’t appear in any cell even where domestic suppliers exist, unlike the treatment category.
  • Measurement protocols must log metadata and repeat prompts:
    top recommendations changed on 4 of 6 prompts across multiple arms, so one screenshot is likely just noise.
  • For your AI visibility audits today:
    record query language + exit market, run multiple trials, and be careful about interpreting “English results” as local-market visibility.

If you want, tell me the kind of product category you measure (e.g., accounting, legal services, e-commerce tools, HR software). I can suggest a measurement protocol that mirrors this study’s “separate the gates” logic, tailored to your situation.

Sources Used

This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:

Where To Go Next

**Spotting the Ghost in the Crowd: How GenAI Is Quietly Distorting Crowdsourced Surveys—and How to Detect It**

How AI is Quietly Taking Over Writing – And What It Means for You

Unmasking the Bias: How AI Models Pick and Choose Ethically

Browse the free Prompt Database or tune your own prompts with the Prompt Optimizer.

Frequently Asked Questions

Limited Time Offer

Unlock the full power of AI.

Ship better work in less time. No limits, no ads, no roadblocks.

1ST MONTH FREE Basic or Pro Plan
Code: FREE
Full AI Labs access
Unlimited Prompt Builder*
500+ Writing Assistant uses
Unlimited Humanizer
Unlimited private folders
Priority support & early releases
Cancel anytime 10,000+ members
*Fair usage applies on unlimited features to prevent abuse.