The Short Answer
BYOAI creates unmanaged, visibility-limited risk pathways that aren’t solved by policy or security controls alone; the study models residual risk as governance maturity improves layered technical coverage and security outcomes.
Practically, you should operationalize BYOAI governance as a ladder that measures and improves enforceable coverage—starting with the dominant risk category: data exposure and compliance—rather than relying on blanket bans.
A key caveat is that framework engagement is inconsistent across sources, so governance maturity and control coverage must be assessed for the specific unmanaged personal-account scenario, not assumed from enterprise-only AI controls.
On this page
- Introduction
- Why This Matters
- Technical + Governance + Human: The Three-Pillar Model Behind the Ladder
- What the Evidence Looked Like: How the Authors Built Their BYOAI Risk Picture
- Frameworks Don’t Fully “Plug In” for BYOAI—So the Model Computes a Framework Gap Index
- The Parameterized Maturity Ladder: How Controls Reduce Residual Risk (Without Pretending Data Exists)
- Retrospective Incidents: Where Each Pillar Would Have Caught the Failure
- Key Takeaways
BYOAI Governance Ladder: Cutting Shadow AI Residual Risk
Introduction
Employees are increasingly doing work with personal generative AI tools—things like ChatGPT, Gemini, or Claude—even when those tools aren’t approved for company use. This practice is called Bring Your Own AI (BYOAI), and it’s a pretty direct cousin of Shadow AI. The twist is that BYOAI often uses employee-authenticated personal accounts, which means the activity can happen outside enterprise identity, logging, and security controls. New research from the original paper digs into what that changes—and how organizations can govern it without just shouting “don’t use it.”
A key idea in the study is that older AI governance frameworks were designed with the assumption that the organization owns the AI system. With BYOAI, the organization often can’t see what’s going in, what’s coming out, or what’s getting stored by the vendor. The paper addresses those visibility gaps by building an evidence-based risk taxonomy and a framework-engagement profile from a systematic literature review, then turning that into a parameterized maturity model that estimates how governance maturity reduces residual risk.
Why This Matters
This is significant right now because BYOAI isn’t a theoretical problem anymore—it’s operational. Your best employees will still want to move fast, fix bugs, summarize docs, draft emails, and generate analysis. If approved enterprise AI is slow, hard to access, or doesn’t match their workflow, many people will “solve the problem” by using a personal account. That’s not a character flaw; it’s a system design mismatch.
Here’s a scenario that could apply today in most organizations: a product engineer uses a personal AI chat to draft code and paste in internal snippets and architecture details. Even if the engineer isn’t trying to leak secrets, the tool may store inputs or use them to improve services (depending on settings), and the output can influence decisions. The paper’s model treats this as a predictable risk pathway—especially around data exposure and privacy leakage, which the authors found dominated the literature.
What’s new in this research is the quantification style. Prior BYOAI governance work often reports governance ratings as asserted values. This paper instead builds an artifact that produces scores from a reproducible chain: governance maturity affects what technical controls are active, coverage affects security outcomes, and those outcomes determine how much framework coverage remains enforceable. In other words: it tries to answer “what’s the risk left after we do governance,” not just “what should we do.”
And it also builds on earlier AI governance research by acknowledging that no single framework is complete for BYOAI. The paper compares how different frameworks show up in the literature and then shows why that unevenness matters—particularly when the model is not organization-owned. If you’ve felt like governance guidelines don’t “snap into place” for BYOAI, this study gives a structured explanation for why.
Technical + Governance + Human: The Three-Pillar Model Behind the Ladder
The paper’s central claim is simple but important: BYOAI can’t be governed by policy alone, and it can’t be solved by security controls alone. The authors structure safeguards into three mutually supporting pillars:
- Technical pillar: controls that reduce data leaving your visibility and shrink the attack surface.
- Governance pillar: rules, contracts, accountability, and workflows (including human verification).
- Human pillar: training and enablement that remove the incentive to “just use the personal tool anyway.”
Think of this like a fire safety system.
- Technical controls are your sprinklers and alarms: they reduce the chance of a big spread and help detection.
- Governance controls are your building rules and responsibilities: they ensure someone is accountable and that responses happen consistently.
- Human controls are the evacuation training and drills: if people don’t know what to do (or don’t believe it’s safe), sprinklers won’t be enough.
The seven technical control families they use as the measurable backbone
The model’s numeric scoring focuses on the technical pillar, using seven canonical control families. These map directly to the kinds of problems BYOAI creates—mainly around data exposure and access outside visibility.
DSPM(Data Security Posture Management)CASB(Cloud Access Security Broker)DLP(Data Loss Prevention)SSPM(SaaS Security Posture Management)IAM(Identity and Access Management)CIEM(Cloud Infrastructure Entitlement Management)SWG(Secure Web Gateways)
In the paper, these aren’t just listed—they’re positioned as progressively enabling enforceable coverage as maturity rises.
What the Evidence Looked Like: How the Authors Built Their BYOAI Risk Picture
This isn’t just a conceptual framework. The researchers built it from a systematic literature review plus a design-science style governance artifact.
Their systematic review: 30 records total, with 24 risk-eligible studies
They searched IEEE Xplore, Scopus, and Google Scholar, using a structured query across three blocks:
1) the phenomenon (e.g., shadow AI, BYOAI, unapproved AI),
2) the technology (e.g., generative AI, LLM), and
3) the governance lens (e.g., governance, risk management, compliance, data protection).
Inclusion criteria were set before selection, and records had to be 2018–2026 and English, and address BYOAI or governance/controls directly applicable to it. Their curated corpus ended up with:
- 30 included records total
- 24 research studies (risk-eligible)
- 6 framework/standard documents (used for framework engagement)
They also used explicit coding rules so the results are, in principle, reproducible:
- Risk categories are coded only when substantively examined (not mentioned in passing).
- Frameworks are coded only when substantively engaged (not merely cited).
Risk categories: what shows up most in the literature
Across the 24 risk-eligible studies, the authors report that the most prominent categories identified are:
- Data exposure and privacy leakage (dominates well over half the corpus)
- A cluster around compliance, governance drift, and hallucination risks
Bias and intellectual property show up, but less frequently, often bundled inside broader data protection or ethics discussions.
One nuance: the authors are careful that these frequencies reflect research emphasis, not a claim about real-world incident rates. Still, it tells you where the community’s attention has naturally landed—and where governance models need to be grounded.
Frameworks Don’t Fully “Plug In” for BYOAI—So the Model Computes a Framework Gap Index
A big frustration in BYOAI governance is that existing AI risk frameworks are built for an assumption that breaks immediately in BYOAI: organization-owned AI with known visibility and controls. The paper tests how consistently those frameworks show up for BYOAI.
Framework engagement is uneven (and AI TRiSM is least represented)
From the full set of 30 records, they coded how many substantively engage each framework family. The authors report:
- Technical controls and ethical/responsible-AI frameworks are engaged most often
- AI TRiSM is engaged comparatively least
- NIST AI RMF shows moderate engagement
Here’s the pattern as a qualitative comparison (since the paper describes the relative engagement magnitude rather than giving exact counts in the excerpt we have):
| Framework family | How prominently it shows up in the corpus (from the paper’s findings) |
|---|---|
| Technical controls (incl. control architectures) | Most engaged |
| Ethical / responsible-AI frameworks | Most engaged |
| NIST AI RMF | Moderate engagement |
| AI TRiSM | Least engaged |
Why that matters: the enforceable share collapses under BYOAI
The paper introduces the Framework Gap Index (FGI) to capture a key mismatch:
Even if a framework claims “nominal” coverage, under BYOAI only a fraction remains enforceable—because the organization doesn’t own the AI tool, doesn’t control identity, and doesn’t necessarily see the interactions.
They model an effectiveness factor α (alpha) representing the enforceable fraction under employee-owned account use. The baseline (unmanaged) posture—Level 1—still has a large gap:
FGI0 ≈ 0.77, meaning only about ~23% enforceable framework coverage at Level 1
That’s the paper’s central pushback against a certain kind of governance optimism: having a framework on paper doesn’t mean you can enforce it when BYOAI is happening in personal accounts.
If you want the “why,” it’s this: BYOAI breaks the observation and control assumptions embedded in most AI governance frameworks.
The Parameterized Maturity Ladder: How Controls Reduce Residual Risk (Without Pretending Data Exists)
Now we get to the model the authors emphasize: a five-level maturity ladder connected to a technical control architecture, with scoring produced by a deterministic chain.
The maturity ladder levels (what changes as you climb)
The ladder advances across all three pillars—but only the technical pillar becomes numeric.
- Level 1: essentially prohibition / informal posture
- Technical coverage is minimal
- Human motivation drivers aren’t addressed
- Levels 2–3: begin stacking technical controls
- Key data-centric controls start appearing
- Governance starts adding accountability workflows
- Levels 3–4: stronger combination
- Human-in-the-loop verification and enablement become part of the posture
- Identity and cloud-access controls integrate
- Level 5: continuous oversight + learning culture
- Everything—technical + governance + human—works as a system
A key detail: earlier levels rely heavily on policy or guidance, which the authors treat as insufficient because they don’t change the adoption incentives.
The measurable chain: from control coverage to security outcomes to framework gap
The model’s numeric logic runs like this (simplified):
- Governance maturity determines which technical control families are active.
- That creates a
TCCSscore (Technical Control Coverage Score). - Increased
TCCSimproves security outcomes (e.g., lower exfiltration risk and better detection dynamics). - Better outcomes reduce the
FGI(Framework Gap Index). - Finally, the model computes a composite score
CGS(Composite Governance Score).
The paper uses a coverage score where TCCS is essentially the share of the seven technical families (DSPM, CASB, DLP, SSPM, IAM, CIEM, SWG) that are active at a maturity level.
What the model predicts across the five levels (direction and magnitude)
The authors say stacking control families raises TCCS from 0% to 100% across the ladder by design. Then the output follows.
They report these structural outputs:
CGSrises from ~1.3to ~4.3(on a 1–5 scale)- Residual
FGIfalls from ~0.77to ~0.08 - Put differently: across the ladder, enforceable framework coverage improves from roughly a quarter of nominal coverage toward much more enforceable alignment
The paper also states the steepest transitions happen:
- entering Level 3 (data-centric controls like DSPM, DLP come online)
- entering Level 4 (identity and cloud-access controls integrate)
Prohibition alone doesn’t magically “solve it”
One of the most practical conclusions is built into those numbers: at Level 1, the model leaves residual risk close to baseline.
The authors back this by arguing that prohibition doesn’t address why people adopt personal tools:
- speed
- availability
- workflow fit
- perceived organizational resource inadequacy
So the “tool ban” can turn visible experimentation into concealed use—without improving the underlying data pathway risks.
Sensitivity analysis: does the conclusion depend on fragile assumptions?
They vary:
- α across [0.15, 0.40]
- governance score weights across a valid simplex (including extremes)
But in all combinations, the key relationship stays the same:
- FGI decreases monotonically with maturity
- CGS increases monotonically with maturity
So the direction isn’t tuned; it’s structurally built into the model’s chain.
Retrospective Incidents: Where Each Pillar Would Have Caught the Failure
The paper maps four documented incidents and a healthcare-sector pattern onto the taxonomy and ladder. Importantly, it’s not claiming “this framework would have prevented these events.” It’s mapping where safeguards would need to sit based on what the incidents show.
Case 1: Samsung engineers submitted proprietary code to ChatGPT → data egress + IP risk
Samsung’s 2023 incident: engineers used (even if authorized in that example) ChatGPT and submitted proprietary semiconductor source code, defect-detection algorithms, and a confidential transcript within about 20 days; Samsung responded with a blanket ban.
Mapped risks:
- Data Exposure/Privacy
- IP/Confidentiality
Mapped maturity:
- Level 1–2 posture (ban aligns with early posture)
What would help from the model:
- as maturity rises: DSPM/DLP/SWG/then CASB type controls would help flag or block similar transfers
But the paper emphasizes a limitation:
- technical controls alone may be incomplete because employees can find alternative tools unless the human pillar and sanctioned pathway reduce the “task completion” incentive.
Case 2: Mata v. Avianca (2023) → hallucinations with no data egress
In Mata v. Avianca, a brief contained fabricated judicial decisions produced by ChatGPT, and the court sanctioned the attorney.
Mapped risk:
- Hallucination / Output Reliability
Why technical controls struggle here:
- there’s no necessarily “sensitive data leaving the org” to catch with DLP or CASB.
What helps:
- governance-layer workflows requiring human verification (later ladder levels, especially Level 4–5).
Case 3: Heppner (2026) → input confidentiality and discoverability
A ruling found that case-strategy materials entered into a consumer AI assistant weren’t protected by attorney-client privilege/work-product doctrine.
Mapped risk:
- Confidentiality and discoverability exposure
Lesson:
- in this kind of scenario, the harm is the act of input, so governance rules and training must cover both inputs and outputs.
Case 4 + healthcare pattern: banking BYOAI upload of customer identifiers → compliance + data exposure
A 2026 community bank incident: an employee used a personal account/device to upload customer names, Social Security numbers, and dates of birth to an unapproved AI application. The institution treated it as material and filed an SEC Form 8-K, then blocked unapproved AI domains and tightened access.
Mapped risks:
- Compliance/Legal
- Data Exposure
What the paper draws from it:
- you need domain blocking (CASB/SWG) plus access tightening (IAM)
- and you need a governance pathway that creates a compliant alternative
Then for healthcare:
- the paper notes consumer AI tools often don’t sign agreements like a HIPAA BAA, so governance must enforce BAA-gated approval by combining:
- classification and DLP (technical)
- with BAA gating and process controls (governance)
Key Takeaways
- BYOAI is different from “generic Shadow AI” because employees use personal accounts outside enterprise identity/security controls, which breaks assumptions in many existing AI governance frameworks.
- A systematic review of 30 records (including 24 risk-eligible studies) shows the most prominent BYOAI risk categories are data exposure/privacy leakage, followed by compliance, governance drift, and hallucination.
- Frameworks don’t automatically translate into enforceable safeguards for BYOAI. The paper’s model quantifies this with a Framework Gap Index, estimating baseline Level 1 enforceability at about ~23% (
FGI0 ≈ 0.77). - Governance maturity helps because it stacks technical controls (measured by
TCCS) and reduces residual risk through a deterministic chain to security outcomes and framework gap. - Prohibition alone (Level 1) leaves residual risk close to baseline because it doesn’t address the reasons people adopt personal AI tools.
- The best protection is layered across technical + governance + human pillars:
- technical controls help with data egress
- governance/human workflows help with hallucination and confidentiality failures
- human enablement and sanctioned pathways reduce the incentive to use personal tools
If you want to use this in your own organization, the practical starting point is to treat BYOAI governance like a portfolio, not a checkbox: figure out which technical control families you can realistically deploy first (DSPM/DLP/CASB/SWG are early candidates), but pair them immediately with accountability workflows and a sanctioned “fast path” so employees don’t work around your rules.
Sources Used
This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:
- Governing Bring Your Own AI: A Parameterized Maturity Model — arXiv
- Authors: Authors: Dare Bello, John Hastings