11% of Everything Your Team Pastes Into ChatGPT Is Confidential
By the AuditSentinel Governance Intelligence Unit
In 2023, Cyberhaven analyzed 1.6 million knowledge workers and found 11% of the data employees paste into ChatGPT is confidential. Not ambiguous. Not borderline. Eleven percent — source code, client data, regulated information, internal financials — material that triggers mandatory breach notification under GDPR, CCPA, and sector-specific regulation.
It didn’t stay at 11%.
The Trajectory
Cyberhaven’s 2024 study (7 million workers): 27.4% of corporate data pasted into AI tools is sensitive — nearly 3x the 10.7% baseline from the prior year. 71% of AI tools classified as high risk.
Your company is almost certainly running AI tools that would fail a standard security review — and nobody has conducted that review because nobody knows the tools exist.
Who’s Actually Doing the Leaking
Less than 1% of workers are responsible for 80% of incidents. This isn’t workforce-wide — it’s concentrated exposure. Cyberhaven’s Howard Ting: “There are two forms of education. The classroom education during onboarding, and the ‘you just did something wrong’ education.” The first has demonstrably failed.
The 1% aren’t malicious. They’re engineers debugging at 11 PM, analysts under deadline, PMs summarizing research.
What “Confidential” Actually Means
- Source code — entire codebases, not snippets
- Client data — names, emails, transactions, support tickets
- Regulated information — HIPAA, SEC, GDPR-triggering data
- Internal financials — projections, cap table, M&A discussions
- API keys and tokens — direct access to production systems
The LayerX Finding
77% of employees have shared sensitive data via AI tools.
Cyberhaven measures proportion of data (27.4%); LayerX measures proportion of employees (77%). Both are correct. Together: most employees have done it at least once, and a smaller number of heavy users are doing it repeatedly.
The Harmonic Security Confirmation
A 2025 study found sensitive information in over 4% of all AI prompts. Three independent research efforts — Cyberhaven, LayerX, Harmonic — all confirm the same pattern.
Real-World Evidence
An r/cybersecurity post: “Employee pasted our customer database schema into ChatGPT. How do you prevent this?” The consensus: detection and real-time monitoring is the only approach that scales.
Why This Matters for Valuation
1. Due diligence discoverability. If an acquirer demonstrates code or customer data exposure through public LLMs, that finding sits in the diligence report and affects the price. It can’t be explained away.
2. Regulatory exposure. GDPR fines reach 4% of turnover. The EU AI Act adds €35M or 7%. The fine isn’t the killer — the enterprise sales pipeline freeze that follows is.
3. Competitive exposure. Code pasted into ChatGPT is not retrievable. It exists in training infrastructure. Competitors can prompt-engineer around it.
The Governance Framework That Works
Visibility — complete AI tool inventory, not just IT-approved ones. Detection — real-time monitoring of data flows into public LLMs. Not blocking — detecting. Education — contextual, at the moment of risk. Pasting source code into ChatGPT is functionally equivalent to posting it on a public website.
The 11% from 2023 was a warning. The 27.4% from 2024 is a trend. The question for 2026: will your company be on the right side before an investor, acquirer, or regulator forces the issue?
Sources: Cyberhaven 2023 (1.6M workers) and 2024 (7M workers) studies, LayerX Security, Harmonic Security 2025. Samsung incident: Forbes, CIO Dive, DarkReading, AI Incident Database.