← Back to Insights Vault
Governance & Security

Shadow AI's Dirty Dozen: The 8 Data Types Your Employees Are Leaking Right Now

Published on Jun 16, 2026  •  6 min read  •  Intel Source: Financial Intelligence Desk

Shadow AI's Dirty Dozen: The 8 Data Types Your Employees Are Leaking Right Now

By the AuditSentinel Governance Intelligence Unit


When Samsung's semiconductor division discovered in April 2023 that three engineers had leaked proprietary data into ChatGPT within 20 days, the headlines focused on the source code. But the incident report revealed a more troubling pattern: the leaked data spanned three entirely different categories — source code, defect algorithms, and meeting transcripts.

This wasn't a single-point failure. It was a categorical failure. Samsung's engineers treated ChatGPT as a universal productivity tool, and in doing so, they demonstrated that every type of confidential data a company possesses has a plausible use case inside a public LLM interface.

Cyberhaven's 2024 analysis of 7 million workers found that 27.4% of corporate data pasted into AI tools qualifies as sensitive — up from 10.7% the previous year. LayerX found that 77% of employees have shared sensitive data via AI tools at least once.

The data isn't abstract. It's catalogued. Here are the eight specific categories, with documented incident data for each.


1. Proprietary Source Code

Incident: Samsung Semiconductor, March 2023. Engineer A pasted the entire semiconductor measurement database source code into ChatGPT to identify bugs and optimize performance.

Scale: Cyberhaven found that source code represents the largest single category of sensitive data pasted into LLMs. Entire codebases — not snippets — are being uploaded for debugging, optimization, and documentation generation.

Risk: Ingested code enters the LLM provider's training infrastructure and is not retrievable. For a startup whose entire valuation rests on proprietary technology, this represents an uninsurable competitive exposure.


2. Customer PII and Database Schemas

Incident: Documented on r/cybersecurity, 2024. An employee pasted the company's complete customer database schema into ChatGPT — including field names, data types, and relationship mappings — to get help writing a migration script.

Scale: Cyberhaven identifies customer data (names, emails, transaction histories, support tickets, account configurations) as the second most-leaked category. Triggering mandatory breach notification under GDPR, CCPA, and sector-specific regulation.

Risk: A single paste containing EU customer data creates GDPR Article 33 notification obligations. For startups without a breach response plan, the process of determining exposure and notifying regulators consumes weeks of executive attention.


3. API Keys and Authentication Tokens

Incident: Persistent across the Cyberhaven dataset. API keys are often embedded in code snippets pasted for debugging. A single exposed key can provide access to production databases, cloud infrastructure, payment processors, and third-party integrations.

Scale: Hard to quantify because keys are often embedded in larger code pastes rather than pasted independently — but the concentration risk is extreme. One key can compromise dozens of systems.

Risk: Unlike source code (which requires reverse-engineering to exploit), an API key is immediately actionable. The mean time to exploitation for an exposed cloud credential is measured in minutes.


4. Internal Financial Projections and Cap Table Data

Incident: No single high-profile case — but the Cyberhaven data confirms financial data is a persistent leakage category. Employees paste revenue projections, cap tables, and board deck materials into LLMs for formatting, summarization, and scenario modeling.

Scale: Financial data represents a smaller proportion of total leaks than source code or customer data, but the materiality is higher. Leaked financial projections create securities law exposure (Reg FD for public companies, insider trading risk for private companies approaching exit).

Risk: Material non-public information in a public LLM is a disclosure event. For companies approaching IPO or acquisition, this creates legal exposure that diligence will uncover.


5. Confidential Meeting Transcripts

Incident: Samsung Semiconductor, March 2023. Employee C recorded a confidential internal meeting — including strategy, product roadmap, and competitive positioning — and fed the full transcript into ChatGPT for meeting notes generation.

Scale: This is the category security teams most consistently underestimate. Employees perceive meeting transcripts as "administrative" rather than "sensitive" — despite containing the highest-concentration strategic information in the company.

Risk: A single leaked strategy discussion can reveal M&A intentions, competitive responses, and product timelines to anyone capable of prompt-engineering the training data.


6. Patent-Pending R&D Materials

Incident: The Samsung defect detection algorithm leak (Incident B) falls into this category. The engineer pasted proprietary test sequences used to identify defective chips — decades of manufacturing process refinement encoded in algorithms.

Scale: R&D leakage is concentrated in engineering and product teams, correlating with the "1% of employees causing 80% of incidents" finding from Cyberhaven. The damage per incident is disproportionately high.

Risk: Patent-pending material disclosed through a third-party AI system may compromise patentability. The US Patent Office's standard for "public disclosure" has not been fully tested against LLM ingestion, creating legal uncertainty at the worst possible moment.


7. Patient Medical Records and Healthcare Data

Incident: Cyberhaven's 2024 study specifically identified healthcare data — patient records, treatment plans, diagnostic information — in the sensitive data stream. HIPAA-covered entities are routing protected health information through tools that have no HIPAA compliance framework.

Scale: Healthcare AI adoption is accelerating faster than healthcare compliance can adapt. The Harmonic Security 2025 study found sensitive information in over 4% of all AI prompts — and healthcare prompts are disproportionately represented.

Risk: HIPAA violations carry penalties of $50,000–$1.5 million per violation category per year. A single employee's ChatGPT usage pattern could trigger multiple categories simultaneously.


8. Equipment Specifications and Manufacturing Algorithms

Incident: Beyond Samsung, manufacturing companies across sectors are experiencing the same pattern. Engineers paste equipment configurations, calibration data, and process optimization algorithms into LLMs for troubleshooting.

Scale: Less visible than consumer-facing data leaks, but higher-value per incident. Manufacturing IP represents decades of capital investment compressed into process parameters that fit in a single ChatGPT prompt.

Risk: For hardware and manufacturing startups, the process parameters are the moat. Leaking them through a public LLM is functionally equivalent to publishing the company's competitive advantage on a public website.


The Concentration Factor

Cyberhaven's finding that less than 1% of workers are responsible for 80% of incidents changes the mitigation calculus. You're not trying to prevent 100% of employees from ever pasting sensitive data. You're trying to identify the specific 1% — typically heavy AI users in engineering, product, and data roles — and intervene before their productivity workflow becomes a data exposure pipeline.


The Detection Imperative

Samsung had a policy. It said: don't enter confidential information into ChatGPT. The policy was communicated. It was acknowledged. It failed.

The lesson across all eight categories is the same: policy without detection is aspiration. The 1% of employees generating 80% of the risk are not reading your acceptable use policy before they paste. The only intervention that works is the one that happens at the moment of risk.


Data sources: Cyberhaven 2023 study (1.6M workers, 11% confidential paste rate), Cyberhaven 2024 study (7M workers, 27.4% sensitive data rate, 71% high-risk tools, 1%/80% concentration finding), LayerX Security (77% employee exposure), Harmonic Security 2025 (4%+ prompt sensitivity rate), Samsung incident reporting (Forbes, CIO Dive, DarkReading, The Economist Korea, RedTeams.ai, AI Incident Database #768).

Suspect unvetted AI tools are hitting your cloud ledger?

Upload a standard CSV corporate card statement to get an instant, confidential corporate governance report.

Run Free AI Audit Now