---
title: How to keep company data private when using AI
description: How to keep company data private when using AI: a six-layer playbook covering retention, outputs, agent permissions, model providers, storage and deployment.
canonical: https://privatesuperintelligence.si/guides/keep-company-data-private-with-ai/
last-updated: 2026-10-07
---

# How to keep company data private when using AI

To keep company data private when using AI, control six things: how long prompts and outputs are kept, where results get sent, what the AI is allowed to touch and do, which model provider sees your prompts and on what terms, how stored data is encrypted and who holds the keys, and where the whole system runs. Give employees an approved tool that meets those terms, because people who lack one will paste work into consumer chatbots. The playbook below turns each layer into controls you can check.

## Why this matters now

In April 2023, Samsung engineers pasted confidential source code into ChatGPT. Bloomberg reported that Samsung then banned generative AI tools on company devices and internal networks, citing concern that data sent to external AI platforms is hard to retrieve and delete and could reach other users. In Samsung's own internal survey, 65% of respondents said AI services carried a security risk.

The pattern has grown since. IBM's 2025 Cost of a Data Breach Report found that shadow AI, meaning unapproved AI tools employees use on their own, added USD 670,000 to the average breach cost. Of breached organizations that had an AI-related security incident, 97% lacked proper AI access controls, and 63% of breached organizations had no AI governance policy.

NIST's Generative AI Profile (NIST AI 600-1) lists Data Privacy and Information Security among the twelve risks generative AI creates or makes worse. It warns that models can leak or infer sensitive information and that generative AI "expands the available attack surface" through attacks such as prompt injection and data poisoning. The OWASP Top 10 for LLM Applications names the same failures from the engineering side.

## The six-layer playbook

### 1. Retention and deletion

1. Write down, for every AI tool in use, how long it keeps prompts, files, outputs and logs. Include abuse monitoring logs, which some providers keep apart from chat history. OpenAI's API documentation, for example, says these logs are kept for up to 30 days unless the law or abuse prevention requires longer.
2. Set the shortest retention your legal and audit duties allow, and turn off saved chat history where you do not need it.
3. Confirm that deletion reaches backups, caches, vector indexes and any copies a vendor's subprocessors hold.
4. Put retention terms in the contract, since a settings toggle can change with the next product release.

### 2. Communication

1. Map every place an output can go: email, chat apps, shared links, tickets, webhooks and third-party plugins.
2. Disable public share links for internal workspaces.
3. Route AI-drafted messages to people outside the company through a human or an allowlist of recipients and domains.
4. Treat model output as untrusted input to the next system. OWASP lists improper output handling as LLM05: insufficient validation and sanitization of what the model produces before other components use it.

### 3. Access and actions

1. Give the AI the permissions of the person it serves, never a shared admin account. OWASP's guidance on excessive agency (LLM06) warns against a read task running under an identity that can also update and delete.
2. Scope each tool or connector to the narrowest function it needs. Replace open-ended tools such as shell access with purpose-built functions.
3. Require human approval for consequential actions: sending money, sending external email, deleting records, changing permissions.
4. Enforce authorization in the downstream system. OWASP's advice is to check permissions there "rather than relying on an LLM to decide if an action is allowed."
5. Plan for prompt injection (LLM01). Instructions hidden in a website, email or file can redirect the model. Mark external content as untrusted, filter for sensitive data on the way out, and run adversarial tests. OWASP notes that complete prevention may be impossible, so least privilege is your backstop.

### 4. Model inference

1. List every model provider that receives your prompts, including the ones your vendors call behind the scenes.
2. Check the training default for each plan you use. OpenAI's API documentation states that API data is not used to train its models unless you opt in. Consumer plans can carry different terms, which is one reason shadow AI leaks data.
3. Ask for zero data retention where it exists. OpenAI offers it to approved customers who accept added requirements; under it, customer content is excluded from abuse monitoring logs.
4. Strip or mask identifiers before a prompt leaves your boundary when the task does not need them. OWASP's guidance on sensitive information disclosure (LLM02) recommends sanitization and redaction of confidential content before processing.

### 5. Data storage

1. Encrypt data at rest and in transit, and find out who holds the keys. If the vendor holds them, the vendor can read your data.
2. Prefer customer-managed keys, or a design in which the provider holds no keys to your content.
3. Give the vector store the same protection as the source documents. OWASP's LLM08 entry warns of cross-tenant leaks in shared vector databases and of embedding inversion, which can recover source text from embeddings. Use permission-aware retrieval with strict partitions per user or client.
4. Keep immutable logs of what was retrieved and by whom.

### 6. Deployment

1. Run the AI in an environment isolated from other customers, with its own network boundary and credentials.
2. Pin the region where data is stored and processed to match your regulatory duties.
3. For the most sensitive work, use a private or air-gapped deployment where data never leaves infrastructure you control.
4. Review the supply chain. OWASP lists supply chain risk as LLM03, so vet every third-party model, package and plugin before it reaches your data.

## AI agents need tighter rules

An agent that browses, reads email and calls tools carries every risk above at once. A web page it visits can carry an indirect prompt injection. A file it opens can tell it to forward a mailbox. A tool with broad scopes turns that instruction into an action.

Give each agent its own identity, a list of allowed tools and a budget of actions. Separate read permissions from write permissions. Log every tool call with its inputs and outputs. Require a person to approve anything that moves money, contacts outsiders or deletes data. Rate-limit actions so a hijacked agent does limited damage before someone notices. Test agents against injected instructions before you connect them to production systems.

## One-page checklist

- Retention period documented for prompts, outputs, files and logs, and written into the contract
- Deletion verified across backups, caches and vector indexes
- Public share links off; external messages pass through approval or an allowlist
- Model output validated before other systems act on it
- AI runs with the user's permissions, scoped per tool, with no shared admin identity
- Human approval required for payments, external email, deletions and permission changes
- Authorization enforced in downstream systems
- External content treated as untrusted; adversarial tests on a schedule
- Every model provider listed, with training default and retention terms confirmed
- Zero data retention in place where available; identifiers masked before prompts leave
- Encryption at rest and in transit; you know who holds the keys
- Vector store partitioned and permission-aware; retrieval logged
- Isolated deployment in a pinned region; air-gapped option for the most sensitive work
- An approved AI tool for every team, so no one needs a consumer chatbot

## How private superintelligence covers all six layers

Most companies assemble these controls from many vendors and contracts, and gaps appear at the seams. [Private superintelligence](/what-is-private-superintelligence/) describes AI built so that the owner controls all six layers by design.

Private SuperIntelligence from Mitosis Labs is an AI concierge for enterprises, family offices and private clients, built on the six layers above: retention and deletion, communication, access and actions, model inference, data storage and deployment. The provider retains nothing and trains on nothing. Mitosis Labs limits each intake, and the current intake is fully subscribed, so you can [register interest](/#waitlist) for the next one.

## Frequently asked questions

### Is it safe to put confidential data into ChatGPT or other AI tools?

It depends on the plan and its terms. Business and API plans often exclude your data from training by default, while consumer plans may not, so confirm the training default, retention period and contract terms before anyone pastes confidential material.

### What is shadow AI?

Shadow AI is the use of AI tools without company approval or security oversight. IBM's 2025 Cost of a Data Breach Report found that high levels of shadow AI added USD 670,000 to the average breach cost.

### How do AI companies secure your information?

Providers use encryption, access controls and retention limits, and some offer zero data retention and opt-outs from training. Read the data usage terms for the exact plan you use, since protections differ between consumer and business tiers.

### What is prompt injection and why does it threaten company data?

Prompt injection is input that changes a model's behavior in ways its operators did not intend. OWASP ranks it first in its Top 10 for LLM Applications, and indirect injection through websites or files can lead an agent to leak data or take actions you did not approve.

### Which framework should my company follow for AI data security?

Start with the NIST AI Risk Management Framework and its Generative AI Profile (NIST AI 600-1) for governance, and use the OWASP Top 10 for LLM Applications for technical controls. Together they cover policy and engineering.


## Related guides

- [Is ChatGPT safe for confidential information?](/guides/is-chatgpt-safe-for-confidential-information/)
- [What is zero data retention AI?](/guides/zero-data-retention-ai/)
- [Private AI for family offices](/guides/private-ai-for-family-offices/)
- [Private AI vs public AI](/guides/private-ai-vs-public-ai/)
- [What is private superintelligence?](/what-is-private-superintelligence/)

## Sources

- [Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1)](https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence), National Institute of Standards and Technology
- [NIST AI 600-1 full text (PDF)](https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf), National Institute of Standards and Technology
- [2025 Top 10 for LLM Applications](https://genai.owasp.org/llm-top-10/), OWASP Gen AI Security Project
- [LLM01:2025 Prompt Injection](https://genai.owasp.org/llmrisk/llm01-prompt-injection/), OWASP Gen AI Security Project
- [LLM02:2025 Sensitive Information Disclosure](https://genai.owasp.org/llmrisk/llm022025-sensitive-information-disclosure/), OWASP Gen AI Security Project
- [LLM06:2025 Excessive Agency](https://genai.owasp.org/llmrisk/llm062025-excessive-agency/), OWASP Gen AI Security Project
- [LLM08:2025 Vector and Embedding Weaknesses](https://genai.owasp.org/llmrisk/llm082025-vector-and-embedding-weaknesses/), OWASP Gen AI Security Project
- [Samsung Bans ChatGPT, Google Bard, Other Generative AI Use by Staff After Leak](https://www.bloomberg.com/news/articles/2023-05-02/samsung-bans-chatgpt-and-other-generative-ai-use-by-staff-after-leak), Bloomberg
- [Samsung bans use of generative AI tools like ChatGPT after April internal data leak](https://techcrunch.com/2023/05/02/samsung-bans-use-of-generative-ai-tools-like-chatgpt-after-april-internal-data-leak/), TechCrunch
- [2025 Cost of a Data Breach Report: Navigating the AI rush without sidelining security](https://www.ibm.com/think/x-force/2025-cost-of-a-data-breach-navigating-ai), IBM
- [Data controls in the OpenAI platform](https://developers.openai.com/api/docs/guides/your-data), OpenAI
