Enterprise code privacy depends on three things: whether your code is used for training, how long prompts are retained, and what each AI tool sends to the model. Vendor policies vary by tier, endpoint, account type, and settings, so “enterprise” does not automatically mean zero exposure.
The risk is practical. Engineers may paste stack traces, configuration files, or code into AI tools without knowing which policy governs that data.
This guide explains how proprietary code reaches external models, what vendor policies state as of September 2026, and six controls CTOs can use to govern AI-assisted development.
What Do Enterprise Code Privacy AI Models Need to Cover?
A complete answer has three parts, and each has a different owner.
| Leak path | Question to answer | Who controls it | Primary mitigation |
|---|---|---|---|
| Training | Can code sent to the tool update model weights? | Vendor contract and account tier | Commercial terms with a no-training clause |
| Retention | How long are prompts and outputs stored, and who can read them? | Vendor policy and any zero data retention (ZDR) agreement | ZDR or short retention, plus your own audit records |
| The request itself | What left your network in the context window, and which tools saw it? | Your engineering configuration | Scoping, scanning and controlled egress |
A privacy position that answers only the training question is incomplete. A vendor can state that it does not train on your data and still hold prompts in logs for 30 days. An agent can run on a no-training commercial tier and still send production database credentials to the model because they sat in a file it read. Any assessment of enterprise code privacy AI models therefore has to cover all three paths, and the controls later in this guide are organised around them.
Where Does Proprietary Code Actually Leak?
Proprietary code can reach external systems through several paths, even when developers have no intention of sharing confidential information. Three paths deserve particular attention: consumer accounts and default settings, agent context and tool access, and data retention and model memorisation.
Path 1: Consumer tiers and default settings
The best-known public example is Samsung. In April 2023, sensitive internal data, reported to include source code, was submitted to ChatGPT by employees. By 1 May the company had restricted generative AI tools on company-owned devices while it worked on safeguards, citing the difficulty of retrieving and deleting data held on external servers.
Commercial-tier terms are clearer than they were in 2023, but the pattern has not changed. Whatever enterprise code privacy AI models provide on a commercial tier applies only to the accounts provisioned under it. An engineer debugging a failing deploy will tend to use the account already signed in, and on a personal account the terms are whatever that account’s settings say.
Anthropic’s consumer plans, for example, use chats to improve models only if the user allows it, a conversation is flagged for safety review, or the user opts in to a programme such as a tester program. Commercial products follow separate terms.
Path 2: Context window data leakage in agent workflows
Context window data leakage is the path most security reviews miss, because nobody has to paste anything. A coding agent assigned a task builds context by reading the repository and running tools. Depending on its permissions, that can include environment files, secrets committed by mistake, test fixtures containing real customer records, internal architecture documents and issue-tracker text. Whatever the agent reads can be serialised into the request sent to the model. If MCP servers or plugins are attached, the same content can also pass through a third-party server with its own retention and security posture.
The path also runs in reverse through prompt injection. Instructions planted in a file, web page, or ticket that an agent reads can steer it to include sensitive content in an output or to call a tool that sends data elsewhere. The OWASP Top 10 for LLM Applications 2025 lists prompt injection first, sensitive information disclosure second and excessive agency sixth.
Vendor-side exclusion features do not always cover agentic modes. GitHub’s documentation states that content exclusion is not supported in Edit and Agent modes of Copilot Chat in Visual Studio Code and other editors.
Treat any tool’s ignore feature as something to verify per surface, and enforce the boundary with permissions and network policy wherever you can. Map each agent’s context as a data flow with a source, a route, a destination and a retention period, the same way you would map any other integration.
Read: The OWASP LLM Top 10 Checklist Before You Ship an AI Feature
Turn the ten risks into release gates. The checklist covers each 2026 category with what to verify before shipping and what evidence to keep.
Path 3: Retention, logs and memorisation
Even when nothing is trained on, prompts are retained for a period. OpenAI’s documentation states that data sent to its API is not used for training since 1 March 2023 unless the customer opts in, and that abuse-monitoring logs, which can contain prompts and responses, are retained for up to 30 days by default.
Customers can be approved for Zero Data Retention or Modified Abuse Monitoring. Anthropic states that API inputs and outputs are deleted from its backend within 30 days of receipt or generation, with exceptions for ZDR agreements, usage policy enforcement, legal requirements and services where you control retention.
Memorization is the reason training and retention need to be assessed separately. Nasr et al. showed that an adversary can extract gigabytes of training data from open, semi-open and closed models, including ChatGPT.
Their divergence attack made ChatGPT emit training data at a rate 150 times higher than normal, and they concluded that current alignment techniques do not eliminate memorization. The attack targeted a 2023 model, and it does not show that any particular piece of proprietary code will be recovered. It does show that extractable memorisation is real.
Training and retention create different risk profiles. LLM training on proprietary code can create exposure that is hard to reverse, because removing specific learned information from model weights remains an open research problem.
Retention creates a time-bounded but still material exposure in stored prompts and outputs, and it matters just as much when those prompts contain source code or credentials.
What Do the Vendor Policies Actually Say?
The table below records what the vendors most enterprise teams use state on their own pages as of September 2026. Policies change, and they differ by tier and by product surface, so treat this as the checklist of questions to re-ask at each contract renewal.
It is a snapshot of what vendors currently commit to on paper for enterprise code privacy AI models. Confirm current terms with the vendor and with counsel before relying on them.
| Vendor, tier and surface | Used for training by default? | Retention stated | Source |
|---|---|---|---|
| OpenAI API (commercial endpoints) | No, since 1 March 2023, unless opted in | Abuse-monitoring logs up to 30 days. Zero Data Retention on eligible endpoints with approval. Conversations, threads, files and vector stores kept until deleted | OpenAI API docs |
| Anthropic API | No by default | Inputs and outputs normally deleted from Anthropic’s backend within 30 days, with exceptions for ZDR agreements, usage policy enforcement and legal requirements | Anthropic privacy center |
| Anthropic Claude for Work and Claude Gov | No by default | Products that save conversations retain chats to provide that functionality. Conversations explicitly submitted as feedback can be kept up to 5 years, de-linked before any use | Anthropic privacy center |
| Anthropic consumer (Free, Pro, Max) | Only if the user turns on model improvement; incognito chats excluded | Per user setting | Anthropic privacy center |
| GitHub Copilot Business and Enterprise: IDE Chat and code completions | No; business and enterprise data is not used to train GitHub’s models | Prompts and suggestions not retained | GitHub Copilot documentation |
| GitHub Copilot Business and Enterprise: other Copilot surfaces | No; business and enterprise data is not used to train GitHub’s models | Prompts and suggestions can be retained for 28 days | GitHub Copilot documentation |
Read together, these policies say something useful and something limiting. The useful part is that LLM training on proprietary code is off by default on the commercial tiers listed.
The limiting part is that the commitment is tied to the tier, the endpoint, the product surface and, in the consumer case, a toggle. The GitHub rows show the point clearly: the same Copilot plan carries different retention depending on how it is accessed.
What enterprise code privacy AI models deliver in practice is therefore a procurement and configuration outcome, and the engineering controls below exist to make sure the right tier and surface are the ones your code reaches.
What Are the Six Controls That Reduce Code Reaching External LLMs?
These are the controls we recommend putting in place before an AI tool is allowed near a production repository. They are ordered by dependency. The first two establish the foundational controls, and they are where the level of protection enterprise code privacy AI models offer is largely decided. Not every control can be applied to every tool, so each one notes where it fits and where it does not.
1. Commercial-tier accounts with the contract terms in writing
Every engineer uses an organization-owned seat on a commercial tier, provisioned through single sign-on and revoked with offboarding. The data processing agreement, the no-training clause, the retention period and any Zero Data Retention approval are captured in the contract rather than inferred from a help-centre page.
This is the control that converts a vendor’s current policy into an obligation, and it is where the promise of enterprise code privacy AI models becomes enforceable. Single sign-on and provisioning establish organization-owned accounts. Where consumer access needs to be restricted, supplement them with managed-device policy, CASB or SSE tooling, DNS and web filtering, proxy controls and endpoint controls, backed by written policy and monitoring. An identity provider alone cannot stop someone opening a consumer AI service in a browser.
2. An LLM gateway for API-controlled workflows
For AI workflows that call model APIs, such as internal agents, CI jobs and custom tooling, route model traffic through a managed gateway where practical. The gateway holds the vendor keys, meters usage, records request metadata, and applies redaction rules before a request leaves. Engineers do not hold vendor API keys directly, and a new tool is registered before it can reach any model.
Many SaaS coding assistants cannot be routed through a custom gateway, because the vendor controls the network and model path. For those, use the vendor’s enterprise controls plus identity, endpoint and network policies appropriate to that surface. A gateway only governs traffic that passes through it, so it narrows shadow AI without eliminating it. Unmanaged browsers, personal devices, mobile apps, browser extensions and unsanctioned endpoints need the device and network controls described in control 1. The same pattern for pipeline policy is covered in our guide to DevSecOps guardrails.
3. Context scoping rules for every agent
Each agent runs with an explicit ignore list and an allow list. Secrets files, environment configuration, anything under a fixtures or data directory, customer exports and internal-only documentation are excluded from the context the agent may read. Where the tool supports it, permission prompts are turned on for file reads outside the working directory and for any network call. Our guide to using Claude Code for coding shows how a permission model can do much of this work.
This is the control that addresses context window data leakage directly, and it is the one many teams have not configured. Because exclusion features vary by tool and surface, verify each one against the mode your engineers actually use.
4. Pre-prompt secrets and sensitive-data scanning
Two different jobs sit under this control, and they should not be treated as equal in reliability.
- Deterministic secret scanning looks for API keys, connection strings, private keys and tokens using known patterns and entropy checks. It is comparatively tractable and belongs in both the pre-request path and a pre-commit hook, so secrets rarely reach the repository in the first place.
- Sensitive-data and PII detection is harder. It depends on the organisation’s own data classification, and it faces false positives, false negatives, unstructured sensitive information and contextual identifiers that no pattern catches. Define what counts as sensitive first, then choose redaction or blocking for each class.
Running a secrets scan across an existing codebase is usually revealing, because it shows what an agent could have sent. No scanner guarantees clean egress, so treat scanning as one layer alongside scoping and contract terms.
5. Private deployments for crown-jewel code
Some code should not reach a shared endpoint regardless of policy: proprietary algorithms, pricing engines, security tooling, anything under a customer’s contractual data residency clause. For those repositories, route to a model deployed in the organisation’s own environment, whether a self-hosted open-weight model or a hyperscaler-hosted frontier model.
Provider-hosted options through Azure, Bedrock, Vertex and similar services involve provider-managed infrastructure, so the boundary is set by architecture and contract, not by the phrase “private”. Choose a provider deployment with private networking and contractual data-handling controls appropriate to the organisation’s residency and security requirements.
The build-versus-buy trade-offs are in our guide to custom AI model development. The aim is to know which repositories qualify and to make the routing automatic, so that the protection enterprise code privacy AI models provide matches the sensitivity of each codebase.
6. Prompt-level audit records with retention you control
An audit trail is what a security team needs when a vendor reports an incident, when a customer asks what of their data touched an AI system, and when a regulator asks the same. It also shows what agents actually read, which is the evidence base for tuning the scoping rules in control 3.
Full prompt logging creates its own risk. If prompts contain source code, customer information, regulated data or credentials that were not caught, the log becomes a second sensitive-data repository. A safer default is to record the model or tool used, the requesting identity, the timestamp, policy decisions, redaction metadata, content hashes, token counts and classification labels. Retain full prompt and response content only where it is justified and permitted. Where full logs are kept, specify:
- encryption at rest and in transit
- least-privilege access
- data classification
- a defined retention period and deletion process
- purpose limitation, so the logs are used only for security and compliance work
| Control | Leak path addressed | Typical owner | Relative effort |
|---|---|---|---|
| Commercial-tier accounts and contract terms | Training, retention | Procurement, IT | Low; mostly contractual |
| LLM gateway for API-controlled workflows | Request, retention (where traffic passes through it) | Platform engineering | Medium |
| Context scoping rules per agent | Request (context window) | Engineering leads | Low per agent, ongoing |
| Pre-prompt secrets and sensitive-data scanning | Request, retention | Security engineering | Low for secrets, higher for PII |
| Private deployments for crown-jewel code | Training, retention, request | Platform, security | Medium to high; per repository |
| Audit records with controlled retention | Evidence across all three | Security engineering | Medium; depends on logging scope |
What Do Teams Get Wrong About Enterprise Code Privacy?
The first mistake is banning without an alternative. A ban that arrives with no usable sanctioned option and no technical enforcement can leave an organisation with poor visibility into how AI is actually being used. IBM’s 2025 report found that 63% of breached organisations either had no AI governance policy or were still developing one. The better pattern is to make the approved path easy to use and to back it with controls that show where AI traffic goes.
The second mistake is stopping at the training question. A vendor’s no-training clause is necessary and covers one of three leak paths. Teams that have signed the enterprise agreement and then let agents read whole repositories have addressed LLM training on proprietary code and left the context window exposed.
The third mistake is reading the wrong policy. The enterprise tier’s terms are quoted in the security review while engineers work on a different tier or surface. What enterprise code privacy AI models commit to on paper only applies to the accounts and surfaces that were actually provisioned, which is why control 1 starts with account ownership and not with the policy document.
When Is Self-Hosting the Wrong Answer?
Self-hosting every model is a response worth questioning. Running open-weight models inside your own environment removes vendor retention and training from the picture. It also adds a new operational surface, covering model updates, GPU capacity, evaluation, and a security team now responsible for the inference stack.
Self-hosting earns its place for a small set of repositories: proprietary algorithms, regulated data with residency clauses, and anything a customer contract names. For many lower-risk repositories, governed commercial services may provide an appropriate balance of privacy, capability and operational cost without the organisation taking on responsibility for an inference stack.
The right answer depends on workload, volume, model choice, GPU utilization, staffing, compliance obligations and deployment architecture. Identify the repositories that need private routing, route them privately, and let the rest use the governed commercial path.
Ariel’s View on Enterprise Code Privacy for AI Models
Three practices are widely regarded as sound discipline for any team using AI coding tools, and they are the ones Ariel recommends when advising on AI-assisted development.
- Code runs on accounts the code owner holds: The organisation that owns the repository provisions commercial-tier seats or a private deployment, holds the contract terms, and can revoke access when a project or engagement ends.
- Every agent is scoped before it reads anything: Ignore lists, permission prompts and scanning are configured up front, and early audit data is reviewed with the security lead to tune what the agent may see.
- The data flow is documented like an integration: Source, route, destination and retention for every AI tool sit in the architecture record next to every other third-party dependency, the same discipline applied to data privacy in analytics work.
These practices apply whether the AI in question is a coding assistant, a retrieval system over internal documents or an agent with tool access, and they sit naturally within AI development services. The same gateway and scoping model applies to enterprise agent patterns, because an agent that can call tools is an agent that can move data.
Own the Account, Scope the Context, Log With Care
How far enterprise code privacy AI models can be trusted is decided in three places, and the vendor controls only one of them. The contract closes the training path. Scoping rules and, where practical, a gateway reduce what reaches the model from the context window. Carefully designed audit records help a security team show what touched an AI system when a customer or regulator asks. None of it requires abandoning the tools, and all of it gives a more reliable footing than a ban.
Start with the foundational controls: organization-owned commercial seats backed by device and network policy, and a gateway for the API-controlled workflows that can use one. Add scoping, scanning, private routing and logging from there.
Get a Clear View of Where Your Code Goes
Ready to understand how AI tools touch your engineering environment? Talk to Ariel about reviewing AI data flows across your tools, accounts and agents, so the controls above can be prioritised for your own risk profile.
Frequently Asked Questions
1. What do enterprise code privacy AI models need to provide?
Enterprise code privacy AI models should protect proprietary code from model training, unacceptable retention, and unauthorized exposure through requests, context windows or third-party integrations.
2. Do AI coding tools train on my company’s code?
Commercial AI coding tools generally exclude customer data from training by default. OpenAI’s API, Anthropic’s commercial products, and GitHub Copilot Business and Enterprise follow this approach, while consumer settings may differ.
3. What is context window data leakage?
Context window data leakage occurs when an AI assistant includes sensitive content in model requests after reading it. Permissions, secrets scanning, context scoping and network controls help reduce it.
4. How does LLM training on proprietary code differ from retention risk?
The two carry different risk profiles. Training can create exposure that is hard to reverse, because removing specific learned information from model weights remains difficult. Retention is time-bounded, but stored prompts and outputs can still contain sensitive code or credentials.
5. Should we ban ChatGPT and similar tools for developers?
A ban without a usable alternative and enforceable controls may leave an organisation with poor visibility into real AI use. Approved commercial tools, scoped agents, device and network policy, and carefully designed audit records provide a managed alternative.
6. When should an enterprise self-host an AI model for code?
Consider self-hosting or private deployment for repositories that cannot reach shared endpoints, such as proprietary algorithms or residency-restricted code. Many other repositories may use governed commercial services.