ARTICLE 05 · PRIVATE AI · 2026-08-30

When employees paste customer data into ChatGPT: private AI options for companies whose data can't leave the building

Employees are already using AI for their work, whether the company allows it or not. The question executives have to answer is which kinds of data each AI may see. This article compares three options, from a ban to an AI provider's enterprise plan to a model installed inside the organization, along with a way of sorting data into tiers that employees can actually follow.

When employees paste customer data into ChatGPT: private AI options for companies whose data can't leave the building
The short version
  • Banning AI rarely works. Employees keep using it on their personal phones, where the company can't see it.
  • Different kinds of data carry different levels of risk. Sort them into tiers, then decide which tier can be used with which AI.
  • Companies with many confidential documents, such as bid prices, contracts or customer data, should have an internal AI that reads documents according to each person's access rights.

What happens to data when someone pastes it into an AI chat

A purchasing officer pastes every supplier's price table into an AI chat to get help comparing them. An HR officer pastes a list of names and salaries to make a chart. Both just want to get their work done faster, and nobody in the company knows that the data has already left.

Text pasted into a public AI service is sent for processing on the provider's machines and stored according to that service's policy. Some types of personal account allow the provider to use conversations to improve its models unless the user switches that setting off. On the company's side, there is no record at all of who sent what, or when. When a customer or an auditor asks, the company has no answer.

Data pasted into a public AI chat leaves the company without anyone noticing
Data pasted into a public AI chat leaves the company without anyone noticing

In May 2023, Samsung banned employees from using public AI tools on company devices after finding that engineers had pasted internal source code and meeting notes into ChatGPT. This kind of AI use, which the company knows nothing about, is called Shadow AI, and it happens in every company whose employees carry phones.

Three options, and what each one costs you

A company has three options: ban AI, buy an enterprise plan from an AI provider, or install a model on its own infrastructure. Each gives you something different at a different price, and most companies end up using more than one.

OptionWhat you getWhat you give upBest suited to
Ban AINo investmentEmployees keep using it through personal channels, and the company loses both productivity and visibilitySpecific work that law or contract strictly prohibits
An AI provider's enterprise planThe most capable models on the market, a contract stating your data won't be used to train models, and management of employee accountsData is still processed outside the organization, and you need to read the terms and check where data is locatedMost companies whose data isn't subject to location restrictions
Private LLM on-premise or in a private cloudData stays in the company's own infrastructure, and you control permissions and usage logs yourselfHardware and maintenance costs, and self-hosted models usually lag behind the latest releases from the major providersOrganizations with data that is confidential under contract, sensitive data, or requirements from customers and regulators

A common misconception is that installing a model in-house makes data safe by itself. If the internal AI can read documents in every folder, an intern can ask it for executives' salaries. The risk simply moves from leaking outside the company to leaking across departments. So an internal AI must see only the documents the person asking is entitled to see.

The internal AI assistant answers each person according to their own access rights, and data they aren't entitled to see stays hidden
The internal AI assistant answers each person according to their own access rights, and data they aren't entitled to see stays hidden

An engineering consultancy and a draft bid that ended up in a chat

A 120-person engineering consultancy bids for work regularly. One engineer pasted a draft proposal containing unit costs into a public AI chat to polish the language. A manager walking past happened to see the screen. When the matter reached the executives, their first question was how many times this had happened before, and nobody could answer.

The executives decided against a ban, because they knew the bid team had to write a large volume of documents and AI genuinely helped. The company sorted its data into four tiers using the method below, bought an enterprise plan for internal-tier work, and built an internal assistant for the confidential tier. This assistant searches past proposals, work standards and contracts, and shows only documents in folders the person asking has permission to open. The company piloted it with the bid department alone first.

Four data lanes

  1. Public: information already on the company's website. It can be used with any AI.
  2. Internal: manuals, work procedures and general meeting notes. These can be used with AI the company has an enterprise contract for.
  3. Confidential: prices, costs, contracts and customer data. These can be used with an internal AI that controls access and keeps usage logs.
  4. Restricted: health data, individual salaries and documents a contract forbids disclosing. These must not be fed to any AI until the data owner and legal counsel approve it case by case.

To put this into practice, have each department head take the ten types of document their team uses most and assign each one a tier, label the folders with their tier, and then teach employees using examples from that department's real work. The evidence to keep is a list of document types with their tiers, plus the number of documents not yet classified. This method can fail in two ways. The first is making the tiers so detailed that employees can't remember them. The second is defining a confidential tier without providing any AI that is approved for it, in which case employees will go back to public tools just as before.

What an internal AI assistant needs before people can trust it

A chatbot that answers from documents can be built with vibe code in a day. What takes time is making it answer each person according to their access rights, cite the source documents and keep logs for later review. Those three things are what get legal and IT to agree to switch it on.

Permissions should be tied to the employee account system the company already uses. Employees log in through SSO with the same account they use for other systems, and the assistant searches only the documents that account can open. When an employee changes departments or leaves, their access in the assistant changes immediately.

Every answer should come with the documents it is based on, so the person asking can open the source and see which version the answer came from. A system that searches documents first and then has the model answer is called RAG. Another advantage is that when a document is updated, the answers change with it, without retraining the model.

The company has two more decisions to make: how long to keep conversations, and who can view the usage logs. If documents contain customers' personal data, sending that data to an external provider for processing falls within the scope of Thailand's Personal Data Protection Act (PDPA). The company should have a data processing agreement with the provider and have its legal counsel or data protection officer review the setup before going live.

An AI model installed inside the company: documents come in for the AI to read without leaving the organization
An AI model installed inside the company: documents come in for the AI to read without leaving the organization
For the development team · Technical detail

Filter search results by permission at the vector search layer by storing an ACL with each document's embedding. Connect accounts through SSO and sync user groups automatically. Record the question, the retrieved documents and the answer in an audit log. Use model routing so the confidential tier goes to the internal model while the internal tier uses the provider's model. Mask PII before anything leaves the organization, and keep an evaluation set built from each department's real questions to measure accuracy and citation quality.

You don't need to invest in a private LLM yet if these apply

  • Most of the company's data is in the public and internal tiers
  • No contract or requirement specifies where data must be located
  • Nobody on the team is available yet to look after the hardware and the model

For companies in this group, an enterprise plan combined with a data tier policy is enough. Once a large amount of confidential-tier work starts to build up, add an internal assistant for that part only.

During the pilot, measure four things: the share of employees using company-approved AI, the number of questions per week, the share of answers that cite the correct version of a document, and the number of times confidential-tier data is found to have been sent outside the company. If employees are still using public AI for confidential work after three months, the internal assistant isn't yet handling real work well enough. Fix the quality of its answers before expanding to other departments.

DNA MAKER · PRIVATE AI

Let your team use AI on company documents while the documents stay inside the company

Your IT team knows the company's permission structure. Your legal counsel and data protection officer can say which kinds of data are restricted, and each department head knows which documents are confidential. DNA Maker runs workshops with all three groups to classify the data and map where documents live, who can open them, and what employees want to ask of those documents.

We then install a language model on your infrastructure or in a private cloud and build an assistant that searches internal documents and answers with citations, filters results by the asker's permissions, keeps an audit log and passes an answer-quality evaluation before going live. The work starts as a pilot with one department, measuring whether people actually use it and how accurately it answers, and only then expands. If you aren't sure yet which option suits your company, bring us a list of the document types your employees most often use AI with. We'll help you sort them into tiers and show which parts need an internal system.

Software engineering glossary

These terms come up when management, IT and legal need to agree on how AI is used with company data.

TermWhat it isA simple exampleWhat executives should ask the development team
Private LLMA language model running on infrastructure the company controls itself, so the data fed into it doesn't go to an outside providerA contract search assistant running on machines in the company's own server roomWhat do the hardware and maintenance cost per year, and how does answer quality compare with outside services?
On-premiseInstalling systems on machines located at the company's own premises, instead of renting them in the cloudThe server running the model sits in the head office's server roomWho looks after the machines, updates the systems and takes responsibility when hardware fails?
Data ClassificationSorting data into tiers by sensitivity, to set how each tier may be stored, sent and usedA cost price table is classified as confidential, so it can only be used with the internal AIWhich document types don't have a tier yet, and who decides when it's unclear?
Shadow AIEmployees using AI tools the company hasn't approved and can't seeAn employee photographs a contract on a personal phone and has an AI app summarize itHow do we know which AI tools employees are using right now, and for what work?
DLPA tool that detects and blocks important data from being sent outside the organization; short for Data Loss PreventionThe system raises a warning when someone pastes a large number of national ID numbers into an outside websiteWhich data tiers do the rules cover, and how often do they raise false alarms?
Data ResidencyA requirement specifying the country or region where data must be stored and processedA customer contract states that project data must stay on machines in ThailandIn which countries is our data processed, including when it is sent to AI?
SSOLogging in once with a company account to access several systems; short for Single Sign-OnEmployees access the AI assistant with their company email account, and when they leave, their access is shut off across every system at onceWhen an employee changes departments, how quickly do their permissions in the assistant change?
Try this tomorrow: Ask three department heads what their teams used AI for last week and what kinds of documents they pasted into it. Their answers will tell you whether confidential data is already flowing out of the company.