HireWebDeveloper.net

AI data privacy: what leaves your company when you use AI tools

What actually happens to data pasted into AI tools: training use, retention, the enterprise tiers that stop it, self-hosting, and the allowlist policy every business needs.

The uncomfortable mechanics, plainly

When an employee pastes text into a consumer AI tool, that text leaves the company: it crosses to the vendor's servers, is processed there, and, depending on the tier and settings, may be retained and reviewed, and in some tiers used to improve the models. Most major vendors now offer enterprise agreements with no-training commitments and defined retention, but the consumer tier most staff default to carries the weak protections. The exposure is not hypothetical: business plans pasted into tools, customer PII in "just summarize this" requests, and source code with credentials have all ended up places their owners did not intend. The AI-assisted build process states this practice's own version of the rule: client data, credentials and anything covered by an NDA never enter an AI tool.

The defense layers, from strongest to cheapest

  • Enterprise tiers with no-training contracts. The major vendors sell business agreements where inputs are not used for training and retention is bounded. This is the baseline for any business where staff use AI daily, it is a checkbox and a fee, not a project.
  • Self-hosted models for the sensitive tier. Running open-weight models on your own infrastructure keeps data inside the walls entirely, viable for classes of work where even enterprise terms are insufficient, at the cost of hardware and weaker model quality at the top end. The the self-hosted vs APIs decision (queued guide, ask directly) covers that decision honestly.
  • The allowlist policy, the cheapest control, today. A one-page policy: what may be pasted (public marketing copy, generic code, published policies), what requires sanitization (strip names, emails, keys, identifiers), and what is prohibited outright (credentials, customer records, unreleased financials). Most AI incidents are policy failures, not technology failures.
  • Sanitization habits for the middle tier. Replace identifiers before pasting; the AI does not need real names to improve a paragraph or debug a query structure.

For builds: the same rule engineered in

AI features on your site inherit this discipline: user data flowing to AI endpoints is disclosed in your privacy policy, enterprise no-training settings are configured on the API account, and retrieval systems are built so sensitive classes never reach the model, the RAG architecture can filter at retrieval time. GDPR adds teeth in the EU: disclosing AI processing and honoring deletion duties are legal requirements, not preferences (the the GDPR-and-AI section of the readiness checklist covers the build-side specifics).

The policy is one page. Writing it is an hour. The alternative, discovering your customer list in a vendor's training set, is not a reversible conversation. The brief starts the readiness audit if you want it done properly; the AI-readiness checklist is the free version.

Quarterly, and only when the numbers move

Get the rate report before you negotiate.

Updated rate bands across the major stacks and regions, plus what changed and why. No other email.

Read the current edition →

Ready to put this guide to work?

Six-question brief, scoped quote within two business days, and every term from the contract guide, in the actual contract.