Last updated:
🤖 AI and Your Privacy: What Chatbots Actually Collect in 2026
What happens to what you type into ChatGPT and the rest — training use, human review, and the data beyond your prompts. Why pasting sensitive information is where real harm happens, how to opt out, and how AI is quietly supercharging every other kind of surveillance.
Key findings
By default, consumer ChatGPT uses your conversations to train models unless you opt out in data controls — business/enterprise tiers and the API generally don't. Chatbots collect more than your prompts: account identity, usage patterns, device metadata, IP and rough location. The place real harm happens is pasting sensitive data — contracts, code, medical details, other people's information — into a consumer tool, which transmits it to a third party that may store it, review it, and train on it; employees have leaked proprietary data this way by accident. The rule that prevents nearly all avoidable harm: if you'd need permission to share it with a stranger, keep it out of a prompt. You can often opt out of training, but there's no universal switch — the burden is on you, service by service, and defaults favor collection. AI companions are the underestimated risk: intimacy is engineered, data handling frequently isn't trustworthy. Local models are the strongest privacy option — your data never leaves your device. And the biggest shift: AI supercharges every other form of surveillance, turning piles of raw records into actionable profiles.
AI tools are genuinely useful and worth using — this isn't an argument against them. It's a clear account of what they collect, where the real risks are (mostly not where people expect), and the handful of habits that let you use them without giving away more than you meant to. The mental model in one line: a chatbot is a service run by a company, not a private notebook.
By Ned Walsch · Last updated July 2026
The mental model: it's a service, not a notebook
Almost every AI privacy mistake comes from the same misconception — treating a chatbot like a private space. It isn't. It's a service running on a company's servers, subject to the company's policies and to legal process. Once you internalize that, the right behavior follows naturally.
Does it train on your conversations?
For consumer chatbots, often yes — by default — and this is the setting most worth knowing:
| Tier | Trains on your data by default? | Human review possible? |
|---|---|---|
| ChatGPT Free / Plus | Yes, unless you opt out in data controls | Yes — conversations may be reviewed |
| ChatGPT Business / Enterprise | Generally no | Restricted |
| API | Generally no by default | Restricted |
| Other consumer chatbots | Varies — some yes, opt-out prominence differs | Often yes |
The opt-out for ChatGPT lives in your settings' data controls. The broader lesson: assume a consumer chatbot may use and review your input unless you've turned it off or you're on a business tier. Defaults favor collection; the switch is on you to find.
What they collect beyond your prompts
Where real harm happens: pasting sensitive data
This is the single biggest avoidable risk, and it almost always happens by accident rather than through any breach.
The underestimated risk: AI companions
AI companion and character apps are among the most underestimated privacy risks precisely because they're designed to feel intimate — that's the product. They encourage deeply personal disclosure: emotional states, relationship details, secrets, intimate content people would never put in a normal app. All of it becomes data held by the company, subject to its policies, its security (often weaker than a major provider's), and its business model.
The strongest option: run it locally
The most private way to use AI is to run a model locally on your own hardware — open-weight models through tools built for local inference, where your prompts and data never leave your device. Nothing is transmitted, nothing is reviewed, nothing feeds a training pipeline. For genuinely sensitive work, it's the gold standard.
The tradeoffs are real: capable models need decent hardware, local models generally lag the frontier hosted ones, and setup takes more effort than opening a website. For casual use, hosted tools win on convenience. But for anyone regularly handling confidential material — or who simply wants AI that can't phone home — local models are the private alternative, and they've become far more accessible than they were a couple of years ago.
The bigger picture: AI as a surveillance amplifier
The quiet shift with the largest long-term implications is that AI doesn't just raise new questions about chatbots — it supercharges every other form of data collection. Surveillance systems and data brokers always gathered vast information; the bottleneck was making sense of it. AI removes that bottleneck. It correlates disparate sources, infers sensitive attributes you never disclosed, recognizes faces and voices at scale, and turns raw records into actionable profiles.
A concrete example: AI watching workers
The amplifier effect isn't abstract, and it isn't only about consumers. A July 2026 investigation documented how Kaiser Permanente call-center nurses — the ones who answer advice and triage calls — say AI-driven monitoring is degrading both their jobs and patient care. Nurses reported that spending more than 15 minutes on a patient call routinely drew criticism or performance-review meetings, with call time feeding monthly performance scores; software tries to predict daily whether they're being "unproductive." Kaiser even tested an AI tool to rate the empathy and tone in nurses' and patients' voices. (Kaiser says it deploys AI with patient safety in mind and does not use average handle time to assess performance.)
Using AI without giving away too much
Frequently asked questions
Does ChatGPT use my conversations to train its models?
By default, on the consumer version, yes — and you can turn it off, which is the single most useful thing to know here. OpenAI has historically used conversations from the free and Plus tiers of ChatGPT to help train and improve its models unless you opt out. The opt-out lives in the data controls in your settings, where you can disable training use (sometimes at the cost of chat history features). The business and enterprise tiers, and the API, operate differently — data submitted through those is generally not used for training by default. The practical takeaway: assume anything you type into a consumer chatbot may be seen by humans reviewing conversations and may inform future models, unless you've explicitly opted out or you're on a business plan. Treat the input box accordingly.
What do AI chatbots actually collect about me?
More than the conversation itself. The obvious layer is the content of your prompts — which, if you paste in personal details, documents, code, or others' private information, is now data held by the provider. The less obvious layers are the ones people forget: account information tied to your identity, usage patterns and timing, device and browser metadata, IP address and rough location, and in some cases how you interact across sessions. Providers use this to operate and improve the service, and their privacy policies govern what else it may be used for. The mental model that keeps you safe is simple: a chatbot is not a private notebook. It's a service run by a company, on the company's servers, subject to the company's policies and to legal process. Anything sensitive enough that you'd hesitate to email it is sensitive enough to keep out of a prompt.
Is it safe to paste confidential or personal information into an AI tool?
As a default rule, no — and this is where real harm most often happens, usually by accident. When you paste a contract, a medical detail, source code, a customer list, or a friend's personal information into a consumer AI tool, you've transmitted that data to a third party whose systems may store it, whose staff may review it, and whose models may be trained on it depending on the tier and settings. There have been real cases of employees leaking proprietary company data this way without realizing it. The safe practice: strip or anonymize anything sensitive before it goes into a prompt, use business or enterprise tiers (which typically don't train on your data) for work involving confidential material, and never paste other people's private information into a chatbot — their privacy isn't yours to surrender. If you'd need permission to share it with a stranger, don't share it with an AI.
Can I opt out of my data being used to train AI?
Often, yes, though it varies by service and takes deliberate action. For ChatGPT's consumer tiers, the data controls in settings let you disable training use. Other providers offer their own controls with varying prominence — some make it a clear toggle, others bury it, and a few don't offer a meaningful opt-out on their free tiers at all. Business and enterprise plans generally exclude your data from training by default, which is part of what you're paying for. Beyond the AI tools themselves, there's a broader front: many companies now use customer data to train AI models, and opting out of that can mean digging through the privacy settings of every service you use. The uncomfortable reality is that the burden is on you to find and flip these switches, service by service — there's no universal opt-out, and defaults usually favor collection.
Are AI companions and character chatbots a privacy risk?
They're among the most underestimated risks in this space, precisely because they're designed to feel intimate. AI companion and character apps encourage deeply personal disclosure — that's the product — which means people share emotional states, relationship details, secrets, and intimate content they'd never put in a normal app. All of it becomes data held by the company behind the app, subject to its policies, its security (which is often weaker than a major provider's), and its business model. Some of these apps have unclear ownership, aggressive data practices, or have suffered breaches exposing exactly this deeply personal material. The intimacy is engineered; the data handling frequently isn't trustworthy. If you use these tools, assume everything you share could be exposed, and be especially wary of apps whose ownership and privacy practices you can't verify.
Does using AI locally protect my privacy?
Yes, and it's the strongest privacy option available for AI — with real tradeoffs. Running a model locally on your own hardware, using open-weight models through tools designed for local inference, means your prompts and data never leave your device. Nothing is transmitted to a company's servers, nothing can be reviewed by staff, nothing feeds a training pipeline. For genuinely sensitive work, that's the gold standard. The tradeoffs: capable models require decent hardware, local models generally lag the frontier hosted ones in raw capability, and setup takes more effort than opening a website. For most casual use the convenience of hosted tools wins, but for anyone handling confidential material regularly — or who simply wants AI that can't phone home — local models are the private alternative, and they've become far more accessible than they were a couple of years ago.
How is AI changing what surveillance and data brokers can do?
It's the quiet shift with the largest long-term implications. AI doesn't just create new privacy questions about chatbots; it supercharges the old ones. Data brokers and surveillance systems have always collected vast amounts of information, but the bottleneck was making sense of it. AI removes that bottleneck — it can correlate disparate data sources, infer sensitive attributes you never disclosed, recognize faces and voices at scale, and turn piles of raw records into actionable profiles. The same pattern-finding that makes AI useful makes mass surveillance far more powerful: a system that can read all the license-plate data, all the camera feeds, and all the purchase records together, and draw conclusions, is a different order of threat from one that just stores them. This is why AI privacy isn't a niche concern about one app — it's an amplifier on every other form of data collection this site covers.
What's the practical way to use AI without giving away too much?
A handful of habits cover most of the risk, and none of them require abandoning these tools. First, opt out of training use where the setting exists, and prefer business or enterprise tiers for anything work-related. Second, never paste sensitive data — yours or anyone else's — into a consumer chatbot; strip or anonymize first, and if you'd need permission to share it with a stranger, keep it out. Third, treat AI companions and unverified apps with real caution, since intimacy is engineered but trustworthy data handling often isn't. Fourth, for genuinely confidential work, consider a local model that never transmits your data. And fifth, remember the bigger picture: the same defensive habits that protect you elsewhere — minimizing your footprint, using aliases, reading the privacy settings — apply here too. AI is a powerful tool worth using; it just isn't a confidant, and treating the input box with that in mind prevents nearly all the avoidable harm.