Learn
What Data Do AI Tools Collect?

Every time you type a prompt, upload a document, or ask an AI tool to summarise a meeting, data changes hands. Some of that data is obvious: the words you typed, the file you attached. Some of it is less visible: your IP address, the time you sent the request, the device you used, the pattern of how often you come back. Understanding what AI tools collect, and what they do with it afterwards, is a basic requirement for making informed decisions about which tools to trust with your work.
This is not a scare piece. Many AI tools handle data responsibly, and some have built their entire product around privacy. But the range of practices across the industry is wide, and the defaults are not always in your favour.
The types of data AI tools collect
The most obvious category is the content you provide directly. This includes the prompts and queries you type into a chat interface, the files you upload for processing or analysis, and any text you paste from other sources. If you ask an AI tool to summarise a PDF, that PDF's contents become data the tool has access to, at least temporarily.
Beyond your direct inputs, AI tools collect metadata. This includes timestamps showing when you used the service, your IP address and approximate location, device information such as your operating system and browser version, and session identifiers that link your activity together over time. Individually, these data points seem innocuous. Collectively, they build a detailed picture of your behaviour.
Usage patterns form another layer. How often you use the tool, which features you rely on, how long your sessions last, what types of files you upload most frequently: all of this is typically logged. Some services track which responses you regenerate or edit, giving them signal about output quality.
Finally, there is account information: your name, email address, payment details if you are on a paid plan, and any profile data you provide. This category is common to nearly all software, not just AI tools, but it matters because it ties all the other data back to a specific person.
How AI companies use your data
The uses fall into several broad categories, and the boundaries between them are not always clearly drawn.
The most contentious use is training. Many AI companies feed user inputs back into their model training pipelines. When you use such a tool, your prompts, documents, and interactions may be used to improve the model's capabilities for all users. The benefit is real: models get better with more data. The cost is that your private documents could subtly influence future model outputs, and you typically have no way to verify what was absorbed or to retract it later.
Product improvement is a related but distinct use. Even companies that do not train on your data may analyse usage patterns to understand which features work well, where users struggle, and what to build next. This kind of analytics is standard across the software industry, though the sensitivity of the data involved with AI tools makes it worth paying attention to.
Some companies share data with third parties. This might mean sharing aggregated, anonymised usage statistics with partners, or it might mean more direct sharing with advertisers or data brokers. The specifics vary enormously between providers, and the only reliable way to know is to read the privacy policy.
Business intelligence is another common use. Companies study how their tools are used to inform pricing, marketing, and strategic decisions. This is generally the least concerning category, provided the data is properly anonymised, but it is worth knowing about.
The training data question
The distinction between "we use your data to improve our product" and "we use your data to train our AI model" is significant and often poorly communicated. When a company trains its model on your inputs, the information you provided becomes, in a statistical sense, part of the model itself. It is not stored as a retrievable file somewhere; it is diffused into the model's weights, influencing how the model responds to future queries from other users.
This matters for several reasons. If you upload a confidential business document and the tool trains on it, fragments of that document's content or style could surface in responses to other users. The risk is low for any single document, but it is not zero, and it scales with the volume of sensitive material processed. For professionals handling client data, intellectual property, or trade secrets, this is a meaningful concern.
Some AI tools have taken a firm position on this. They explicitly do not train on user data, and they make this commitment a central part of their product. Others train on everything by default but offer an opt-out mechanism, sometimes buried in settings. Still others are vague about their practices, using language in their terms of service that reserves broad rights without specifying exactly how data will be used. When evaluating tools for sensitive work, a clear no-training policy matters. Platforms designed around privacy and security tend to make this commitment explicit rather than leaving it ambiguous.
Data retention: how long and in what form
Retention policies determine how long a company keeps your data after you provide it. This is distinct from whether they train on it: a company might not train on your files but still retain them on their servers for months or years.
The questions worth asking are specific. How long are your prompts and uploaded files stored? What happens to your data if you delete your account? Is there a difference between deleting data from the visible product (removing a conversation from your history) and deleting it from the company's servers and backups? If the company trained on your data before you opted out or deleted your account, does the opt-out apply retroactively, or only going forward?
Most companies retain data for at least some period after you stop using the service. Backups, legal holds, and technical constraints mean that "deleted" in the user interface does not always mean "erased from all systems." Understanding the specifics of what happens to your data when you delete an app requires reading the fine print, which most people understandably do not do.
The best practice from a privacy standpoint is a clear, short retention window, with deletion from all systems (including backups) within a defined timeframe after account closure. Some services offer this; many do not.
Reading privacy policies with purpose
Privacy policies are long, dense, and written by lawyers. Reading them cover to cover is unrealistic for most people, but you can learn a lot by scanning for specific sections.
Look for language about data usage for training. Phrases like "improve our services," "enhance our models," or "develop new features" often indicate that your data may be used for training purposes. More privacy-conscious companies will explicitly state that they do not use customer data for model training.
Check for third-party sharing clauses. Who does the company share data with, and under what circumstances? "Service providers" is normal (every company uses infrastructure partners). "Business partners," "affiliates," or "advertising partners" should prompt closer reading.
Examine retention periods. A good privacy policy will state how long data is kept and what triggers deletion. Vague language like "as long as necessary" without further specifics is a warning sign.
Look for opt-out mechanisms. If the company does train on user data, can you opt out? How? Is the opt-out retroactive? Some tools offer granular controls; others offer only an all-or-nothing choice. Understanding data minimisation principles can help you evaluate whether a service collects more than it needs.
Finally, check the effective date and change notification policy. Companies update their privacy policies, sometimes in ways that materially affect your rights. Knowing whether you will be notified of changes, and whether continued use constitutes acceptance, is worth understanding upfront.
The spectrum of data practices
AI tools fall along a wide spectrum when it comes to data handling. At one end are services that operate on a zero-knowledge basis: they cannot see your data even if they wanted to, because it is encrypted in a way that only you can decrypt. Tools that support end-to-end encryption or zero-knowledge architectures fall into this category, though the trade-off is that server-side AI processing becomes more complex when the service cannot read your data in plain text.
In the middle are services that can see your data for processing purposes but commit to not training on it, not sharing it, and deleting it within a defined window. This is a practical compromise that allows AI features to work while limiting exposure.
At the other end are services that train on everything, retain data indefinitely, and share it broadly. Some of these are free tools, where the implicit bargain is that your data is the product. Others are paid services with surprisingly permissive data clauses.
The right position on this spectrum depends on what you are doing. Asking an AI to help you write a birthday card is a different risk profile from uploading a confidential merger document. Workspace tools that keep your cloud storage under your control, rather than routing everything through third-party AI services, can help manage this balance. A tool like Fabric, for instance, provides AI assistant capabilities within a workspace where your files remain yours, rather than becoming training data for someone else's model.
How to protect yourself
The most effective protection is awareness combined with deliberate tool selection.
Start by reading the terms of service for any AI tool you use regularly, particularly the sections on data usage, training, and retention. This is tedious but important, especially if you handle sensitive or confidential information.
Choose tools that explicitly do not train on your data, particularly for professional use. The comparison of sustainable AI apps can help you evaluate options that respect both your data and broader environmental considerations.
Be cautious about what you upload. Before sending a file to an AI service, consider whether the contents are sensitive and whether you are comfortable with the service's data practices. For particularly sensitive work, consider whether self-hosted or local AI models might be more appropriate.
Understand your opt-out settings. If you use a tool that trains on user data, check whether you can opt out, and do so if the training does not benefit you. Keep in mind that opting out may not be retroactive.
Consider data portability when choosing tools. A service that makes it easy to export your data gives you more control and makes it easier to leave if the company's practices change. Tools that support connections to your existing services without requiring you to re-upload everything to a new platform reduce the amount of data you need to expose.
Finally, think about your data holistically. Your files, notes, and knowledge are valuable. Treating them as something worth protecting, rather than something to scatter across dozens of free services, is a reasonable starting point. A personal data vault approach, where your information lives in a place you control and AI tools access it on your terms, is increasingly viable.
Frequently asked questions
Do all AI tools train on my data?
No. Some AI tools explicitly do not train on user data. Others train on data by default but offer an opt-out. The only way to know for certain is to check the provider's privacy policy and terms of service. Free-tier tools are more likely to train on user inputs than paid services, though this is not a universal rule.
What is the difference between data collection and model training?
Data collection refers to gathering and storing information you provide. Model training means using that data to adjust the AI model's parameters so it performs differently in the future. A company can collect your data without training on it, for example by storing it for your own access or for analytics. Training is a specific, more consequential use of collected data.
Can I get my data removed from a trained AI model?
In practice, no. Once data has been used to train a model, it is absorbed into the model's weights in a statistical form. There is no reliable way to extract or remove a specific individual's contribution from a trained model. This is why prevention, choosing tools that do not train on your data, is more effective than trying to undo training after the fact.
What metadata do AI tools collect beyond my prompts?
Typical metadata includes your IP address, device type and operating system, browser information, timestamps of your interactions, session duration, feature usage patterns, and account identifiers. Some tools also collect location data, referral sources, and information about how you arrived at the service.
Are paid AI tools safer than free ones?
Generally, paid tools have less incentive to monetise your data through training or third-party sharing, since you are already paying for the service. However, paying for a tool does not guarantee good data practices. Always check the specific privacy policy rather than assuming that a subscription means your data is protected.
What should I do before uploading sensitive documents to an AI tool?
Read the tool's privacy policy, specifically the sections on data usage and training. Check whether the tool trains on uploaded files. Consider whether the document contains confidential, proprietary, or personally identifiable information. If it does, either use a tool with a verified no-training policy or redact sensitive portions before uploading.
How do I know if an AI tool is using my data for training?
Check the privacy policy and terms of service. Look for phrases like "improve our models," "train our systems," or "enhance our services using customer data." If the policy is ambiguous, contact the company directly. Some tools also provide a settings page where you can see and control whether your data is used for training.
What is zero-knowledge encryption in the context of AI tools?
Zero-knowledge encryption means the service provider cannot access your data, even on their own servers, because it is encrypted with a key that only you hold. This is the strongest form of data protection, though it limits what the service can do with your data server-side. Not all AI features are compatible with zero-knowledge encryption, so there is often a trade-off between privacy and functionality.
Do AI tools keep my data after I delete my account?
This varies by provider. Some delete all user data within a specified period after account closure. Others retain data in backups or aggregated datasets for longer. If the tool trained on your data before deletion, that training effect persists regardless of account status. Check the provider's data retention policy for specifics.
How can I audit what data an AI tool has collected about me?
Many jurisdictions give you the right to request a copy of the data a company holds about you (for example, under GDPR or CCPA). Most AI tool providers have a process for data access requests, usually through their privacy or support pages. The response may take several weeks, and the completeness of the data provided varies.