Blog
What the AI Boom Means for Your Personal Files

Your photos, documents, and creative work are now valuable in ways they weren't a few years ago. Here's what changed and what you can do about it.
For most of the history of personal computing, your files were inert. A photo sat on your hard drive until you opened it. A document lived in a folder until you searched for it by name. Your files mattered to you, but they had no particular value to the companies whose software you used to store them.
That has changed. In the age of large language models, image generators, and AI assistants, your personal files have become something new: potential training data. The photos you upload, the documents you write, the notes you take, the creative work you produce. All of it can be used to train AI models, to provide context for AI agents, or to feed recommendation engines that shape what you see and do online.
This shift has happened quickly, and much of it has happened quietly. Understanding what it means, and what choices you have, does not require technical expertise. It requires attention.
Your files as training data
Large AI models need enormous quantities of data to train on. Text, images, audio, video. The companies building these models have consumed most of the publicly available data on the internet, and many are now looking for new sources. Your personal files, stored on platforms that these companies operate or partner with, are one such source.
In 2024 and 2025, several major technology companies updated their terms of service to include language permitting the use of user-uploaded content for AI training. The specifics vary. Some companies allow training on your data unless you opt out. Some allow it by default with no opt-out. Some use your data to improve "product features" or "user experience," language that can encompass model training without saying so explicitly.
The practical effect is that a photo you uploaded to a cloud service three years ago, a document you drafted in an online editor, or a voice memo you saved to a synced folder may now be part of the training data for an AI model. Once your data enters a model's training set, it becomes part of the model's weights. It cannot be individually removed. The model has, in a sense, learned from your files permanently.
This is not illegal in most jurisdictions. The terms of service you agreed to when you signed up for these platforms generally grant broad licences over your content. But the fact that it is legal does not mean it is what you expected or wanted.
AI assistants and the context problem
Beyond training data, your files are increasingly used as context for AI assistants. This is, in many cases, a useful feature. An AI assistant that can read your documents and answer questions about them is more helpful than one that cannot. The ability to ask "What did the contract say about termination?" and get an answer drawn from your own files is a genuine convenience.
But context access creates new privacy considerations. When an AI assistant reads your files to answer a question, where does that processing happen? Is it on your device, on the provider's servers, or on a third party's infrastructure? Is the content of your query, and the files referenced in it, retained after the answer is delivered? Could it be used to improve the assistant's underlying model?
These questions have different answers depending on which AI assistant you use. Some process queries on-device. Some send your files to remote servers for processing. Some retain query logs. Some do not. The differences are significant, and they are rarely surfaced clearly to the people using these tools.
The distinction matters because AI assistants are becoming more capable and more deeply integrated into daily workflows. The more you rely on an assistant to help with your files, the more of your files it needs access to. And the more access it has, the more important it becomes to understand what happens to that data.
Agents acting on your behalf
AI agents represent the next step beyond assistants. Where an assistant answers questions, an agent takes actions. It can send emails, schedule meetings, organise files, make purchases, and interact with other services on your behalf.
To do any of this, an agent needs access to your accounts, your files, and your data. A scheduling agent needs your calendar and contacts. A research agent needs your documents and notes. A file management agent needs access to your entire storage system.
This access is the point. Without it, agents cannot function. But every grant of access is also a grant of trust. You are trusting the agent to use your data only for the purposes you intended. You are trusting the agent's provider to secure that access properly. You are trusting that the agent's behaviour will match what you expect, even when it is operating autonomously.
The security implications of agent proliferation are covered elsewhere. From a privacy perspective, the key issue is simpler: the more agents you authorise, the more copies of your data, or access paths to your data, exist in systems you do not control. Each agent connection is a relationship of trust with a company whose incentives may not align perfectly with yours.
What your files are worth now
It is worth being specific about what has changed. Your personal files are now valuable in at least four distinct ways.
As training data, they improve AI models. Companies that can train on large, diverse, high-quality datasets build better products. Your files contribute to that quality and diversity.
As context, they make AI assistants more useful. The more an assistant knows about your specific situation, the more relevant its answers become. Your files are the raw material for that relevance.
As behavioural signals, they feed recommendation engines. The documents you create, the photos you take, the notes you write all reveal patterns about your interests, habits, and intentions. These patterns are commercially valuable.
As targets for agents, they enable automation. Agents that can read and act on your files can save you time. They can also expose your files to new risks if the agent or its provider is compromised.
None of this makes your files less useful to you. But it does mean that the platforms storing your files have incentives that go beyond providing you with storage.
What you can do about it
The good news is that meaningful choices exist. You do not need to stop using cloud services or AI tools. But you can be more deliberate about which ones you use and how.
Reading terms of service is the most direct step, particularly the sections on data usage, AI training, and content licences. Look for language about "improving our services," "training our models," or "enhancing user experience." These phrases often indicate that your data may be used for purposes beyond basic storage and retrieval. If a service's terms allow AI training on your content, and you are not comfortable with that, look for an opt-out mechanism. If none exists, consider whether that service is the right place for your files.
Choosing tools that do not train on your data is one of the most effective protections available. Not all platforms treat your files the same way. Some have explicit commitments not to use user data for model training. Some, like Fabric, are built on the principle that your data stays yours and is never used to train AI models. When you use Fabric's AI features, they work on your files without sending them to third parties for training or processing.
Keeping your files in systems with clear data policies reduces ambiguity. If a platform's privacy policy is vague about how your data is used, that vagueness is itself informative. Platforms that are clear about their data handling, that say plainly "we do not train on your data" and mean it, are making a commitment you can hold them to.
Controlling where your files live physically adds another layer of protection. Fabric's bring-your-own-storage model lets you connect your own Google Drive, Dropbox, or other storage. Your files remain in the storage you choose, under the policies you selected when you set up that storage. The workspace layer sits on top without requiring you to surrender physical control of your data.
Consolidating your files into fewer systems reduces the number of policies you need to track and the number of relationships of trust you need to maintain. When your files are scattered across a dozen services, you are subject to a dozen different terms of service, a dozen different privacy policies, and a dozen different security postures. Bringing them together in one cloud workspace with smart organisation simplifies your exposure without sacrificing accessibility.
Being selective about agent permissions is increasingly important. When you authorise an AI agent, consider what access it needs and whether the provider's data handling practices meet your standards. Grant the minimum permissions necessary. Review and revoke agent access periodically, just as you would for any other app authorisation.
A note on creative work
For people who create things, whether writing, photography, illustration, music, or video, the implications of the AI boom are particularly acute. Creative work has clear commercial and personal value. It is also exactly the kind of high-quality data that AI companies need for training.
If you are a photographer, the photos you upload to a cloud service could end up in the training data for an image generation model. If you are a writer, your drafts and published work could be used to train a language model. The legal landscape around this is still evolving, but the technical reality is that once your work enters a training set, it is effectively irrecoverable.
Storing creative work in systems that respect your ownership is not paranoia. It is basic professional practice. Using a platform that encrypts your files, does not train on them, and gives you control over where they live physically is a reasonable safeguard for work that has both personal and commercial value. Tools like Fabric's annotations and publish features let you collaborate on and share creative work without sacrificing control over the underlying files.
The quiet shift
What has changed about your personal files is not dramatic in the way a data breach is dramatic. There is no single moment of exposure, no notification that your data has been compromised. Instead, there is a gradual shift in the terms under which your files exist on the platforms you use.
Your files are still yours in the legal sense. You can still access them, move them, delete them. But the platforms storing those files have found new uses for them, uses that were not part of the original bargain when you signed up for cloud storage or a note-taking app.
Being aware of this shift is the first step. Making deliberate choices about where your files live, which platforms you trust, and how much access you grant is the second. The tools to protect your data exist. The question is whether you use them.
Fabric is built for people who want the benefits of AI without giving up control of their files. Your data stays yours. It is encrypted, it is not used for training, and you decide where it lives. That is not a radical position. It is simply the way things should work.
Frequently asked questions
Are cloud services training AI on my files?
Some are. Several major cloud and productivity platforms updated their terms of service in 2024 and 2025 to permit the use of user-uploaded content for AI model training. The specifics vary by service. Some allow you to opt out; others do not. Check the terms of service and privacy policy for each service you use, paying particular attention to language about "improving services" or "training models."
Can I remove my data from an AI model once it has been trained on it?
In practice, no. Once data is incorporated into a model's training set and the model is trained, the data becomes part of the model's parameters. Individual data points cannot be extracted or removed. Some companies offer processes for requesting data deletion from future training runs, but this does not undo training that has already occurred.
Does Fabric train AI on my files?
No. Fabric does not use your files to train AI models. When you use Fabric's AI assistant, it processes your files to answer your questions, but the content is not retained for training purposes. Your data stays yours.
What is bring-your-own-storage?
Bring-your-own-storage means you can connect your existing storage services, such as Google Drive or Dropbox, to Fabric rather than uploading your files to Fabric's servers. Your files remain in the storage you choose, under the policies of that storage provider. Fabric's connections feature lets you work with these files in a unified workspace without moving them.
How do AI agents access my files?
AI agents typically request permission to access your accounts and files through OAuth or similar authorisation mechanisms. When you authorise an agent, you grant it access to specific services and data. The scope of that access varies. Some agents request broad access; others are more narrowly scoped. It is worth reviewing what permissions each agent has and revoking access for agents you no longer use.
What should I look for in a privacy policy?
Look for clear statements about whether your data is used for AI training, how long your data is retained, whether you can delete your data, who has access to it, whether it is encrypted, and whether the company shares your data with third parties. Vague language about "improving services" without specifics is a signal that your data may be used in ways you would not expect.
Is local storage safer than cloud storage?
Local storage avoids the risks associated with third-party data handling, but it introduces other risks: hardware failure, theft, and the inability to access files from other devices. A well-designed cloud platform with strong encryption and privacy practices can offer better security than a local hard drive, along with the convenience of access from anywhere. The key is choosing a platform whose data handling practices align with your expectations.
How can I protect my creative work from being used as training data?
Store your work in platforms that explicitly commit to not training on user data. Be selective about where you publish or share work online. Review the terms of service of any platform you upload creative work to. Consider using platforms that offer secure sharing and publishing without requiring you to grant broad content licences. For professional work, keeping files in a system with clear ownership and access controls is a reasonable baseline.
