Learn

How to Use AI Without Losing Your Data


The promise of AI tools is that they make you more productive: they summarise documents, draft emails, organise information, answer questions about your files. The less-discussed reality is that using these tools often means handing your data to companies whose incentives may not align with yours. Your prompts become training data. Your files sit on servers you do not control. Your notes, research, and intellectual output become inputs for someone else's model.

This does not mean you should avoid AI. It means you should use it deliberately, with an understanding of where your data goes and what happens to it once it gets there. Keeping control of your information while benefiting from AI capabilities is not a contradiction. It just requires some care.


What "losing your data" means in practice

When people talk about losing data, they usually mean deletion: a file disappears, a hard drive fails, a service goes offline. But in the context of AI tools, data loss is more often a loss of control rather than a loss of access. Your files still exist; you just no longer determine how they are used.

This can take several forms. Your data might be used to train AI models, meaning your private documents influence a model that serves millions of other users. It might be shared with third parties for analytics, advertising, or other purposes outlined somewhere deep in a terms-of-service document. It might be retained indefinitely, long after you stop using the service, sitting on servers in jurisdictions with different data protection standards. Or it might become practically inaccessible if a service shuts down or changes its export capabilities, leaving your work trapped in a proprietary format.

Each of these represents a different kind of loss, and protecting against them requires different strategies. Understanding the full picture of what AI tools collect is a good starting point.


Choosing AI tools carefully

The most consequential decision you make about your data happens before you type your first prompt. Choosing the right tool eliminates most of the risks; choosing the wrong one creates problems that are difficult to undo.

The first thing to check is whether the tool trains on your inputs. This information should be in the privacy policy, but some companies make it more accessible than others. Look for explicit statements: "We do not use customer data to train our models" is clear. "We may use data to improve our services" is vague enough to cover almost anything. If you cannot find a clear answer, treat that as a warning.

Data residency matters more than many people realise. Where your data is physically stored determines which laws govern its protection. A tool that stores data in a country with strong privacy regulations offers different protections than one that routes data through jurisdictions with weaker standards. For professionals working with client data or regulated information, this is not a theoretical concern.

Encryption is another differentiator. Some tools encrypt data in transit (between your device and their servers) but not at rest (on their servers). Others offer end-to-end encryption, where data is encrypted before it leaves your device and can only be decrypted by you. The strongest approach is zero-knowledge encryption, where even the service provider cannot access your data. Each level offers progressively more protection, with some trade-offs in functionality.

Consider also whether the tool is designed as a standalone AI chatbot or as a workspace that integrates AI capabilities. A chatbot typically requires you to send your data to the AI service for processing. A workspace that keeps your files in your own cloud storage and brings AI to the data rather than sending data to the AI, gives you a fundamentally different relationship with your information.


Practical habits for protecting your data

Even with a well-chosen tool, habits matter. Small decisions about what you share and how you share it compound over time.

Do not upload sensitive documents to AI tools you have not vetted. This sounds obvious, but the convenience of pasting a contract into a chat window and asking for a summary is powerful, and it is easy to do without thinking about where that contract ends up. Before uploading anything sensitive, confirm that the tool does not train on your inputs and that its retention policy is acceptable to you.

When working with confidential material, consider whether you can use an anonymised or redacted version. If you need an AI to analyse the structure of a legal agreement, you may not need to include the parties' names, financial figures, or other identifying details. Removing sensitive specifics before sharing reduces your exposure without eliminating the utility.

Be mindful of what you paste into chat interfaces. It is common to copy text from emails, documents, or internal communications and paste it into an AI tool for processing. Each paste is a data transfer, and the cumulative effect of many small transfers can be significant. Understanding how LLMs process your inputs can help you appreciate what happens once that data reaches the model.

Remember that "deleted" does not always mean erased. When you delete a conversation in an AI tool, it typically disappears from your view, but the underlying data may persist in backups, logs, or training datasets. Understanding what happens when you delete an app is worth the time, particularly for tools where you have shared substantial amounts of information.


Data portability: can you leave?

A tool's value is diminished if it traps your data. Before committing to any AI tool or workspace, check whether you can export your data, and in what form.

The best tools support export in open, standard formats that other tools can read. Markdown, PDF, CSV, and common file formats are good signs. Proprietary formats that only work within the original tool are a red flag. Data portability is not just a nice-to-have; it is your insurance policy against a service changing its terms, raising its prices, or shutting down.

Pay attention to what survives the export. Can you take your files? Your notes? Your folder structure and organisation? What about metadata like tags, timestamps, and relationships between documents? Some tools export the raw content but strip away the context that makes it useful. Tools built around smart organisation should preserve that organisation when you export.

Also consider whether the tool supports interoperability while you are using it, not just when you leave. Services that offer connections to other platforms let you work across tools without duplicating data or creating inconsistencies. This reduces lock-in and keeps your information accessible regardless of which tool you are using at any given moment.


Self-hosting versus cloud AI

Running AI models locally, on your own hardware, is the most privacy-preserving option available. When the model runs on your machine, your data never leaves your device. There is no third-party server, no privacy policy to parse, no risk of training data leakage.

The trade-offs are significant, though. Local models require capable hardware, particularly a GPU with enough memory to run useful models at reasonable speeds. The models available for local use are generally less capable than the largest cloud-hosted models, though the gap is narrowing. Setup and maintenance require technical knowledge, and you are responsible for keeping the software updated and secure.

Cloud AI services offer more powerful models, easier setup, and no hardware requirements beyond a web browser. The cost is that your data passes through someone else's infrastructure, subject to their policies and practices.

For many people, the practical answer is a hybrid approach: use cloud AI for non-sensitive tasks where the convenience outweighs the privacy cost, and use more controlled environments for sensitive work. A workspace that provides AI capabilities within a privacy-respecting framework can offer a middle path, giving you cloud convenience with stronger data protections than a generic AI chatbot. Some tools, including Fabric, offer AI agents that work within your workspace rather than requiring you to send documents to an external service.


Organisational policies for AI use

If you work within a team or organisation, individual caution is not enough. One person's careful habits can be undermined by a colleague who uploads a client database to an unvetted AI tool.

Organisations need clear AI usage guidelines that specify which AI tools are approved for use with company data, what types of data can and cannot be shared with AI tools, how to evaluate new AI tools before adopting them, and what to do if sensitive data is inadvertently shared with an unapproved tool.

These policies should be practical, not prohibitive. Banning AI tools entirely is unrealistic and counterproductive; people will use them regardless, just without guidance. A better approach is to provide approved tools that meet the organisation's security requirements and make it easy for people to use them. Workspaces designed for team use with built-in AI capabilities reduce the temptation to reach for unvetted alternatives.

For startups and smaller organisations, the policy does not need to be elaborate. A simple document that lists approved tools, explains the reasoning, and provides basic guidelines for handling sensitive data is a solid foundation. As the organisation grows and its data practices mature, the policy can evolve with it.


The role of workspace tools

The tension between AI capability and data control is not inherent to AI itself. It is a product of how most AI tools are designed: as standalone services that require you to bring your data to them.

An alternative model is the AI-native workspace, where AI capabilities are built into the environment where your files already live. Instead of uploading a document to a chatbot for analysis, you ask the workspace's AI to analyse a document that is already stored in your cloud. Instead of pasting text into an external service, you use an AI assistant that can access your notes and files within a controlled environment.

This approach has several advantages. Your data stays in one place rather than being scattered across multiple AI services. The AI can draw on your full context, including your files, notes, and organisational structure, through search and retrieval rather than requiring you to manually provide context with each query. And the privacy boundary is clearer: your data lives in your workspace, and the AI accesses it there rather than ingesting it into an external system.

Fabric takes this approach, providing AI capabilities through self-writing documents, search, and an assistant that works within your workspace. The intent is that you get the benefits of AI, including knowledge retention and intelligent organisation, without the data trade-offs that come with sending your files to standalone AI services.


Frequently asked questions

Can AI tools see the files on my computer?

Only if you give them access. Most AI tools can only see data you explicitly provide, either by uploading files, pasting text, or granting access to cloud storage. However, some desktop applications may request broader file access permissions. Always review what permissions an AI tool requests before granting them.

Is it safe to use AI for work documents?

It depends on the tool and the document. For non-sensitive work, most reputable AI tools are fine. For confidential, proprietary, or regulated documents, use only tools with explicit no-training policies, strong encryption, and clear data handling commitments. When in doubt, redact sensitive details before sharing.

What is the safest way to use AI with sensitive data?

The safest approach is to use a tool that does not train on your data, offers end-to-end encryption, stores data in a jurisdiction with strong privacy laws, and allows you to delete your data completely when you choose. Running a local model on your own hardware is the most private option, though it comes with capability and convenience trade-offs.

How do I know if an AI tool is training on my data?

Check the privacy policy and terms of service. Look for explicit statements about whether user data is used for model training. Some tools provide a settings page where you can opt out of training. If the policy is unclear, contact the company directly. The absence of a clear no-training statement should be treated as a presumption that training may occur.

What happens to my data if an AI company goes out of business?

This depends on the company's data handling practices and the laws of its jurisdiction. In some cases, user data may be sold as an asset during bankruptcy proceedings. In others, it may be deleted. The best protection is to use tools that support data export, so you always have a copy of your data regardless of what happens to the service.

Should my company have an AI usage policy?

Yes. Even a simple policy that lists approved tools, defines what data can be shared with AI services, and provides guidelines for evaluating new tools is better than no policy. Without guidelines, employees will make individual decisions with varying levels of care, and the organisation has no consistent standard for data protection.

Is cloud AI less secure than running models locally?

Cloud AI introduces risks that local models do not, such as data transmission over networks, storage on third-party servers, and potential use in model training. However, cloud providers often have more robust security infrastructure than individual users. The right choice depends on the sensitivity of your data, your technical capabilities, and your tolerance for risk.

Can I use AI tools and still comply with data regulations like GDPR?

Yes, but it requires care. You need to ensure that the AI tools you use have appropriate data processing agreements in place, that data transfers comply with relevant regulations, and that you can fulfil data subject requests (such as deletion requests) across all tools in your workflow. Using fewer tools with strong data practices is simpler than managing compliance across many services.

What is retrieval-augmented generation, and why does it matter for data control?

Retrieval-augmented generation, or RAG, is a technique where an AI retrieves relevant documents at query time rather than relying solely on its training data. It matters for data control because your documents are searched and referenced but not absorbed into the model's training. This means the AI can use your information to give you better answers without that information becoming part of the model itself.

How do I export my data from AI tools?

Most AI tools offer some form of data export, usually through account settings or a dedicated export feature. Check whether the export includes all your data (conversations, files, settings) and whether it uses open formats you can import elsewhere. If a tool makes export difficult or exports in a proprietary format, that is a significant lock-in risk.


Related pages

The workspace that thinks with you.

Ready when you are.

The workspace that thinks with you.

Ready when you are.

The workspace that thinks with you.

Ready when you are.