Learn
How LLMs Use Your Files

When you drag a PDF into ChatGPT, attach a spreadsheet to Claude, or upload an image to any AI assistant, a series of steps takes place between your action and the model's response. Understanding what those steps are, what data the model sees, and what happens to your file afterwards, is essential for making informed decisions about which documents you share with AI tools and which you keep to yourself.
The process is less mysterious than it might seem, but the details vary between providers and between pricing tiers in ways that matter.
The upload pipeline
The basic sequence is consistent across most AI tools, even if the implementation details differ.
When you upload a file, the first step is parsing. The system extracts the content from whatever format you have provided. For a text-based PDF, this means extracting the text layer. For a Word document, it means reading the document structure and pulling out the text, headings, and any embedded metadata. For an image, a vision model processes the visual content. For a spreadsheet, the cell contents are converted into a text representation, typically something resembling CSV or a simple table format.
The extracted content is then added to the conversation context. In most AI tools, the conversation context is the full set of information the model can see when generating its response: the system prompt, the conversation history, and any attached files or retrieved data. Your uploaded file becomes part of this context, sitting alongside your question in the model's input.
The model processes the entire context and generates its response. It does not "open" the file in the way a human would open a document in a word processor. It receives the extracted text (or the image pixels, in the case of vision models) as part of a single input, processes it through its neural network, and produces output tokens in sequence.
This means the model's interaction with your file is bounded by the conversation. In most implementations, once the conversation ends or the context window is filled, the model no longer has access to what you uploaded. Whether the provider retains the file on their servers after this point is a separate question, and an important one.
Processing versus training
This is perhaps the most important distinction in the entire topic.
When an AI tool processes your file to answer your question, it is performing inference. Your file is part of the input, the model generates an output, and the transaction is complete. This is analogous to a search engine indexing a web page to include it in search results; the content is used to provide a service, not to improve the underlying system.
When an AI tool uses your file for training, it incorporates the content into the dataset used to adjust the model's weights. The model learns from your data, and that learning persists in future versions of the model. Your document's content, in some transformed form, becomes part of what the model knows.
These two activities are fundamentally different in their privacy implications. Processing is transient: the model uses your data for the immediate task and (in principle) moves on. Training is persistent: your data contributes to a model that will be used by others.
Most major AI providers now separate these activities, but the distinction has not always been clear. In the early days of consumer AI tools, some providers included user inputs in their training data by default, with an opt-out buried in settings. Public attention and regulatory pressure have shifted the norm towards clearer disclosure and explicit consent for training use. Paid tiers of most major AI services now state that user inputs are not used for training, while free tiers may still reserve that right.
Reading the data usage policy before uploading sensitive files is not overly cautious; it is basic due diligence. The policies differ between providers, between pricing tiers within the same provider, and between API access and consumer product access.
What the model sees
The model does not see your file as you see it. It sees the extracted representation of your file's content, and that representation varies by file type and by the quality of the extraction process.
For text-based PDFs, the extraction is usually reliable. The text layer is read directly, and the model receives the text content in sequence. Formatting, however, is often lost or simplified. Bold text, italics, headers, and other visual formatting may not be preserved. Tables can be particularly problematic: a table that is visually clear in a PDF may be extracted as a confusing sequence of text fragments, with rows and columns losing their spatial relationships.
For scanned PDFs and images of documents, optical character recognition (OCR) is applied first. The quality of OCR varies with the quality of the scan, the clarity of the text, and the language. Handwritten text, unusual fonts, or low-resolution scans can produce errors in the extracted text, which then propagate into the model's understanding.
Images that are not documents, such as photographs, diagrams, or screenshots, are processed differently. A vision model analyses the image directly, interpreting its visual content. The model can describe what it sees, answer questions about the image, and extract text from screenshots. But it is interpreting pixels, not reading a structured document, and its understanding is limited by what the vision model can discern.
Spreadsheets are typically converted to a text representation. A complex spreadsheet with multiple sheets, formulas, conditional formatting, and pivot tables will be reduced to a flat text representation of cell values. The formulas themselves may not be visible; the model sees the computed values. Relationships between sheets may be lost.
Code files are treated as text, which is one of the areas where LLMs perform well. The model can read, understand, and reason about code with relatively little information loss, since code is already a text format.
Audio and video files require an additional preprocessing step: transcription. The audio is converted to text by a speech recognition model, and the resulting transcript is what the language model sees. The quality depends on the transcription model, the audio quality, the number of speakers, and the language.
The common thread is that your file undergoes a transformation before the model encounters it. What the model "sees" is a processed version of your file, and some information is inevitably lost in that processing. Understanding this helps set realistic expectations for what AI can do with your documents.
Token limits and truncation
Every language model has a context window: a maximum number of tokens it can process at once. A token is roughly three-quarters of a word in English, though the exact ratio varies by model and language.
When you upload a large document, the extracted text may exceed the model's context window. In this case, the system must decide what to do. Some tools truncate the file, processing only the first portion that fits within the context limit. Others use a retrieval-augmented approach, breaking the document into chunks and selecting the chunks most relevant to your question. Some refuse the upload and ask you to provide a shorter document.
This is relevant because it means the model may not see your entire file. If you upload a 200-page report and ask a question about page 180, a system that truncates at the context limit may never reach that section. Workspace tools that implement intelligent search across your files often handle this more gracefully, using retrieval techniques to surface the relevant sections regardless of document length. The approach described in detail in the concept of retrieval-augmented generation is how many systems address this limitation.
Data retention after processing
What happens to your file after the model has generated its response? This is where provider policies diverge most significantly.
Some providers delete your uploaded file immediately after processing. The file is used for the current inference request and then discarded. No copy is retained on the provider's servers.
Others retain uploaded files for a period, typically to support conversation history features. If you can scroll back through a conversation and see your uploaded file, it is being stored somewhere. Retention periods vary: some providers keep files for 30 days, others for the duration of your account's existence, and others until you manually delete the conversation.
A further consideration is whether conversations are logged for quality assurance or safety monitoring. Even if your file is not used for training, the conversation in which you discussed it may be reviewed by human evaluators or stored in logs. This is more common with free tiers and less common with enterprise agreements, but it is worth checking.
Enterprise tiers often come with explicit data handling commitments. Providers may guarantee that enterprise data is not retained beyond the inference request, is not used for training, is not accessible to other customers, and is stored in specific geographic regions. These commitments are typically documented in data processing agreements and are subject to audit.
For anyone working with sensitive data, whether that means client work for consultants, proprietary code for developers, or confidential business plans for founders, understanding the retention policy of your AI tool is not optional.
File types and how they are handled
A brief summary of how common file types are processed by AI tools.
PDFs come in two varieties. Text-based PDFs, where the text was generated digitally, are parsed by extracting the text layer directly. This is fast and usually accurate, though formatting and layout information may be lost. Scanned PDFs, where the content is an image of text, require OCR before the text can be extracted. OCR accuracy depends on scan quality, and errors in OCR become errors in the model's understanding.
Images are processed by vision models. These models can identify objects, read text, describe scenes, and answer questions about visual content. They work well for photographs, screenshots, diagrams, and charts, though their accuracy varies with image complexity and quality.
Word documents and other rich-text formats are parsed to extract text content, headings, and basic structure. Embedded images may or may not be processed, depending on the tool. Track changes, comments, and other collaborative features are typically stripped during extraction.
Spreadsheets are converted to a text representation. For simple data tables, this works reasonably well. For complex workbooks with formulas, charts, and multiple sheets, significant information is lost in the conversion.
Code files are handled as plain text, which preserves their content fully. This is one area where the file-to-model pipeline introduces minimal distortion.
Audio files are transcribed to text using speech recognition. Video files may be transcribed (for the audio track) and have keyframes extracted for visual analysis, depending on the tool's capabilities.
Privacy-preserving alternatives
If the privacy implications of uploading files to cloud AI services concern you, several alternatives exist.
On-device processing keeps your files on your hardware. Some AI tools run locally, either using smaller models optimised for personal devices or connecting to local model servers. The trade-off is typically performance: local models are smaller and less capable than their cloud counterparts.
Local models, run through tools like Ollama or LM Studio, give you full control over the processing environment. Your files never leave your machine. The models available locally are improving rapidly, but for complex tasks, cloud models still tend to produce better results.
Workspace-integrated AI offers a middle path. Rather than uploading individual files to a general-purpose AI chat, you can work within a cloud workspace that processes your files within a controlled environment. A workspace that is designed to be private and secure can provide AI features, including search, summarisation, and question-answering, while maintaining clear boundaries around data handling.
Self-hosted solutions give organisations full control over the infrastructure. By running models on their own servers or in their own cloud tenancy, organisations can ensure that no data leaves their environment. This requires significant technical investment but provides the highest level of control.
The right approach depends on the sensitivity of the data, the capabilities you need, and the resources available. For most personal and professional use, the key is choosing tools with clear, auditable data handling policies and using them in accordance with the sensitivity of the material.
The enterprise angle
For businesses, the stakes around file privacy are higher and more varied. A startup founder uploading a pitch deck faces different risks than a law firm uploading client contracts or a hospital system processing patient records.
Enterprise AI deployments typically involve data processing agreements that spell out retention, access, training use, geographic storage, and breach notification requirements. These agreements go well beyond the consumer terms of service and provide enforceable commitments.
Many enterprise AI tools offer features specifically designed for sensitive environments: data isolation between tenants, encryption at rest and in transit, audit logging of all model interactions, and compliance certifications (SOC 2, ISO 27001, HIPAA, and others).
Some organisations choose to bring AI capabilities into their existing infrastructure rather than sending data to external providers. This is more complex to set up but provides stronger guarantees about data handling.
For teams that work with sensitive material, choosing a workspace with a clear security posture and documented data handling practices is a foundational decision. The capabilities of the AI matter, but so does the environment in which it operates. An AI assistant embedded in a workspace that already manages your files and access controls can offer a more integrated and secure experience than uploading files to a separate AI service.
Practical guidance
A few principles can guide decisions about which files to share with AI tools.
Read the data usage policy. Specifically, look for answers to: Are my inputs used for model training? How long are uploaded files retained? Are conversations reviewed by human evaluators? Do the policies differ between free and paid tiers? These questions have concrete, discoverable answers for every major provider.
Understand the tier you are on. Free tiers of AI services often have more permissive data usage policies than paid tiers. The free version of a tool might use your inputs for training, while the paid version might not. The API might have different terms than the consumer product. Check the specific tier and access method you are using.
Consider what needs AI processing and what does not. Not every file benefits from AI analysis. A short document that you can read in a few minutes may not need to be uploaded. Reserving AI file processing for tasks where it provides clear value, like searching across a large collection or summarising lengthy material, reduces unnecessary exposure.
Use tools that keep your files in context. Rather than uploading files one at a time to a generic chat interface, consider working within a system that already holds your files and can make them available to AI as needed. A workspace where your notes and documents live alongside an AI assistant means fewer uploads to external services and more control over the data environment.
Know what you are comfortable with. Some people are comfortable uploading any document to any AI tool. Others prefer to keep sensitive material local and use AI only for non-confidential work. Both approaches are reasonable, and the right choice depends on your circumstances, your data, and your risk tolerance.
For those who want AI capabilities without the privacy trade-offs of uploading to external services, workspace tools that process files within a secure, private environment offer a practical path forward. The goal is to get the benefits of AI-assisted work, from search to summarisation to question-answering, without giving up control over your data in the process.
Frequently asked questions
Does the AI keep my file after I upload it?
It depends on the provider and your pricing tier. Some delete files immediately after processing. Others retain them for conversation history, typically for 30 days or the life of the conversation. Enterprise agreements often specify shorter retention periods. Check your provider's data retention policy.
Is uploading a file the same as giving it for training?
No. Processing a file to answer your question (inference) is different from using it to improve the model (training). Most major AI providers now separate these activities, but the distinction depends on the provider and tier. Paid tiers generally do not use inputs for training; free tiers may.
What happens to formatting when I upload a PDF?
Text content is extracted, but visual formatting (bold, italics, colours, layout) is often lost or simplified. Tables can be particularly problematic, as their spatial structure may not survive extraction. Scanned PDFs require OCR, which introduces additional potential for errors.
Can the AI see images inside my documents?
Some tools process embedded images using vision models; others extract only the text. If your document contains important diagrams, charts, or images, check whether your AI tool supports multimodal processing. Uploading images separately is sometimes more reliable.
Are free and paid tiers different in how they handle my data?
Often, yes. Free tiers may use your inputs for model training, retain data longer, or allow human review of conversations. Paid tiers typically offer stronger privacy commitments. API access may have different terms again. Always check the specific terms for the tier you are using.
What if my file is too large for the AI to process?
If the extracted text exceeds the model's context window, the tool may truncate the file, retrieve only the most relevant sections, or refuse the upload. For very large documents, tools that use retrieval-augmented generation handle this more gracefully than those that try to process the entire file at once.
Is it safe to upload confidential business documents to AI tools?
It depends on the tool, the tier, and the data handling policy. Enterprise agreements typically provide stronger guarantees. For highly sensitive material, consider workspace tools with explicit privacy commitments, self-hosted models, or on-device processing. Never upload material that would cause serious harm if exposed without first understanding the provider's retention and usage policies.
How do AI tools handle spreadsheets?
Spreadsheets are converted to a text representation of cell values. Formulas are typically evaluated to their results, and the formulas themselves may not be visible to the model. Complex features like conditional formatting, pivot tables, and cross-sheet references are usually lost in the conversion.
Can AI tools process audio and video files?
Yes, though they require an additional step. Audio is transcribed to text using speech recognition, and the transcript is what the language model processes. Video may be transcribed for its audio track and have keyframes analysed for visual content. The quality depends on the transcription and vision models used.
What is the safest way to use AI with sensitive files?
Use paid tiers with clear no-training policies. Prefer tools that operate within a private workspace rather than generic chat interfaces. Consider local models for the most sensitive material. Review data retention policies. Avoid uploading files that do not need AI processing. If working with regulated data (legal, medical, financial), ensure the tool meets the relevant compliance requirements.