Learn

What Is AI Memory?


When you have a conversation with an AI tool, it seems to understand you. It follows your thread of thought, refers back to things you said earlier, and builds on previous answers. But this understanding has limits that are easy to overlook. Ask the same tool a question tomorrow, and it may have no idea who you are, what you discussed, or what you were trying to accomplish. The sense of continuity within a conversation and the lack of it between conversations is one of the most important things to understand about how AI tools work today.

AI memory, in the broadest sense, refers to the ability of an AI system to retain information from previous interactions and use it in future ones. This is distinct from the model's training data, which represents everything the model learned during its training process. Memory is about what the system knows about you specifically, rather than what it knows about the world in general.


Context windows: what the AI can see right now

Every AI conversation operates within what is called a context window. This is the amount of text the model can consider at once, measured in tokens (roughly, chunks of words). When you are chatting with an AI, the context window contains the system instructions the tool uses to shape the AI's behaviour, the conversation history so far, and any documents or files you have provided for the current session.

The context window has a fixed size. Older models had windows of a few thousand tokens; current models range from tens of thousands to over a million. But regardless of size, the window is finite. Once a conversation exceeds the window's capacity, older parts of the conversation are dropped or compressed. The model does not choose what to forget; it simply cannot hold everything at once.

This is why AI tools sometimes seem to "forget" things you told them earlier in a long conversation. If you provided important context at the beginning of a discussion and the conversation has grown long enough to push that context out of the window, the model no longer has access to it. It is not that the model decided the information was unimportant; it is that the information is no longer visible to it.

Working within context limits is a practical skill. Restating important context periodically, summarising earlier discussion when shifting topics, and keeping conversations focused rather than sprawling all help maintain quality. Some tools manage this automatically by summarising earlier turns; others leave it to you. Understanding these mechanics helps explain why AI assistants sometimes give inconsistent answers in long sessions.


Persistent memory: remembering between conversations

Context windows handle the short term. Persistent memory addresses a different problem: retaining information across separate conversations, so the AI can build a cumulative understanding of you and your work over time.

Several AI tools have introduced persistent memory features. The general approach is that the system identifies key facts from your conversations, such as your preferences, your role, projects you are working on, or decisions you have made, and stores them in a separate memory layer. In future conversations, these stored facts are injected into the context window alongside your current prompt, giving the model background knowledge about you without your having to repeat it.

This is useful. If you tell an AI tool that you prefer British spelling, that you work in consulting, and that you are preparing a report for a specific client, having the tool remember these details across sessions saves time and produces better results. The alternative is re-establishing context from scratch every time you start a new conversation.

But persistent memory raises questions that are worth thinking through. The information stored in memory is typically controlled by the AI tool, not by you. You may be able to view and delete individual memories, but the system decides what to remember and how to phrase it. There is an asymmetry: you may not remember what you told the tool three months ago, but the tool does, and it uses that information to shape every response it gives you.

The privacy implications are significant. Persistent memory means the tool is building a profile of you: your interests, your work, your opinions, your habits. This profile persists on the provider's servers, subject to their data practices. If you are concerned about what data AI tools collect, persistent memory is an area that deserves particular attention, because it represents an explicit, structured record of your interactions over time.


Retrieval-augmented generation: grounding AI in your documents

Retrieval-augmented generation, commonly called RAG, takes a different approach to memory. Rather than trying to remember facts about you, RAG systems search through a corpus of documents at query time and pull in relevant information to inform the AI's response.

When you ask a RAG-enabled tool a question, the system first searches your files, notes, or other stored documents for passages relevant to your query. It then includes those passages in the context window alongside your question, giving the model specific, sourced information to work with. The model's response is grounded in your own data rather than relying solely on what it learned during training.

This matters for several reasons. RAG does not require the model to have been trained on your data. Your documents are retrieved and referenced, not absorbed into the model's weights. This makes RAG more privacy-friendly than training, because your information is used to answer your questions without becoming part of a model that serves other users.

RAG also reduces hallucination, the tendency of AI models to generate plausible-sounding but incorrect information. When the model has access to your specific documents, it can cite and draw from them rather than generating answers from its general training. For professionals working with precise information, such as researchers or consultants, this grounding in real documents is valuable.

The quality of a RAG system depends on the quality of its retrieval. If the search component fails to find the right documents, the AI cannot use them, no matter how capable the model is. This is why search quality in a workspace matters so much. A tool that can find the right document across your files, notes, and connected services gives the AI much better material to work with than one that searches poorly or narrowly.

Fabric uses retrieval-augmented generation to let its AI assistant draw on your stored files and notes when answering questions. Because the retrieval happens within your cloud workspace, your documents do not need to be sent to a separate AI service for processing.


The forgetting problem

There is a fundamental distinction between an AI model and the application built around it. The model itself, the large language model, has no memory between calls. Each time it processes a prompt, it starts from its training and the context window it is given. It does not know what it said five minutes ago unless that previous conversation is included in the current context.

Everything that looks like memory in an AI tool is provided by the application layer. The chat interface stores your conversation history and feeds it back into the model with each new message. Persistent memory features extract and store facts, then inject them into future conversations. RAG systems search your documents and include relevant passages. The model itself is stateless; it is the software around it that creates the illusion of continuity.

This matters because it means AI memory is a design choice, not an inherent capability. Different tools make different choices about what to remember, how to remember it, and who controls that memory. A tool that stores your conversation history on its servers creates a different privacy profile than one that keeps history only on your device. A tool that extracts and stores "memories" about you creates a persistent profile that you may or may not be comfortable with.

Understanding this architecture helps demystify AI behaviour. When an AI tool forgets something you told it, the explanation is usually mundane: the conversation exceeded the context window, or the application did not store or retrieve the relevant information. When a tool remembers something surprising, the explanation is similarly mechanical: a persistent memory feature stored it, or a RAG search retrieved it.


How different tools handle memory

The landscape of AI memory implementations is varied and evolving.

Chat-based AI tools like ChatGPT, Claude, and Gemini primarily rely on conversation context, sending the recent history back to the model with each message. Some have added persistent memory features that store facts between sessions. These features are typically opt-in or on by default with the ability to opt out, and they usually provide a way to view and manage stored memories.

AI-powered workspaces take a different approach. Rather than building memory into a chat interface, they provide memory through the workspace itself. Your files, notes, and documents serve as the AI's long-term memory, accessed through retrieval at query time. This is conceptually closer to how human memory works with external aids: you do not need to remember every detail of a project if you can look it up in your notes. Tools with smart organisation and strong search make this kind of retrieval-based memory more effective.

Some tools combine approaches. An AI assistant within a workspace might use the context window for the current conversation, RAG to pull in relevant documents, and a lightweight persistent memory for user preferences. The key question is not which approach a tool uses, but how transparent and controllable the memory is from the user's perspective.

AI agents introduce additional memory considerations. An agent that performs tasks on your behalf, such as researching a topic across your files or drafting a document based on your past work, needs access to substantial context. How that context is gathered, stored, and used after the task is complete is worth understanding, particularly for agents that operate through external connections or across multiple data sources.


Memory and privacy

The more an AI tool remembers about you, the more useful it can be and the more sensitive the information it holds becomes. This is a genuine tension with no perfect resolution, only trade-offs.

Persistent memory creates a detailed profile of your interactions. If this data is stored on a provider's servers, it is subject to their security practices, their data retention policies, and potentially their use for model training. Understanding who controls your data in this context is important.

RAG-based memory is generally less concerning from a privacy standpoint, because the AI accesses your documents at query time without absorbing them into a model. However, the retrieval process still involves processing your data, and the documents being searched need to be indexed in a way the system can access. The privacy characteristics depend on where that index lives and who can access it.

The ability to delete memories is a basic control that not all tools provide equally well. Can you see what the AI has stored about you? Can you delete individual memories? Can you wipe all stored memory? What happens to memories when you delete your account? These questions are worth asking before investing significant personal or professional context in a tool.

Tools that approach memory through your own workspace, where your files and notes are the memory and the AI accesses them under your control, offer a model where the answer to "who controls the memory?" is clearer. If your documents live in your cloud storage and the AI searches them when you ask a question, the memory is your data, managed and deletable by you.


Implications for how you work

Understanding AI memory changes how you use these tools. If you know the context window has limits, you can structure conversations to work within those limits rather than being surprised when the AI loses track. If you know persistent memory is building a profile of you, you can review and curate that profile rather than leaving it to the system's defaults.

For professionals who work with sensitive information, memory architecture should be a factor in tool selection. A developer who discusses proprietary code with an AI assistant, or a founder who uses AI to brainstorm strategy, should understand where those conversations end up and how long they persist.

For teams, memory has additional dimensions. Collaborative workspaces where AI can draw on shared documents create a form of collective memory. This is powerful for knowledge retention, ensuring that institutional knowledge survives personnel changes and that context is not lost when someone leaves. But shared AI memory also means that what one person shares with the AI may inform responses given to others on the team, raising questions about access control and information boundaries.

The landscape is evolving rapidly. As context windows grow larger, the boundary between "in-session" and "persistent" memory blurs. As RAG implementations improve, the distinction between remembering and retrieving becomes less relevant to the end user. What remains constant is the need to understand, and control, what the AI knows about you and where that knowledge lives.


Frequently asked questions

Does AI remember my previous conversations?

It depends on the tool. Within a single conversation, AI models work with a context window that includes the current discussion. Between conversations, most AI tools do not remember anything unless they have a specific persistent memory feature enabled. Some tools, such as ChatGPT, offer opt-in memory that retains key facts across sessions.

What is a context window?

A context window is the total amount of text an AI model can consider at once, measured in tokens. It includes the system prompt, conversation history, and any provided documents. When the conversation exceeds the window size, older content is dropped. Context windows range from tens of thousands to over a million tokens, depending on the model.

Can I delete what an AI remembers about me?

For tools with persistent memory features, most providers offer a way to view and delete stored memories, typically through a settings or privacy page. For conversation history, you can usually delete individual conversations. However, if the tool has already used your data for model training, deleting memories or conversations does not remove that influence from the trained model.

What is the difference between AI memory and AI training?

AI memory refers to information retained about you specifically, across your conversations, to personalise future interactions. Training refers to the process of adjusting a model's parameters using large datasets, which may or may not include user data. Memory affects how the tool responds to you; training affects how the model responds to everyone.

How does retrieval-augmented generation differ from training?

In training, data is absorbed into the model's parameters during a lengthy training process, becoming part of the model itself. In retrieval-augmented generation, documents are searched and referenced at query time, included in the context window alongside your question. The model uses the retrieved information for that specific response but does not incorporate it into its permanent knowledge. RAG is more privacy-preserving because your data is referenced, not absorbed.

Why does AI forget things I told it earlier in the same conversation?

This happens when the conversation exceeds the model's context window. The model can only process a fixed number of tokens at once. In long conversations, earlier messages are dropped or summarised to make room for more recent ones. The model does not choose to forget; it simply cannot hold the entire conversation at once.

Is AI memory a privacy risk?

It can be. Persistent memory creates a stored profile of your interactions, preferences, and personal information. This data is typically held on the provider's servers and subject to their security and data handling practices. The risk depends on what information is stored, how securely it is kept, and whether the provider uses it for purposes beyond personalising your experience.

Can AI memory be shared between team members?

Some tools allow shared memory or shared context, where information from one user's interactions informs responses to others on the same team. This is useful for maintaining consistent organisational knowledge, but it raises questions about what is shared, who controls it, and whether individuals can opt specific information out of the shared pool.

How do AI workspaces handle memory differently from chatbots?

AI chatbots typically rely on conversation history and optional persistent memory features. AI workspaces use your stored files and documents as the AI's knowledge base, accessed through retrieval rather than memorisation. This means the "memory" is your own data, stored and controlled by you, rather than a profile built and maintained by the AI provider.

Will AI memory improve over time?

Yes. Context windows are growing larger, retrieval systems are becoming more accurate, and memory management is becoming more sophisticated. The trend is toward AI tools that maintain richer, longer-term context about your work. As these capabilities improve, the questions about privacy, control, and data ownership become more important, not less.


Related pages

The workspace that thinks with you.

Ready when you are.

The workspace that thinks with you.

Ready when you are.

The workspace that thinks with you.

Ready when you are.