Blog

Your research is your moat


Large pharmaceutical companies have spent tens of billions of dollars over decades running experiments, generating data, and building up proprietary understanding of how specific molecules behave in specific contexts. That accumulated knowledge is the core of what makes each company competitive. When AI labs recently began approaching these companies asking them to contribute their proprietary data to train specialised models, offering early access and some vague proprietary advantage in return, nearly all of them said no.

The reasoning was clear once someone laid it out: contributing your proprietary data to a shared model that also ingests your competitors' data commoditises the exact thing that differentiates you. Your years of expensive, hard-won understanding becomes a feature of a general-purpose tool that anyone can buy. The edge disappears, and you've paid for the privilege of watching it go.

The pharma companies can see this dynamic clearly because the stakes are measured in billions and they have legal teams whose entire job is to think about intellectual property. But the same logic applies, at every scale, to anyone whose accumulated knowledge is what makes their work valuable.


The personal version

A consultant's competitive advantage is the analytical frameworks they've developed through years of client engagements, the pattern recognition that lets them see what's happening in a new organisation within the first two weeks, and the institutional knowledge they've built up about how specific industries actually work. That knowledge took years to accumulate and can't be replicated by reading the same books and articles the consultant read, because so much of it comes from experience, from seeing what happens when you apply theory to practice and noticing where the theory breaks.

A researcher's edge is similar: the unpublished hypotheses, the experimental results that didn't make the paper, the research log of what was tried and why it didn't work, the intuitions about which directions are promising that come from years of being close to a problem. The published papers are the tip. The iceberg underneath, the accumulated context of a research career, is what produces the next insight.

A lawyer's case strategy, a founder's understanding of a market gap, an investor's pattern recognition across hundreds of deals, a writer's accumulated notes and half-formed ideas: all of these are proprietary knowledge assets that compound over time. Each year of work makes the next year's work better, because the context gets richer and the pattern matching gets sharper.

This is the same compounding dynamic that makes personal AI context valuable: the knowledge on day 180 is qualitatively different from the knowledge on day one. The question is who benefits from that compounding.


Where the risk lives

When you pour years of accumulated thinking into a tool controlled by a company that might productise your category, you're doing a version of what the pharma companies were asked to do. You're contributing your proprietary knowledge to a system that may use it to build products that compete with you, or at minimum, that can provide a similar quality of analysis to anyone who pays for the same subscription.

The platform predation pattern makes this concrete. Cursor became one of Anthropic's biggest API customers, proving the coding assistant market was enormous. Anthropic launched Claude Code. The developers who had built their workflows on Cursor had, in effect, been demonstrating the commercial opportunity to their infrastructure provider. The same pattern is playing out across design, legal, financial analysis, and scientific research, as frontier labs launch vertical products in each category.

For individual knowledge workers, the risk is less dramatic but structurally identical. Your consultant who has built up a rich library of case analyses, strategic frameworks, and client insights inside a frontier lab's tool has given that lab a detailed understanding of what consulting looks like at a granular level, which use cases are most valuable, and what a productised version of that work might contain. They won't build a product targeting that specific consultant, but they'll build one targeting the category, informed by the aggregate of what thousands of consultants have poured in.


The knowledge audit

It's worth asking yourself a version of the question the pharma companies asked themselves: if you extracted all the proprietary thinking you've accumulated in your AI tools, your note-taking apps, your research databases, and put it in a pile, how much of your competitive advantage would be sitting in that pile?

For many knowledge workers, the answer is: most of it. The degree or the credentials got you into the room. The accumulated knowledge of how to do the actual work is what keeps you there. And a meaningful portion of that knowledge is now stored in systems you don't control, governed by terms of service you haven't read, and potentially accessible to companies whose next product announcement might be aimed squarely at your line of work.

This isn't a reason to stop using AI. AI makes every form of knowledge work dramatically more productive, and the compounding benefits of a rich personal context are real and significant. The question is where that context lives and who controls the terms of access.


Building a moat that stays yours

The structural answer is the same one that works at the enterprise scale, adapted for individuals and small teams.

Store your research in infrastructure you control. Your notes, annotations, research logs, literature reviews, and accumulated thinking should live in storage you own rather than in the database of an AI provider. Bring-your-own-storage means the data sits in your S3 bucket or R2 account, and the tools you use to interact with it are a layer on top rather than the container itself.

Use models that serve you without absorbing you. Open source models can process your data without it flowing to a third-party training pipeline. Choosing your model means you're not dependent on a single provider, and your data isn't a side payment for the service you're receiving.

Let AI amplify your context without giving it away. A knowledge library with semantic search lets you get the full benefit of AI-powered retrieval and synthesis across your accumulated research, without that research becoming the property of the system that searches it. The AI reads your library on your terms through controlled protocols rather than ingesting it as training data.

Make the compounding work in your favour. Every article you annotate, every note you write in your own words, every research log entry that documents what you predicted versus what you found, is a data point that makes your personal context richer and your AI interactions more useful. That compounding should benefit you, and the way to ensure it does is to keep the context in a system you control.


The time advantage

There's a timing dimension to this that's easy to overlook. The pharma companies who said no to contributing their data did so because they could see the trade clearly: decades of accumulated advantage being commoditised in exchange for early access to a tool. The knowledge workers who see the same trade clearly now and choose to build their context in systems they control are making a similar calculation.

The context compounds every day. A researcher who starts building their personal knowledge library in controlled infrastructure today will have six months of compounded context that can't be fast-forwarded, bought, or replicated. That context makes every AI interaction more useful, every search more specific, every synthesis more grounded in their actual work. And because it lives in their storage, under their terms, the compounding benefits flow entirely to them.

The alternative, building the same depth of context in a system controlled by a company whose interests may diverge from yours, is a bet on trust that grows riskier exactly in proportion to how valuable the accumulated context becomes.


Frequently asked questions

I'm just one person. Does this really apply to me? If your accumulated knowledge is what makes your work valuable, yes. The pharma example is dramatic in scale but identical in principle to a freelance consultant's client insights, a doctoral researcher's literature notes, or an independent analyst's industry understanding. The question is the same: is the system you're building up your knowledge in one that serves your interests, or one that might eventually serve someone else's?

What if I've already put years of work into a provider's tools? Start going forward rather than trying to extract everything retrospectively. New captures, new notes, new research go into a system you control. The existing material is still useful where it is, and much of it can be connected to your new library through integrations without needing a painful manual migration.

Does owning my data mean worse AI performance? Open source models have reached the point where the performance difference for most tasks is negligible. And when your AI has access to rich personal context (your research, your notes, your annotations), the relevance of the output often exceeds what a frontier model produces from a standing start with no context at all.


Related reading: The AI advantage isn't the model, it's the memory, Don't rent your intelligence, Where does your knowledge live?, How to actually do research, The research log. Related guides: Research workflow, Literature review, How people use Fabric.

The workspace that thinks with you.

Ready when you are.

The workspace that thinks with you.

Ready when you are.

The workspace that thinks with you.

Ready when you are.