The Machine is Killing the Data Tollbooth
The way we consume data is changing. Data providers need to evolve... fast.
There’s an argument circulating in the investment community. Every time I hear it, there’s a small part of me thinking something about it doesn’t hold up.
The argument goes like this: AI is not really eating software because the value still sits in the underlying data. Software companies remain the system of record. AI is just a layer on top, dependent on whatever sits beneath it.
The logic sounds reasonable. But I think it misses something important.
I ask you a simple question: “What’s eight multiplied by nine?”
You don’t calculate it. You don’t reach for a calculator. Your brain immediately says “seventy-two”. At some point in childhood, repetition burned that pattern into your memory. It stopped being arithmetic and became a reflexive recall function.
Large language models work in a surprisingly similar way.
When you ask an LLM to solve ‘eight multiplied by nine’, it also answers instantly.
If you picture the AI furiously multiplying eight by nine every time you or someone else asks, you’d be wrong. That would be absurdly inefficient. During training, the model has encountered ‘8 × 9 = 72’ so many times that the association becomes embedded in its weights. It becomes part of its neural fabric. When you ask the question, it remembers. Kind of the same way we do.
The model isn’t calculating. It’s recalling.
That distinction matters more than people realize.
Before AI, computing was algorithmic. A program running on a CPU required the computer to fetch data from source every time it ran. Data retrieval was part of the process. It would reach out across the internet, knock on an API’s door, and say, “Give me that piece of data, please.” And the data provider would say, “Certainly. That’ll be a fraction of a penny.” If the same fact was requested a million times, the provider monetized it a million times.
That model worked because software needed continuous access to the source of truth.
Now comes the era of the large language model. An LLM operates entirely differently. To be clear, it still burns an immense amount of computational power (FLOPs) during the query phase; every single token it generates requires billions of matrix multiplications across its weights. The AI then compresses information directly into its billions of algebraic parameters during training.
When a factual query is made, the answer emerges deterministically from those weights. It is statistical recall, not an API lookup loop. The model doesn’t know it has stored a discrete fact. It has simply seen the sequence so many times that the probability of “72” being the correct output has converged close to ‘1’.
Now imagine licensing a massive archive of historical stock prices to train a model. After training, the model can answer questions about Apple’s closing price on January 15, 2020 without querying a financial API. The information has effectively been absorbed into the network, synthetically memorized.
No API request. No toll booth. No recurring monetization.
The old model monetized a dumb algorithm that needed to look up the answer time and time again. AI learns and artificially remembers.
This explains that little voice casting a doubt in the back of my mind. AI is not as dependent on historic data as many imagine.
That is why I think “the data remains the moat” narrative is incomplete. In many cases, AI weakens the value of static data precisely because repeated retrieval is no longer necessary.
For companies built on charging for the same fact millions of times, that changes everything. It’s an existential crisis. It directly attacks the economic foundation of their business model. The uncomfortable question is simple: what exactly are you selling once the machine remembers?
Increasingly, the value is moving toward freshness, verification, and low latency.
An AI’s memory is a lossy compression library. You and I remember that eight times nine is 72, but if I asked you for the menu prices of a restaurant we visited in Omaha back in 2012, you’d look at me blankly. LLMs are also great at “8 x 9” because it’s ubiquitous, but, like us, they run into serious difficulty with rare or highly precise data.
Consider financial statement analysis. When trained on millions of rows of precise but similar numbers, details often blur. Because LLMs rely on statistical probabilities rather than exact database lookups, they can conflate digits or flat-out hallucinate.
That limitation matters enormously in fields like finance, law, medicine, and enterprise infrastructure, where a “near 1” statistical probability isn’t good enough. That lossiness is exactly why data providers still have some leverage.
That is why Retrieval-Augmented Generation (RAG) matters. In a RAG system, the model actively queries external sources during inference to retrieve verified or current information. The value is no longer the raw fact itself. The value is confidence that the answer is current and correct. It’s analogous to a lawyer checking a law book prior to offering counsel, rather than simply relying on memory.
The best run data businesses have recognized this stark new reality and are already pivoting.
Some are shifting toward large one-time licensing agreements for training data. Others are focusing on real-time information that cannot simply be memorized: live prices, inventory levels, sports scores, logistics data, market feeds and such like.
So, as an investment community, we should probably spend more time asking a simple question: which companies are adapting with humility, and which still believe the world will keep rewarding yesterday’s playbook?
Some management teams look at AI, changing consumer behaviour, shifting economics, and new competitive dynamics with genuine curiosity. They are willing to rethink assumptions, cannibalize parts of their own business, and accept that prior success does not guarantee future relevance. Those companies tend to evolve before they are forced to.
Others behave very differently. Years of dominance create institutional certainty. Processes harden. Incentives calcify. Management starts defending the existing model rather than questioning it. The company mistakes market leadership for permanence. That is usually when hubris begins to creep in. Management continues talking as though the old rules still apply.
And hubris is dangerous because it often looks rational in the moment. The incumbent still has scale. Still has margins. Still has customers. The decline rarely starts as a collapse. It starts as complacency. It happens slowly at first, then the cliff edge comes into view and it’s invariably far too late to change course.
AI has introduced an environmental shift. No question about it. Now, software companies need to evolve. It’s a Darwinian case of survival of the fittest; those that don’t evolve are likely to die out.
The difficult part for investors is separating temporary noise from genuine adaptability. Every company claims it’s embracing AI change. Yet very few actually restructure incentives, products, or business models for fear it will threaten their current cash flows.
The approaches of financial data businesses differ enormously. If you’re a thin UI layer on top of a system of record, you’re going to have to earn your keep. Dow Jones / Factiva is clearly adapting, packaging its content for GenAI use and licensing it for AI models and solutions. In stark contrast, many others still seem focused on defending legacy walled gardens. Bloomberg, FactSet, and LSEG are responding by embedding AI into their own product stacks rather than fully leaning into data licensing as a broader AI distribution opportunity.
The companies most likely to survive major transitions are often the ones most willing to disrupt themselves early. The ones most at risk are usually the ones still explaining why disruption supposedly doesn’t apply to them.
So I’m curious,
Which software/data companies do you think are showing real foresight and humility today?
And which ones seem increasingly trapped by the assumptions created by their own past success?
Please leave your answers in the comments section below.





$CSGP
it has tons of well-protected in house very niche RE industry data moat..
you can not google it and GET IT FREE
YOU need PAY ...
If the company owns the data and the llm e.g. RELX, then they have a clear moat and are likely to become net beneficiaries of AI. If the data has to be precise e. g. Sage, again there is a clear moat. If neither are true, and especially if the application is easy to reproduce, e.g. Salesforce, Outsystems, then they are in trouble IMO. BTW the LLM providers themselves fall into this category - it is increasingly easy to transfer work and apps between, say Claude and Chatgpt. That's one of the key reasons they are having to invest so much. If they had a moat they wouldn't need to.