·
Opinion
The Centralisation Pattern
AI Is Incredible Right Now. What Comes Next Will Be Everywhere.
Every dominant tech platform in history evolved into something more distributed, more accessible, more embedded. AI is following the same curve.
Before you read on — quick question
Where do you honestly see AI in your organisation 5 years from now?
I’ve spent the last decade building real-time systems for capital markets. Order books, event streams, distributed trade blotters. The kind of work where you live and die by latency and you learn pretty quickly what architectural decisions cost you two years down the line. So when I look at where AI is right now, I see something I recognise. Not the technology itself, that part is genuinely new and genuinely impressive. But the shape of it. The way it’s organised. The stage it’s at.
I’ve seen this stage before. And history gives us a pretty clear view of what tends to come next.
There is a pattern that keeps showing up across every major shift in computing. Something powerful gets built. People get tremendous value from it. Then over time, someone figures out how to deliver the same capability cheaper, smaller, more distributed. The original technology doesn’t disappear, it evolves and finds its right level. The companies that built it adapt and keep being relevant. But the centre of gravity shifts. And the ecosystem that grows up around the next phase is usually much bigger than the one before it.
AI is at that early centralised phase right now. Which is genuinely exciting, because if the pattern holds, what comes next will be even bigger.
Section I
This Has Happened Before. Repeatedly.
Start with something simple. Lighting. You probably don’t spend much time thinking about the history of how humans got light into rooms, but the progression is worth a quick look:
Lighting
Same outcome the whole way through: light in a room. Each generation does it with less energy, more control, smaller hardware, lower cost. The complexity gets buried deeper each time until it just disappears. Nobody building a modern stadium is thinking about filaments. They’re thinking about lumens and power draw.
Data went through the exact same compression:
Data
And programming languages, which is where the parallel to AI gets really interesting:
Programming Languages
Every time someone moved up that ladder, the people on the rung below said the same things. “You lose control.” “The performance isn’t there.” “Real work needs assembly / C / a proper type system.” They were sometimes right about the technical trade-offs. They were wrong about what the market actually needed. The market needed the next rung, not a better version of the current one.
“Each generation of abstraction didn’t make the technology worse. It made it invisible. And invisible is the only form that scales to everyone.”
Section II
We Watched This Happen to Software Architecture and Still Didn’t Learn
Between 2010 and 2020, software architecture went through a version of this that I lived through professionally. The monolith vs microservices debate sounds academic until you’ve been in the room where a team can’t ship a feature for two months because it touches code owned by four other teams. Then it becomes very practical very fast.
Monolith
Everything lives together. One deploy, one team’s mess tangled with every other team’s mess. You want to scale one feature, you scale everything and pay for all of it.
Microservices
Each service owns its domain. Ships independently, scales independently, can be swapped out without touching anything else. Complexity managed at the boundary, not dissolved into a single pile.
Centralised Compute
One mainframe, everyone queuing for time on it. Expensive to access, controlled by whoever owns the machine, single point of everything including failure.
Distributed Ledger
No central authority, no single point of control, no single point of failure. Strip the speculation out and what’s left is just distributed computing, which predates Bitcoin by decades.
The interesting thing about distributed ledger tech isn’t cryptocurrency. It’s the underlying idea: when you don’t need a central authority to validate state, you remove the central cost too. That principle keeps showing up, in different forms, across different domains. It showed up in peer-to-peer file sharing. It showed up in edge computing. It’s about to show up in AI in a very big way.
Section III
What AI Looks Like Right Now
The current setup is remarkable. A handful of organisations have done something genuinely hard: trained models at a scale that produces real intelligence, made them accessible via API, and built products that millions of people use every day. Claude, ChatGPT, and Gemini are legitimately capable tools. I use them daily. Most engineers I know do.
And structurally, right now, they work the same way the mainframe did. Powerful centralised hardware, accessed remotely, paying for compute time. That’s not a criticism. It’s just where the technology is at this point in the curve. The mainframe was also a genuinely useful tool in its era. The interesting question is what the next phase looks like.
AI Platform Evolution
MCP (Model Context Protocol) is interesting as a signal. It’s the industry’s first serious attempt at a standardised interface between AI models and the tools they use. That matters not because it fixes the centralisation problem, but because standardised interfaces are historically what make decentralisation possible. You can’t distribute something that nothing can consistently talk to. MCP might be the TCP/IP moment for AI interoperability. Or it might not. But someone building something like it was inevitable.
The question the big labs are not asking out loud is: what happens when the standardised interface means models become interchangeable?
“The companies selling AI API access today built something genuinely remarkable. The interesting question is what they build next, because the pattern says the next phase is always bigger than the current one.”
Section IV
It’s Already Running in a Browser Tab
I want to be specific here because vague predictions about “the future of AI” are everywhere and most of them are useless. So let me tell you what is happening right now, today, on consumer hardware.
A 3 billion parameter language model runs entirely inside a browser tab. No API. No server. No data leaving the device. It streams tokens in real time via WebGPU, handles multi-turn conversation, and can be interrupted mid-generation without losing what it had already produced. The whole thing downloads once, lives in browser cache, and runs for free forever after that.
Two years ago, 3B parameters meant serious research infrastructure. Today it’s a weekend project that runs on a mid-range laptop. The hardware keeps getting faster. The models keep getting more efficient per parameter. The tooling keeps improving. At some point these lines cross and the API model stops being the obvious default.
The model is the new runtime. ONNX quantised weights are the new bytecode. Transformers.js is close enough to what the JVM was for Java that the comparison holds. The abstraction ladder is being built, one rung at a time, right now.
Section V
The Infrastructure Is Already There. It Just Doesn’t Have a Name Yet.
There won’t be a big announcement. This transition is already happening in the background, quietly, the same way most real infrastructure shifts happen.
Look at what already exists today. Hugging Face hosts over half a million models. ONNX is a standardised model format that runs across hardware. Transformers.js brings inference to the browser. Ollama lets you run models locally with one command. llama.cpp runs quantised models on CPUs that were never designed for AI workloads. These things aren’t prototypes. They’re in production use right now.
The current integration pattern still looks like this:
const response = await fetch('https://api.openai.com/v1/chat/completions', {
headers: { 'Authorization': `Bearer ${process.env.OPENAI_API_KEY}` },
// your data just left the building
});
But the building blocks for something very different already exist:
import { pipeline } from '@huggingface/transformers';
// downloads once, cached in browser, runs local forever after
const classifier = await pipeline('sentiment-analysis',
'Xenova/distilbert-base-uncased-finetuned-sst-2-english'
);
That’s not pseudocode. That’s real code running in production today. The model ID is a bit verbose, the packaging is rough around the edges, and the developer experience still needs work. But the core of it is there. Someone classifying text in a browser, with no API call, no cost per request, no data leaving the device.
What’s missing isn’t the technology. It’s the convention layer on top of it. A shared understanding of how these components get versioned, how they get composed, how you swap one model for a better one without changing the rest of your code. npm didn’t invent the idea of reusable JavaScript modules. It just made the convention clear enough that everyone adopted it. That’s what AI tooling still needs. And when it arrives, it won’t feel like a revolution. It’ll feel obvious in hindsight, the way npm feels obvious now.
Section VI
What Happens to the Big Labs
The honest answer is: some of them adapt and some of them don’t. The ones that built the original mainframes didn’t all disappear. IBM is still around. But IBM stopped being the centre of the computing universe a long time ago, and the companies that built their entire strategy around “IBM will always be the centre” mostly aren’t around anymore.
The frontier models will survive. GPT-5-class, Claude-class capability for genuinely hard problems, multi-modal reasoning, research-grade tasks, the stuff that actually needs a 70B parameter model running on a cluster. That market is real and probably grows.
But 70% of what people currently pay per-token API costs for? Sentiment analysis, summarisation, classification, code suggestions, entity extraction, question answering over a document, customer support routing. All of that is going to move local. The models are already good enough. The economics will force it. The privacy regulations will accelerate it.
-
1
Routine tasks go local within 3-4 years. Classification, summarisation, code completion. Small purpose-built models, no API call, no cost per use. The 0.5B model that today people dismiss as a toy will be the workhorse of half the business software market. -
2
Frontier models stay relevant but get repositioned. Complex reasoning, synthesis, genuinely novel tasks. The same way mainframes kept running after the PC arrived, just no longer as the default computing platform for everyone. -
3
Someone builds the npm for AI models and whoever does it first with the right licensing model will reshape this industry the way npm reshaped web development. It’s a land-grab waiting for the right timing and the right execution. -
4
Compliance forces the issue in regulated sectors. Healthcare, legal, finance will get pushed to local inference by data sovereignty requirements before economics get them there. Regulators will move faster than the technology roadmap. -
5
The device layer is the real frontier. Apple’s Neural Engine is the canary. Dedicated on-device AI accelerators will make capable local inference a standard feature of every laptop and phone within a few product cycles.
This isn’t hypothetical — it’s running in a browser tab today
LiveLens is a small proof of where this is going: object detection on TensorFlow.js, real Python analytics via Pyodide, and a language model running fully offline via Transformers.js and WebGPU, all in one browser tab. No server, no API cost, nothing leaving the device. It’s not a production AI system. It’s a demonstration that the pattern already works on hardware you already own. Read the technical breakdown here.
Conclusion
The Abstraction Always Wins
Computing history is really just the history of complexity being buried in layers until it stops being visible. Fire became LED. Assembly became WASM. Mainframes became microservices. The technology that once required a specialist and a data centre becomes a utility that anyone can run, anywhere, for close to nothing.
AI follows the same pattern.
The large language model sitting on a remote server, billed per token, controlled by a company whose roadmap you cannot predict and whose pricing you cannot negotiate, is the penultimate step. Not the final one. The final one is the model running on your machine, owned by nobody, imported like a package, doing its specific job quietly and efficiently while your application gets on with doing whatever it’s actually for.
The companies that see this evolution coming and invest ahead of it will shape what the next phase looks like. The good ones always have. IBM didn’t disappear when the PC arrived. They shifted. AWS didn’t disappear when containers arrived. They built the best container platform. The pattern isn’t about disruption in the dramatic sense. It’s about the centre of gravity moving, and the smart organisations moving with it.
The shift doesn’t happen because someone writes a think piece about it. It happens because a relatively small number of people build the framework foundations that make the new paradigm actually usable for everyone else. Someone built the JVM before Java went mainstream. Someone built Webpack before modern frontend became possible. Someone built React before half the web’s UI ran on it. The ideas came first, but the infrastructure had to follow.
That infrastructure work for AI is what’s genuinely interesting right now. Not which model wins, not which API is cheapest. The architectural patterns that let AI components compose cleanly, integrate with existing systems, scale from local dev to distributed production. That’s the still-open problem. And it looks a lot like the problems that got solved in previous transitions, if you’ve been in the industry long enough to see the shape of it.
My answer to the question at the top, by the way: C. Hybrid. Frontier models for the genuinely hard problems where the cost is worth it, small purpose-built local models for the 70% of tasks that don’t need a 70B parameter model sitting in a data centre. That’s where I see this landing. It’s an architecture problem as much as a model problem.
The abstraction always wins. The only question is who builds the layer it runs on.