Rajnish Noonia

Category: Architecture

  • AI Is Incredible Right Now. What Comes Next Will Be Everywhere.

    Pixytech
    ·
    Opinion

    The Centralisation Pattern

    AI Is Incredible Right Now. What Comes Next Will Be Everywhere.

    Every dominant tech platform in history evolved into something more distributed, more accessible, more embedded. AI is following the same curve.

    Before you read on — quick question

    Where do you honestly see AI in your organisation 5 years from now?

    • A Still primarily cloud APIs (OpenAI, Claude, Gemini) — more capable, better priced
    • B Mostly local or embedded models running on-device, no per-token cost
    • C Hybrid — frontier models for genuinely complex work, local for everything routine
    • D Something we haven’t fully figured out yet

    I’ve spent the last decade building real-time systems for capital markets. Order books, event streams, distributed trade blotters. The kind of work where you live and die by latency and you learn pretty quickly what architectural decisions cost you two years down the line. So when I look at where AI is right now, I see something I recognise. Not the technology itself, that part is genuinely new and genuinely impressive. But the shape of it. The way it’s organised. The stage it’s at.

    I’ve seen this stage before. And history gives us a pretty clear view of what tends to come next.

    There is a pattern that keeps showing up across every major shift in computing. Something powerful gets built. People get tremendous value from it. Then over time, someone figures out how to deliver the same capability cheaper, smaller, more distributed. The original technology doesn’t disappear, it evolves and finds its right level. The companies that built it adapt and keep being relevant. But the centre of gravity shifts. And the ecosystem that grows up around the next phase is usually much bigger than the one before it.

    AI is at that early centralised phase right now. Which is genuinely exciting, because if the pattern holds, what comes next will be even bigger.

    Section I

    This Has Happened Before. Repeatedly.

    Start with something simple. Lighting. You probably don’t spend much time thinking about the history of how humans got light into rooms, but the progression is worth a quick look:

    Lighting

    Fire
    Gas Lamp
    Bulb
    Tube Light
    LED
    Micro LED

    Same outcome the whole way through: light in a room. Each generation does it with less energy, more control, smaller hardware, lower cost. The complexity gets buried deeper each time until it just disappears. Nobody building a modern stadium is thinking about filaments. They’re thinking about lumens and power draw.

    Data went through the exact same compression:

    Data

    Flat Files
    SQL + Norms
    Document DB
    Data Lakes
    Lakehouses

    And programming languages, which is where the parallel to AI gets really interesting:

    Programming Languages

    Assembly
    C / C++
    Java / C#
    JavaScript
    WASM
    Runtime Anywhere

    Every time someone moved up that ladder, the people on the rung below said the same things. “You lose control.” “The performance isn’t there.” “Real work needs assembly / C / a proper type system.” They were sometimes right about the technical trade-offs. They were wrong about what the market actually needed. The market needed the next rung, not a better version of the current one.

    “Each generation of abstraction didn’t make the technology worse. It made it invisible. And invisible is the only form that scales to everyone.”

    Section II

    We Watched This Happen to Software Architecture and Still Didn’t Learn

    Between 2010 and 2020, software architecture went through a version of this that I lived through professionally. The monolith vs microservices debate sounds academic until you’ve been in the room where a team can’t ship a feature for two months because it touches code owned by four other teams. Then it becomes very practical very fast.

    Monolith

    Everything lives together. One deploy, one team’s mess tangled with every other team’s mess. You want to scale one feature, you scale everything and pay for all of it.

    Microservices

    Each service owns its domain. Ships independently, scales independently, can be swapped out without touching anything else. Complexity managed at the boundary, not dissolved into a single pile.

    Centralised Compute

    One mainframe, everyone queuing for time on it. Expensive to access, controlled by whoever owns the machine, single point of everything including failure.

    Distributed Ledger

    No central authority, no single point of control, no single point of failure. Strip the speculation out and what’s left is just distributed computing, which predates Bitcoin by decades.

    The interesting thing about distributed ledger tech isn’t cryptocurrency. It’s the underlying idea: when you don’t need a central authority to validate state, you remove the central cost too. That principle keeps showing up, in different forms, across different domains. It showed up in peer-to-peer file sharing. It showed up in edge computing. It’s about to show up in AI in a very big way.

    Section III

    What AI Looks Like Right Now

    The current setup is remarkable. A handful of organisations have done something genuinely hard: trained models at a scale that produces real intelligence, made them accessible via API, and built products that millions of people use every day. Claude, ChatGPT, and Gemini are legitimately capable tools. I use them daily. Most engineers I know do.

    And structurally, right now, they work the same way the mainframe did. Powerful centralised hardware, accessed remotely, paying for compute time. That’s not a criticism. It’s just where the technology is at this point in the curve. The mainframe was also a genuinely useful tool in its era. The interesting question is what the next phase looks like.

    AI Platform Evolution

    Research Labs
    OpenAI API
    ChatGPT / Claude
    MCP + Agents
    Offline Models
    Embedded / Local

    MCP (Model Context Protocol) is interesting as a signal. It’s the industry’s first serious attempt at a standardised interface between AI models and the tools they use. That matters not because it fixes the centralisation problem, but because standardised interfaces are historically what make decentralisation possible. You can’t distribute something that nothing can consistently talk to. MCP might be the TCP/IP moment for AI interoperability. Or it might not. But someone building something like it was inevitable.

    The question the big labs are not asking out loud is: what happens when the standardised interface means models become interchangeable?

    “The companies selling AI API access today built something genuinely remarkable. The interesting question is what they build next, because the pattern says the next phase is always bigger than the current one.”

    Section IV

    It’s Already Running in a Browser Tab

    I want to be specific here because vague predictions about “the future of AI” are everywhere and most of them are useless. So let me tell you what is happening right now, today, on consumer hardware.

    A 3 billion parameter language model runs entirely inside a browser tab. No API. No server. No data leaving the device. It streams tokens in real time via WebGPU, handles multi-turn conversation, and can be interrupted mid-generation without losing what it had already produced. The whole thing downloads once, lives in browser cache, and runs for free forever after that.

    Two years ago, 3B parameters meant serious research infrastructure. Today it’s a weekend project that runs on a mid-range laptop. The hardware keeps getting faster. The models keep getting more efficient per parameter. The tooling keeps improving. At some point these lines cross and the API model stops being the obvious default.

    The model is the new runtime. ONNX quantised weights are the new bytecode. Transformers.js is close enough to what the JVM was for Java that the comparison holds. The abstraction ladder is being built, one rung at a time, right now.

    Section V

    The Infrastructure Is Already There. It Just Doesn’t Have a Name Yet.

    There won’t be a big announcement. This transition is already happening in the background, quietly, the same way most real infrastructure shifts happen.

    Look at what already exists today. Hugging Face hosts over half a million models. ONNX is a standardised model format that runs across hardware. Transformers.js brings inference to the browser. Ollama lets you run models locally with one command. llama.cpp runs quantised models on CPUs that were never designed for AI workloads. These things aren’t prototypes. They’re in production use right now.

    The current integration pattern still looks like this:

    const response = await fetch('https://api.openai.com/v1/chat/completions', {
      headers: { 'Authorization': `Bearer ${process.env.OPENAI_API_KEY}` },
      // your data just left the building
    });

    But the building blocks for something very different already exist:

    import { pipeline } from '@huggingface/transformers';
    
    // downloads once, cached in browser, runs local forever after
    const classifier = await pipeline('sentiment-analysis',
      'Xenova/distilbert-base-uncased-finetuned-sst-2-english'
    );

    That’s not pseudocode. That’s real code running in production today. The model ID is a bit verbose, the packaging is rough around the edges, and the developer experience still needs work. But the core of it is there. Someone classifying text in a browser, with no API call, no cost per request, no data leaving the device.

    What’s missing isn’t the technology. It’s the convention layer on top of it. A shared understanding of how these components get versioned, how they get composed, how you swap one model for a better one without changing the rest of your code. npm didn’t invent the idea of reusable JavaScript modules. It just made the convention clear enough that everyone adopted it. That’s what AI tooling still needs. And when it arrives, it won’t feel like a revolution. It’ll feel obvious in hindsight, the way npm feels obvious now.

    Section VI

    What Happens to the Big Labs

    The honest answer is: some of them adapt and some of them don’t. The ones that built the original mainframes didn’t all disappear. IBM is still around. But IBM stopped being the centre of the computing universe a long time ago, and the companies that built their entire strategy around “IBM will always be the centre” mostly aren’t around anymore.

    The frontier models will survive. GPT-5-class, Claude-class capability for genuinely hard problems, multi-modal reasoning, research-grade tasks, the stuff that actually needs a 70B parameter model running on a cluster. That market is real and probably grows.

    But 70% of what people currently pay per-token API costs for? Sentiment analysis, summarisation, classification, code suggestions, entity extraction, question answering over a document, customer support routing. All of that is going to move local. The models are already good enough. The economics will force it. The privacy regulations will accelerate it.

    • 1
      Routine tasks go local within 3-4 years. Classification, summarisation, code completion. Small purpose-built models, no API call, no cost per use. The 0.5B model that today people dismiss as a toy will be the workhorse of half the business software market.
    • 2
      Frontier models stay relevant but get repositioned. Complex reasoning, synthesis, genuinely novel tasks. The same way mainframes kept running after the PC arrived, just no longer as the default computing platform for everyone.
    • 3
      Someone builds the npm for AI models and whoever does it first with the right licensing model will reshape this industry the way npm reshaped web development. It’s a land-grab waiting for the right timing and the right execution.
    • 4
      Compliance forces the issue in regulated sectors. Healthcare, legal, finance will get pushed to local inference by data sovereignty requirements before economics get them there. Regulators will move faster than the technology roadmap.
    • 5
      The device layer is the real frontier. Apple’s Neural Engine is the canary. Dedicated on-device AI accelerators will make capable local inference a standard feature of every laptop and phone within a few product cycles.

    This isn’t hypothetical — it’s running in a browser tab today

    LiveLens is a small proof of where this is going: object detection on TensorFlow.js, real Python analytics via Pyodide, and a language model running fully offline via Transformers.js and WebGPU, all in one browser tab. No server, no API cost, nothing leaving the device. It’s not a production AI system. It’s a demonstration that the pattern already works on hardware you already own. Read the technical breakdown here.

    Conclusion

    The Abstraction Always Wins

    Computing history is really just the history of complexity being buried in layers until it stops being visible. Fire became LED. Assembly became WASM. Mainframes became microservices. The technology that once required a specialist and a data centre becomes a utility that anyone can run, anywhere, for close to nothing.

    AI follows the same pattern.

    The large language model sitting on a remote server, billed per token, controlled by a company whose roadmap you cannot predict and whose pricing you cannot negotiate, is the penultimate step. Not the final one. The final one is the model running on your machine, owned by nobody, imported like a package, doing its specific job quietly and efficiently while your application gets on with doing whatever it’s actually for.

    The companies that see this evolution coming and invest ahead of it will shape what the next phase looks like. The good ones always have. IBM didn’t disappear when the PC arrived. They shifted. AWS didn’t disappear when containers arrived. They built the best container platform. The pattern isn’t about disruption in the dramatic sense. It’s about the centre of gravity moving, and the smart organisations moving with it.

    The shift doesn’t happen because someone writes a think piece about it. It happens because a relatively small number of people build the framework foundations that make the new paradigm actually usable for everyone else. Someone built the JVM before Java went mainstream. Someone built Webpack before modern frontend became possible. Someone built React before half the web’s UI ran on it. The ideas came first, but the infrastructure had to follow.

    That infrastructure work for AI is what’s genuinely interesting right now. Not which model wins, not which API is cheapest. The architectural patterns that let AI components compose cleanly, integrate with existing systems, scale from local dev to distributed production. That’s the still-open problem. And it looks a lot like the problems that got solved in previous transitions, if you’ve been in the industry long enough to see the shape of it.

    My answer to the question at the top, by the way: C. Hybrid. Frontier models for the genuinely hard problems where the cost is worth it, small purpose-built local models for the 70% of tasks that don’t need a 70B parameter model sitting in a data centre. That’s where I see this landing. It’s an architecture problem as much as a model problem.

    The abstraction always wins. The only question is who builds the layer it runs on.

    Rajnish Noonia

    Lead Architect  ·  Full Stack Engineer  ·  London

    Over 24 years I’ve had one recurring job: build the platform layer that the next generation of applications stands on. That meant WPF frameworks when the industry was moving from desktop to web. It meant micro-frontend and SDK architecture when monoliths were breaking apart into services. It meant real-time event-driven platforms for trading systems at hedge funds, tier-one investment banks and industry bodies, where the margin for getting the architecture wrong is measured in actual money, in real time.

    Each of those transitions looked different on the surface. Underneath they had the same shape: a powerful centralised thing was getting replaced by something distributed and composable, and the teams that would build on top of the new thing needed foundations they could trust. That’s the work I find interesting, and it keeps showing up in different forms across every generation of the industry.

    The AI transition is the same shape again, just bigger and faster. I’ve been building toward it for the last year and I’ll keep writing about it here. If you’re part of an organisation thinking seriously about the architectural foundations for what AI looks like when it stops being an API and starts being infrastructure, I’m happy to have that conversation.

  • LiveLens: Two Languages, One Browser Tab, Zero Servers — Now with In-Browser AI Chat

    A real-time trading blotter and a webcam object detector have more in common than they look like they should. Both are event-driven: a stream of updates arrives continuously, something needs to render the current state without falling behind, and something else needs to watch that stream for patterns worth flagging – a position that’s moved too far, a count that’s spiked. I’ve spent most of the last decade building the first kind of system for capital markets. LiveLens is what happens when you point the same architectural instincts at a webcam instead of a market feed, and run the whole thing client-side, in a single browser tab, with no backend at all.

    Since the first version the app has grown quite a bit. What started as an object detection demo with a Python analytics layer now also includes a full AI chat interface that runs either entirely in the browser (no API key, no server) or connects to Claude, OpenAI, or Gemini via their streaming APIs. All in the same tab.

    What it does

    Grant camera access and LiveLens detects objects in the live feed, tracks each one’s identity across frames, and continuously computes rolling statistics: per-class counts, average dwell time, flagging anything that looks like a genuine anomaly rather than normal frame-to-frame noise. None of the video, and none of the analysis, ever leaves the tab.

    Flip to the Chat tab and you get a conversational AI interface that lets you run language models completely offline in the browser, or switch over to a cloud provider if you need more horsepower. The two modes share the same UI, same settings panel, same streaming experience. It’s one interface with a source selector, not two separate features bolted together.

    Three layers, three tools, on purpose

    It would be tempting to reach for one language and use it for everything. I split LiveLens into three layers instead, each running the tool that’s actually good at that job:

    • Perception: TensorFlow.js runs a MobileNet-based object detector (coco-ssd) directly in the browser, WebGL-accelerated.
    • Identity: a small TypeScript tracker turns a stream of unlabelled per-frame detections into persistent tracked objects.
    • Statistics: a genuine Python module, running in-browser via Pyodide, does the counting, aggregation, and anomaly detection with pandas and numpy.
    • Chat: Transformers.js v3 runs ONNX-quantised language models directly in the browser via WebGPU or WASM, no API key required. Optionally routes to Claude, OpenAI or Gemini instead.

    Why the neural net runs in JavaScript, not Python

    This is worth being upfront about, because “Python in the browser” is the headline feature and it would be easy to assume the model runs there too. It doesn’t, deliberately.

    In-browser neural-net inference is a mature, well-supported path in the JavaScript ecosystem. TensorFlow.js and transformers.js both have WebGL/WebGPU-accelerated runtimes, wide model support, and years of production hardening. Pyodide, which compiles the CPython interpreter and the scientific Python stack to WebAssembly, doesn’t have an equivalent neural-net execution engine yet. Asking it to run the detector would mean either a much slower pure-Python inference path or shipping a second runtime just to reinvent what TensorFlow.js already does well.

    So the detector runs where the ecosystem is strongest, and Python gets used for what it’s actually strong at: data manipulation and statistics. That division of labour is the whole point of the architecture, not an incidental detail.

    The tracker: turning detections into objects

    coco-ssd gives you a fresh, unlabelled set of bounding boxes on every frame. It has no concept of “the same coffee cup as last frame” – as far as the model is concerned, every detection is a new observation with no history.

    To compute dwell time, or to feed any kind of coherent time series into the analytics layer, that identity has to come from somewhere. LiveLens uses a lightweight IoU-based tracker: on each frame, every new detection is matched to the closest existing track of the same class by bounding-box overlap, using greedy matching against an intersection-over-union threshold.

    function intersectionOverUnion(a, b) {
    const [ax, ay, aw, ah] = a;
    const [bx, by, bw, bh] = b;
    const x1 = Math.max(ax, bx);
    const y1 = Math.max(ay, by);
    const x2 = Math.min(ax + aw, bx + bw);
    const y2 = Math.min(ay + ah, by + bh);
    const interArea = Math.max(0, x2 - x1) * Math.max(0, y2 - y1);
    const unionArea = aw * ah + bw * bh - interArea;
    return unionArea <= 0 ? 0 : interArea / unionArea;
    }

    Tracks that go unmatched for more than 1.5 seconds are dropped. Long enough to bridge a missed frame or brief occlusion without inventing a new identity for the same physical object, short enough that a track doesn’t linger indefinitely after something actually leaves the frame. It’s a simplified relative of the greedy-matching idea behind SORT-style trackers, without a Kalman filter. Accurate enough for a browser demo, not for a production tracking system.

    Pyodide: real pandas, real numpy, zero server

    This is the part that made the project worth writing up. Pyodide loads the full CPython interpreter, plus numpy and pandas, compiled to WebAssembly, and runs it in the same tab as the React app:

    def analyze(window_json: str) -> str:
    objects = json.loads(window_json)
    df = pd.DataFrame(objects)
    counts = df.groupby("class")["id"].nunique().to_dict()
    df["dwell"] = df["last_seen"] - df["first_seen"]
    dwell_ms = df.groupby("class")["dwell"].mean().to_dict()
    anomalies = []
    for cls, count in counts.items():
    hist = _history[cls]
    if len(hist) >= 5:
    mean, std = np.mean(hist), np.std(hist)
    z = (count - mean) / std if std > 0 else 0
    if abs(z) >= 2.5:
    anomalies.append({"class": cls, "count": count, "z": round(z, 2)})
    hist.append(count)
    return json.dumps({"counts": counts, "dwell_ms": dwell_ms, "anomalies": anomalies})

    _history is module-level state inside the Pyodide runtime, so it persists across calls for the life of the tab without React having to manage a rolling buffer itself. React just calls analyze() on an interval with the current tracked-object snapshot and gets back a JSON summary, exactly as if it were talking to a REST endpoint. The difference is there’s no network hop, no serialization over the wire beyond a function call boundary, and no server to deploy or pay for.

    The anomaly check itself is deliberately simple: a rolling per-class count history, a z-score against the mean and standard deviation of the last ~60 windows, and a threshold. It won’t catch anything subtle, but it’s the same basic shape as the alerting logic behind far more sophisticated systems, computed live, over a real stream, with real statistics.

    The AI Chat layer: running language models in the browser

    The newer addition to LiveLens is a chat tab that lets you talk to a language model without leaving the browser. There’s no proxy, no serverless function, nothing phoning home. The model weights download once, get cached by the browser, and run locally on WebGPU or fall back to WASM if WebGPU isn’t available.

    Under the hood it’s Transformers.js v3 running ONNX-quantised models in a Web Worker so the inference doesn’t block the UI thread. The worker loads the pipeline, receives messages, streams tokens back, and the main thread just renders whatever it receives. It’s a clean separation: the camera detection and Python analytics keep running even while the language model is thinking.

    Which models are available offline

    The model picker offers two groups, general purpose and coding, so you can pick based on what you’re actually doing:

    • SmolLM2 135M / 360M / 1.7B: HuggingFace’s small but surprisingly capable series, good for quick questions and demos. The 135M model loads in seconds even on slower connections.
    • Qwen 2.5 0.5B / 1.5B: strong multilingual reasoning from Alibaba, fits comfortably in browser memory.
    • Llama 3.2 1B: Meta’s smallest Llama, good general-purpose model, 2GB download.
    • Phi 3.5 Mini: Microsoft’s model that punches above its weight for reasoning and code, 2.2GB.
    • Qwen2.5-Coder 0.5B / 1.5B / 3B: dedicated code models that are genuinely useful for explaining snippets, debugging, and answering tech questions without sending your code to a third-party server.

    Switching to a coding model automatically swaps the system prompt to something more appropriate. It makes a noticeable difference in the quality of responses for technical questions. The system prompt is fully editable anyway, so you can tune it however you want.

    Cloud API fallback: Claude, OpenAI, Gemini

    The offline models are genuinely useful but they have obvious limits. A 1B parameter model is not going to match GPT-4o on complex reasoning tasks. So the same chat interface also supports switching to a cloud provider, using the exact same streaming UX.

    All three providers stream tokens over server-sent events. There’s a readSSE async generator that handles the data: line parsing, [DONE] termination, and abort signals, and each provider has its own thin wrapper on top that handles the auth header and extracts the right field from the delta object:

    • Claude: API key stored in localStorage, anthropic-dangerous-direct-browser-access: true header to allow direct browser calls, extracts content_block_delta.text_delta from the SSE stream.
    • OpenAI: API key in localStorage, standard chat/completions stream, extracts choices[0].delta.content.
    • Gemini: uses OAuth 2.0 via Google Identity Services rather than an API key, because Gemini supports it and it’s a nicer UX than asking users to generate and paste service account credentials. You provide an OAuth client ID, click Sign in with Google, and it handles the token flow. Roles get mapped to Gemini’s convention (assistant becomes model) before the request goes out.

    API keys and the client ID are stored in localStorage only. They never leave the browser except in the Authorization header of the request to the respective API. No telemetry, no backend logging, nothing persisted anywhere else.

    Settings UX: floating panels, not a modal

    The settings live behind two icon buttons in the chat topbar: a gear icon for model and provider config, and a pencil icon for the system prompt. Clicking either one opens a floating panel that overlays the messages area without pushing any content around or blocking interaction with the rest of the page. The panel is non-modal, so you can still scroll through the conversation and type in the input while it’s open.

    Keeping them separate felt right. Model selection is something you do once per session. The system prompt is something you might actually want to tweak mid-conversation, so giving it its own dedicated panel makes it feel less buried.

    Interrupt and send

    One thing that bothered me about a lot of chat demos is that the Send button disappears while the model is responding and you can’t send anything new until it’s done. If you realise mid-response that your question was badly phrased, or you want to follow up immediately, you have to wait.

    In LiveLens you can send a new message at any point. If generation is in progress, the current response gets aborted, the partial text gets saved into the conversation history, and the new request starts immediately. The button label changes to “Interrupt & Send” during generation so it’s clear what’s going to happen. There’s also a separate Stop button if you just want to halt the current response without sending anything new.

    For offline models the abort signal goes to the Web Worker which terminates the generation loop. For API sources it cancels the fetch via AbortController. Either way the partial text is already in state when the abort fires so nothing gets lost.

    Context management

    A small token usage indicator in the input bar shows roughly what percentage of the model’s context window is consumed. It turns amber at 80%. At that point you probably want to either clear the context or be a bit more concise. “Clear context” is one click and it’s right there next to the usage pill, not buried in a settings menu somewhere.

    The app also remembers the last selected tab, model, and provider across reloads. The offline model restarts loading automatically on page load if that was what you had selected. It’s already cached locally after the first download so it’s fast.

    Why this matters beyond a toy demo

    This is the same shape as a real-time trading blotter watching market data for anomalous price moves, or an operations dashboard flagging unusual order flow: a continuous event stream, an identity/state layer sitting between raw events and anything downstream, and a statistics layer watching for deviation from a rolling baseline. The domain changed from financial instruments to webcam detections. The architecture didn’t. Being able to build that shape entirely client-side, with a real statistical computing stack and no backend at all, says something about how far browser runtimes have come.

    The chat layer adds something different: a language model that runs in the same tab as the rest of the app, that can be pointed at what the detection pipeline is seeing, and that doesn’t require an account or API key to use. Whether that’s useful in production depends on the use case, but the fact that it’s a viable option at all is worth noting. A 1B parameter model that runs locally in a browser tab and streams tokens in real time without a server would have been a strange claim to make even two or three years ago.

    Practical implications

    Detection quality depends on lighting and the model’s speed/accuracy trade-off. lite_mobilenet_v2 is tuned to keep up with live video, not to win accuracy benchmarks. Expect it to miss small or partially occluded objects.

    WebGL matters. TensorFlow.js prefers a WebGL backend; without it the app falls back to WASM, which is noticeably slower. Most modern browsers and GPUs handle this fine, but it’s worth knowing if performance looks off on an older machine.

    Pyodide’s first load has a real cost. Fetching the CPython interpreter plus numpy and pandas as WASM packages is a few megabytes of one-time download, cached by the browser afterwards. The analytics panel shows “Loading Python…” for exactly this reason – it’s not a bug, it’s an honest reflection of what’s happening.

    The tracker is intentionally simple. Greedy IoU matching with no motion prediction means fast-moving objects or heavy occlusion can split one physical object into two tracked IDs. A production system would add a motion model; a demo doesn’t need one to make the point.

    Offline model quality scales with size. The 135M SmolLM2 is genuinely fast but don’t expect it to write production code or reason through anything complicated. The 3B Qwen-Coder is a lot more capable but it’s a 3GB download, so pick based on what you actually need. For anything serious the API providers are the more pragmatic choice anyway. The offline path is there for when you want zero data leaving the machine.

  • From Staff Engineer to Engineering Director: What Actually Changes

    Many years as a lead architect building financial services platforms, delivering for large institutions with distributed teams across several countries and time zones. Hard integrations, architectures that had outgrown their original shape, problems where the right answer wasn’t technically obvious.

    Moving into an engineering director role, the first surprise is that last part describes the org problems too.

    The architecture parallel

    Take a document comparison (blacklining) platform. It tracked changes across large volumes of legal and financial documents: master agreements, contract amendments, standards documents. Per-client deployed. Each client ran their own instance. The core logic was identical across all of them, but nothing was shared. Every update meant coordinating across separate environments. Clients had diverged slightly from each other. Operational overhead scaled with the number of clients, not the volume of work.

    Before — per-client deployment: each client runs a full separate stack, updates and versions drift independently
    Client A Gateway Identity Compare Orchestrator Converter Logging v1.1 · own server Client B Gateway Identity Compare Orchestrator Converter Logging v1.3 · config drifted Client C Gateway Identity Compare Orchestrator Converter Logging v1.0 · update backlog ↔ ↔ isolated isolated

    The redesign moved it to a shared serverless SaaS platform. Two core services, independently deployable. A comparison service accepts two HTML documents and returns <ins>/<del> markup, with parallel processing and long-running execution for very large legal documents. A document orchestrator manages the multi-stage pipeline: fetching and caching source documents, dispatching comparison pairs via a message queue, receiving results via callback, and handing off to downstream consumers. A conversion service handles upstream document conversion. Centralised logging and metrics cover observability; an identity provider handles auth and an API gateway handles routing. Infrastructure is defined as code, so identical configuration is deployed to every environment with zero manual provisioning. CI/CD uses zero-downtime deployments. The platform scales out under queue pressure and drops to zero at idle.

    After — shared SaaS: one platform, multiple environments reproduced identically via infrastructure as code
    CONSUMERS Client A Client B Client C Support UI API Gateway routing · rate limiting · TLS Identity / auth Comparison Service <ins>/<del> comparison parallel · long-running serverless functions Orchestrator enrichment pipeline queue orchestration progress monitoring Converter document → HTML full fidelity serverless functions Observability structured logging metrics & tracing per-service views Queue Callback Cloud storage queues · blobs · tables Document store document source Infrastructure as code · CI/CD Dev → Test → Prod · identical config · zero-downtime deploys · scale to zero auto-scaling
    Architecture in motion — client traffic converges to a shared platform; services scale out on demand

    B E F O R EA F T E RClient AClient BClient CAPI GatewayCompareOrchestratorConverterIdentityObservabilitySaaS PlatformPer-client deploys → shared SaaS · scales out on demand, to zero at idle · identical environments via IaC

    It wasn’t designed just to solve the immediate problem. It was designed so the next engineer on the next document-heavy async workflow could follow the same pattern: queue topology, service boundaries, infrastructure module structure, callback model between services, poison queue and retry handling. When the team needed a similar capability for a different document type later, the shape was already there.

    Good org design works the same way. Build the operating structure once, make it followable, and the next person in a similar situation doesn’t reinvent from scratch. The difference is verification. In software you run tests and find out quickly. In an engineering organisation, feedback takes quarters and it’s noisy.

    What changes when the unit of output shifts

    As a lead architect, the value you deliver is legible. You write the design, the team builds from it, you can point to the system and attribute specific decisions. The feedback loop is tight. Did the queue-backed processing hold up under the document volume we projected? Did the infrastructure module composition cause problems when we added the conversion service? You find out quickly.

    As a director, the best work you do is nearly invisible. A team that ships consistently, that surfaces real problems early, that makes sound architectural decisions without needing sign-off on every one: that’s the output. It doesn’t show up in any particular artefact. It shows up as the absence of the problems that would otherwise be present.

    That makes the feedback loop much harder to close. In architecture, you know whether your comparison service handles very large documents within timeout bounds. In leadership, the equivalent signal arrives slowly and with a lot of noise: does this team have the right capability mix for what’s coming in the next two quarters? The technical habits that help most here are the ones about holding problems at the right level of abstraction. Not diving into implementation detail when the real question is about interface design. Not fixating on a specific solution when the actual problem is a constraint you haven’t identified yet. Lead architecture work builds that instinct through iteration. It applies directly when the system is a team and the constraints are skills, relationships, and delivery risk.

    Technical credibility has a different role

    Staying technically credible in a director role is less about making good technical decisions yourself and more about maintaining the environment where other people make them well.

    That means being precise enough in design discussions that engineers can’t paper over complexity with confident language. It means understanding the failure modes of the patterns the team uses. Queue-backed serverless architectures have specific operational characteristics around poison messages, scaling lag, and storage cost that matter in production and need to be understood, not just referenced. It means being able to read an infrastructure-as-code plan and understand what’s actually being proposed, not just whether the CI checks pass.

    What it doesn’t mean, and this is the trap, is using that credibility to make the calls yourself. The job is to raise the quality of the decisions, not to be the source of them. A director who consistently overrides architectural choices in design reviews has optimised for short-term correctness at the cost of long-term ownership. A few months of that and engineers stop proposing designs and start waiting for instructions. The technical quality stays roughly the same. The team’s ability to function without you degrades steadily.

    In practice: less “have you considered X” and more “what’s driving the tradeoff between X and Y?” The first substitutes your judgment. The second develops theirs.

    Hiring is the hardest architectural decision

    The hardest architecture decisions aren’t about technology choices. They’re about what to build versus buy, what to build now versus defer, what to design for extensibility versus what to deliberately dead-end because the requirements aren’t stable enough yet. The cost of a wrong call is proportional to how deeply it’s been built in and how long it’s been there.

    The same logic applies to team composition. A team’s capability is the result of who is on it, and hiring decisions compound over years. A well-placed senior hire in the right gap can unblock six months of stuck delivery. Slow backfilling, or consistently hiring for available rather than needed, degrades what the team can build in ways that don’t show up clearly until the gap is already expensive to close.

    Working across distributed teams in several countries adds a layer of complexity the org chart doesn’t capture. The skills market is genuinely different in each location. The availability of engineers with specific domain backgrounds varies significantly. The timing of a hire matters relative to which team is carrying which delivery risk. You can’t treat it as one global hiring problem with uniform parameters. Model it the way you’d model a distributed system with different latency and throughput characteristics in each region: understand the local constraints, plan accordingly, and don’t assume what works in one location translates to another.

    The actual transition

    The shift isn’t about personality. It’s about what you’re optimising for. As a lead architect, you’re optimising for the quality of what gets built. As a director, you’re optimising for the quality of the team that builds things: what it can do now, its capability six months from now, its risk profile, how much delivery depends on a small number of individuals, whether the architectural patterns it’s following will hold under changing requirements.

    Both need technical depth. Neither requires abandoning it. The depth operates at a different resolution: wider and shallower across the whole surface, deliberately deep in the areas that most affect the team’s next quarter of work. The instincts from years of architecture work, finding where complexity hides, separating fundamental constraints from implementation choices, remain the most useful tools in the kit. Same tools, different system.

  • Building a TypeScript-First Frontend SDK: An Architectural Journey

    Over the past few years I have spent much of my time on one recurring problem: how do you let many teams build many data-intensive web applications without each of them reinventing the same foundations? This post walks through the architectural journey of designing a TypeScript-first frontend SDK to answer that question. It focuses on the patterns and trade-offs rather than any particular implementation, because the ideas travel well beyond the project that prompted them.

    Application layer Individual applications built on the platform Data grid apps Dashboards Workflow tools Admin tools Others… Framework layer Native React and Angular — thin bindings over the core Angular bindings Components · Directives · DI bridge React bindings Components · Hooks · DI bridge Core Framework-agnostic TypeScript — no UI framework dependency MVVMViewModel · State IoC / DIContainer · Discovery Data gridConfig · Streaming FormsRules · Validation ShellConfig-driven LayoutPanels · Persist ThemingTokens · Dark mode MessagingEvents · Notifications VisualisationCharts · Maps StreamingObservables ObservabilityLogging · Tracing · APM SecurityAuth · Permissions AI-ready contextContext files · Agent rules Scaffolding CLI + delivery pipeline CI/CD · Containers · AWS · dev / test / prod environments
    A layered SDK: framework-agnostic core, thin React/Angular bindings, and applications on top

    The Starting Point: Why a Component Library Is Not Enough

    The usual answer to “we need consistency across teams” is a component library: styled buttons, inputs and grids published as a package. In my experience that solves perhaps a fifth of the real problem. It says nothing about where state lives, how view logic is tested without rendering, how the same feature behaves identically in two UI frameworks, how dependent form fields validate each other, or how any of it stays coherent across years of changing teams.

    So the goal shifted from “shared components” to “shared architecture”: a TypeScript-first, framework-agnostic SDK in which React and Angular are equal, native rendering targets. The choice of view framework becomes a team preference, not an architectural fork.

    Principle 1: MVVM Keeps the View Thin

    The most important decision was to adopt Model–View–ViewModel as the backbone. The view is a template that binds to properties and commands. The ViewModel is a plain TypeScript class that owns state and behaviour. The model is data and services. Nothing else is allowed to blur those lines.

    Three benefits follow. Testability: a ViewModel has no JSX and no decorators, so business logic is tested as ordinary unit tests instead of by mounting components and poking the DOM. Reuse across frameworks: the same ViewModel drives a React view and an Angular view, so behaviour is written once. Clarity: engineers always know where logic belongs, which removes a whole category of code-review debate.

    Principle 2: Dependency Injection Owned by the Core

    MVVM only works well when ViewModels can ask for what they need without knowing where it comes from. Neither framework’s native mechanism fits: Angular’s injector only exists in Angular, and React context is not really DI. The answer was an IoC container that lives in the core, independent of any UI framework, with thin bridges into each one – a hook-based resolver for React and an injector bridge for Angular.

    Services are registered against interfaces, and can be discovered by capability rather than concrete type. This is where SOLID stops being a slide and becomes the path of least resistance: new behaviour arrives as a new registration, not an edit to existing classes, and every dependency can be swapped for a test double.

    Principle 3: Layers With Enforced Boundaries

    The SDK is organised as a monorepo with three layers:

    • Core: pure TypeScript – MVVM primitives, the container, state, messaging, streaming, configuration, theming tokens and platform services. No UI framework dependency.
    • Framework bindings: thin React and Angular packages that render the core’s concepts natively and bridge the container into each framework.
    • Applications: products built on a shared, configuration-driven shell, deployable to the browser or to desktop containers.

    Build-time boundary rules stop applications from reaching into the core directly and reject circular dependencies. Architecture that is only documented erodes; architecture that fails the build stays intact.

    Principle 4: Solve Cross-Cutting Concerns Once

    Most of the value of a platform is in the unglamorous capabilities that every application needs and no team wants to build. Rather than leave them to each project, the SDK provides them as shared services:

    • Configuration-driven shell and layout: navigation, panels and workspaces declared in configuration, with dockable panels whose arrangement is persisted per user.
    • State persistence: any screen can opt in to having column widths, filters, sort order and layout restored on reload.
    • Data grid abstraction: grids configured declaratively, with streaming updates delivered through observables, so teams write configuration and a ViewModel rather than grid plumbing.
    • Forms with a rule engine: field dependencies, visibility and validation expressed as rules; client and server validation surface through one pipeline.
    • Design tokens: semantic colours, spacing and typography exposed as typed constants and CSS variables, with dark mode as a first-class theme.
    • Observability and security: structured logging, tracing and application performance monitoring wired in by default, alongside authentication and permission handling.

    Principle 5: Start Every Project From Working, Not Blank

    A scaffolding CLI turns the SDK into a starting point rather than a set of instructions. A few commands produce an application that already has the shell, theming, authentication, observability, CI/CD pipelines, container configuration and cloud deployment templates for AWS across dev, test and production environments. The decisions that usually consume a project’s first sprint are already made, and every project starts in the same shape – which makes moving between them much easier for engineers.

    Principle 6: Design for AI-Assisted Development

    A newer lesson: an SDK now has two audiences, human developers and AI coding agents. Each package carries a concise context file describing its patterns, naming conventions and rules, and a routing structure points an agent at only the context relevant to its task. Generated projects inherit these files automatically.

    Treating that context as a maintained deliverable – updated alongside code and checked in CI – keeps agents inside the architecture instead of inventing their own. The result is AI assistance that accelerates teams without quietly eroding standards.

    Principle 7: Governance That Scales

    • Enforced module boundaries at build time
    • Semantic versioning with deprecation windows, so nothing disappears without warning
    • Living documentation: component catalogues and visual regression tests in the pipeline
    • Context files as a contract for both humans and agents

    What I Took Away

    The measurable outcome was that applications which once took months could be delivered in weeks, with a consistent look, feel and behaviour. But the more lasting lessons were architectural:

    • Separate behaviour from rendering early; it is very hard to retrofit.
    • Own your dependency injection in the core if you want framework independence.
    • Encode rules in the build, not in wiki pages.
    • Invest in the boring cross-cutting services – that is where teams lose the most time.
    • Treat AI context as part of the product, not an afterthought.

    A good platform is measured by how rarely the teams building on it need to think about it.

  • The JavaScript Execution Engine: Event Loop, Microtasks, Macrotasks, Web Workers and Service Workers

    JavaScript is single-threaded. There is one call stack, one thread of execution, and at any given moment exactly one thing is running. Yet it handles timers, network requests, user events, animations, background computation, and offline support without blocking. Understanding how it works changes how you write async code, debug race conditions, and reason about performance. This post covers the full picture: the event loop, the queues, Web Workers, and Service Workers.

    JS engine (single thread) Call stack Executes synchronous code console.log(‘hello’) main() Heap Object memory allocation { name: ‘obj’, val: 42 } Web APIs / Node.js setTimeout Timer (0ms+ delay) fetch / XHR Network I/O DOM events click, scroll, resize rAF / Promises Async coordination Microtask queue Drains completely before next task Promise.then() / .catch() / .finally() queueMicrotask() / MutationObserver Macrotask queue One task per loop iteration setTimeout / setInterval I/O callbacks / UI events Event loop offload Priority: (1) sync (2) microtask – drains fully (3) macrotask – one per iteration
    JavaScript execution engine: call stack, Web APIs, microtask queue, macrotask queue and event loop

    The Moving Parts

    Four components interact to make async JavaScript work on the main thread.

    The call stack

    The call stack is where synchronous code executes. When you call a function, a frame is pushed onto the stack. When it returns, the frame is popped. The engine can only execute code at the top of the stack. If a function takes 500ms of CPU work, nothing else can happen during those 500ms – the stack is blocked. This is why long-running synchronous operations freeze the browser: the event loop cannot process anything while the stack is occupied.

    Web APIs (or Node.js equivalents)

    The browser provides APIs that live outside the JavaScript engine: timers, network I/O, DOM events, file system access. When you call setTimeout, the JS engine hands the callback and delay to the browser’s timer mechanism and immediately returns. The actual waiting happens in a separate system thread managed by the browser. When the timer fires, the browser posts the callback into a queue.

    This is the key insight: JavaScript itself never does I/O. It delegates to the runtime, registers a callback, and moves on. The runtime notifies JavaScript when the work is done by placing callbacks into one of two queues.

    The microtask queue

    Microtasks are high-priority callbacks that run immediately after the current task completes, before the event loop picks up any macrotask and before the browser renders the next frame. Sources of microtasks include:

    • Promise.then(), .catch(), .finally()
    • queueMicrotask()
    • MutationObserver callbacks
    • await continuations (which desugar to Promise .then())

    The critical rule: the microtask queue drains completely before the event loop moves on. If a microtask callback queues another microtask, that microtask also runs before any macrotask. This can theoretically starve the macrotask queue – and the rendering pipeline – if microtasks keep spawning more microtasks.

    The macrotask queue (task queue)

    Macrotasks are lower-priority callbacks scheduled by the runtime. Sources include:

    • setTimeout and setInterval callbacks
    • I/O callbacks (file reads, network in Node.js)
    • UI event callbacks (click, keydown, scroll)
    • setImmediate in Node.js
    • MessageChannel port messages

    The event loop picks exactly one macrotask per iteration. After that one macrotask runs, it drains the entire microtask queue again before picking the next macrotask.

    The event loop algorithm

    while (true) {
      executeCurrentTask();
      while (microtaskQueue.length > 0) {
        const task = microtaskQueue.shift();
        task();
      }
      maybeRender();
      if (macrotaskQueue.length > 0) {
        const task = macrotaskQueue.shift();
        task();
      }
    }

    Why Promise.then() always beats setTimeout(fn, 0)

    setTimeout(fn, 0) does not mean “run immediately”. It means “run as soon as possible, but only after the current task and all microtasks have finished, and only as a macrotask”. A Promise.then() registered at the same time will always run first because it goes into the microtask queue, which drains before any macrotask is picked.

    setTimeout(() => console.log('timeout'), 0);
    Promise.resolve().then(() => console.log('promise'));
    
    // Output:
    // promise  (microtask - runs first)
    // timeout  (macrotask - runs after all microtasks)

    How fetch actually works

    When you call fetch('/data'):

    1. fetch() returns a Promise immediately. The browser’s networking layer starts the HTTP request in a separate thread.
    2. JavaScript continues executing synchronously – nothing waits.
    3. When the response arrives, the browser resolves the promise, queuing any .then() handlers as microtasks.
    4. The next time the microtask queue drains, those handlers run.
    5. Each .then() in a chain is a separate microtask, queued lazily when the previous one resolves.

    How async/await desugars

    async/await is syntactic sugar over promises. Every await suspends the async function (pops its frame off the call stack) and queues the continuation as a microtask when the awaited promise resolves.

    // async/await
    async function load() {
      const res = await fetch('/data');
      const data = await res.json();
      console.log(data);
    }
    
    // Is approximately:
    function load() {
      return fetch('/data')
        .then(res => res.json())
        .then(data => { console.log(data); });
    }

    A worked example: what does this log?

    console.log('A');
    setTimeout(() => console.log('B'), 0);
    Promise.resolve()
      .then(() => console.log('C'))
      .then(() => console.log('D'));
    console.log('E');
    
    // Output: A E C D B

    Sync runs first: A, then E. Microtask queue drains: C resolves and queues D, so C then D. Finally the macrotask runs: B. The second .then() was not enqueued until the first one ran – promise chains are lazy microtask sequences, not a pre-loaded batch.

    Main thread UI + event loop Call stack Sync execution DOM access Microtask queue Promise callbacks Macrotask queue setTimeout, events postMessage() Send / receive data Render pipeline Layout – Paint Blocked by JS Workers never block this Web Worker Separate OS thread Own call stack CPU-heavy work No DOM access Own microtask queue Promises work here Own macrotask queue setTimeout works here postMessage() Send results back Service Worker Network proxy thread Lifecycle events install – activate idle / terminated fetch event Intercepts all requests respondWith(promise) Cache API Serve offline / stale Cache-then-network Push + sync events Background wakeup No DOM / no window postMessage to page send result fetch() intercepted by SW cached / network response All three have their own event loop – only the main thread can access the DOM
    Main thread, Web Worker, and Service Worker – three independent event loops with distinct responsibilities

    Web Workers: true parallelism for JavaScript

    The event loop handles async waiting by delegating I/O to the browser and resuming when ready. But it does not solve CPU-intensive computation. If you need to parse a 50MB JSON file, run a physics simulation, encrypt a large payload, or process an image, doing that on the main thread blocks the call stack and freezes the UI regardless of how cleverly you structure your promises.

    Web Workers give you a genuine second OS thread. A Worker runs in a completely separate execution context with its own call stack, its own microtask and macrotask queues, its own memory heap, and its own event loop. The main thread and workers run in parallel on multi-core hardware.

    What Workers can and cannot do

    A Worker has access to: fetch, setTimeout, setInterval, Promise, IndexedDB, WebSockets, Cache API, crypto, console, and most Web APIs that don’t require a rendering context. What it does not have access to is the DOM – no document, no window, no direct manipulation of page elements. The rendering pipeline lives exclusively on the main thread, and the browser enforces this hard boundary.

    Communication via postMessage

    The main thread and workers communicate exclusively through postMessage(). Data is serialised (structured clone algorithm) and deserialised on the other side; each side gets its own copy, not a shared reference. For large data (images, audio buffers, typed arrays) you can use Transferable Objects to hand ownership to the other thread without copying, which is zero-copy and fast.

    // main.js
    const worker = new Worker('worker.js');
    worker.postMessage({ action: 'process', data: largeArray });
    worker.onmessage = (event) => {
      console.log('Result:', event.data);
    };
    
    // worker.js
    self.onmessage = (event) => {
      const result = heavyComputation(event.data.data);
      self.postMessage(result);
    };

    When a worker calls postMessage(), the message arrives on the main thread as a macrotask – it queues a message event on the worker object. The main thread’s event loop picks it up in the normal way: after all current microtasks have drained. This means worker communication is non-blocking in both directions.

    Transferable Objects for zero-copy transfer

    const buffer = new ArrayBuffer(100 * 1024 * 1024); // 100MB
    worker.postMessage({ buffer }, [buffer]);
    // buffer is now detached in main - worker owns it
    
    self.onmessage = (e) => {
      processBuffer(e.data.buffer);
      self.postMessage({ buffer: e.data.buffer }, [e.data.buffer]);
    };

    Shared memory with SharedArrayBuffer

    For high-frequency communication (game engines, audio processing, real-time data pipelines), copying data on every message is too expensive. SharedArrayBuffer creates a memory region that both the main thread and workers can read and write without copying. You coordinate access using Atomics – atomic operations that prevent race conditions by guaranteeing visibility and mutual exclusion.

    const shared = new SharedArrayBuffer(4);
    const view = new Int32Array(shared);
    worker.postMessage({ shared });
    Atomics.store(view, 0, 1);
    Atomics.notify(view, 0, 1);

    Worker types

    There are three kinds of Worker: Dedicated Workers (owned by one page, created with new Worker()), Shared Workers (shared across multiple pages from the same origin, accessible via new SharedWorker()), and Service Workers (covered in the next section). Dedicated Workers are by far the most common.

    Service Workers: the network proxy

    A Service Worker is a type of worker with a fundamentally different purpose. Where a Dedicated Worker offloads CPU computation, a Service Worker sits between your application and the network, intercepting and handling every fetch request the page makes. It is the foundation of Progressive Web Apps (PWAs) – offline support, background sync, push notifications, and fine-grained caching strategies all live here.

    The Service Worker lifecycle

    A Service Worker has a distinct lifecycle that separates it from ordinary workers:

    1. Registration: the page calls navigator.serviceWorker.register('/sw.js'). The browser downloads and parses the worker script.
    2. Install: the browser fires the install event. This is where you pre-cache static assets. You call event.waitUntil(promise) to tell the browser not to proceed until your caching is complete. If the promise rejects, the install fails and the worker is discarded.
    3. Activate: once installed, the worker waits to activate. A new worker only activates when no existing controlled pages are open (or when you call self.skipWaiting()). The activate event is where you clean up old caches from previous versions.
    4. Idle: the worker is now active and controlling pages, but the browser may terminate it at any time to save memory. It is relaunched on demand when a controlled page makes a fetch or a push notification arrives.
    const CACHE = 'v1';
    const STATIC = ['/index.html', '/app.js', '/style.css'];
    
    self.addEventListener('install', event => {
      event.waitUntil(
        caches.open(CACHE).then(cache => cache.addAll(STATIC))
      );
      self.skipWaiting();
    });
    
    self.addEventListener('activate', event => {
      event.waitUntil(
        caches.keys().then(keys =>
          Promise.all(keys.filter(k => k !== CACHE).map(k => caches.delete(k)))
        )
      );
      self.clients.claim();
    });

    Intercepting fetch requests

    The fetch event is the core of a Service Worker. Every network request made by a controlled page – including fetch(), XMLHttpRequest, CSS imports, image loads, script tags – passes through this handler. You use event.respondWith(promise) to provide the response, which can come from the cache, the network, or be constructed programmatically.

    self.addEventListener('fetch', event => {
      event.respondWith(
        caches.match(event.request).then(cached => {
          if (cached) return cached;
          return fetch(event.request).then(response => {
            const toCache = response.clone();
            caches.open(CACHE).then(cache => cache.put(event.request, toCache));
            return response;
          });
        })
      );
    });

    Caching strategies

    • Cache first: serve from cache if present, otherwise fetch from network and cache the result. Best for assets that change infrequently (fonts, versioned JS bundles).
    • Network first: try the network, fall back to cache if offline. Best for frequently updated content (API responses, news feeds).
    • Stale while revalidate: serve from cache immediately (fast), then fetch from network in the background to update the cache for the next request. Best for content where showing slightly stale data is acceptable.
    • Network only: never use the cache. For analytics, payment flows, anything where stale data is dangerous.
    • Cache only: never hit the network. For pre-cached static assets in a fully offline app.

    Background sync and push notifications

    Service Workers can be woken up by the browser even when no page is open. The sync event fires when connectivity is restored after a period offline – you can queue writes made while offline and flush them here. The push event fires when a push message arrives from your server, allowing you to display a notification even if the user doesn’t have your site open.

    self.addEventListener('sync', event => {
      if (event.tag === 'submit-form') {
        event.waitUntil(flushPendingSubmissions());
      }
    });
    
    self.addEventListener('push', event => {
      const data = event.data.json();
      event.waitUntil(
        self.registration.showNotification(data.title, {
          body: data.body,
          icon: '/icon.png',
        })
      );
    });

    Communicating with the page

    A Service Worker communicates with its controlled pages via postMessage, just like a Dedicated Worker. The worker can send messages to all controlled clients via self.clients.matchAll(), and individual pages can send messages to the worker via navigator.serviceWorker.controller.postMessage(). As with Dedicated Workers, these messages arrive as macrotasks on the receiving end.

    self.clients.matchAll().then(clients => {
      clients.forEach(client => client.postMessage({ type: 'CACHE_UPDATED' }));
    });
    
    navigator.serviceWorker.addEventListener('message', event => {
      if (event.data.type === 'CACHE_UPDATED') {
        showUpdateBanner();
      }
    });

    Web Worker vs Service Worker: when to use which

    The two are often confused because both are “workers that run off the main thread”. The distinction is purpose, not mechanism.

    Web Worker Service Worker
    Primary purpose CPU-intensive computation Network proxy and caching
    Lifetime As long as the page holds a reference Persists independently, browser controls termination
    Scope Single page All pages on the same origin
    fetch access Can make fetch calls Intercepts all fetch calls from the page
    DOM access No No
    Offline support No Yes, via Cache API
    Push notifications No Yes
    Background sync No Yes
    Typical use cases Image processing, data parsing, encryption, physics, ML inference Caching strategy, offline PWA, push, background sync

    Practical implications

    Long microtask chains can block rendering

    The browser cannot render a new frame until the microtask queue is empty. For large data processing, break work into chunks using setTimeout to yield to the rendering pipeline, or use a Web Worker to move the work off the main thread entirely.

    Service Workers require HTTPS

    Service Workers can intercept and modify any network request, which makes them powerful – and dangerous if compromised. Browsers only register Service Workers on HTTPS origins (and localhost for development). There are no exceptions.

    Service Workers are version-sensitive

    When you deploy a new Service Worker, users with the old version keep it until all their tabs are closed and reopened. skipWaiting() + clients.claim() force immediate takeover – useful during development but potentially disruptive if the new worker uses an incompatible cache structure. Design your activate handler to clean up old caches explicitly.

    Workers have their own event loops

    Both Web Workers and Service Workers have their own complete event loop – their own call stack, microtask queue, and macrotask queue. Promise, setTimeout, fetch, and async/await all work inside a worker the same way they work on the main thread. The only difference is the absence of the DOM and window APIs.

    Summary

    JavaScript’s concurrency model has three layers. The event loop handles async waiting on the main thread – sync code runs first, then all microtasks drain (Promises, await continuations), then one macrotask runs, then microtasks drain again. Web Workers handle CPU parallelism – genuine OS threads running their own event loops, communicating with the main thread via postMessage without blocking the UI. Service Workers handle network and persistence – a long-lived proxy thread that intercepts fetch requests, implements caching strategies, enables offline use, and receives push notifications and background sync events independently of any open page.

    Together they give JavaScript – a language with one main thread – a complete answer to async I/O, CPU parallelism, and network resilience.