Many years as a lead architect building financial services platforms, delivering for large institutions with distributed teams across several countries and time zones. Hard integrations, architectures that had outgrown their original shape, problems where the right answer wasn’t technically obvious.
Moving into an engineering director role, the first surprise is that last part describes the org problems too.
The architecture parallel
Take a document comparison (blacklining) platform. It tracked changes across large volumes of legal and financial documents: master agreements, contract amendments, standards documents. Per-client deployed. Each client ran their own instance. The core logic was identical across all of them, but nothing was shared. Every update meant coordinating across separate environments. Clients had diverged slightly from each other. Operational overhead scaled with the number of clients, not the volume of work.
The redesign moved it to a shared serverless SaaS platform. Two core services, independently deployable. A comparison service accepts two HTML documents and returns <ins>/<del> markup, with parallel processing and long-running execution for very large legal documents. A document orchestrator manages the multi-stage pipeline: fetching and caching source documents, dispatching comparison pairs via a message queue, receiving results via callback, and handing off to downstream consumers. A conversion service handles upstream document conversion. Centralised logging and metrics cover observability; an identity provider handles auth and an API gateway handles routing. Infrastructure is defined as code, so identical configuration is deployed to every environment with zero manual provisioning. CI/CD uses zero-downtime deployments. The platform scales out under queue pressure and drops to zero at idle.
It wasn’t designed just to solve the immediate problem. It was designed so the next engineer on the next document-heavy async workflow could follow the same pattern: queue topology, service boundaries, infrastructure module structure, callback model between services, poison queue and retry handling. When the team needed a similar capability for a different document type later, the shape was already there.
Good org design works the same way. Build the operating structure once, make it followable, and the next person in a similar situation doesn’t reinvent from scratch. The difference is verification. In software you run tests and find out quickly. In an engineering organisation, feedback takes quarters and it’s noisy.
What changes when the unit of output shifts
As a lead architect, the value you deliver is legible. You write the design, the team builds from it, you can point to the system and attribute specific decisions. The feedback loop is tight. Did the queue-backed processing hold up under the document volume we projected? Did the infrastructure module composition cause problems when we added the conversion service? You find out quickly.
As a director, the best work you do is nearly invisible. A team that ships consistently, that surfaces real problems early, that makes sound architectural decisions without needing sign-off on every one: that’s the output. It doesn’t show up in any particular artefact. It shows up as the absence of the problems that would otherwise be present.
That makes the feedback loop much harder to close. In architecture, you know whether your comparison service handles very large documents within timeout bounds. In leadership, the equivalent signal arrives slowly and with a lot of noise: does this team have the right capability mix for what’s coming in the next two quarters? The technical habits that help most here are the ones about holding problems at the right level of abstraction. Not diving into implementation detail when the real question is about interface design. Not fixating on a specific solution when the actual problem is a constraint you haven’t identified yet. Lead architecture work builds that instinct through iteration. It applies directly when the system is a team and the constraints are skills, relationships, and delivery risk.
Technical credibility has a different role
Staying technically credible in a director role is less about making good technical decisions yourself and more about maintaining the environment where other people make them well.
That means being precise enough in design discussions that engineers can’t paper over complexity with confident language. It means understanding the failure modes of the patterns the team uses. Queue-backed serverless architectures have specific operational characteristics around poison messages, scaling lag, and storage cost that matter in production and need to be understood, not just referenced. It means being able to read an infrastructure-as-code plan and understand what’s actually being proposed, not just whether the CI checks pass.
What it doesn’t mean, and this is the trap, is using that credibility to make the calls yourself. The job is to raise the quality of the decisions, not to be the source of them. A director who consistently overrides architectural choices in design reviews has optimised for short-term correctness at the cost of long-term ownership. A few months of that and engineers stop proposing designs and start waiting for instructions. The technical quality stays roughly the same. The team’s ability to function without you degrades steadily.
In practice: less “have you considered X” and more “what’s driving the tradeoff between X and Y?” The first substitutes your judgment. The second develops theirs.
Hiring is the hardest architectural decision
The hardest architecture decisions aren’t about technology choices. They’re about what to build versus buy, what to build now versus defer, what to design for extensibility versus what to deliberately dead-end because the requirements aren’t stable enough yet. The cost of a wrong call is proportional to how deeply it’s been built in and how long it’s been there.
The same logic applies to team composition. A team’s capability is the result of who is on it, and hiring decisions compound over years. A well-placed senior hire in the right gap can unblock six months of stuck delivery. Slow backfilling, or consistently hiring for available rather than needed, degrades what the team can build in ways that don’t show up clearly until the gap is already expensive to close.
Working across distributed teams in several countries adds a layer of complexity the org chart doesn’t capture. The skills market is genuinely different in each location. The availability of engineers with specific domain backgrounds varies significantly. The timing of a hire matters relative to which team is carrying which delivery risk. You can’t treat it as one global hiring problem with uniform parameters. Model it the way you’d model a distributed system with different latency and throughput characteristics in each region: understand the local constraints, plan accordingly, and don’t assume what works in one location translates to another.
The actual transition
The shift isn’t about personality. It’s about what you’re optimising for. As a lead architect, you’re optimising for the quality of what gets built. As a director, you’re optimising for the quality of the team that builds things: what it can do now, its capability six months from now, its risk profile, how much delivery depends on a small number of individuals, whether the architectural patterns it’s following will hold under changing requirements.
Both need technical depth. Neither requires abandoning it. The depth operates at a different resolution: wider and shallower across the whole surface, deliberately deep in the areas that most affect the team’s next quarter of work. The instincts from years of architecture work, finding where complexity hides, separating fundamental constraints from implementation choices, remain the most useful tools in the kit. Same tools, different system.