Several teams consuming a single SDK. Each one built their own AI-GUIDE.md. Each one drifted. The fix wasn’t better documentation. It was shipping agents as part of the package itself.
// I
The Problem Nobody Talks About
Every enterprise SDK team eventually faces the same version of this. You’ve built solid components. You’ve documented them properly. ADRs, component catalogues, generated API docs, the lot. And then you watch several different teams use those components in several different ways, most of them wrong.
AI coding tools make it worse. The model doesn’t read your ADRs. It reads whatever’s in the conversation window, which in most projects means a README and whatever the current file contains. So it guesses. Sometimes well. Usually not.
The first fix teams reach for is an AI-GUIDE.md. A file written for the model, not humans. Describes the architecture, naming conventions, which components are for what, what patterns to follow. Better than nothing. But it stays behind in the project. Each team maintains their own. None of them stay in sync with the SDK.
The natural next step is shipping context with the SDK itself. Put an AGENT-GUIDE.md in the npm package. Postinstall copies it to the project. Now the SDK team maintains it, and each consuming team gets the right context automatically on every upgrade.
That works. But context isn’t behaviour. A context file tells the model what things exist. It doesn’t tell it what to do: which questions to ask first, what to produce before writing any code, which rules can’t be broken. What you actually want is an agent that knows your SDK the way a senior engineer does. Context files don’t do that. Agents do.
// II
Three-Layer Architecture
The design is three files. Each one has a distinct job and a distinct owner.
Trigger. A slash command in .claude/commands/. One file per workflow: create-grid.md, create-chart.md, create-form.md. This is the developer’s entry point. They type /create-grid and the agent starts. The command file is thin: it names the agent, states the intent, hands off to PLAYBOOK.
Behaviour. PLAYBOOK.md. This is what the agent actually does. A deterministic contract: which questions to ask, in what order, what to produce before writing any code, what rules hold regardless of instruction. The SDK team writes it. The SDK ships it. Developers don’t touch it.
Memory. JOURNAL.md. Project-specific knowledge that accumulates: patterns that work in this codebase, API quirks, decisions the team made and why. The project owns it. SDK upgrades never touch it.
The SDK owns PLAYBOOK. The project owns JOURNAL. That separation is the whole thing. Without it, every upgrade either breaks project customisations or falls behind on SDK changes.
// III
Anatomy of PLAYBOOK.md
PLAYBOOK.md isn’t a prompt. It’s a contract. Ten named sections, each with a specific job, grouped into three categories: identity, protocol, and standards.
Section 04 (SPECIFY Protocol) defines what the spec document looks like: components, configuration options, data shape, SDK calls. The agent produces this in markdown before touching any code. The developer reads it and approves it, or corrects it.
Section 06 (Hard Rules) is the most important one to get right. Things like: never install a dependency that isn’t already in package.json, never write a component that bypasses the SDK’s data service layer, never make direct API calls from a UI component. If these aren’t explicit, the model ignores them. Stated explicitly, they hold.
// IV
Three-Phase Protocol
The protocol is the agent’s runtime behaviour. Three phases, and one of them is a hard stop.
DISCOVER exists because the agent needs to know things you haven’t written down. What data source is this grid pointing at, does it need grouping, what’s the expected row count, does it need real-time updates. Without those answers it will guess. The questions in PLAYBOOK are chosen specifically to prevent the most common wrong guesses.
SPECIFY produces a markdown document the developer can read and push back on. Not a comment in code, not something buried in a terminal. An actual spec. The agent doesn’t start building until the developer approves it. “Looks good” is sufficient. Silence is not.
IMPLEMENT follows a fixed build order: infrastructure first (data services, state management), then layout components, then configuration-heavy elements, then integration wiring. That order prevents circular dependencies. It also means you can stop the agent at any point and the code that exists is valid, not half-finished.
The mandatory human checkpoint is what makes developers trust the agent. Without it, they watch code appear and feel nervous. With it, they approved the plan. They’re responsible for it. They rewrite far less.
// V
The Sync Mechanism
Getting agents into projects is the part most write-ups skip. The install experience determines whether developers actually use the agents or quietly work around them.
The approach is a postinstall hook. When a developer installs your SDK package, the hook runs and syncs agent files into .claude/agents/. No manual setup. No documentation to follow. The agents are just there when the SDK is there.
Three merge rules:
- PLAYBOOK.md is overwritten on every sync. The SDK team owns it entirely.
- JOURNAL.md is never touched. The project owns it entirely.
- PROJECT-RULES.md is merged. The SDK ships a base set of rules; the project adds custom ones and the merge preserves both. On conflict, SDK rules win.
The active CLI matters for mid-sprint SDK upgrades. If a new agent version ships while a developer is in the middle of work, they don’t need to reinstall. One command brings PLAYBOOK up to date without touching JOURNAL.
Test the postinstall hook in CI. Not every npm client runs it. Some teams disable it. Some CI environments strip it. If the hook fails silently, developers get an outdated agent and have no idea why it’s behaving differently from the team next door.
// VI
What It Looks Like in Practice
The developer types:
/create-grid
Four questions follow. Data source, grouping requirements, expected row count, real-time or static. The agent synthesises the answers and produces a spec:
## Grid Specification: Orders Overview
Components:
DataGrid (container, data-service mode)
DataGridColumn x6 (id, name, category, quantity, status, updatedAt)
DataGridGroupRow (group by status)
Data service: OrdersDataService (existing)
Row model: server-side (volume: large)
Update mode: streaming
Grouping: status column, expandable
SDK calls:
useGridDataService(OrdersDataService)
configureServerRows({ groupBy: ['status'] })
withLiveUpdates(ordersStream)
Ready to implement?
The developer reads it. Says yes. Three minutes later, a fully configured data grid is in the editor: correct imports, correct SDK calls, correct data binding, error handling, TypeScript types. No code the developer wrote by hand.
The important detail: the agent didn’t guess. PLAYBOOK told it which components exist, which data service patterns are correct, what server-side row configuration looks like. The model’s job was filling in the values, not inventing the structure.
// VII
What I Learned
The agent is only as good as PLAYBOOK. If it’s vague, the agent is vague. Write it the way a senior engineer would brief a junior one who asks no questions. Specific enough that there’s no ambiguity about what to do next.
The human checkpoint makes the difference. Without it, developers watch code appear and feel like they’re not in control. With it, they approved the plan. They caught problems before the code was written. They rewrite far less.
Shipping agents with the SDK transfers responsibility correctly. The team that knows the API best is now the team responsible for AI quality. Each consuming team doesn’t need to figure out the right way to prompt the model. The SDK team did it once.
JOURNAL.md compounds. After six months, a project’s JOURNAL has real knowledge: specific data patterns, API quirks, decisions the team made and why they made them. That context doesn’t exist if every SDK upgrade clears it. The separation between PLAYBOOK and JOURNAL is what lets the memory accumulate safely.
The sync mechanism is where things break. Postinstall hooks are fragile. Test them in CI. Add the CLI as a fallback. Document what to run if the hook didn’t fire. Don’t assume npm install always ran the hook.
// VIII
Where This Is Heading
JOURNAL.md currently gets written by developers. They add entries when they make decisions, when they find patterns that work, when they hit problems. That’s fine and it works.
The logical extension is an agent that annotates JOURNAL automatically. When a SPECIFY output gets corrected, the correction gets recorded. When a pattern the agent chose works well and the developer says so, that gets recorded. Over time, JOURNAL captures real correction signals from the people who actually know the codebase.
That’s not fine-tuning. It’s a memory system. You get an agent that improves without touching the model. The SDK team ships base behaviour. The project’s history makes it better for that project specifically.
It raises a question about ownership, though. Right now the agent is Claude Code with a slash command. As JOURNAL grows, it accumulates context that’s genuinely specific to one team. At some point it isn’t really a generic SDK tool anymore. It’s a team-specific tool built on SDK primitives. What’s the SDK team’s responsibility when that agent makes a bad call?
Worth thinking about before the JOURNAL gets long.
Leave a Reply