Rajnish Noonia

Your SDK Should Ship Its Own AI Agents. Here’s How I Built It.

Several teams consuming a single SDK. Each one built their own AI-GUIDE.md. Each one drifted. The fix wasn’t better documentation. It was shipping agents as part of the package itself.


// I

The Problem Nobody Talks About

Every enterprise SDK team eventually faces the same version of this. You’ve built solid components. You’ve documented them properly. ADRs, component catalogues, generated API docs, the lot. And then you watch several different teams use those components in several different ways, most of them wrong.

AI coding tools make it worse. The model doesn’t read your ADRs. It reads whatever’s in the conversation window, which in most projects means a README and whatever the current file contains. So it guesses. Sometimes well. Usually not.

The first fix teams reach for is an AI-GUIDE.md. A file written for the model, not humans. Describes the architecture, naming conventions, which components are for what, what patterns to follow. Better than nothing. But it stays behind in the project. Each team maintains their own. None of them stay in sync with the SDK.

The natural next step is shipping context with the SDK itself. Put an AGENT-GUIDE.md in the npm package. Postinstall copies it to the project. Now the SDK team maintains it, and each consuming team gets the right context automatically on every upgrade.

That works. But context isn’t behaviour. A context file tells the model what things exist. It doesn’t tell it what to do: which questions to ask first, what to produce before writing any code, which rules can’t be broken. What you actually want is an agent that knows your SDK the way a senior engineer does. Context files don’t do that. Agents do.

README humans API Docs humans Storybook humans machine-first AI-GUIDE context only Shipped Agents context + behaviour
FIG 1 — Documentation evolution: from human readers to shipped behaviour

// II

Three-Layer Architecture

The design is three files. Each one has a distinct job and a distinct owner.

Trigger. A slash command in .claude/commands/. One file per workflow: create-grid.md, create-chart.md, create-form.md. This is the developer’s entry point. They type /create-grid and the agent starts. The command file is thin: it names the agent, states the intent, hands off to PLAYBOOK.

Behaviour. PLAYBOOK.md. This is what the agent actually does. A deterministic contract: which questions to ask, in what order, what to produce before writing any code, what rules hold regardless of instruction. The SDK team writes it. The SDK ships it. Developers don’t touch it.

Memory. JOURNAL.md. Project-specific knowledge that accumulates: patterns that work in this codebase, API quirks, decisions the team made and why. The project owns it. SDK upgrades never touch it.

The SDK owns PLAYBOOK. The project owns JOURNAL. That separation is the whole thing. Without it, every upgrade either breaks project customisations or falls behind on SDK changes.

SDK ships Trigger .claude/commands/create-grid.md SDK ships Behaviour PLAYBOOK.md · protocol, rules, pattern library, component registry SDK ships Memory JOURNAL.md · project patterns, codebase decisions, accumulated knowledge Project owns
FIG 2 — Three-layer architecture: Trigger / Behaviour / Memory

// III

Anatomy of PLAYBOOK.md

PLAYBOOK.md isn’t a prompt. It’s a contract. Ten named sections, each with a specific job, grouped into three categories: identity, protocol, and standards.

IDENTITY PROTOCOL STANDARDS 01 Agent Identity Domain, role, what it covers and what it doesn't 02 Guards Pre-flight checks before any work starts 03 DISCOVER Protocol Questions to ask before touching any code 04 SPECIFY Protocol Spec document format, explicit approval required 05 IMPLEMENT Protocol Build order, dependency rules, integration sequence 06 Hard Rules What the agent never does, regardless of instruction 07 Pattern Library Canonical patterns to copy, not invent 08 Component Registry Which SDK components exist and what they're for 09 Review Checklist What to verify before done 10 Error Catalogue Known failure modes and recovery paths
FIG 3 — PLAYBOOK.md: ten sections across three categories

Section 04 (SPECIFY Protocol) defines what the spec document looks like: components, configuration options, data shape, SDK calls. The agent produces this in markdown before touching any code. The developer reads it and approves it, or corrects it.

Section 06 (Hard Rules) is the most important one to get right. Things like: never install a dependency that isn’t already in package.json, never write a component that bypasses the SDK’s data service layer, never make direct API calls from a UI component. If these aren’t explicit, the model ignores them. Stated explicitly, they hold.


// IV

Three-Phase Protocol

The protocol is the agent’s runtime behaviour. Three phases, and one of them is a hard stop.

DISCOVER Structured questions before any code SPECIFY Spec document in markdown HUMAN REVIEW IMPLEMENT Deterministic build order Data source? Volume? Groups? Real-time? Components, config, data shape, SDK calls Approve or correct the spec Infra first, layout next, integration last
FIG 4 — Three-phase protocol with mandatory human checkpoint

DISCOVER exists because the agent needs to know things you haven’t written down. What data source is this grid pointing at, does it need grouping, what’s the expected row count, does it need real-time updates. Without those answers it will guess. The questions in PLAYBOOK are chosen specifically to prevent the most common wrong guesses.

SPECIFY produces a markdown document the developer can read and push back on. Not a comment in code, not something buried in a terminal. An actual spec. The agent doesn’t start building until the developer approves it. “Looks good” is sufficient. Silence is not.

IMPLEMENT follows a fixed build order: infrastructure first (data services, state management), then layout components, then configuration-heavy elements, then integration wiring. That order prevents circular dependencies. It also means you can stop the agent at any point and the code that exists is valid, not half-finished.

The mandatory human checkpoint is what makes developers trust the agent. Without it, they watch code appear and feel nervous. With it, they approved the plan. They’re responsible for it. They rewrite far less.


// V

The Sync Mechanism

Getting agents into projects is the part most write-ups skip. The install experience determines whether developers actually use the agents or quietly work around them.

The approach is a postinstall hook. When a developer installs your SDK package, the hook runs and syncs agent files into .claude/agents/. No manual setup. No documentation to follow. The agents are just there when the SDK is there.

SDK Package PLAYBOOK.md commands/ adapters/ AGENT-GUIDE.md PASSIVE npm install runs postinstall hook automatically ACTIVE npx sync-agents on-demand, mid-sprint .claude/agents/ PLAYBOOK.md SYNC JOURNAL.md SKIP PROJECT-RULES.md MERGE
FIG 5 — Sync mechanism: what gets overwritten, skipped, and merged

Three merge rules:

  • PLAYBOOK.md is overwritten on every sync. The SDK team owns it entirely.
  • JOURNAL.md is never touched. The project owns it entirely.
  • PROJECT-RULES.md is merged. The SDK ships a base set of rules; the project adds custom ones and the merge preserves both. On conflict, SDK rules win.

The active CLI matters for mid-sprint SDK upgrades. If a new agent version ships while a developer is in the middle of work, they don’t need to reinstall. One command brings PLAYBOOK up to date without touching JOURNAL.

Test the postinstall hook in CI. Not every npm client runs it. Some teams disable it. Some CI environments strip it. If the hook fails silently, developers get an outdated agent and have no idea why it’s behaving differently from the team next door.


// VI

What It Looks Like in Practice

The developer types:

/create-grid

Four questions follow. Data source, grouping requirements, expected row count, real-time or static. The agent synthesises the answers and produces a spec:

## Grid Specification: Orders Overview

Components:
  DataGrid (container, data-service mode)
  DataGridColumn x6 (id, name, category, quantity, status, updatedAt)
  DataGridGroupRow (group by status)

Data service:  OrdersDataService (existing)
Row model:     server-side  (volume: large)
Update mode:   streaming
Grouping:      status column, expandable

SDK calls:
  useGridDataService(OrdersDataService)
  configureServerRows({ groupBy: ['status'] })
  withLiveUpdates(ordersStream)

Ready to implement?

The developer reads it. Says yes. Three minutes later, a fully configured data grid is in the editor: correct imports, correct SDK calls, correct data binding, error handling, TypeScript types. No code the developer wrote by hand.

The important detail: the agent didn’t guess. PLAYBOOK told it which components exist, which data service patterns are correct, what server-side row configuration looks like. The model’s job was filling in the values, not inventing the structure.


// VII

What I Learned

The agent is only as good as PLAYBOOK. If it’s vague, the agent is vague. Write it the way a senior engineer would brief a junior one who asks no questions. Specific enough that there’s no ambiguity about what to do next.

The human checkpoint makes the difference. Without it, developers watch code appear and feel like they’re not in control. With it, they approved the plan. They caught problems before the code was written. They rewrite far less.

Shipping agents with the SDK transfers responsibility correctly. The team that knows the API best is now the team responsible for AI quality. Each consuming team doesn’t need to figure out the right way to prompt the model. The SDK team did it once.

JOURNAL.md compounds. After six months, a project’s JOURNAL has real knowledge: specific data patterns, API quirks, decisions the team made and why they made them. That context doesn’t exist if every SDK upgrade clears it. The separation between PLAYBOOK and JOURNAL is what lets the memory accumulate safely.

The sync mechanism is where things break. Postinstall hooks are fragile. Test them in CI. Add the CLI as a fallback. Document what to run if the hook didn’t fire. Don’t assume npm install always ran the hook.


// VIII

Where This Is Heading

JOURNAL.md currently gets written by developers. They add entries when they make decisions, when they find patterns that work, when they hit problems. That’s fine and it works.

The logical extension is an agent that annotates JOURNAL automatically. When a SPECIFY output gets corrected, the correction gets recorded. When a pattern the agent chose works well and the developer says so, that gets recorded. Over time, JOURNAL captures real correction signals from the people who actually know the codebase.

That’s not fine-tuning. It’s a memory system. You get an agent that improves without touching the model. The SDK team ships base behaviour. The project’s history makes it better for that project specifically.

It raises a question about ownership, though. Right now the agent is Claude Code with a slash command. As JOURNAL grows, it accumulates context that’s genuinely specific to one team. At some point it isn’t really a generic SDK tool anymore. It’s a team-specific tool built on SDK primitives. What’s the SDK team’s responsibility when that agent makes a bad call?

Worth thinking about before the JOURNAL gets long.

Leave a Reply

All posts

Discover more from Pixytech

Subscribe now to keep reading and get access to the full archive.

Continue reading