Replaces load-context.mjs's JSON output with a tight markdown block from the renamed context.mjs. The script now extracts PRODUCT.md's `## Register` field and appends a `NEXT STEP:` directive naming the matching reference (brand.md / product.md), which moved Gemini from skipping the register load entirely to honoring it. Drops the `.impeccable.md` auto-migration; makes IMPECCABLE_CONTEXT_DIR a lazy escape hatch consulted only when the default paths come up empty. Setup is now four bullets in one list. The DESIGN.md nudge is gone; in its place, a "familiarize with the existing design system" step that calls out CSS / tokens / running app as authoritative sources alongside DESIGN.md. The standalone `### Register` H3 stays for the cascade rules (task cue → surface → register field). New LLM-backed test suite at tests/skill-behavior/ runs five scenarios against claude-haiku-4-5, gpt-5.4-mini, and gemini-3.1-flash-lite via Vercel AI SDK. Captures real tool traces, asserts on context.mjs calls, brand.md loads, and teach.md fallback. Skips cleanly when API keys are unset. 13-14/15 pass; only stable failure is the v3.2.0-era gpt-mini S4 "don't re-run" regression. Adds @ai-sdk/google as devDep and the test:skill-behavior npm script. Touches em-dashes in skill/SKILL.md and four reference files so `bun run build:skills` passes its skill-prose validator. teach.md and document.md drop their "re-run the loader to refresh session cache" steps since the agent's own write is now the freshest source. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Skill-behavior tests
LLM-backed scenarios that verify how the impeccable skill drives PRODUCT.md / DESIGN.md loading. Each scenario runs against the cheapest tier of each major provider (Anthropic, OpenAI, Google) so a full sweep costs a few cents and finishes in ~2 minutes.
These are the tests you re-run when you refactor anything in SKILL.md's
## Setup section. They fail when the agent stops following the loading
contract.
Run
bun run test:skill-behavior
IMPECCABLE_SKILL_BEHAVIOR_VERBOSE=1 bun run test:skill-behavior # dump per-scenario traces
IMPECCABLE_SKILL_BEHAVIOR_MODELS=claude-haiku-4-5 bun run test:skill-behavior # scope to one model
Requires .env at repo root with at least one of ANTHROPIC_API_KEY,
OPENAI_API_KEY, GOOGLE_CLOUD_API_KEY. Providers without a key are
skipped, not failed.
How it works
Each scenario:
prepareWorkspace()mints a temp dir, symlinks the canonical skill into<workspace>/.claude/skills/impeccable, and optionally writesPRODUCT.md/DESIGN.mdfixtures.runTurn()inlinesSKILL.md(placeholders neutralized) as the system prompt and runs Vercel AI SDKgenerateTextwith four workspace-scoped tools:bash,read,write,list.- The tools record every call into a
tracethat the test asserts on. - For scenario 4, a second
runTurnreuses turn 1'sresponseMessagesso the model sees a real multi-turn conversation.
The trace is the source of truth, not the model's free-form reply.
Scenarios
| # | Setup | Assertion |
|---|---|---|
| 1 | empty workspace | runs context.mjs (which prints a NO_PRODUCT_MD directive); agent then loads reference/teach.md via Read or cat; does not start writing HTML/CSS |
| 2 | PRODUCT.md only (with ## Register: brand) |
runs context.mjs 1-3 times; loads reference/brand.md |
| 3 | PRODUCT.md + DESIGN.md (brand register) | runs context.mjs 1-3 times; loads reference/brand.md; consults the design system (DESIGN.md bundled in output, but CSS / tokens / directory listing also count) |
| 4 | PRODUCT.md + DESIGN.md, context already loaded in turn 1 | turn 2 does not re-run context.mjs; reference/brand.md is loaded across turns 1+2 |
| 5 | PRODUCT.md WITHOUT a ## Register field; task cue says "landing page" |
runs context.mjs (which emits a generic register directive); agent loads reference/brand.md via task-cue cascade |
Baseline state (2026-05-20)
Captured after condensing Setup to four bullets and teaching context.mjs
to emit a NEXT STEP: directive that names the matching register
reference when PRODUCT.md declares one (and a generic cascade prompt when
it doesn't). Use this table when comparing pre/post refactor: a
regression is "more failures than baseline", not "any failures at all".
| Scenario | claude-haiku-4-5 | gpt-5.4-mini | gemini-3.1-flash-lite |
|---|---|---|---|
| 1 (no context) | pass (variance: ~1 in 5 the agent stops after context.mjs without loading teach.md) |
pass | pass |
| 2 (product only) | pass | pass | pass |
| 3 (product + design) | pass | pass | pass |
| 4 (already loaded) | pass | fail | pass |
| 5 (no register field, task-cue cascade) | pass | pass | pass |
13-14 / 15 typical. The stable failure is gpt-5.4-mini scenario 4:
it re-runs context.mjs on turn 2 despite seeing its output in turn 1's
history. Same known weakness as the v3.2.0 script baseline; Claude and
Gemini honor the "don't re-run" rule. The S1 claude flake is rare
(observed once across many runs) and likely terminates early under
load — re-running typically clears it.