From 9e637a3732cdd143ab42b2964eedc14183a83ff4 Mon Sep 17 00:00:00 2001 From: Paul Bakaus Date: Wed, 20 May 2026 22:56:41 -0700 Subject: [PATCH] skill: add sub-command + existing-project scenarios; move sub-command load to step 2 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Adds three new LLM-backed scenarios to tests/skill-behavior: - S6: `/impeccable polish` → loads polish.md - S7: `/impeccable audit` → loads audit.md - S8: existing SvelteKit project (PRODUCT.md + DESIGN.md + src/app.css + src/lib/components/*.svelte + src/routes/+page.svelte) → agent reads at least one project code file to understand the existing design system S6/S7 surface a real model-floor: gpt-5.4-mini reads brand.md, reads the target index.html, and just does the polish/audit without ever loading the sub-command reference. Stronger SKILL.md wording didn't move it. Captured in the README baseline as a known weakness. Claude and Gemini honor the load reliably. To fix Gemini on S6/S7, sub-command reference loading is now Setup step 2 (right after context.mjs), not step 4 — placing it before the model gets focused on "doing the work". Step 3 (design-system familiarization) is tightened to require at least one project code read even when a sub-command reference loads in step 2, so Claude doesn't laser-focus on the sub-command flow and skip the broader exploration. Two new fixtures: MINIMAL_LANDING_HTML (a tiny static landing page for S6/S7) and SVELTE_PROJECT_FILES (a minimal SvelteKit scaffold with tokens, components, and a routes/+page.svelte for S8). Both designed to look real enough that agents treat them as production code. Suite is now 24 tests across three providers; baseline is 21-22/24, with the stable failures being gpt-5.4-mini scenarios 6 and 7. Co-Authored-By: Claude Opus 4.7 (1M context) --- skill/SKILL.md | 6 +- tests/skill-behavior/README.md | 36 ++++---- tests/skill-behavior/fixtures.mjs | 106 ++++++++++++++++++++++++ tests/skill-behavior/scenarios.test.mjs | 93 ++++++++++++++++++++- 4 files changed, 223 insertions(+), 18 deletions(-) diff --git a/skill/SKILL.md b/skill/SKILL.md index 1369b0aea..35d153904 100644 --- a/skill/SKILL.md +++ b/skill/SKILL.md @@ -15,9 +15,9 @@ Designs and iterates production-grade frontend interfaces. Real working code, co You MUST do these steps before proceeding: 1. Run `node {{scripts_path}}/context.mjs` once per session. If you've already seen its output in this conversation, do not re-run it. The script either prints the project's PRODUCT.md (and DESIGN.md when present) as a markdown block, or tells you it's missing. Follow whatever it prints. **If it reports `NO_PRODUCT_MD`, stop and follow `reference/teach.md` before doing anything else.** -2. Familiarize yourself with any existing design system, conventions, and components in the code. Don't reinvent the wheel; use what's there when it works, branch out when the UX wins. -3. Read the matching register reference. **This is non-optional; skipping it produces generic output.** If the project is marketing, a landing page, a campaign, long-form content, or a portfolio (design IS the product), read `reference/brand.md`. If it is app UI, admin, a dashboard, or a tool (design SERVES the product), read `reference/product.md`. Pick by first match: (1) task cue ("landing page" vs "dashboard"); (2) surface in focus (the page, file, or route being worked on); (3) `register` field in PRODUCT.md. -4. If the user invoked a sub-command (`craft`, `shape`, `audit`, ...), load its reference too. +2. If the user invoked a sub-command (`craft`, `shape`, `audit`, `polish`, ...), you MUST read `reference/.md` next. Non-optional. The reference defines the command's flow; without it you will skip steps the user expects. +3. Familiarize yourself with any existing design system, conventions, and components in the code. Read at least one project file (CSS / tokens / theme / a representative component or page). **Required even when you've loaded a sub-command reference in step 2.** Don't reinvent the wheel; use what's there when it works, branch out when the UX wins. +4. Read the matching register reference. **This is non-optional; skipping it produces generic output.** If the project is marketing, a landing page, a campaign, long-form content, or a portfolio (design IS the product), read `reference/brand.md`. If it is app UI, admin, a dashboard, or a tool (design SERVES the product), read `reference/product.md`. Pick by first match: (1) task cue ("landing page" vs "dashboard"); (2) surface in focus (the page, file, or route being worked on); (3) `register` field in PRODUCT.md. ## Design guidance diff --git a/tests/skill-behavior/README.md b/tests/skill-behavior/README.md index cab227664..8fd29c060 100644 --- a/tests/skill-behavior/README.md +++ b/tests/skill-behavior/README.md @@ -46,26 +46,34 @@ The trace is the source of truth, not the model's free-form reply. | 3 | PRODUCT.md + DESIGN.md (brand register) | runs `context.mjs` 1-3 times; loads `reference/brand.md`; consults the design system (DESIGN.md bundled in output, but CSS / tokens / directory listing also count) | | 4 | PRODUCT.md + DESIGN.md, context already loaded in turn 1 | turn 2 does **not** re-run `context.mjs`; `reference/brand.md` is loaded across turns 1+2 | | 5 | PRODUCT.md WITHOUT a `## Register` field; task cue says "landing page" | runs `context.mjs` (which emits a generic register directive); agent loads `reference/brand.md` via task-cue cascade | +| 6 | PRODUCT.md + DESIGN.md + a minimal `index.html`; prompt is `/impeccable polish` | loads `reference/polish.md` | +| 7 | same fixture; prompt is `/impeccable audit` | loads `reference/audit.md` | +| 8 | PRODUCT.md + DESIGN.md + a SvelteKit scaffold (`src/app.css`, components, `+page.svelte`); prompt is `/impeccable polish src/routes/+page.svelte` | reads at least one project code file (CSS / component / page) — not just the skill's reference files | ## Baseline state (2026-05-20) -Captured after condensing Setup to four bullets and teaching `context.mjs` -to emit a `NEXT STEP:` directive that names the matching register -reference when PRODUCT.md declares one (and a generic cascade prompt when -it doesn't). Use this table when comparing pre/post refactor: a -regression is "more failures than baseline", not "any failures at all". +Captured after moving sub-command reference loading from step 4 to step 2 +of Setup (so the agent loads `reference/.md` right after +`context.mjs`, before "doing the work" preempts it), and tightening +step 3 to require at least one project code read even when a sub-command +reference loads first. Use this table when comparing pre/post refactor: +a regression is "more failures than baseline", not "any failures at all". | Scenario | claude-haiku-4-5 | gpt-5.4-mini | gemini-3.1-flash-lite | |---|---|---|---| -| 1 (no context) | pass (variance: ~1 in 5 the agent stops after `context.mjs` without loading `teach.md`) | pass | pass | +| 1 (no context) | pass (rare flake — agent stops after `context.mjs` without loading `teach.md`) | pass | pass | | 2 (product only) | pass | pass | pass | -| 3 (product + design) | pass | pass | pass | -| 4 (already loaded) | pass | **fail** | pass | +| 3 (product + design) | pass | pass | pass (rare flake — sub-command ref loads but register ref doesn't) | +| 4 (already loaded) | pass | pass | pass | | 5 (no register field, task-cue cascade) | pass | pass | pass | +| 6 (`polish` routing) | pass | **fail** | pass | +| 7 (`audit` routing) | pass | **fail** | pass | +| 8 (existing project, explore design system) | pass | pass | pass | -13-14 / 15 typical. The stable failure is gpt-5.4-mini scenario 4: -it re-runs `context.mjs` on turn 2 despite seeing its output in turn 1's -history. Same known weakness as the v3.2.0 script baseline; Claude and -Gemini honor the "don't re-run" rule. The S1 claude flake is rare -(observed once across many runs) and likely terminates early under -load — re-running typically clears it. +21-22 / 24 typical. The stable failures are gpt-5.4-mini scenarios 6 and 7: +the model reads `index.html` (the target file), recognizes "polish" or +"audit" as a familiar action, and proceeds with the work without ever +loading the sub-command reference. Stronger SKILL.md wording (MUST, +"non-optional", reordered earlier) didn't move it; this looks like a +model-floor behavior rather than a skill ambiguity. Claude and Gemini +honor the load. diff --git a/tests/skill-behavior/fixtures.mjs b/tests/skill-behavior/fixtures.mjs index dc96e1dc2..beb1ebf07 100644 --- a/tests/skill-behavior/fixtures.mjs +++ b/tests/skill-behavior/fixtures.mjs @@ -72,6 +72,112 @@ landing page. - Never lead with screenshots. Lead with the idea. `; +/** + * Tiny static landing page fixture for scenarios that invoke sub-commands + * (polish, audit) without standing up a full framework project. Gives the + * agent something concrete to inspect so it doesn't bail with "what + * should I work on?" before completing Setup. + */ +export const MINIMAL_LANDING_HTML = ` + + + + Acme Notes + + + +

For those who build their own index.

+

Acme Notes is not a productivity tool. It is an archive, designed for the researcher who treats their notes as an external brain.

+ Begin Archive + + +`; + +/** + * Minimal SvelteKit project. Exercises Setup step 2 ("familiarize + * yourself with any existing design system, conventions, and components"): + * the agent should explore at least one of these code files before + * producing a polish or craft pass. + */ +export const SVELTE_PROJECT_FILES = { + 'package.json': `${JSON.stringify( + { + name: 'acme-notes', + type: 'module', + dependencies: { svelte: '^4.0.0', '@sveltejs/kit': '^2.0.0' }, + scripts: { dev: 'vite dev', build: 'vite build' }, + }, + null, + 2, + )}\n`, + 'svelte.config.js': `import adapter from '@sveltejs/adapter-auto'; + +export default { + kit: { adapter: adapter() }, +}; +`, + 'src/app.css': `:root { + --ink: oklch(0.16 0.02 250); + --paper: oklch(0.98 0.01 90); + --accent: oklch(0.55 0.18 28); + --hairline: oklch(0.16 0.02 250 / 0.08); +} + +body { + background: var(--paper); + color: var(--ink); + font-family: 'Inter', sans-serif; + line-height: 1.55; +} +`, + 'src/lib/components/Button.svelte': ` + + + + +`, + 'src/lib/components/Card.svelte': `
+ +
+ + +`, + 'src/routes/+page.svelte': ` + + + Acme Notes + + +
+

For those who build their own index.

+

Acme Notes is not a productivity tool. It is an archive.

+ + Some featured content. +
+`, +}; + export const DESIGN_MD_SAMPLE = `# Acme Notes — Design System ## Colors diff --git a/tests/skill-behavior/scenarios.test.mjs b/tests/skill-behavior/scenarios.test.mjs index 84547f46b..2050a202d 100644 --- a/tests/skill-behavior/scenarios.test.mjs +++ b/tests/skill-behavior/scenarios.test.mjs @@ -24,7 +24,13 @@ import { summarizeTrace, } from './harness.mjs'; import { detectProvider, getModel, hasKey, resolveModelList, PROVIDERS } from './providers.mjs'; -import { PRODUCT_MD_SAMPLE, PRODUCT_MD_SAMPLE_NO_REGISTER, DESIGN_MD_SAMPLE } from './fixtures.mjs'; +import { + PRODUCT_MD_SAMPLE, + PRODUCT_MD_SAMPLE_NO_REGISTER, + DESIGN_MD_SAMPLE, + MINIMAL_LANDING_HTML, + SVELTE_PROJECT_FILES, +} from './fixtures.mjs'; const CRAFT_PROMPT = '/impeccable craft a landing page for the project in this workspace'; const PRIMER_PROMPT = @@ -250,5 +256,90 @@ for (const modelId of resolveModelList()) { cleanupWorkspace(workspace); } }); + + it('scenario 6: sub-command routing (`/impeccable polish` loads polish.md)', async () => { + const workspace = prepareWorkspace({ + files: { + 'PRODUCT.md': PRODUCT_MD_SAMPLE, + 'DESIGN.md': DESIGN_MD_SAMPLE, + 'index.html': MINIMAL_LANDING_HTML, + }, + }); + try { + const { trace, text } = await runTurn({ + workspace, + model, + userPrompt: '/impeccable polish index.html', + maxSteps: 6, + }); + logTrace('S6', 'polish-routing', modelId, trace, { textSample: text.slice(0, 300) }); + assert.ok( + fileLoaded(trace, 'polish.md'), + `agent should load polish.md when /impeccable polish is invoked.\n` + + `Trace: ${JSON.stringify(summarizeTrace(trace), null, 2)}`, + ); + } finally { + cleanupWorkspace(workspace); + } + }); + + it('scenario 7: sub-command routing (`/impeccable audit` loads audit.md)', async () => { + const workspace = prepareWorkspace({ + files: { + 'PRODUCT.md': PRODUCT_MD_SAMPLE, + 'DESIGN.md': DESIGN_MD_SAMPLE, + 'index.html': MINIMAL_LANDING_HTML, + }, + }); + try { + const { trace, text } = await runTurn({ + workspace, + model, + userPrompt: '/impeccable audit index.html', + maxSteps: 6, + }); + logTrace('S7', 'audit-routing', modelId, trace, { textSample: text.slice(0, 300) }); + assert.ok( + fileLoaded(trace, 'audit.md'), + `agent should load audit.md when /impeccable audit is invoked.\n` + + `Trace: ${JSON.stringify(summarizeTrace(trace), null, 2)}`, + ); + } finally { + cleanupWorkspace(workspace); + } + }); + + it('scenario 8: existing SvelteKit project (agent explores design system)', async () => { + const workspace = prepareWorkspace({ + files: { + 'PRODUCT.md': PRODUCT_MD_SAMPLE, + 'DESIGN.md': DESIGN_MD_SAMPLE, + ...SVELTE_PROJECT_FILES, + }, + }); + try { + const { trace, text } = await runTurn({ + workspace, + model, + userPrompt: '/impeccable polish src/routes/+page.svelte', + maxSteps: 8, + }); + logTrace('S8', 'existing-project', modelId, trace, { textSample: text.slice(0, 400) }); + // Setup step 2: familiarize with existing design system. The + // agent should read at least one project code file (CSS / tokens / + // component / page), not just the skill's PRODUCT.md / DESIGN.md + // / reference files. + const projectReads = trace.readPaths.filter((p) => + /\.(css|svelte|tsx?|jsx?|astro)$/i.test(p) && !p.includes('.claude/skills/'), + ); + assert.ok( + projectReads.length >= 1, + `agent should read at least one project code file to understand the existing design system.\n` + + `readPaths: ${JSON.stringify(trace.readPaths, null, 2)}`, + ); + } finally { + cleanupWorkspace(workspace); + } + }); }); }