mirror of
https://github.com/pbakaus/impeccable.git
synced 2026-09-19 17:46:36 +03:00
The critique report and the AskUserQuestion call shipped in one assistant message, so the report stayed hidden until the user answered the picker and the command read as if it had never run. Reorder critique's persistence steps so the temp-file cleanup runs after the report and trend line are sent. That cleanup now ends the message carrying the report, leaving the questions to open a fresh one. Both critique.md and overdrive.md state the constraint and why it exists, so the ordering is not an unexplained sequence a model can optimize away. Overdrive additionally moves its direction descriptions inside the question options, where the user is actually reading them. Also fix the ask_instruction splices. The placeholder is a complete sentence, but five call sites spliced it mid-sentence and shipped text like "stop and STOP and call the AskUserQuestion tool to clarify. before expanding it". Every call site is now sentence-initial and the twelve lowercase provider values are capitalized to match, with a comment in utils.js pinning the contract. Record a workflow-contract baseline for the current model lineup. The two failures seen while validating this change are pre-existing: bolder refinement fails on deepseek-v4-flash identically with bolder.md reverted to HEAD, and redesign replaces DESIGN is flaky on assertions driven by new-work.md, which this change does not touch. Prepared with AI assistance (Claude Code). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
158 lines
9.5 KiB
Markdown
158 lines
9.5 KiB
Markdown
# Skill-behavior tests
|
|
|
|
LLM-backed scenarios that verify how the impeccable skill drives context,
|
|
command-reference, new-work, and native-platform loading. Each scenario runs
|
|
against one current model from each supported provider (Anthropic, OpenAI,
|
|
Google, DeepSeek).
|
|
|
|
These are the tests you re-run when you refactor anything in SKILL.md's
|
|
`## Setup` section. They fail when the agent stops following the loading
|
|
contract.
|
|
|
|
## Run
|
|
|
|
```bash
|
|
bun run test:skill-behavior
|
|
IMPECCABLE_SKILL_BEHAVIOR_VERBOSE=1 bun run test:skill-behavior # dump per-scenario traces
|
|
IMPECCABLE_SKILL_BEHAVIOR_MODELS=claude-sonnet-5 bun run test:skill-behavior # scope to one model
|
|
```
|
|
|
|
Requires `.env` at repo root with at least one of `ANTHROPIC_API_KEY`,
|
|
`OPENAI_API_KEY`, `GOOGLE_CLOUD_API_KEY`, `DEEPSEEK_API_KEY`. Providers without a key are
|
|
skipped, not failed.
|
|
|
|
## How it works
|
|
|
|
Each scenario:
|
|
|
|
1. `prepareWorkspace()` mints a temp dir, symlinks the canonical skill
|
|
into `<workspace>/.claude/skills/impeccable`, and optionally writes
|
|
`PRODUCT.md` / `DESIGN.md` fixtures.
|
|
2. `runTurn()` inlines `SKILL.md` (placeholders neutralized) as the
|
|
system prompt and runs Vercel AI SDK `generateText` with four
|
|
workspace-scoped tools: `bash`, `read`, `write`, `list`, and a fake
|
|
provider-neutral `ask_user_question` backed by a deterministic simulated user.
|
|
3. The tools record every call into a `trace` that the test asserts on.
|
|
4. For scenario 4, a second `runTurn` reuses turn 1's `responseMessages`
|
|
so the model sees a real multi-turn conversation.
|
|
|
|
The trace is the source of truth, not the model's free-form reply.
|
|
|
|
## Scenarios
|
|
|
|
| # | Setup | Assertion |
|
|
|---|---|---|
|
|
| 1 | empty workspace | runs `context.mjs`; loads `reference/init.md` before implementation; automation is not an init bypass |
|
|
| 2 | PRODUCT.md only | runs `context.mjs` 1-3 times; loads `reference/new-work.md` to resolve visual authority, establish a world when needed, and develop the surface |
|
|
| 3 | PRODUCT.md + DESIGN.md | runs `context.mjs` 1-3 times; receives the committed design system and loads `reference/new-work.md` for the task-scoped concept |
|
|
| 4 | PRODUCT.md + DESIGN.md, context already loaded in turn 1 | turn 2 does **not** re-run `context.mjs` |
|
|
| 5 | PRODUCT.md without the legacy `## Register` field and no DESIGN.md | runs `context.mjs`; greenfield craft loads `reference/new-work.md`, not init, to establish the missing world |
|
|
| 6 | PRODUCT.md + DESIGN.md + a minimal `index.html`; prompt is `/impeccable polish` | loads `reference/polish.md` |
|
|
| 7 | same fixture; prompt is `/impeccable audit` | loads `reference/audit.md` |
|
|
| 8 | PRODUCT.md + DESIGN.md + a SvelteKit scaffold (`src/app.css`, components, `+page.svelte`); prompt is `/impeccable polish src/routes/+page.svelte` | reads at least one project code file (CSS / component / page) — not just the skill's reference files |
|
|
| 9 | PRODUCT.md + `index.html` + a seeded update cache with a newer version (`skillVersion` copy-mode so `context.mjs` has a `SKILL.md` to version-check against); prompt is `/impeccable polish index.html` | `context.mjs` runs and its output carries the `UPDATE_AVAILABLE` directive (proven via captured bash output); the agent does **not** auto-run `npx impeccable update` (it must ask first) |
|
|
| 10 | no PRODUCT.md + a minimal `index.html`; prompt is `/impeccable polish index.html` | runs `context.mjs`, loads `reference/polish.md`, and does **not** divert into `reference/init.md` |
|
|
| 11 | empty workspace; prompt is `/impeccable shape ...` | runs `context.mjs`; resolves `reference/init.md` before planning the surface |
|
|
| 12 | empty workspace; prompt is natural-language build intent with no command word | runs `context.mjs`; resolves `reference/init.md` before implementation |
|
|
| 13 | empty workspace; prompt is `/impeccable teach` | runs `context.mjs` and diverts into `reference/init.md` because `teach` aliases `init` |
|
|
| 14 | PRODUCT.md with `## Platform: ios` (native iOS app); prompt is `/impeccable craft a tide detail screen` | `context.mjs` runs and emits the contents of `reference/ios.md` directly, placing native conventions in context without a second model-directed read |
|
|
| 15 | same iOS fixture; prompt is `/impeccable audit` | agent loads `reference/audit.native.md` (the Commands-table native variant, routed instead of `audit.md`) |
|
|
|
|
The workflow-contract file adds end-to-end assertions for attended fresh init,
|
|
an initialized natural build request, replacement-world redesign, and scope-preserving bolder
|
|
refinement. It checks question order and context/artifact writes rather than
|
|
only reference-file loading.
|
|
|
|
## Workflow-contract baseline (2026-08-13, current lineup)
|
|
|
|
Measured on `claude-sonnet-5`, `gpt-5.6-luna`, `gemini-3.5-flash`, and
|
|
`deepseek-v4-flash` while checking whether an `{{ask_instruction}}` rewrite had
|
|
regressed anything. Two of the four workflow-contract scenarios fail for reasons
|
|
that predate that change. Treat both as the known floor; a regression is a
|
|
failure beyond these.
|
|
|
|
| Scenario | claude-sonnet-5 | gpt-5.6-luna | gemini-3.5-flash | deepseek-v4-flash |
|
|
|---|---|---|---|---|
|
|
| attended fresh init | not measured | not measured | not measured | not measured |
|
|
| initialized natural build | not measured | not measured | not measured | not measured |
|
|
| redesign replaces DESIGN | flaky | not measured | not measured | not measured |
|
|
| bolder refinement | not measured | pass | pass | **fail** |
|
|
|
|
`not measured` means exactly that: the cell was never run in isolation on this
|
|
lineup. Only the two failing scenarios were scoped per model, because the
|
|
investigation was about whether a specific edit had regressed them. The rows are
|
|
worth keeping anyway, since a scenario absent from the table is easy to mistake
|
|
for a scenario that passed.
|
|
|
|
**`bolder refinement`, deepseek-v4-flash.** The model runs `context.mjs`, reads
|
|
`bolder.md`, `craft-floor.md`, and `current.html`, then ends its turn without
|
|
editing anything: empty `writePaths`, no `ask_user_question` call, well short of
|
|
the 16-step cap. Confirmed identical on HEAD with `bolder.md` reverted, so it is
|
|
not a skill-text problem. Same shape as the gpt-5.4-mini scenario 6/7 failures
|
|
below: the model consumes the references and then declines to act.
|
|
|
|
**`redesign replaces DESIGN`, flaky.** It has failed on two different assertions
|
|
across runs (`designWrite > question` and `implementation > designWrite`), and on
|
|
one run claude-sonnet-5 exhausted the 300s per-test timeout instead of asserting.
|
|
The traces never load `document.md`; the ordering under test comes from
|
|
`new-work.md`. Re-run before believing a single red result here. Which model
|
|
produced which failure was not pinned down, so the row records only that the
|
|
scenario is unstable.
|
|
|
|
The `bolder` claude-sonnet-5 cell is unmeasured for a specific reason: the scoped
|
|
run that produced this table used a 180s cap, which sonnet exceeded. That is a
|
|
timeout, not a failure, and it is why the guidance below insists on 300000.
|
|
|
|
### Scoping a run while investigating
|
|
|
|
Both files honor `--test-name-pattern`, which is much cheaper than a full sweep
|
|
when bisecting one scenario:
|
|
|
|
```bash
|
|
IMPECCABLE_QUESTION_DISABLED=1 CI=1 IMPECCABLE_SKILL_BEHAVIOR_MODELS=deepseek-v4-flash \
|
|
node --test --test-timeout=300000 --test-force-exit \
|
|
--test-name-pattern="bolder refinement" tests/skill-behavior/workflow-contract.test.mjs
|
|
```
|
|
|
|
Keep `--test-timeout` at 300000. A tighter cap turns claude-sonnet-5's slower
|
|
runs into timeouts that look like failures. Set `IMPECCABLE_QUESTION_DISABLED=1`
|
|
and `CI=1` so `serve-question.mjs` cannot open a browser window on the host. Pipe
|
|
to a file rather than `tail`; node prints the failing-test summary at the end,
|
|
and truncating it costs you the per-model attribution.
|
|
|
|
## Baseline state (2026-05-20, previous cheap tier)
|
|
|
|
> **Historical record.** The default models are now `claude-sonnet-5`,
|
|
> `gpt-5.6-luna`, `gemini-3.5-flash`, and `deepseek-v4-flash`. The table below
|
|
> was measured on an older cheap tier
|
|
> (`claude-haiku-4-5` / `gpt-5.4-mini`) and is kept as the historical record.
|
|
> Re-measure on the current lineup and update this section; the stronger
|
|
> models are expected to clear the scenario 6/7 routing failures that the old
|
|
> gpt tier showed.
|
|
|
|
Captured after moving sub-command reference loading from step 4 to step 2
|
|
of Setup (so the agent loads `reference/<command>.md` right after
|
|
`context.mjs`, before "doing the work" preempts it), and tightening
|
|
step 3 to require at least one project code read even when a sub-command
|
|
reference loads first. Use this table when comparing pre/post refactor:
|
|
a regression is "more failures than baseline", not "any failures at all".
|
|
|
|
| Scenario | claude-haiku-4-5 | gpt-5.4-mini | gemini-3.1-flash-lite |
|
|
|---|---|---|---|
|
|
| 1 (no context) | pass (rare flake — agent stops after `context.mjs` without loading `init.md`) | pass | pass |
|
|
| 2 (product only) | pass | pass | pass |
|
|
| 3 (product + design) | pass | pass | pass (rare flake — sub-command ref loads but world ref doesn't) |
|
|
| 4 (already loaded) | pass | pass | pass |
|
|
| 5 (no register field, task-cue cascade) | pass | pass | pass |
|
|
| 6 (`polish` routing) | pass | **fail** | pass |
|
|
| 7 (`audit` routing) | pass | **fail** | pass |
|
|
| 8 (existing project, explore design system) | pass | pass | pass |
|
|
|
|
21-22 / 24 typical. The stable failures are gpt-5.4-mini scenarios 6 and 7:
|
|
the model reads `index.html` (the target file), recognizes "polish" or
|
|
"audit" as a familiar action, and proceeds with the work without ever
|
|
loading the sub-command reference. Stronger SKILL.md wording (MUST,
|
|
"non-optional", reordered earlier) didn't move it; this looks like a
|
|
model-floor behavior rather than a skill ambiguity. Claude and Gemini
|
|
honor the load.
|