diff --git a/AGENTS.md b/AGENTS.md index dabda041f..65b7a95b3 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -16,6 +16,7 @@ If you load the Impeccable skill from `.agents/skills/impeccable/SKILL.md` (or a - `bun test tests/build.test.js` - run a focused Bun test. - `bun run test` - run the full Bun + Node test suite. - `bun run test:live-e2e` - opt-in live-mode E2E against framework fixtures (~2 min; needs `npx playwright install chromium` once). +- `bun run test:skill-behavior` - opt-in LLM-backed checks that the SKILL.md Setup flow actually drives the agent (~5 min; runs claude-haiku-4-5 / gpt-5.4-mini / gemini-3.1-flash-lite, costs cents per run, needs `.env` with provider keys). - `bun run build:browser` / `bun run build:extension` - rebuild browser-specific bundles. Run `bun run build` after changing anything in `source/`, transformer code, or user-facing counts. @@ -41,6 +42,8 @@ For changes to `skill/scripts/live-*.{mjs,js}`, also run `bun run test:live-e2e` Set `IMPECCABLE_E2E_AGENT=llm` to swap the deterministic fake agent for a Claude-backed one (`tests/live-e2e/agents/llm-agent.mjs`, default Haiku 4.5, override via `IMPECCABLE_E2E_LLM_MODEL`). Requires `ANTHROPIC_API_KEY`; tests skip cleanly when it's unset. This path hits the API — use it for verification, not CI. +For changes to `skill/SKILL.md`'s Setup section, `skill/scripts/context.mjs`, or any Setup-touching reference file (`teach.md`, `document.md`, `brand.md`, `product.md`, sub-command refs), also run `bun run test:skill-behavior`. The suite spawns real LLMs (claude-haiku-4-5, gpt-5.4-mini, gemini-3.1-flash-lite — all three, every run) with the source SKILL.md inlined as system prompt and a workspace-scoped tool set, then asserts on the tool-call trace. Provider keys live in repo-root `.env`; missing keys skip cleanly. Scope to one provider with `IMPECCABLE_SKILL_BEHAVIOR_MODELS=`; add `IMPECCABLE_SKILL_BEHAVIOR_VERBOSE=1` to dump per-scenario traces. Baseline (21-22/24) and per-scenario assertions live in `tests/skill-behavior/README.md`. + ## Anti-pattern detection rules `cli/engine/detect-antipatterns.mjs` is the source of truth for the rule engine. It feeds the CLI, the site overlay (`cli/engine/detect-antipatterns-browser.js`, regenerated by `bun run build:browser`), the Chrome extension (`extension/detector/`, regenerated by `bun run build:extension`), and the homepage `DETECTION_COUNT` in `site/public/js/generated/counts.js` (regenerated by `bun run build`). After any rule change run all three builds plus `bun run test` so nothing drifts. diff --git a/CLAUDE.md b/CLAUDE.md index 06312d3be..cd8127797 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -118,8 +118,9 @@ Local state files inside harness directories (e.g. `.claude/scheduled_tasks.lock ## Testing ```bash -bun run test # Default suite: unit + static framework fixtures -bun run test:live-e2e # Opt-in: full-cycle live-mode E2E across framework fixtures +bun run test # Default suite: unit + static framework fixtures +bun run test:live-e2e # Opt-in: full-cycle live-mode E2E across framework fixtures +bun run test:skill-behavior # Opt-in: LLM-backed checks that the skill text actually drives the agent's setup flow ``` Unit tests (build orchestration, detector logic) run via `bun test`. Fixture tests (jsdom-based HTML detection) run via `node --test` because bun is too slow with jsdom. The `test` script handles this split automatically. @@ -146,6 +147,38 @@ The agent is pluggable via a one-method interface in `tests/live-e2e/agent.mjs`: Adding a new fixture is a matter of cloning a directory under `tests/framework-fixtures/`, swapping the source files, and writing a `fixture.json`. See `tests/framework-fixtures/README.md` for the full schema. +### Skill-behavior tests + +`tests/skill-behavior/scenarios.test.mjs` is the LLM-backed safety net for edits to `skill/SKILL.md` and the Setup-adjacent reference files (`teach.md`, `document.md`, `brand.md`, `product.md`, sub-command refs). It inlines the source `skill/SKILL.md` into the system prompt of a real LLM, gives the agent `bash` / `read` / `write` / `list` tools scoped to a temp workspace, and asserts on the tool-call trace — not on the model's free-form output. The trace is the source of truth. + +```bash +bun run test:skill-behavior # full suite (24 tests, ~5 min, ~$0.10 across providers) +IMPECCABLE_SKILL_BEHAVIOR_MODELS=gemini-3.1-flash-lite bun run test:skill-behavior # scope to one provider +IMPECCABLE_SKILL_BEHAVIOR_VERBOSE=1 bun run test:skill-behavior # dump per-scenario trace JSON to stderr (use when iterating) +``` + +**Three providers per run, every run.** The suite always exercises `claude-haiku-4-5`, `gpt-5.4-mini`, and `gemini-3.1-flash-lite` — the cheapest tier of each major coding-agent provider. The point of running the cheap tier is to catch instruction-following weaknesses early. If a SKILL.md edit relies on a tightening that only Sonnet / GPT-5.5 / Gemini Pro can follow, the cheap tier will flag it. **Don't substitute Claude alone** — many of the most useful findings come from divergence between providers. + +**Auth** lives in repo-root `.env` (copied from `~/code/impeccable-evals/.env`, gitignored). Providers skip cleanly when their key is unset; they don't fail. + +**Eight scenarios:** +1. empty workspace → agent loads `reference/teach.md` +2. PRODUCT.md only → loads `brand.md` +3. PRODUCT.md + DESIGN.md → loads `brand.md` + consults the design system +4. context already loaded in turn 1 → turn 2 does **not** re-run `context.mjs` +5. PRODUCT.md without `## Register` field → agent infers `brand` from task cue +6. `/impeccable polish` → loads `reference/polish.md` +7. `/impeccable audit` → loads `reference/audit.md` +8. existing SvelteKit project → agent reads at least one project code file + +**Baseline** (post-refactor): 21-22 / 24. Stable failures live on gpt-5.4-mini scenarios 6 and 7 — it just doesn't load sub-command references regardless of MUST language. Captured in `tests/skill-behavior/README.md` as a known weakness. + +**Cost.** Each run is real LLM calls, billed to the keys in `.env`. ~$0.05-0.15 per full sweep. Keep it out of CI unless you really want it there. + +**Adding a scenario.** Write the fixture in `tests/skill-behavior/fixtures.mjs`, add the `it()` block in `scenarios.test.mjs` (the harness uses the source `skill/` dir via a symlink, so no rebuild needed), and update the baseline table in the suite's README. The harness's `fileLoaded(trace, filename)` helper checks both `read` and bash `cat` — different models prefer different tools. + +**The harness symlinks source, not built output.** This is deliberate so SKILL.md / reference / `scripts/context.mjs` edits show up immediately without `bun run build:skills`. The trade-off: reference files surface their raw `{{placeholders}}`, but the assertions key on tool calls rather than content, so it doesn't matter for correctness. + ## CLI The CLI lives in this repo under `cli/`: `cli/bin/` (entry + sub-commands), `cli/engine/` (the detect-antipatterns rule engine + browser variant), `cli/lib/` (helpers shared by CLI and Cloudflare Pages Functions). Published to npm as `impeccable`. diff --git a/DEVELOP.md b/DEVELOP.md index 2bd922b0c..2e0537210 100644 --- a/DEVELOP.md +++ b/DEVELOP.md @@ -131,6 +131,16 @@ scripts/ - `replacePlaceholders()`: Substitutes `{{model}}`, `{{config_file}}`, etc. per provider - `generateYamlFrontmatter()`: Serializes objects to YAML frontmatter (auto-quotes values starting with `[` or `{`) +## Testing + +```bash +bun run test # Default suite — unit + static fixtures (no API keys needed) +bun run test:live-e2e # Opt-in — full-cycle live-mode E2E across framework fixtures (~2 min, needs `npx playwright install chromium` once) +bun run test:skill-behavior # Opt-in — LLM-backed checks that the SKILL.md Setup flow actually drives the agent (~5 min, costs cents, needs `.env`) +``` + +The skill-behavior suite runs three providers (claude-haiku-4-5, gpt-5.4-mini, gemini-3.1-flash-lite — the cheapest tier of each, every run) with the source `skill/SKILL.md` inlined as the system prompt and a workspace-scoped `bash`/`read`/`write`/`list` tool set. It then asserts on the tool-call trace, not on free-form output. Use it whenever you edit `skill/SKILL.md`'s Setup section, `skill/scripts/context.mjs`, or any Setup-touching reference (`teach.md`, `document.md`, `brand.md`, `product.md`, sub-command refs). Per-scenario assertions and the current baseline (21-22/24) live in `tests/skill-behavior/README.md`. Provider keys live in repo-root `.env` (gitignored); missing keys skip cleanly. + ## Best Practices ### Skill Writing @@ -140,7 +150,7 @@ scripts/ 3. **Clear instructions**: LLM should understand exactly what to do 4. **Include examples**: Where they clarify intent 5. **State constraints**: What NOT to do as clearly as what to do -6. **Test across providers**: Verify it works in multiple contexts +6. **Test across providers**: Verify it works in multiple contexts. For Setup-related edits to `skill/`, `bun run test:skill-behavior` automates this across three providers. ## Reference Documentation