Files
pbakaus_impeccable/CLAUDE.md
T
9ffd3211d5 Neo Kinpaku design system + Live Mode v3 (#169)
* Add neo kinpaku design system page

* skill: rip out baked-in category recipes and saturated-default motion tropes

Programmatic bias mining (impeccable-evals) traced four major defects
back to specific lines in this skill that contradicted SKILL.md's own
first-order-reflex warning:

- brand.md "Pairing and voice" prescribed four category→aesthetic
  recipes (editorial → serif+sans, tech/dev/fintech → tight tracking,
  consumer/food/travel → script/display serif, creative → rule-break).
  These directly drove OpenAI's 76% extreme-negative letter-spacing
  on tech briefs and Anthropic/Google's 28-34% italic-serif-display
  slop on editorial/food briefs. Replaced with one sentence: the
  shape depends on the brand, not on the brand's category.
- brand.md "Brand permissions" had "Typographic risk. Enormous
  display type, unexpected italic cuts, mixed cases, hand-drawn
  headlines, a single oversize word as a hero." — a four-for-one
  slop driver behind 97% OpenAI comically-large H1, 42% bad-SVG
  illustration, and the editorial-italic slop. Deleted outright.
- typeset.md and teach.md repeated the same category recipes;
  trimmed to the principle without the recipe.
- SKILL.md Typography: added a hard hero-H1 ceiling (clamp() max
  ≤ 6rem ≈ 96px), with a <codex> block to make it explicit since
  OpenAI over-indexes here (97% ≥128px vs 24% for Anthropic).
- animate.md, bolder.md, brand.md: removed "staggered reveals" and
  "scroll-triggered transitions" as the prescribed default ambitious
  motion. By 2026 that's the saturated AI tell, not a choreography.
  Reserved stagger for legitimate list-sibling rhythm.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* skill: anti-cream + codex-specific defect bans + universal slop bans

Second pass after measuring more biases against the eval corpus.

- SKILL.md Color: explicit "cream/sand/beige body bg is the saturated
  AI default of 2026" rule. Tone down the "tint every neutral" line so
  it doesn't read as "default to warm-tinted near-white" (which OpenAI
  hits at 74% and Anthropic at 31%-47%).
- SKILL.md Absolute bans: add universal bans for two slop patterns
  detected at 55-95% across providers — tiny uppercase tracked eyebrow
  above every section (the 2023-era kicker that's now AI grammar) and
  numbered section markers (01/02/03). Also explicit "text that
  overflows its container is the universal defect on tablet/mobile."
- SKILL.md Absolute bans → <codex> block: ban the GPT-specific defects
  Paul annotated repeatedly — `border:1px solid` + soft-wide-shadow
  (≥16px blur) "ghost cards", `border-radius:32px+` over-rounding,
  hand-drawn/sketchy SVG illustrations (loose-sketch / *-sketch classes,
  feTurbulence paper-grain filters), repeating-linear-gradient stripes,
  "X theater" AI-slop copy phrases.
- SKILL.md Motion → <gemini> block: the image :hover transform tell
  (38% Google skill-on rate). Hover effects on images add no info; the
  image isn't an action target. Animate card chrome, not the image.
- SKILL.md Typography: hard display letter-spacing floor ≥-0.04em
  (OpenAI defaults to -0.075em → cramped). Existing hero ceiling
  <codex> block extended with the letter-spacing rule.
- codex.md Step A example: stop seeding "warm-grounded (deep oxblood +
  cream)" as the warm-palette template, which primes the cream default.
- colorize.md Tinted backgrounds: stop printing the literal cream
  recipe `oklch(97% 0.01 60)`; replace with brand-anchored guidance.
- document.md examples: warm-ash-cream → cool-paper so the example
  doesn't seed cream as the canonical neutral example.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* skill: universal anti-slop bans + contrast/font-count/all-caps-body rules

Third pass after measuring the rest of the cross-provider matrix:

- Color: explicit "Verify contrast" rule. Low-contrast text fires at
  68% across all providers skill-on (90+% off). The most common
  failure is muted gray body on a tinted near-white; light-gray-for-
  elegance is named as the single biggest cause of unreadable AI
  pages.
- Typography: max-3-font-families rule. Overused-fonts (>4 families)
  fires at 28% Anthropic / 36% Google / 0% OpenAI skill-on; >50% off.
  Also: universal "no all-caps body copy" (moved from brand-only ban
  to Shared design laws since product-register also overuses caps).
- Copy: anti-aphoristic-cadence ban targets Anthropic's signature
  "X. No Y." / "X. Just Y." voice (63% skill-on copy-slop rate, 77%
  off — the worst rate in the matrix). Once-is-voice / three-or-more-
  is-tell framing per the runner's copy-slop detector.
- Copy: anti-SaaS-buzzword-string ban with the literal phrase list
  the detector watches for (streamline/empower/supercharge, trusted-
  by-leading, best-in-class/enterprise-grade/cutting-edge, etc).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* skill: strengthen anti-cream rule across full warm-neutral band

Smoke validation showed the cream fix worked for Google + OpenAI but
Anthropic Sonnet italian-restaurant still shipped `--paper: oklch(90%
.018 88)` — cream just outside the L≥95% band the rule cited.

Broaden the rule:
- Band: OKLCH L 0.84-0.97, C < 0.06, hue 40-100 (was 95-97% / 60-95).
- Name the token-name tells explicitly (paper / cream / sand / bone /
  flour / linen / parchment / wheat / biscuit / ivory) — the model
  defaults to one of these regardless of what hex it lands on.
- Call out the specific brief patterns ("warm, traditional, family-
  coastal-Italian" / "editorial-restraint") that the model translates
  into cream by reflex. Then provide three explicit non-cream options:
  saturated brand color, true off-white at C=0, or darker mid-tone.

Warmth in the brand is carried by accent + typography + imagery, not
by body bg.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* v3.2.0: skill bias-fix release

Bumps version from 3.1.1 to mark the four-commit skill cleanup that
rips out baked-in category recipes (brand.md), saturated-default motion
tropes (staggered reveals everywhere), the cream/sand body-bg AI tell,
codex-specific defects (1px+wide-shadow, over-rounding, hand-drawn SVGs,
stripes, X-theater copy), the extreme-letter-spacing default, and
universal slop bans (all-caps eyebrow on every section, numbered-section
markers, all-caps body, font-family-count > 3, aphoristic copy cadence,
SaaS buzzword strings). Plus a hard hero-H1 ceiling (clamp() ≤6rem) and
a Gemini-specific image:hover transform block.

Validated against ~190 post-fix samples — see impeccable-evals
biases tab for per-provider deltas.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* drop "no pure black/white" rule entirely

The rule was contested in the design world and causing more damage than
good — pushing every page into the tinted-near-white default which is
the cream/sand AI tell we already explicitly ban elsewhere. Vercel,
SVKMS, Brutalist sites, et al. use pure black/white successfully; the
skill shouldn't second-guess that.

Skill markdown deletions:
- SKILL.md Color: drop the "Never use #000 or #fff" bullet.
- color-and-contrast.md: drop the "Never Use Pure Gray or Pure Black"
  subsection, the "Never pure black" table-row prescription, and the
  "Avoid: Using pure black for large areas" bullet.
- colorize.md: drop the "NEVER use pure black or pure white for large
  areas" bullet.
- polish.md: drop the "Tinted neutrals: No pure gray or pure black"
  half of the bullet (the gray-on-color bullet survives).

Detector code (cli/engine):
- registry/antipatterns.mjs: remove the `pure-black-white` entry.
- rules/checks.mjs: remove the three `findings.push({ id:
  'pure-black-white', ... })` emit points (inline #000 bg, Tailwind
  bg-black class, plain-HTML scan path).
- engines/regex/detect-text.mjs: remove the two pure-black-white regex
  rules (CSS `background: #000…` + Tailwind `bg-black`).
- detect-antipatterns-browser.js: regenerated via
  scripts/build-browser-detector.js.

Tests:
- detect-antipatterns-fixtures.test.mjs: invert the assertion that
  pure-black-white fires; expect it to NOT fire post-v3.2. Drop the
  Tailwind bg-black-opacity edge-case test (no longer relevant).
- detect-antipatterns.test.js: drop the standalone "detects pure-
  black-white in styled-components" test and remove pure-black-white
  from the multi-detector assertions in PricingCard, globals.css, and
  GlobalStyle.tsx tests.

166 bun tests pass; 24 node fixture tests pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* skill: strip example patterns from copy rules, strengthen gemini block

v3.2 rerun validation surfaced two issues:

1. Copy-slop detector fires more on Gemini under v3.2 (48% → 84%) than
   under no-skill baseline. Root cause: the anti-aphoristic-cadence rule
   printed the literal "X. No Y." / "X. Just Y." patterns as examples,
   and Gemini imitated them as the recommended voice. Same recipe-becomes-
   bias trap we hit with brand.md:116's "Enormous display type, unexpected
   italic cuts, mixed cases, hand-drawn headlines" enumeration. Fix:
   describe the cadence as a rhythm ("serious statement, then punchy
   short negation") without printing literal patterns. Buzzword list
   trimmed to a single inline phrase family rather than quoted strings.

2. Gemini image:hover transform Gemini-tell hadn't dropped (31% off →
   32% v3.2). Strengthen the <gemini> block: explicit "Never animate
   <img> elements on hover", call out the Tailwind group-hover:scale /
   group-hover:rotate / group-hover:translate parent-hover patterns by
   name (Gemini was reaching for these via Tailwind even though the
   prior text talked about :hover on the image directly).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* skill: simplify context loading and inline register directive

Replaces load-context.mjs's JSON output with a tight markdown block from
the renamed context.mjs. The script now extracts PRODUCT.md's `## Register`
field and appends a `NEXT STEP:` directive naming the matching reference
(brand.md / product.md), which moved Gemini from skipping the register
load entirely to honoring it. Drops the `.impeccable.md` auto-migration;
makes IMPECCABLE_CONTEXT_DIR a lazy escape hatch consulted only when the
default paths come up empty.

Setup is now four bullets in one list. The DESIGN.md nudge is gone; in
its place, a "familiarize with the existing design system" step that
calls out CSS / tokens / running app as authoritative sources alongside
DESIGN.md. The standalone `### Register` H3 stays for the cascade rules
(task cue → surface → register field).

New LLM-backed test suite at tests/skill-behavior/ runs five scenarios
against claude-haiku-4-5, gpt-5.4-mini, and gemini-3.1-flash-lite via
Vercel AI SDK. Captures real tool traces, asserts on context.mjs calls,
brand.md loads, and teach.md fallback. Skips cleanly when API keys are
unset. 13-14/15 pass; only stable failure is the v3.2.0-era gpt-mini S4
"don't re-run" regression. Adds @ai-sdk/google as devDep and the
test:skill-behavior npm script.

Touches em-dashes in skill/SKILL.md and four reference files so
`bun run build:skills` passes its skill-prose validator. teach.md and
document.md drop their "re-run the loader to refresh session cache"
steps since the agent's own write is now the freshest source.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* skill: merge orphan reference files into command sub-skills + inline S-tier invariants

Two related restructurings:

1. SKILL.md now carries the cross-domain invariants that catch defects in any
   project (contrast/placeholder/gray-on-color, similar-font pairing, text-wrap,
   tabular-nums, centered-stack default, Flex/Grid choice, auto-fit grids,
   semantic z-index, reduced motion, stagger vs section-fade, premium motion
   materials, focus-visible, placeholders-aren't-labels, dropdown overflow trap,
   button/link copy). Greenfield-only rules (theme picking, color strategy,
   tinted neutrals) live under "New projects only".

2. Reference files merged into their command counterparts:
   - spatial-design.md  -> layout.md
   - motion-design.md   -> animate.md
   - color-and-contrast.md -> colorize.md
   - responsive-design.md  -> adapt.md
   - ux-writing.md         -> clarify.md
   - typography.md         -> typeset.md (bolder.md redirected)
   - cognitive-load.md + heuristics-scoring.md + personas.md -> critique.md

   craft.md and shape.md "load references" lists updated to new file homes.
   interaction-design.md stays standalone (no 1:1 command verb).

Net: 36 -> 27 reference files. Same content, fewer files, no orphaned
reference loaded only from craft.md.

Also extends the routing rules: if the user's first word doesn't match a
command but the intent clearly maps to one, load that command's reference
and proceed as if invoked.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* skill: add sub-command + existing-project scenarios; move sub-command load to step 2

Adds three new LLM-backed scenarios to tests/skill-behavior:
- S6: `/impeccable polish` → loads polish.md
- S7: `/impeccable audit` → loads audit.md
- S8: existing SvelteKit project (PRODUCT.md + DESIGN.md + src/app.css +
  src/lib/components/*.svelte + src/routes/+page.svelte) → agent reads
  at least one project code file to understand the existing design system

S6/S7 surface a real model-floor: gpt-5.4-mini reads brand.md, reads the
target index.html, and just does the polish/audit without ever loading
the sub-command reference. Stronger SKILL.md wording didn't move it.
Captured in the README baseline as a known weakness. Claude and Gemini
honor the load reliably.

To fix Gemini on S6/S7, sub-command reference loading is now Setup step 2
(right after context.mjs), not step 4 — placing it before the model gets
focused on "doing the work". Step 3 (design-system familiarization) is
tightened to require at least one project code read even when a
sub-command reference loads in step 2, so Claude doesn't laser-focus on
the sub-command flow and skip the broader exploration.

Two new fixtures: MINIMAL_LANDING_HTML (a tiny static landing page for
S6/S7) and SVELTE_PROJECT_FILES (a minimal SvelteKit scaffold with
tokens, components, and a routes/+page.svelte for S8). Both designed to
look real enough that agents treat them as production code.

Suite is now 24 tests across three providers; baseline is 21-22/24, with
the stable failures being gpt-5.4-mini scenarios 6 and 7.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* skill: add reveal-animation safety rule (must enhance, not gate visibility)

Class-triggered visibility transitions pause on hidden tabs and headless
renderers. The italian-restaurant smoke produced a build where 2 sections
shipped opacity:0 because the CSS transition never advanced past
currentTime=0 (timeline paused). Added one-liner under Motion to prevent
the antipattern: reveals must enhance an already-visible default, never
gate content visibility on a class-triggered transition.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* skill: restore prescriptive cream/sand/beige paragraph

Bisection across 5 historical skill commits on Gemini 3.5 flash fast
lane n=3 found that 0cf2debd was the peak quality state. The regression
between 0cf2debd and HEAD came from simplifying the long anti-cream
paragraph into a one-liner.

Restoring the paragraph (with em-dashes replaced by parens to satisfy
prose lint) recovers ~0.22pt average on Gemini vs HEAD, with the
largest gains on:
- 09-luxury-hotel: +0.50 (restores editorial drama in photo-led briefs)
- 10-food-magazine: +0.67
- 03-italian-restaurant: +0.51

The paragraph's load-bearing parts are the (a)(b)(c) alternatives that
give the model actionable replacements for cream-tinted body bg
("saturated brand color as body", "true off-white at chroma 0",
"darker mid-tone tinted neutral"). Without them, the one-line warning
left the model with no concrete alternative.

Cross-provider validation showed the pattern matches historical
behavior: Gemini benefits from prescriptive scaffold (+0.12 over off),
Sonnet is roughly neutral (+0.01), GPT-5.5 slightly regresses (-0.11
matching the v3.1.0 pattern of -0.11). The skill has never been
uniformly better than skill-off across providers; this is the closest
achievable state without provider-specific rework.

The structural improvements from the prior restructure stay (file
merges, S-tier inlines, routing rule extension, reveal-animation
safety rule).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs: teach CLAUDE.md / AGENTS.md / DEVELOP.md about the skill-behavior tests

Adds the `bun run test:skill-behavior` script to the test commands lists
in all three docs. CLAUDE.md gets a full `### Skill-behavior tests`
subsection paralleling the existing Live-mode E2E one: how the suite
works (inlines source SKILL.md, scoped tools, asserts on the trace),
which providers it always runs (claude-haiku-4-5, gpt-5.4-mini,
gemini-3.1-flash-lite — all three every run), the eight scenarios, the
baseline (21-22/24 with stable gpt-mini sub-command-routing failures),
auth via repo-root `.env`, and how to add a scenario.

AGENTS.md gets the one-liner plus a paragraph in Testing Guidelines that
points contributors at the suite for Setup-touching edits (SKILL.md
Setup section, context.mjs, teach.md, document.md, register / sub-command
refs).

DEVELOP.md gets a short Testing section that didn't exist before, plus a
nudge in the "Test across providers" bullet pointing at the new suite as
the automated way to do that.

No code changes.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* detector: add 5 new antipatterns (em-dash-overuse, broken-image, marketing-buzzword, numbered-section-markers, aphoristic-cadence)

Consolidates eval-side detection logic into the canonical impeccable
detector. Before this change, the eval harness had its own duplicate
implementations of em-dash, copy-slop, and broken-image checks. They
now live alongside the existing 28 antipatterns in the impeccable
registry, available to the CLI, browser extension, critique skill,
and eval (via the existing slop grader child-process call).

New antipatterns:
- em-dash-overuse: 5+ em-dashes in body text content (threshold
  permits legitimate prose use of em-dash; only triggers on AI
  cadence-level density)
- broken-image: <img> with empty src, missing src, or src="#"
- marketing-buzzword: SaaS phrase list (streamline / empower /
  supercharge / enterprise-grade / cutting-edge / etc)
- numbered-section-markers: repeated 01 / 02 / 03 sequence as
  section labels — the AI editorial scaffold one tier deeper than
  tracked eyebrow chips
- aphoristic-cadence: 3+ manufactured-contrast ("Not a X. A Y.")
  or short-rebuttal ("Sentence. No clause." / "Sentence. Just
  clause.") constructions in body text

Engine wiring:
- broken-image runs as a static-html element rule (selector: img)
  and a fallback regex matcher (for non-HTML files)
- em-dash / buzzword / numbered / aphoristic run as regex
  page-analyzers, factored into a new runTextContentAnalyzers()
  helper that both detectText (non-HTML) and detectHtml (HTML)
  call, so .html files get the same coverage as .css/.tsx

Tests: 166 detector + 12 browser + 24 fixture all pass.
Browser detector rebuilt (162.7 KB).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* skill: drop unvalidated anti-centering rule; add image-led hero carve-out

The anti-centering rule ("Don't default to centering everything") was
added without empirical support. We have a detector for it
(everything-centered, threshold ≥70%) that fires on 0 / 998 samples
in the corpus — never validated, never useful.

Meanwhile the rule was almost certainly responsible for collapsing
Gemini 3.5 flash's luxury-hotel skill-on output from the canonical
"full-bleed photo + centered overlay headline" cinematic hero (the
shape skill-off Gemini chooses 67% of the time) to a 50/50
magazine grid (full-bleed rate drops to 18% under skill-on, -49pp).

Changes:
- skill/SKILL.md #### Layout: drop "Don't default to centering..."
- skill/reference/brand.md ## Layout: drop the same rule; replace
  with a positive carve-out — image-led briefs (hotels, restaurants,
  magazines, photography) often want full-bleed hero with overlaid
  menu and centered headline; let the photograph be the design
- skill/reference/layout.md: drop the assessment question and the
  "asymmetric breaks centered-content pattern" framing

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Apply neo-kinpaku design system and improve live picker UX

Restyle the live picker to match the site kinpaku kit, persist pick mode
in localStorage, fix DESIGN.md color swatches in the parser, and land the
neo-kinpaku site refresh with new tokens, assets, palette script, and
detector rules.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Add live Steer end-to-end: poll protocol, browser UI, and E2E harness.

Wire page-level Steer through the live server and agent poll loop with steer_done
unlock semantics, extend live.md for agents, and add smoke tests with LLM
handleSteer plus recovery for hidden heroes, HMR lag, and dev-tool overlays.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Add experimental live-poll --stream mode; keep one-shot default for Cursor.

Stream keeps one process alive with ack-aware resume, but live.md documents
that Cursor should stay on one-shot background notify after testing showed
~5s pickup vs sub-second on exit-based notify.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Sync harness output and fix build validators for poll stream release.

Regenerate provider skills after live-poll --stream work, update homepage
detection counts to 41, and replace em dashes in site/skill copy so
bun run build passes prose and count checks.

Co-authored-by: Cursor <cursoragent@cursor.com>

* homepage: add testimonials marquee section

A two-row testimonial marquee on a tinted graphite plinth, sitting
between the hero and the slop teaser.

29 testimonials sourced via api.fxtwitter.com (lightly cleaned: leading
@-mention reply targets stripped, trailing self-links removed). Avatars
downloaded into site/public/assets/testimonials/ so they're served
locally. Quote order curated for impact — both rows lead with the
punchiest quotes (Ben Davis spotlight, "Impeccable > Claude design",
"THIS. This shit works.", "Uninstall whatever frontend skill you're
using.") so the first viewport is loaded with the most memorable
testimonials.

Engineering notes:
- Section uses width:100vw + margin-left:calc(50% - 50vw) to escape
  main.site-content's max-width + side padding (cards now clip cleanly
  at the actual viewport edges).
- Marquee runs at 110s linear infinite. Both rows share the same
  duration so on-screen speeds match; track is doubled so the loop
  back to 0 reads as continuous.
- Hero min-height reduced from 100svh to calc(100svh - 115px) so the
  dotted divider and top of row A peek above the fold on landing,
  signalling the section is there.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* homepage: keep the hero demo clear of the fixed header on short viewports

The hero centers its content in the full viewport (the site header is a fixed
overlay), so on shorter screens the tall Live Mode demo tucked under the nav.
Raise the hero's top padding above the 97px header (113px wide, 108/92px when
stacked) so content always pins below the header while still centering on tall
viewports, and cap the demo frame to the viewport so the whole demo stays on
screen.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Add steer voice input and refine processing animation.

Wire Web Speech API on the Steer mic with auto-submit, block Cursor's preview browser with a clear message, and replace truncated "Working" text with a dots-only processing state.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Add agent poll connectivity indicator and tighten global bar spacing.

Surface poller state on the Impeccable mark via SSE and /status, with an instant disconnected tooltip, steer timeout failsafe, and matched brand/chat section gaps.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Fix steer focus to allow page text selection without losing type-to-steer.

Blur the hidden steer input on page interaction, pause refocus during selection gestures, and reschedule focus recovery after clicks and cleared selections.

Co-authored-by: Cursor <cursoragent@cursor.com>

* site: rework "Design in production" section glyphs and audience band

Put the three how-it-works steps back into thin-line cards and drop the
overused browser-chrome bars from each glyph. Redraw the step 2 and 3
visuals to mirror the real Live Mode UI: step 2 shows the on-canvas pick
outline with an attached comment bubble, step 3 shows the floating
contextual accept bar plus the source-write confirmation. Re-treat the
audience tiles as verdigris-lined text (no card box) under a "Who it's
for" eyebrow, so each role reads as distinct from the gold step band.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Add live insert mode with HMR-safe placeholder recovery.

Ships insert picking, scaffold helpers, variant cycling fixes for hidden
variants, and placeholder snapshot/recreation so Astro HMR does not drop
the wait-state box or re-anchor to the hero container.

Co-authored-by: Cursor <cursoragent@cursor.com>

* site: mobile pass — hamburger nav + designing hero overflow fix

The header was rendering inline nav links + GitHub button that overflowed
narrow viewports (~363px). Pre-existing display:none hacks hid Designing
and Live to make the row fit, but those items still belonged in the menu.

Header.astro: added a hamburger toggle button + inline script. The right
cluster (nav + GitHub) becomes a collapsible drawer below the header on
mobile, with data-nav-open driving the open/closed state and animating
the two-line glyph into an X.

kinpaku-kit.css: hamburger button (kinpaku-bordered glyph), mobile drawer
panel (solid lacquer-deep bg, hairline separators between rows, full-width
tappable rows), and overrides for the older sub-pages.css mobile rules
(horizontal-scroll mask on the nav, hidden [data-nav="home"] item, hidden
GitHub star label) — all redundant now that the drawer surfaces everything.

home-kinpaku.css: dropped the @media (max-width: 560px) block that hid
Designing / Live / GitHub. The drawer pattern shows them all.

designing-kinpaku.css: hero h1 "Designing with Impeccable" was overflowing
at narrow viewports. Three fixes:
  - grid-template-columns 1fr → minmax(0, 1fr) so the column shrinks to
    fit container instead of growing to "Impeccable"'s 472px intrinsic
    min-content width.
  - mobile h1 size override (clamp(2.2rem, 11vw, 3rem) at <=480px) since
    the display token's 3.4rem minimum is sized for desktop hero impact.
  - hide the decorative loop-wheel SVG below 600px (was overflowing 22px
    past the right edge).

Verified clean at both 363px and 403px viewports across /, /docs,
/docs/animate, /slop, /designing, /live-mode. scrollWidth matches viewport
width on every page (no horizontal scroll).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* detector: refine new rules + run provider tells in browser env

Follow-up to the detector port (rules landed in 7648af00):
- oversized-h1: flag long headlines set at display size, not punchy
  one/two-word heroes (length, not size alone, is the tell)
- provider tells (--gpt/--gemini) now always run in a real browser env
  (detector page, live overlay, extension); gating is a CLI-output
  concern only, applied in the Node engine return paths
- move theater-slop-phrase into checkHtmlPatterns so it runs in the
  bundled browser path, not just CLI/static (browser bundle excludes
  detect-text.mjs)
- hero-eyebrow-chip overlay highlights the eyebrow, not the heading
- gemini-tells fixture: data-URI images so the hover-zoom renders
- rebuild browser bundle

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* site: migrate /detector lab to neo-kinpaku design system

Rebuild the detector lab tool shell on --ks-* tokens (lacquer ground,
gold hairlines, champagne/mono type) instead of the legacy warm-paper
palette. Swap the "/" placeholder for the real carved-tile brand lockup,
restyle the toolbar actions as kinpaku primary/secondary buttons, and
recolor the finding overlay from off-brand magenta to vermilion.

Update the global theme-color from #fafafa to #010101 (the sRGB render
of the lacquer ground) so the browser chrome matches the dark site.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Homepage: hero finalist, compact live demo, real picker bar.

Switch the hero to m-01-v2-01, tighten the in-hero demo layout, and replace
the marketing gbar with a shared LiveDemoGbar that mirrors live-browser.js.
Size the bar with max-content so controls are not clipped inside the capsule.

Co-authored-by: Cursor <cursoragent@cursor.com>

* site: migrate /cases/neo-mirai to neo-kinpaku design system

Rebuild the Neo Mirai case-study page on --ks-* tokens: lacquer ground
(drops the off-brand magenta radial spotlight), Alumni Sans Pinstripe
display headings instead of the banned italic serif, gold eyebrow/labels,
gold hairline image frames, kinpaku primary/secondary buttons, and a
lacquer-deep command panel with a gold-bordered code block.

Opt .neon-case-page into the shared kinpaku site-header/footer chrome in
kinpaku-kit.css (per the "add new kinpaku pages to the selector list"
note) so the global header and footer go dark to match the page.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* site: consolidate kinpaku header+footer into one reusable .kinpaku-chrome class

The dark header/footer were not a reusable unit: the header was scoped to
a per-page selector list, the github star pill was home-only, and the
default footer was copy-pasted into four page stylesheets. Pages not on
the lists (like /cases/neo-mirai) fell back to the legacy light chrome.

Collapse all of it into one `.kinpaku-chrome` block in kinpaku-kit.css —
header, github pill, and default footer — and opt every kinpaku page in
via a single body class. Delete the four duplicated per-page footer
blocks and the home-only github pill. The home page keeps its textured
verdigris footer as a deliberate override, raised to body.home-kinpaku
specificity so it wins regardless of import order. Genuinely light pages
(privacy, tutorials) just omit the class.

Fixes on /cases/neo-mirai: footer and github star now render dark/kinpaku
(were legacy-light), and the content sections are wrapped in the .neon-case
container so they sit in header-aligned gutters instead of bleeding to the
viewport edge.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* site: migrate privacy + tutorials to kinpaku via a reusable surface class

These were the last two light pages. Rather than rewrite their per-rule
styling, add a reusable .kinpaku-surface class that remaps the legacy
--color-* / --font-* tokens to kinpaku values at the body scope, so the
existing legacy-token CSS (sub-pages.css prose, the pages' inline styles)
renders dark for free. Same trick docs-kinpaku/slop-kinpaku use per page,
lifted into one shared class. Pair it with .kinpaku-chrome for header +
footer.

privacy + both tutorials pages now carry both classes. Also force the
sub-1.2rem headings (tutorial card titles, prose h1/h2) back to the
upright body face: the legacy display face was italic serif, and the
kinpaku Pinstripe face reads wrong synthesized-italic at small sizes.

No light pages remain.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* site: re-add Tutorials to the /docs sidebar

Tutorials lost its docs placement across two refactors: the Astro docs
rebuild never carried over the sidebar tutorials list the old generated
pages had, and the kinpaku homepage redesign dropped the "Full
walkthrough" link. It survived only via /designing and /live-mode.

Add a "Tutorials" group at the top of the docs sidebar (matching the
command-category styling) linking the index plus all four tutorials,
restoring the old information architecture.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* site: make kinpaku the default — flip legacy :root tokens to dark (phase 1)

Repoint the legacy design tokens in tokens.css from light-mode to kinpaku:
--font-* now reference the --ks-* brand faces (retiring Cormorant/Instrument/
Space Grotesk), surfaces carry dark-lacquer oklch, and --color-accent is gold
instead of magenta. Values mirror the per-page kinpaku remaps.

Every live page already overrides these at its body-class scope, so this
changes the fallback (any classless/new page now renders kinpaku) without
altering existing pages — verified home, designing, slop, live-mode, docs
unchanged, and the deliberate-light demos (slop specimens, home's Aurelia
mock) still render light via their own colors.

First step toward removing the per-page remaps; those become redundant next.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* detector + slop: cream-palette rule, drop everything-centered, polish catalog

- new deterministic cream-palette rule ("claude beige"): flags warm
  lightly-tinted off-white page backgrounds; wired into static + browser
  engines, with fixture + test
- remove everything-centered rule entirely (no longer in the skill) from
  registry, regex analyzer (+ index-offset fix), checkPageLayout, and tests
- catch Instrument Serif in overused-font (regex + OVERUSED_FONTS)
- /slop: reconcile catalog (cream card in, everything-centered out; counts),
  and fix demo visuals — visible hairline border, gigantic clipped hero,
  more extreme crushed tracking, padded gray-on-color card, uniform-rhythm
  monotonous-spacing, long line-length line, elastic-overshoot dialog for
  bounce easing, real zooming image for image-hover; flip the demo surface
  off warm beige to a cool neutral

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* detector page: add cream-palette fixture to the catalog

Surfaces the new cream/beige palette rule on /detector alongside the
other Color specimens.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* site: shared docs sidebar + tutorial pages join the layout

Extract the /docs section sidebar into a reusable DocsSidebar component
and wire it into all three entry points so the navigation is consistent
across docs index, command pages, and tutorial pages.

site/components/DocsSidebar.astro (new): one source of truth. Loads the
tutorials + skills collections, renders Tutorials → Commands grouped by
category, and highlights the active entry via activeCommand / activeTutorial
props.

site/pages/docs/index.astro: swap the inline sidebar markup for the
component. Drop the "All tutorials" link — the dedicated tutorials
listing page wasn't earning its slot in the rail.

site/layouts/Doc.astro: same swap. Command pages now also see the
Tutorials section above Commands, matching /docs.

site/pages/tutorials/[...slug].astro: rewrite from a standalone page
(custom .tutorial-page wrapper, ad-hoc breadcrumb) to the full
skills-layout shell with DocsSidebar in the left rail. Tutorial content
now reads in the same layout as command reference pages.

site/content/tutorials/brand-vs-product.md (deleted): the skill picks
the register automatically from PRODUCT.md, so a tutorial telling users
to pick it themselves was misleading.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* detector: catch Tailwind warm-light bg utilities in cream-palette

The static engine can't resolve Tailwind classes to computed CSS, so a
`bg-amber-50` on <body> slipped past the cream-palette rule. Add a
class-list fallback that scans body/html for arbitrary `bg-[...]` values
and named warm-light utilities (amber/orange/yellow/stone), each run
through the same isCreamColor test so neutrals and over-saturated shades
drop out. Fixture + test for the class-only case.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* site: drop redundant per-page token remaps (phase 2)

With kinpaku now the :root default, the --color-* / --font-* remap blocks
in docs/slop/designing/live-mode-kinpaku.css re-declared values identical
to :root. Removed them, keeping only the --ks-muted alias (still read by
name in those files) and each page's shell (gradient bg, color, min-height).

home-kinpaku.css keeps its remap: it uses home-specific values (e.g.
--color-charcoal: var(--ks-text), --color-cream: var(--ks-lacquer-raised))
plus the --cat-* gradient overrides, so it is not redundant.

Verified designing (PRODUCT.md viz), slop (specimens stay light), docs,
live-mode unchanged.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* site: drop italic from 15 dead editorial-serif heading rules

Audited every font-style: italic in sub-pages.css and main.css against
the live markup. Removed italic from the 15 rules whose selectors don't
appear in any page/component/content/script:

  sub-pages.css: docs-home-card-title, docs-category-title,
    tutorial-embed-caption, skill-demo-caption, skill-source-card-subtitle,
    skill-references-heading, skill-reference-title
  main.css: hero-title-combined, hero-tagline-combined, impeccable-title,
    loading-state, install-primary-howto .install-path-desc em,
    install-howto-steps > li::before, install-step-status, consulting-title

These were dormant remnants of the retired Cormorant italic-serif look —
the kinpaku Pinstripe face renders them as bad synthesized-italic, but
no markup matches the selectors so nothing rendered. Removed only the
font-style declaration; the rest of each rule stays (whole-rule cleanup
is out of scope).

Kept the 5 live selectors (slop-section-heading, tutorial-card-title,
visual-mode-demo-caption, visual-mode-method-name, gallery-card-title)
per the "if they're not used anywhere" condition, plus .prose em (real
emphasis) and .prose blockquote (conventional blockquote italic).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* skill: brand-seed palette.mjs + Setup step to run it

New-brand color now starts from a curated seed color (129 OKLCH seeds)
instead of the model guessing or defaulting to warm-cream. The script
returns one seed + composition guidance (pure-bg architecture, perceptual
text-on-fill, anti-cliché moods, jewel-tone range), with inverse-frequency
hue weighting for fair rainbow exposure and deterministic --from picking.
SKILL.md Setup step 5 makes it run for greenfield projects. Curation
tooling lives in the impeccable-evals repo (tools/palette/).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Remove accidental live mode inject from Base.astro.

The localhost live.js tag was left in the site layout after a dev session and should never ship in the Astro template.

Co-authored-by: Cursor <cursoragent@cursor.com>

* site: dedicated /changelog + /faq, epic v3.5.0 notes, Live Mode → Beta

Split changelog and FAQ out of the homepage into two standalone kinpaku
pages, linked from the footer (and a quiet hint under the Get-started CTA).

/changelog: every release inline (no collapsible), newest first. The
v3.5.0 entry leads with a one-line summary, a real before/after pair from
the GPT-5.5 eval corpus (luxury-hotel brief, skill off vs on), and a stat
row (74% cream-bg, 76% extreme tracking, 90%+ low-contrast — measured
across ~190 samples). Then five scannable bold-led bullets, biggest
takeaway first: per-provider skill compilation, the bias-fix, Live Mode,
the 7 new detector rules, the tighter skill. Before/after JPGs optimized
to ~470KB total (down from ~2.5MB PNGs).

/faq: the six support questions, each deep-linkable.

Live Mode is now Beta everywhere it surfaces: the /live-mode eyebrow
badge and note, the homepage bento tile badge, and the changelog entry.
The historical v3.0 changelog entry stays "Alpha" — accurate to what
shipped then.

Footer trimmed to the four links not already in the top nav (Changelog,
FAQ, Privacy, GitHub).

Version bumped 3.2.0 → 3.5.0 across the three plugin manifests; the
3.2 bias-fix work folds into this release rather than shipping separately.

astro.config.mjs: disable the dev toolbar.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* site: point /design-system hero at the m-01-v2-01 finalist

design-system.css referenced kintsugi-hero-v2.png, an untracked orphan
that was never committed. Repoint it at the committed m-01-v2-01 finalist
so /design-system and the homepage hero share one image, and the page
no longer depends on a file outside the repo. The v2 orphan moved to tmp/.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* build: sync harness mirrors + green the prose gate

Rebuild propagates the committed skill source (palette.mjs Setup step,
detector rule updates, brand.md) into the 13 harness output dirs and the
plugin subtree, which had drifted from source.

Also fixes the prose validator, which had been red on six pre-existing
hits across committed files:
- Four em dashes in code comments (Testimonials.astro, LiveDemoGbar.astro,
  index.astro) and one in skill/reference/live.md — reworded to colons/commas.
- Two in the slop catalog (an em-dash-overuse specimen and the
  marketing-buzzword rule naming "empower"). Those are intentional: the
  slop page documents every antipattern by example, so it must contain
  them. Exempted site/pages/slop from validateProse rather than neutering
  the specimens.

`bun run build` is now green end to end: counts validate, prose passes,
site builds.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* skill: rewrite no-section-fade rule to fix Gemini zero-motion overcorrection

The old rule ("whole-section fade-on-scroll is the saturated AI motion
reflex") drove Gemini to overcorrect into shipping pages with no motion
at all: motion-variety 39% / zero-motion 12% with the skill on, vs
~74-78% variety and ~3% zero-motion without it.

Rewrite keeps the legitimate-stagger carve-out, names the defect at
shape level (one identical entrance on every section) without
enumerating motion primitives, and adds an explicit clause that
suppressing the reflex is never grounds for a static page.

Validated on Gemini 3.5-flash (n=10, luxury-hotel + infra-platform):
motion-variety 39% -> 70%, zero-motion 12% -> 0%, staggered-reveal
stays 0% (reflex not re-inflated).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* release: bump CLI to 2.2.0 and extension to 1.1.0

Both ship the expanded detector: the 7 new rules (cream-palette,
em-dash-overuse, marketing-buzzword, numbered-section-markers,
aphoristic-cadence, broken-image, italic-serif-display) plus
hero-eyebrow-chip, with everything-centered removed. 41 rules total.

The extension settings page already supports toggling them: the rule
list renders from detector/antipatterns.json, grouped by category, and
disabledRules flows through chrome.storage.sync into the scan config,
which detect.js honors by rule id. New rules are toggleable with no UI
change.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* release: fix release.mjs for the moved changelog + add CLI/ext entries

The changelog moved from site/pages/index.astro to its own
site/pages/changelog.astro with new markup (cf-version / cf-entry /
cf-items), which left release.mjs reading the wrong file with the old
selectors. All three release commands would have failed at note
extraction. Point it at changelog.astro, match cf-version, and scope
notes to the <ul class="cf-items"> bullet list — that also skips the
lead paragraph, before/after figure, and stat row on the v3.5.0 entry,
keeping release notes to clean bullets.

Add CLI v2.2.0 and Extension v1.1.0 changelog entries (the shared
detector update: 7 new rules, everything-centered removed, 41 total;
plus the extension's per-rule toggles) so release:cli and release:ext
have notes to extract.

Verified extraction for all three labels: v3.5.0 (5 bullets),
CLI v2.2.0 (3), Extension v1.1.0 (2).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix: correct dev server port to 4321 and drop stale pnpm-lock

Astro serves on 4321, not 3000 as the docs claimed; update CLAUDE.md,
AGENTS.md, and screenshot-antipatterns.js. Remove the leftover
pnpm-lock.yaml from the Astro migration so Cloudflare's frozen install
uses the maintained, in-sync bun.lock instead of a drifted pnpm lockfile.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* site: rework /designing flow, rhythm, Live Mode mock, and CTA

Restructure the page so iteration reads as the core value, not net-new.
The four loop phases are wrapped in a track with a sticky scroll-spy nav
(Start/Iterate/Polish/Maintain) that pins under the header and highlights
the active phase; the surfaces section (skill/CLI/extension) moves out of
the loop into the post-loop context group so the loop runs uninterrupted.

Fix the iterate split: shared subgrid row tracks so the terminal and the
Live Mode mock align on the same baseline regardless of paragraph length,
wider intro measure (52ch, was a crammed 36ch), and a deeper picker stage
so the context and global bars breathe instead of stacking on the card.

Rebuild the Live Mode mock to mirror the real picker: carved-tile mark plus
Pick / Insert / Detect / DESIGN.md controls on lacquer-deep with the gold
border, and a /impeccable live entry line so the reader knows how to start.

Reframe Start as the hard mode, move h3 subheads off the thin display face
onto Albert Sans, and trim Start so it no longer dominates the loop.

Rework the closing CTA into two standalone raised cards (the bento plinth
made them read as boxes nested in a box), and fix the tutorials copy: there
are three walkthroughs now, and the brand-vs-product tutorial is gone, so
drop it from the CTA and remove the dead lane link to it.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* site: reorder Get Started so usage follows setup, link out to more

Move the /impeccable usage examples below the Chrome extension, CLI, and
Stay-updated block. Running a command is the logical next step once the
skill, extension, CLI, and subscriptions are all in place, so the section
now reads install -> set up the extras -> use it. Add a closing "Go deeper"
line linking to the Designing with Impeccable workflow page and the docs.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix: install compiled per-provider skill variants, not uncompiled source

`npx skills add` (and `impeccable skills install`, which wrapped it) installed
the uncompiled skill/ source verbatim: the skills CLI dedupes discovery by name
and picks skill/SKILL.md first, so installs shipped unresolved {{placeholders}}
and no vendored detector (#168).

- Rename skill/SKILL.md -> skill/SKILL.src.md so the skills CLI's discovery
  skips the source and falls through to a compiled .agents variant; update the
  build reader, skill-behavior harness, and docs to match.
- Refactor `impeccable skills install` to copy each harness's compiled variant
  from the universal bundle (real dirs, no npx skills, no symlink), with
  project/global harness detection and a --providers override.
- Fix stale unit tests (replacePlaceholders, readPatterns, transformer
  prefix/summary) that asserted removed pre-v3.0 behavior, and wire the three
  orphaned test files into `bun run test` so the drift can't recur.
- Split skills-cli.test.js: pure blocks run by default, network blocks move
  behind a new `bun run test:cli-e2e`; fix its stale update assertions.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* site: default to `npx impeccable skills install`, restore install-method panel

Get Started recommended `npx skills add`, which installs a single shared build
across harnesses. Make our CLI the default (it installs the build compiled for
each harness) and bring back the "Other install methods" disclosure the
neo-kinpaku redesign dropped.

- Homepage: primary command is now `npx impeccable skills install`; a native
  <details> panel offers the Claude Code plugin and `npx skills` (caveated as
  installing one shared build rather than the per-harness one).
- FAQ: recommend `npx impeccable skills install` to install, `--force` to
  reinstall, and note the npx skills shared-build caveat.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* site: reword craft tagline so it doesn't lead with "Shape"

The craft card's tagline began with the word "Shape", which reads like
the name of the sibling /shape command and made the two cards look
swapped (#166). Reword to "Design it, then build it, all in one flow."
No data was actually swapped; this is a copy collision fix.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* skill: rename teach -> init and expand its setup flow

Rename the `/impeccable teach` command to `/impeccable init` across the
skill, site, CLI, and tests. `teach` stays as a deprecated router alias and
/docs/teach + /skills/teach redirect to /docs/init.

Expand the command beyond writing PRODUCT.md/DESIGN.md: the same codebase
crawl now also pre-configures `.impeccable/live/config.json` (Step 6, with
CSP consent) so live mode boots with no first-time detour, and the flow ends
by recommending the best commands to run next from what the scan surfaced
(Step 7).

Fold two items into the unreleased v3.5.0 changelog entry: the init rename
and the brand-seed palette picker. No version bump.

Regenerates all harness skill output dirs and the _redirects file.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: lead README install + usage with the CLI installer

Add `npx impeccable skills install` as the recommended install option and
update the Usage section to the `/impeccable <command>` form, dropping the
nonexistent `/normalize` example.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(skill-behavior): swap to production-tier models (sonnet + gpt-5.5)

Replace the cheap-tier default lineup (claude-haiku-4-5, gpt-5.4-mini) with
production-tier models (claude-sonnet-4-6, gpt-5.5) so the skill-behavior
suite reflects what users actually run. gemini stays on flash-lite.

Sync the docs (CLAUDE.md, AGENTS.md, tests/skill-behavior/README.md): new
model names, cost estimate raised to ~$0.50-1.50/sweep, and the old 21-22/24
baseline reframed as previous-cheap-tier history pending re-measurement on
the new lineup.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat: self-updating skill via boot-time version check

context.mjs now polls a new lightweight /api/version endpoint at most once
per day (cached globally in ~/.impeccable) and appends an UPDATE_AVAILABLE
directive when a newer skill version has shipped, prompting the agent to
offer `npx impeccable skills update`. Best-effort and silent on any failure;
asks before updating; suppresses re-prompts for a declined version for a
week. Opt out with IMPECCABLE_NO_UPDATE_CHECK=1.

- skill/scripts/context.mjs: version read, throttle + anti-nag cache, directive
- scripts/build.js + _redirects: /api/version endpoint (from plugin.json version)
- skill/SKILL.src.md: document the UPDATE_AVAILABLE boot branch
- tests/context.test.mjs: coverage for cached/newer/suppressed/opt-out paths
- changelog: v3.5.0 entry
- synced harness skill dirs via bun run build

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test: cover the self-update path (network + LLM behavior)

context.test.mjs: add a localhost stub-server integration test for the live
fetch path (poll /api/version, cache a newer version, stay silent on
same-or-older, fail silent + stamp lastCheck when unreachable). Runs against
127.0.0.1 only, never the real site; uses async spawn so the in-process stub
isn't deadlocked by spawnSync blocking the event loop.

skill-behavior: add scenario 9 asserting the agent surfaces UPDATE_AVAILABLE
but never auto-runs `npx impeccable skills update` without asking. New
prepareWorkspace `skillVersion` copy-mode (so context.mjs has a SKILL.md to
version-check), env threading through runTurn -> execBash, and bash-output
capture to prove the agent actually received the directive. Passed on
claude-sonnet-4-6, gpt-5.5, and gemini-3.1-flash-lite.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-28 14:22:22 -07:00

27 KiB
Raw Blame History

Project Instructions for Claude

Skill bootstrap: skip the post-update-cleanup step

The <post-update-cleanup> block at the top of .claude/skills/impeccable/SKILL.md runs scripts/cleanup-deprecated.mjs and then self-deletes. Do not run it in this repo, and do not delete the block. This repo is the source of truth for the Impeccable skill; the cleanup script scans for "leftover" skill files to remove and its judgment doesn't fit a repo where the skill lives in skill/. The block is also regenerated from skill/SKILL.src.md on every bun run build, so any deletion churns back on the next build. Skip the step silently and proceed.

Same rule for AGENTS.md and every other harness-specific instruction file: treat post-update-cleanup as a no-op in this repo.

Architecture (v3.0+)

There is one user-invocable skill, impeccable, with 23 commands underneath it. Users type /impeccable polish, /impeccable audit, etc. The skill is defined in skill/:

  • SKILL.md — frontmatter (with the auto-trigger-optimized description and the allowed-tools list), shared design laws, and the Commands router table.
  • reference/ — one <command>.md per command (audit.md, polish.md, critique.md, etc.) plus the domain reference files (typography.md, color-and-contrast.md, etc.). When a sub-command is matched, the router loads its reference file.
  • reference/brand.md and reference/product.md — the two register references. SKILL.md's Setup section selects one based on the task cue, the surface in focus, or the register field in PRODUCT.md (first match wins).
  • scripts/command-metadata.json — single source of truth for each command's description, argument hint, and (eventually) category. Both the build and pin.mjs read from this.
  • scripts/pin.mjs — creates/removes lightweight redirect shims so users can have /audit as a standalone shortcut that delegates to /impeccable audit.
  • scripts/cleanup-deprecated.mjs — runs once after an update to remove leftover files from renamed/merged commands.

Do not add standalone skills unless there's a strong reason. The consolidation was deliberate: the / menu pollution problem is real and gets worse as users install more plugins.

Register (brand vs product)

Every design task belongs to one of two registers:

  • Brand — design IS the product: marketing, landing pages, brand sites, campaign surfaces, portfolios, long-form content. Distinctiveness is the bar. Spans every visual lane (tech-minimal, luxury, editorial-magazine, consumer-warm, brutalist, etc.) — do not default to only one.
  • Product — design SERVES the product: app UI, admin, dashboards, tools. Earned familiarity is the bar — fluent users of Linear / Figma / Notion / Raycast / Stripe should trust it.

PRODUCT.md at the project root carries a ## Register section with a bare value (brand or product). /impeccable teach asks about register first because it shapes every downstream answer.

Sub-command reference files add a short ## Register section near the top only where the answer diverges between the two. Don't restate the register files' content in sub-commands — link instead. Sub-commands where register meaningfully diverges today: typeset, animate, bolder, delight, colorize, layout, quieter.

a11y lives in audit.md, not in SKILL.md, brand.md, or product.md. Models over-cautious themselves into safe, underdesigned output when reminded about accessibility at design time. The audit command is the dedicated place for that check.

CSS

Plain hand-written CSS, no Tailwind. Imported into Astro pages/layouts via frontmatter import statements; Vite resolves @import chains automatically.

The CSS architecture (under site/styles/):

  • main.css — Main entry point, imports the partials and defines tokens/reset
  • workflow.css — Commands section, glass terminal, magazine spread styles
  • sub-pages.css/docs, /anti-patterns, /tutorials, detail pages
  • tokens.css — OKLCH color tokens (ink, charcoal, ash, mist, cream, accent)
  • footer.css — shared across all pages, imported in Base.astro

Edit any of these directly and the dev server hot-reloads. No rebuild needed for CSS changes.

Color token rule

  • --color-ink (10% lightness) is for body copy. Use it even for small text.
  • --color-charcoal (25% lightness) reads as washed-out gray in small text. Only use for headings or larger body copy at ≥16px.
  • --color-ash (55%) is for secondary labels, captions, relationship meta lines.
  • Never use pure black or pure white. Use the tinted tokens.

Prose: read STYLE.md before writing user-facing copy

Editorial brief is at STYLE.md (root). Read it before editing the homepage, sub-pages, command editorials, tutorials, or READMEs. The site has been called out for AI prose; the rules there exist to keep that from creeping back.

The build's validateProse step (in scripts/build.js) enforces a denylist: em dashes ( and HTML entities), the -- em-dash substitute, load-bearing, highest-leverage, biggest unlock, seamless, robust, delve, elevate, empower, underscore, pivotal, tapestry, data-driven, reflex defaults, collapses into monoculture, in today's, gone are the days, whether you're, let's dive in, in summary, in conclusion, moreover, furthermore. Each rule prints a rationale and a suggested replacement when it fires. Do not silently work around the regex. If a banned word has earned a real meaning here, raise it as a STYLE.md amendment.

The validator scans site/pages/, site/content/, site/components/, site/layouts/, README.md, README.npm.md. It deliberately skips skill/ because LLM-facing reference instructions sometimes need technical phrasings the marketing copy can't.

The deeper structural issues (negation pivot, triadic auto-pilot, uniform paragraph rhythm, hollow confidence) require human judgment. STYLE.md lists them. Use them on every editorial pass.

Editorial content lives under site/content/

Skill editorials and tutorials are read by scripts/build.js (for taglines and downstream tooling) and by Astro's content collection (for what actually renders on the site). One tree, one place to edit:

  • site/content/skills/<id>.md — optional editorial wrapper with frontmatter tagline plus body sections
  • site/content/tutorials/<slug>.md — full tutorial content
  • site/data/anti-patterns-catalog.js — detection-rule catalog (visual examples, gallery items, layer definitions)

Development Server

bun run dev        # Bun dev server at http://localhost:4321
bun run preview    # Build + Cloudflare Pages local preview

The dev server runs Astro (astro dev). Editing files in site/content/skills/, skill/, or scripts/lib/sub-pages-data.js requires a server restart (not just a browser reload) to see the change. CSS, components, and pages hot-reload fine without a restart.

Legacy URL redirects are emitted to _redirects by scripts/build.js (via generateCFConfig); the dynamic /skills/:id → /docs/:id redirect lives in site/public/_redirects (Cloudflare Pages reads both at deploy). Current redirects: /skills/docs, /skills/:id/docs/:id, /cheatsheet/docs, /gallery/visual-mode#try-it-live.

Deployment

Hosted on Cloudflare Pages. Static assets served from build/, API routes handled via _redirects rewrites (JSON) and Pages Functions (downloads).

bun run deploy     # Build + deploy to Cloudflare Pages

Build System

The build system compiles the impeccable skill from skill/ to provider-specific formats in dist/:

bun run build      # Build all providers
bun run rebuild    # Clean and rebuild

Source files use placeholders that get replaced per-provider:

  • {{model}} — Model name (Claude, Gemini, GPT, etc.)
  • {{config_file}} — Config file name (CLAUDE.md, .cursorrules, etc.)
  • {{ask_instruction}} — How to ask user questions
  • {{command_prefix}}/ or $ depending on provider
  • {{available_commands}} — auto-populated list of commands (from IMPECCABLE_SUB_COMMANDS in scripts/lib/utils.js)
  • {{scripts_path}} — provider-aware path to the skill's scripts directory

Harness output directories are tracked

.claude/skills/, .cursor/skills/, .agents/skills/, and the other 8 harness directories are intentionally committed to the repo. npx skills reads them directly from this repo at install time, and they enable clean submodule use. Do not gitignore them. Run bun run build to refresh them after editing skill/.

Local state files inside harness directories (e.g. .claude/scheduled_tasks.lock, .claude/settings.local.json) ARE gitignored.

Generated sub-pages are gitignored

site/public/docs/, site/public/anti-patterns/, site/public/tutorials/, site/public/visual-mode/, site/public/slop/ are gitignored as legacy generator output paths. Astro's content collections drive the live site under site/pages/docs/, site/pages/tutorials/, etc.; nothing reads from those gitignored dirs anymore.

Testing

bun run test                  # Default suite: unit + static framework fixtures
bun run test:live-e2e         # Opt-in: full-cycle live-mode E2E across framework fixtures
bun run test:skill-behavior   # Opt-in: LLM-backed checks that the skill text actually drives the agent's setup flow

Unit tests (build orchestration, detector logic) run via bun test. Fixture tests (jsdom-based HTML detection) run via node --test because bun is too slow with jsdom. The test script handles this split automatically.

Important: tests/build.test.js uses spyOn(transformers, 'transformCursor') with the named exports from scripts/lib/transformers/index.js. Those named exports (transformCursor, transformClaudeCode, etc.) are kept specifically for test spying, even though build.js itself uses createTransformer + PROVIDERS directly. Do not delete them as "dead code" — I made that mistake once and broke 8 tests.

Live-mode E2E

tests/live-e2e.test.mjs drives the entire user flow (handshake → pick → Go → cycle → accept → carbonize cleanup) against every fixture in tests/framework-fixtures/ that declares a runtime block. Each fixture installs real deps, boots its framework dev server (Vite, Next, SvelteKit, Astro, Nuxt static), and runs Playwright Chromium against a deterministic fake agent that produces realistic variants in the exact format reference/live.md describes.

bun run test:live-e2e                                       # full suite, ~2 min, 19 fixtures
IMPECCABLE_E2E_ONLY=vite8-react-modal bun run test:live-e2e # scope to one fixture
IMPECCABLE_E2E_DEBUG=1 bun run test:live-e2e                # dump page DOM + dev-server tail on failure

One-time setup: npx playwright install chromium (the suite uses a specific Chromium build keyed to the bundled Playwright version).

Kept out of the default bun run test because (a) it does real npm install per fixture, (b) it boots framework dev servers, (c) wall time is ~2 minutes, and (d) it requires Playwright's browser cache. Run it locally before shipping changes to anything in skill/scripts/live-*.{mjs,js}.

The agent is pluggable via a one-method interface in tests/live-e2e/agent.mjs: generateVariants(event, context) → { scopedCss, variants[] }. The default fake agent emits canned variants that exercise all three param kinds (range, steps, toggle). The orchestrator (wrap, write, accept, carbonize) is agent-agnostic.

LLM agent (opt-in): set IMPECCABLE_E2E_AGENT=llm to swap the fake agent for tests/live-e2e/agents/llm-agent.mjs, which calls Claude (default Haiku 4.5) via @anthropic-ai/sdk. Requires ANTHROPIC_API_KEY in env; the test runner skips with a clear message when it's unset. Override the model with IMPECCABLE_E2E_LLM_MODEL=claude-sonnet-4-6 if Haiku produces unreliable JSON. Caching is on — live.md is the cacheable prefix, and after the first call subsequent fixtures pay only the cache-read rate. Pass rate on a typical sweep is 18/19; the modal fixture's intrinsic state-loss flake is amplified by LLM latency and may need a re-run. This path hits the API and costs money — keep it out of CI unless you really want it there.

Adding a new fixture is a matter of cloning a directory under tests/framework-fixtures/, swapping the source files, and writing a fixture.json. See tests/framework-fixtures/README.md for the full schema.

Skill-behavior tests

tests/skill-behavior/scenarios.test.mjs is the LLM-backed safety net for edits to skill/SKILL.src.md and the Setup-adjacent reference files (teach.md, document.md, brand.md, product.md, sub-command refs). It inlines the source skill/SKILL.src.md into the system prompt of a real LLM, gives the agent bash / read / write / list tools scoped to a temp workspace, and asserts on the tool-call trace — not on the model's free-form output. The trace is the source of truth.

bun run test:skill-behavior                                              # full suite (27 tests, ~5 min, ~$0.50-1.50 across providers)
IMPECCABLE_SKILL_BEHAVIOR_MODELS=gemini-3.1-flash-lite bun run test:skill-behavior   # scope to one provider
IMPECCABLE_SKILL_BEHAVIOR_VERBOSE=1 bun run test:skill-behavior          # dump per-scenario trace JSON to stderr (use when iterating)

Three providers per run, every run. The suite always exercises claude-sonnet-4-6, gpt-5.5, and gemini-3.1-flash-lite. Sonnet and GPT-5.5 are production-tier, matching what users actually run, so the pass/fail signal reflects real agent behavior rather than a cheap proxy; gemini stays on the flash-lite tier. Don't substitute Claude alone: many of the most useful findings come from divergence between providers.

Auth lives in repo-root .env (copied from ~/code/impeccable-evals/.env, gitignored). Providers skip cleanly when their key is unset; they don't fail.

Nine scenarios:

  1. empty workspace → agent loads reference/teach.md
  2. PRODUCT.md only → loads brand.md
  3. PRODUCT.md + DESIGN.md → loads brand.md + consults the design system
  4. context already loaded in turn 1 → turn 2 does not re-run context.mjs
  5. PRODUCT.md without ## Register field → agent infers brand from task cue
  6. /impeccable polish → loads reference/polish.md
  7. /impeccable audit → loads reference/audit.md
  8. existing SvelteKit project → agent reads at least one project code file
  9. context.mjs emits UPDATE_AVAILABLE (seeded newer version) → agent surfaces it but does not auto-run npx impeccable skills update

Baseline. The 21-22 / 24 baseline (with stable gpt scenario 6/7 failures) was measured on the old cheap tier (claude-haiku-4-5 / gpt-5.4-mini). It needs re-measuring on the current claude-sonnet-4-6 / gpt-5.5 lineup; the production-tier models are expected to do better on the sub-command routing scenarios the old gpt tier failed. See tests/skill-behavior/README.md.

Cost. Each run is real LLM calls, billed to the keys in .env. Production-tier models put a full sweep around $0.50-1.50. Keep it out of CI unless you really want it there.

Adding a scenario. Write the fixture in tests/skill-behavior/fixtures.mjs, add the it() block in scenarios.test.mjs (the harness uses the source skill/ dir via a symlink, so no rebuild needed), and update the baseline table in the suite's README. The harness's fileLoaded(trace, filename) helper checks both read and bash cat — different models prefer different tools.

The harness symlinks source, not built output. This is deliberate so SKILL.md / reference / scripts/context.mjs edits show up immediately without bun run build:skills. The trade-off: reference files surface their raw {{placeholders}}, but the assertions key on tool calls rather than content, so it doesn't matter for correctness.

CLI

The CLI lives in this repo under cli/: cli/bin/ (entry + sub-commands), cli/engine/ (the detect-antipatterns rule engine + browser variant), cli/lib/ (helpers shared by CLI and Cloudflare Pages Functions). Published to npm as impeccable.

npx impeccable detect [file-or-dir-or-url...]   # detect anti-patterns
npx impeccable detect --fast --json src/         # regex-only, JSON output
npx impeccable live                              # start browser overlay server
npx impeccable skills install                    # install skills
npx impeccable --help                            # show help

The browser detector (cli/engine/detect-antipatterns-browser.js) is generated from the main engine. After changing cli/engine/detect-antipatterns.mjs, rebuild it:

bun run build:browser

IMPORTANT: Always use node (not bun) to run the detect CLI. Bun's jsdom implementation is extremely slow and will cause scans with HTML files to hang for minutes.

Versioning

There are three independently versioned components. Only bump the one(s) that actually changed:

CLI (npm package):

  • package.jsonversion
  • Bump when: CLI code changes (cli/bin/, cli/engine/detect-antipatterns.mjs, etc.)

Skills (Claude Code plugin / skill definitions):

  • .claude-plugin/plugin.jsonversion
  • .claude-plugin/marketplace.jsonplugins[0].version
  • Bump when: skill content changes (skill/, reference files, command metadata, etc.)

Chrome extension:

  • extension/manifest.jsonversion
  • Bump when: extension code changes (extension/)

Website changelog (site/pages/index.astro):

  • Hero version link text + new changelog entry in the changelog section
  • Update for user-facing changes only, not internal build/tooling details
  • Use the most prominent version that changed (skills version is usually the right one)

After bumping, see Releases below for how to tag and publish.

Releases

GitHub releases are tagged per-component, not per-version, since the three components ship independently. Tag prefixes: skill-v, cli-v, ext-v.

Workflow for any component:

  1. Bump the manifest version (see Versioning above).
  2. Add a changelog entry to site/pages/index.astro. Skill entries use a bare vX.Y.Z label; CLI and extension entries use the prefixed forms CLI vX.Y.Z and Extension vX.Y.Z. The release script extracts notes by matching this label, so the prefix matters.
  3. Commit and push to main.
  4. Run bun run release:<skill|cli|ext>. Preview first with node scripts/release.mjs <component> --dry-run.

The script refuses to run if: the working tree is dirty, HEAD is ahead of origin, the tag already exists, the matching changelog entry is missing, or (for skill/extension) bun run build / bun run build:extension produces uncommitted changes — meaning the harness output dirs or extension/detector/ files weren't refreshed before the bump was committed.

Skill releases attach dist/universal.zip. Extension releases run bun run build:extension first and attach dist/extension.zip. CLI releases print a reminder to run npm publish separately; extension releases print a reminder to upload the zip to the Chrome Web Store dashboard.

If you need to fix release notes after the fact (typo, missing thank-you, formatting bug): gh release edit <tag> --notes-file <md>. The release script's htmlToMarkdown function is the cleanest source for regenerating notes from the changelog.

Adding New Commands

All commands live under /impeccable. To add a new one:

  1. Create skill/reference/<command>.md with the command's instructions (this is what the LLM loads when the command is invoked)
  2. Add a row to the Sub-command reference table in skill/SKILL.src.md
  3. Add an entry to the Command menu section in the same file
  4. Add the command name to IMPECCABLE_SUB_COMMANDS in scripts/lib/utils.js
  5. Add it to VALID_COMMANDS in skill/scripts/pin.mjs
  6. Add its metadata (description + argumentHint) to skill/scripts/command-metadata.json
  7. Add its category to SKILL_CATEGORIES in scripts/lib/sub-pages-data.js
  8. Add its relationships (leadsTo / pairs / combinesWith) to COMMAND_RELATIONSHIPS in the same file
  9. Add the same category entry to site/scripts/data.js commandCategories and commandProcessSteps (for the homepage carousel)
  10. Add symbol + number to commandSymbols and commandNumbers in site/scripts/components/framework-viz.js (periodic table)
  11. Optional: write an editorial wrapper at site/content/skills/<command>.md with a short tagline and expanded body (When to use it / How it works / Try it / Pitfalls)

The build system counts commands from the router table automatically. Update the command count in all of these locations when the total changes:

  • site/pages/index.astro — meta descriptions, hero box, section lead
  • /cheatsheet redirects to /docs (no standalone page)
  • README.md — intro, command count, commands table
  • NOTICE.md — command count
  • AGENTS.md — intro command count
  • .claude-plugin/plugin.json — description
  • .claude-plugin/marketplace.json — metadata description + plugin description

The build validator (generateCounts in scripts/build.js) checks these files for stale numeric counts and fails the build if any disagree with the router table.

Adding editorial content for existing commands

Editorial files live at site/content/skills/<command>.md and have a tagline frontmatter plus a body with the standard four sections:

  • When to use it — the specific scenarios this command owns
  • How it works — the internal process, phases, or approach
  • Try it — one or two concrete examples with expected output
  • Pitfalls — real failure modes, with alternatives to reach for instead

The tagline is used by UI surfaces (magazine spread, docs cards) that need a short human-friendly label. The long description in command-metadata.json stays optimized for auto-trigger keyword matching in the AI harness.

Every command should have an editorial file eventually, but the build does not require one: commands without editorials fall back to the frontmatter description.

Adding or modifying anti-pattern detection rules

cli/engine/detect-antipatterns.mjs is the source of truth for the rule engine. It powers the CLI, the public-site overlay, the Chrome extension, and the homepage rule count. Five places stay in sync:

Where How it stays in sync
cli/engine/detect-antipatterns.mjs (ANTIPATTERNS array + checkXxx logic) Hand-edited
cli/engine/detect-antipatterns-browser.js bun run build:browser
extension/detector/detect.js + extension/detector/antipatterns.json bun run build:extension
site/public/js/generated/counts.js (DETECTION_COUNT) bun run build
skill/SKILL.src.md and reference/*.md Hand-edited if the rule introduces new design guidance

Always run all three builds and the test suite after a rule change:

bun run build && bun run build:browser && bun run build:extension && bun run test

TDD order (non-negotiable)

  1. Fixture at tests/fixtures/antipatterns/{rule-id}.html with two columns (should-flag / should-pass), each case identified by a unique heading. Cover ≥4 flag cases and ≥5 false-positive shapes. Use explicit pixel dimensions in CSS because jsdom does no layout.
  2. Failing test in tests/detect-antipatterns-fixtures.test.mjs using the snippet-substring pattern (regex /"([^"]+)"/ against SHOULD_FLAG / SHOULD_PASS lists). Run it and watch it fail before implementing.
  3. Rule entry in the ANTIPATTERNS array: id, category (slop for AI tells, quality for real design or a11y issues), name, description, optional skillSection and skillGuideline.
  4. Pure check function checkXxx(opts) returning [{ id, snippet }]. No DOM access in the pure function.
  5. Two adapters: checkElementXxxDOM(el) for the browser (getComputedStyle + getBoundingClientRect) and checkElementXxx(el, tag, window) for jsdom (parseFloat(style.width) instead of layout). Wire both into both element loops in cli/engine/detect-antipatterns.mjs — the browser loop (~line 1837) and the jsdom loop in detectHtml (~line 2058). Forgetting one is the most common mistake; symptom is "test passes, live page silent" or vice versa.
  6. Verify on a live page: http://localhost:4321/fixtures/antipatterns/{rule-id}.html and the homepage (no false positives). The two adapter paths can disagree, so manual browser checks catch what the fixture test can't.

Conventions and jsdom gotchas

  • Snippet format: wrap the identifying heading text in straight double quotes (e.g. 'icon tile above h3 "Lightning Fast"') so the fixture test can extract it. For rules not anchored to a heading, pick another stable identifier.
  • jsdom doesn't lay out: getBoundingClientRect() returns 0×0. Read parseFloat(style.width) and parseFloat(style.height) from explicit CSS instead.
  • background: shorthand isn't decomposed in jsdom: use the existing resolveBackground() and resolveGradientStops() helpers (~line 631 / 670).
  • Computed colors aren't normalized in jsdom: parseGradientColors() handles both hex and rgb forms.

Reference rules to copy from: side-tab (border, ~line 312), low-contrast (color + gradient, ~line 339), icon-tile-stack (sibling relationship, ~line 425), flat-type-hierarchy (page-level, ~line 1080).

Evals Framework (separate private repo)

The eval framework lives in a separate private repo at ~/code/impeccable-evals/. It measures whether the /impeccable skill improves or harms AI-generated frontend design by running the same brief through a model with and without the skill loaded.

If you're picking up eval work, switch to that repo and read its AGENT.md first. It captures model choices, sample size policy, lessons learned, common workflows, and gotchas.

cd ~/code/impeccable-evals
bun run serve            # dashboard on http://localhost:8723

The eval runners read this repo's skill from ../impeccable/skill/ and staged provider skills from ../impeccable/build/_data/dist/*. Run bun run build in this repo before an eval sweep if you want the Claude/Gemini staged skills to reflect your latest edits.

After structural skill changes, update inline-skill.ts in the evals repo

The harness inlines SKILL.md into the system prompt for "skill-on", stripping sections irrelevant to an API-driven craft run. The stripped list in runner/inline-skill.ts needs to stay in sync with SKILL.md's top-level ## headings. As of v3.0, it should strip ## Setup (non-optional) (was ## Context Gathering Protocol), ## Commands (was ## Command Router), and ## Pin / Unpin. Keep ## Shared design laws. If you add or rename a top-level section, update the strip list there.