mirror of
https://github.com/pbakaus/impeccable.git
synced 2026-09-12 06:06:37 +03:00
0ee3e80aec32f12e731ee83073ada461f4da1f1b
38
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
79573ce55b |
refinement scope: keep content + media footprint; recompose for emphasis (codex+gemini consult)
Replaces the a14/a15 attempts (both deleted). Diagnosis: incentive stacking; the placeholder-completion MUST plus the image tool turned 'bolder' into full-bleed photo insertion. Scope preservation is the missing rule, not imagery policy. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
6f3076051f |
persuade: scope the imagery MUST to new surfaces; existing systems decide their own vocabulary
x02 a14 rerun: 3/3 samples still imported photos — the unscoped MUST in the Persuade mode block overrode the existing-worlds principle. Scoping keeps the greenfield ablation win, frees iteration asks. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
5674a94114 |
existing-worlds: boldness from committed materials; new medium = redesign, not refinement
x02-tidewater-bolder eval: 3/3 skill-on samples imported photography into a photo-free seed system (0% arena vs competitor, which amplified the seed's own vocabulary instead). One sentence, shape-level, no examples. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
073e17e180 |
three-directions sketch + the scene decides the theme
Paul's a11 review: heroes are safe SaaS viewports, everything predictable; mobile Operate ships dark despite a brief that specifies outdoors-in-motion use. Decide-then-build now opens with three one-line directions differing in concept (the instinctive pick that any studio would reach for is the default wearing your name); the Operate mode adds: the usage scene is part of the spec, the theme follows the scene, not the category's habit. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
cfbac54440 |
deprecate craft: the build flow lives in new-work.md, checkpoints are a mode
Per Paul: rather than gating a second file, fold what made the craft path superior into the file both models already read 21/21 through the gate. new-work.md gains 'Decide, then build' (direction as one confirmable paragraph; attended pauses, unattended records-and-goes; codex.md mock flow when image generation exists) and 'Finish like a studio' (inspect, honest critique, patch, detector). craft becomes a deprecated alias like teach: invoking it forces attended checkpoints, nothing else differs; the reference is a redirect stub. codex.md retargeted. Existing-world feature builds remain governed by the core floor (unmeasured path, noted in the plan doc). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
139d69f2b7 |
bare build requests follow the craft orchestration; brief-coverage joins the floor
Invocation A/B on Fable (a9 craft-path vs a9-direct plain): the plain path scored 38% vs the competitor against the craft path's 50%, and brief fidelity collapsed to 14% vs bare — the direct path drops asked- for features that craft's direction step and engineering bar preserve. Routing now sends any build request through the craft orchestration unprompted (its gates pause only when a user can respond), and the craft floor gains a brief-coverage recheck: every requirement the brief names must exist on the page. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
eeff485c20 |
mode belongs to the surface; palette sources are exclusive
a7 transcript evidence: 01-observability samples drew orange-honey and green seeds, recited the color-strategy menu, and shipped dark category-reflex palettes anyway; the model applied the subject's workmanlike grammar to its own landing page. Two generic lines: the mode belongs to the surface, not the subject (a landing page for a dense tool is still Persuade; deciding a page can be plain because its subject is workmanlike is the category error in reverse), and the palette has exactly two legitimate sources (seed or the subject's world; the category's habitual palette is neither). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
a55b162a24 |
routing owns the craft-vs-direct decision; craft.md stops advising its own loading
The when-to-choose guidance sat inside the file that only loads after the choice is made. SKILL.md's routing now says it: bare build requests build directly through the gate and floor; craft is routed only when named or when the user asks for a guided, checkpointed build. The Commands row describes craft by its checkpoints. craft.md's intro just describes the supervised flow it orchestrates. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
b409bedf5d |
drop brand.md/product.md stubs; craft repositioned as the collaborative build
Stubs removed per Paul (register: values remain harmless family hints; nothing points at the files anymore). craft.md now opens by defining itself against plain invocation: a bare build request goes straight through the gate and the craft floor; craft is the supervised path with guaranteed checkpoints and the mock pipeline. One shipping-discipline line joins the core floor (real content, interaction states, respect the build pipeline) so one-shots inherit the bar that previously lived only in craft's Step 4. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
051f856113 |
finish the mode migration: brand.md/product.md become redirect stubs
Answering the obvious question the family-depth framing dodged: with modes derived per task, files named for the old two-register taxonomy had no architectural reason to exist. brand.md's surviving depth (lane test + inverse test, reflex-reject lanes, color discipline, layout moves, permissions) folds into new-work.md, where all of it belonged: it is new-identity Persuade/Experience guidance. product.md's content moves unchanged to operate.md, its true name. Both old files remain as one-line redirect stubs because register: brand|product in existing PRODUCT.md files and older links point there. All cross-references retargeted (SKILL.md modes intro, context.mjs REGISTER hint, live.md, typeset.md); 85 tests green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
0fde0850cf |
skill v4.0.0-alpha.9: daily-driver core + mandatory new-work playbook
Architecture per Paul: impeccable is primarily a daily driver on existing codebases; the always-loaded core should serve that 90% path, not carry the full generative arsenal on every invocation. SKILL.md now holds brief-wins, existing-worlds (the headline path), the four visitor modes, the full craft floor, and a hard gate: new identity work (greenfield, or a redesign discarding the current look) MUST read reference/new-work.md before any design decision. That file carries the generative playbook (seed, subject grounding, plan/self-check/signature, hero-thesis, everything-bold, prove-don't-claim, color commitment, calibration, persuade type/imagery). context.mjs enforces the gate mechanically: NEW_WORK directive when no PRODUCT.md/DESIGN.md exists, and the old mandatory register-file read is replaced by a REGISTER family hint. No surfaces: map anywhere; mode is derived per task. Gate compliance is measurable via skillEvidence.directSkillFileReads. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
bf2dd7ec13 |
skill v4.0.0-alpha.8: four visitor modes replace the brand/product bifurcation
Field report: impeccable SaaS-ified a developer docs page; the Opus galleries showed the same on an album page. Root cause: two registers force every surface into persuade-or-operate grammar. The register section now names the visitor's mode first (Persuade / Operate / Read / Experience) with mode-borrowing called out as the canonical failure, and PRODUCT.md's register field maps as family (brand = Persuade + Experience, product = Operate + Read) for compatibility. Read mode: comprehension deliverable, navigable structure, chrome out of the way. Experience mode: the artifact leads at every screen size. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
2e71facb59 |
skill v4.0.0-alpha.7: cultural surfaces are the work, not a funnel
Paul's Opus gallery observation: every impeccable 05-experimental-album generation reads decidedly SaaS while frontend-design's open with the art itself, especially at narrow viewports. Cause: the brand register prescribed stop-the-scroll/earn-the-click/convert for ALL brand surfaces. Split the register's deliverable by surface: product/service pages convert; cultural surfaces (album, portfolio, publication, body of work) lead with the artifact, recede the interface, and treat conversion grammar as a category error — the visitor meets the work in the first viewport at every screen size. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
02f760fbad |
skill v4.0.0-alpha.6: boldness is page-level commitment, not an element budget
Paul: everything should be bold, nothing bland; bold is neither decoration nor clutter but commitment to the concept, whose form the concept chooses (maximal or severely clean, drenched or monochrome, piercing copy, the product demonstrating itself). Replaces the 'spend your boldness in one place' rule imported from frontend-design, whose one-bold-element-on-a-quiet-page framing pulled pages toward the tasteful softness the galleries showed losing. The signature becomes where the concept peaks rather than the only place it lives. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
56edce955a |
skill v4.0.0-alpha.5: hero-as-thesis + commit-over-refined (distinctiveness push)
Paul's spot-check of the Fable validation galleries: frontend-design's lektor generations read vastly more distinctive and subject-faithful despite losing the overall pairwise verdict on craft. The arena agrees on the axis (distinctiveness 8-31 at n=5). Two additions to the core: the opening viewport is a thesis (open with the most characteristic thing in the subject's world, with a concrete memory test), and an explicit polish-is-the-floor counterweight so the craft floor stops reading as a mandate for quiet. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
50002a9e05 |
skill: rewrite codex block as positive calibration (self-priming fix)
gpt-5.6-sol evals: skill-on lost craft 0-25 to bare gpt-5.6; removing the enumerated codex ban block recovered it to 4-16, confirming the block's literal CSS patterns self-prime the defects they ban (the same mechanism the v2.1 ablation sweep documented). Replaced with three shape-level calibration lines: tracking floor (kept, it's a numeric ceiling), elevation-declared-once + modest container radius, and material honesty (real assets, surfaces not decoration, specific claims). Detector rules continue to enforce the mechanical patterns. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
6b3d174e93 |
skill v4.0.0-alpha.4: the lean core — full design guidance at a quarter the length
Pairwise evals on Fable one-shot (6-task regression set, opus-4-8 judge, position-bias-cancelled): the hand-distilled ~55-line lean core beat the heavy v4 core 66% overall / 67% craft head-to-head, and moved the decisive win-rate vs frontend-design from 13% to 27% (40% with the completion-time QA scan; craft went positive 6-5 for the first time). 18/18 lean samples ran context.mjs + palette.mjs vs a minority under the heavy core: shorter instructions get followed. Context weight itself was suppressing both compliance and boldness. Structure: persona + brief-wins + existing-worlds + subject-grounding + plan/self-check + boldness + prove-don't-claim + commit + calibration + compressed craft floor + two-paragraph registers. Commands table kept; the no-arg context-aware menu logic moved to reference/routing.md (read on demand in the only case that is inherently interactive). Provider blocks and rule anchors preserved. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
1172898020 |
skill v4a3: prove-don't-claim + load-bearing signature
Judge rationales across cand-v4a2 arenas: competitor wins by showing the product working (mix panels, comparison tables, live demos) and by signatures big enough to organize the page; our samples claim, decorate, and sometimes stop at the hero. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
f1078c59b0 |
skill v4a2: seed defers to subject's world; unattended-mode gates for craft/shape; init skip when no user
Eval evidence (cand-v4a1-prose): palette.mjs handed a random violet seed to the Polish-TV lektor brief and the model anchored on it, overriding subject-grounding; craft/shape user gates can't fire in one-shot runs and each model improvises around them. Seed is now a reflex-check that yields to a subject-dictated palette; craft/shape gain an explicit unattended mode (same bar, no waiting); init interview is skipped when no user can respond. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
440348a498 |
skill v4 core: existing-world/new-work gate, register-scoped type rules
Per Paul's guidance: (1) existing committed design systems are the bread-and-butter case and get a first-class core rule (work inside the world, no parallel colors/fonts/styles, no perf regressions); (2) a redesign that discards the current look is new identity work and runs the full concept/tokens/signature process instead of anchoring to the incumbent skeleton (the lektor failure); (3) the reflex-reject font list and physical-object font procedure are brand-register rules, moved out of the universal Commit section — system stacks and workhorse UI faces are legitimate, often correct, for product UI, stated positively in the product register. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
a38a0765a9 |
skill v4.0.0-alpha.1: always-loaded core — brief-wins, subject grounding, token/self-check process, inline craft floor + registers
One-shot evals on Fable 5 (impeccable-evals notes/fable-oneshot-craft-plan.md) showed the reference-file architecture failing: models skip the register reads, so most design guidance never reaches them, and skill-on collapses toward bare-model output (0/9 pairwise wins vs frontend-design on r10). SKILL.md is now self-contained for one-shot work: persona, the-brief-wins rule, ground-it-in-the-subject, a plan/tokens/signature/self-check process gate, commitment guidance, a compact inline craft floor, and distilled brand/product registers. Reference files remain as sub-command flows and optional depth. The enumerated absolute-bans list is retired from prose; mechanical slop enforcement moves to the detector/hook. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
c11cc7b58c |
Route native projects to native command variants (audit, adapt) (#357)
* Route native projects to native command variants for audit and adapt Follow-up to #269. The web audit.md and adapt.md carried "translate this yourself" Platform notes, so a native invocation paid for the full web file (~1.8k / ~2.6k tokens, mostly inapplicable) and did error-prone run-time translation. Authored with AI assistance (Claude Code) under maintainer direction. - New reference/audit.native.md and reference/adapt.native.md: authored native content (VoiceOver/TalkBack, platform conformance, adaptivity dimensions; phone-to-tablet, platform-to-platform, web-to-native strategies). One variant per command covers ios, android, and adaptive; per-OS specifics stay in the platform refs Setup loads regardless. - SKILL.src.md: Commands table lists the variants; Setup step 2 reads the variant instead of the web file when the platform is native. - audit.md / adapt.md: Platform sections replaced with a one-line web-only guard pointing at the variant. - animate.md / layout.md: Platform sections deleted; the Motion and Layout sections of the already-loaded platform refs carry that content. Web users now pay zero tokens for the platform axis in these files. - Skill-behavior scenario 15 pins the route-instead behavior (passes live on claude-sonnet-4-6); CLAUDE.md documents the variant convention. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Phrase command-reference routing as one rule, not rule-plus-exception Copilot review catch: step 2 said "MUST read reference/<command>.md" and then carved out the native variant, which invites loading both files. Now a single rule: read the web reference or the table's native variant, one file, not both. Scenario 15 re-verified live. Applied with AI assistance (Claude Code) under maintainer direction. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Anchor native runs in animate/layout, drop loaded-refs assumption Review-thread fixes, applied with AI assistance (Claude Code) under maintainer direction: - Greptile: deleting the animate/layout Platform sections left native runs alone with web tooling instructions (CSS keyframes, GSAP, Grid, clamp()). Restore a one-line anchor in each pointing at the loaded platform reference's Motion / Layout section (~20 tokens, not the old restatements). - Bugbot: audit.native.md and adapt.native.md asserted the platform refs were "already loaded in Setup", but the command reference loads at step 2, before step 5. Now they instruct: read the platform reference first if Setup hasn't already. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Carry the native-variant rule into routing rules 2 and 3 Bugbot catch: Setup step 2 routed native projects to the variant, but routing rules 2 and 3 (the operative text at command time) still said to load the generic reference file. Both now reference the same one-file variant rule. Applied with AI assistance (Claude Code) under maintainer direction. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Point animate/layout native anchors at the files, not "loaded" refs Bugbot catch, same class as the variant wording fix: the anchor lines said "the loaded platform reference" but command files load at step 2, before the platform refs at step 5. Both anchors now name the files and instruct reading them first if Setup hasn't already. Applied with AI assistance (Claude Code) under maintainer direction. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
3e38e595c7 |
Add platform axis (web / ios / android / adaptive) (#269)
* Add a platform axis (web / ios / android / adaptive) to the skill Orthogonal to register: register decides whether design IS or SERVES the product; platform decides the delivery target and which native conventions apply. Set `## Platform` in PRODUCT.md; a missing field defaults to `web`, so legacy projects are unaffected. - extractPlatform() in skill/scripts/context.mjs (mirrors extractRegister); the CLI appends a NEXT STEP directive to read the native reference(s). `adaptive` (Flutter / RN / KMP shipping both iOS and Android) loads both ios.md and android.md. - New reference/ios.md (Apple HIG distilled) and reference/android.md (Material 3 distilled); reference/web.md is a thin pointer. The native refs frame register's role as narrow: platform conformance is the bar, brand lives in the expressive layer the platform gives you, never by breaking the rails. - Setup step 5 loads the native reference(s) when platform is native. Live mode and the detect CLI stay web-only, gated off ios/android/adaptive. - init asks platform right after register; adapt/audit/animate/layout carry short platform divergence notes; all secondary spots thread `adaptive`. - a11y stays in audit.md (loading it at design time makes output timid), so the native refs carry no Accessibility section; audit.md's Platform section owns native a11y. - Tests: extractPlatform unit coverage + skill-behavior scenario 10 (PRODUCT.md platform ios -> agent loads ios.md). Source-first: only skill/, scripts/, tests/, CLAUDE.md, NOTICE.md, the changelog and version are committed; the sync workflow regenerates the provider trees and ./plugin on merge. ios.md / android.md are distilled from the MIT-licensed ehmo/platform-design-skills; attribution in NOTICE.md. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Address review: gate web tools on native platforms, drop version churn Maintainer-review fixes applied with AI assistance (Claude Code), on top of the rebased platform-axis commit: - Design hook (post-edit and Cursor pre-edit) now resolves the project platform via loadContext + extractPlatform and skips its web rule scan for ios / android / adaptive projects, so React Native / Flutter code never draws web-shaped findings (new hook-lib resolveProjectPlatform / isNativePlatform helpers, covered by unit and subprocess tests). - context.mjs CLI warns on an unrecognized ## Platform value (e.g. a toolchain name like `flutter`) instead of silently defaulting to web; extractRegister / extractPlatform now share extractSectionValue. - Removed reference/web.md: nothing loaded it; CLAUDE.md carries the "web has no extra rulebook" explanation. - init.md: skip live-mode config (Step 6) for native platforms; note the per-app PRODUCT.md pattern for repos shipping web + native. - android.md: Material-everywhere apps that also ship on iPhone still owe iOS OS guarantees (safe areas, Reduce Motion, edge-swipe back). - ios.md: reworded a design-time line that framed Dynamic Type as an accessibility check (a11y stays owned by audit.md). - Renumbered the new skill-behavior scenario to 14 after main's 10-13; updated CLAUDE.md scenario list; added android + unrecognized-value CLI test cases. - No version or changelog changes: versioning happens at release time. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Tighten platform reference prose Editorial pass on the platform-axis text, applied with AI assistance (Claude Code) under maintainer direction: - ios.md / android.md rewritten to house style: single-line paragraphs (no hard wraps), one-sentence scope intro, deduplicated intro/slop-test, register-compression down to two sentences. In-file attribution paragraphs removed (NOTICE.md owns attribution); "read on top of the register reference" cruft removed (SKILL step 5 and the context.mjs directive already say it). Bans sections dropped: they restated the rules above them; the two additive items (tab-bar overload, hover-dependent affordances) folded into rules. ~40% smaller each. - Sub-command Platform sections (adapt, audit, animate, layout), SKILL step 5, init.md platform prose, and the context.mjs directive trimmed the same way. Build (prose validators, counts) and both test runners green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Treat an empty PRODUCT.md section as absent, not the next heading Copilot review catch: extractSectionValue read the next `## ...` heading as the section value when a field was left empty, which made the CLI warn "value `## Product Purpose` is not recognized". Stop at the next heading and return null instead. Regression tests for extractPlatform, extractRegister, and the CLI warning path. Applied with AI assistance (Claude Code) under maintainer direction. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Only read a token list of both native targets as adaptive Bugbot catch: after the exact platform tokens failed, any Platform line containing the words ios and android was classified adaptive, so negated or explanatory prose ("web only, not ios or android") silently loaded both native refs and skipped the hook, with no warning. The combo parse now accepts only list separators and the two platform words; anything else falls through to the CLI's unrecognized-value WARNING. Regression tests added. Applied with AI assistance (Claude Code) under maintainer direction. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Paul Bakaus <paul.bakaus@gmail.com> |
||
|
|
1a46353b29 |
Don't force init on scoped commands when PRODUCT.md is missing (#277)
* Don't force init on scoped commands when PRODUCT.md is missing Setup step 1 told the agent: "If it reports NO_PRODUCT_MD, stop and follow reference/init.md before doing anything else." For a project with no PRODUCT.md, that turned every scoped request (polish, critique, audit, layout, ...) into a full from-scratch init detour. The user asks to polish one button and the skill instead starts writing PRODUCT.md from the beginning. Faced with that gate, agents also frequently abandon the command and do an ad-hoc pass without loading the command reference. Make the gate command-aware. A missing PRODUCT.md still routes into init for the from-scratch build flows where captured product context is the point (init, craft, shape). For any other command, a scoped request against existing code, the code is the context: proceed with the requested command, infer the register from the surface in focus, and offer /impeccable init once as a suggestion rather than a blocker. - skill/SKILL.src.md: rewrite the step 1 NO_PRODUCT_MD rule; reconcile the no-argument routing rule so it leads the menu with init instead of silently jumping into it; extend the craft init-then-resume footnote to cover shape, now also a from-scratch flow. - skill/scripts/context.mjs: soften the NO_PRODUCT_MD message to defer to the step 1 rule instead of "Stop the current task"; refresh the stale file-level JSDoc that still described the old empty-stdout signal. - tests/skill-behavior/scenarios.test.mjs: add scenario 10 (scoped command + no PRODUCT.md proceeds without forcing init) and scenario 11 (shape + no PRODUCT.md still diverts into init). Scenario 1 (craft diverts) stays green and pins the build path. Source-only per repo convention; provider and plugin copies are regenerated by the maintainer's build:skills sync. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Fix missing-context routing for build intent --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Paul Bakaus <paul.bakaus@gmail.com> |
||
|
|
7b2c2a1f23 | Fix Impeccable setup path guidance (#341) | ||
|
|
4ac0348032 | Add Codex grid background slop rule | ||
|
|
b7d2ad5589 |
Fix: allow skill's bundled node helpers under strict-permission harnesses (#301) (#310)
The skill declared only `Bash(npx impeccable *)` in allowed-tools, but Setup and the no-arg menu shell out to `node {{scripts_path}}/*.mjs`. Under a default-deny Claude Code allowlist those calls are blocked, so Setup fails on context.mjs.
Add a provider-aware `Bash(node {{scripts_path}}/*)` entry and resolve {{scripts_path}} in the frontmatter (the build previously substituted it only in the body). Provider-aware rather than the hardcoded `.claude/...` path the issue suggested, since five providers honor allowed-tools with different script dirs.
|
||
|
|
0306b41949 |
Add monorepo context support (#213)
Context files (PRODUCT.md / DESIGN.md) resolve child-first then fall back to the repo root, and /impeccable live lets the user pick a child app in a monorepo. Single-app behavior is unchanged. Closes #202. Co-Authored-By: abdulwahabone |
||
|
|
672517f76e |
Add automatic design hook install and exceptions (#170)
* docs: add PRD for design detector hook integration Plans a PostToolUse hook for Claude Code and Codex that runs the existing design detector after every relevant file write and feeds findings back to the agent as advisory system-reminder context. No implementation in this commit; covers UX, technical design, build pipeline changes, distribution, coverage tradeoffs, and rollout. Co-authored-by: Cursor <cursoragent@cursor.com> * docs: revise hook PRD with best-practices review Folds in the P0/P1/P2 findings from an online best-practices critique against the official Claude Code and Codex hook references plus 10+ 2026 community guides and similar prior-art tools (claw-hooks, claude-code-hooks-mastery). Key changes: - Exec form everywhere (Codex snippet was shell form), with Windows rationale. - Default timeout dropped from 10s to 5s. - Re-entrancy guard (CLAUDE_HOOK_DEPTH) and per-file edit counter. - Session-scoped finding dedup promoted from open question to v1. - Per-language inline-ignore syntax map (HTML/JSX/CSS/JS). - Hard-skip rules for sensitive paths and generated/lock files. - Honest framing about Claude Code lacking per-plugin hook disable. - Honest framing about Bash-written files being invisible in v1. - Codex Windows-not-supported call-out, feature flag note, trust ceremony detail. - Optional NDJSON audit log via IMPECCABLE_HOOK_LOG. - Findings cap lowered 8 → 5 with attention-budget rationale. - Versioned envelope ([impeccable@1]) on rendered template. - Expanded test plan, decision log, and stdin payload appendix. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(hooks): ship the design detector hook for Claude Code and Codex Implements docs/hooks-prd.md: a PostToolUse hook that runs the impeccable design detector after every Edit/Write/MultiEdit on a UI file and pushes findings into the agent's next-turn context as a short system reminder. Silent on clean files. Never blocks an edit. Why this matters: today, design slop (side-tab borders, gradient text, purple/cyan palettes, bounce easing, etc.) only gets caught when a human notices or someone explicitly runs /impeccable audit. The hook closes the loop at the moment slop is written. What ships in v1 - skill/scripts/hook.mjs: PostToolUse entry. Reads stdin, runs the detector in-process (no `npx impeccable` cold start), emits hookSpecificOutput.additionalContext when fresh findings exist. - skill/scripts/hook-lib.mjs: extracted helpers (config, cache, filter, render, audit log, runHook orchestrator). 100% unit-testable. - skill/scripts/hook-session-start.mjs: SessionStart greeting, gated by a project-scannable probe + 30-day throttle. - skill/scripts/hook-admin.mjs: backs /impeccable hooks on/off/status/ignore-rule/ignore-file/reset. Hardening built in - Re-entrancy guard (IMPECCABLE_HOOK_DEPTH) so the hook can never recursively spawn itself. - Hard-skip regexes for sensitive paths (.env, .pem, id_rsa, secrets, credentials, .git) and generated/lock/build output. These fire before the file is even read; cannot be turned off via config. - Path-traversal check on the inbound file_path. - Session-scoped dedup keyed by (session, file, rule, line) so the same finding never lands in context twice. Prevents the ~12.5K wasted tokens per chatty session called out in the PRD. - Per-(session, file) edit counter with a one-shot suppression notice on the 7th edit, silent after. - Fail-open contract: every error path returns exit 0 with no stdout. Optional NDJSON audit log via IMPECCABLE_HOOK_LOG. Three kill switches (precedence high to low): 1. IMPECCABLE_HOOK_DISABLED env var (1/true/yes/on, case-insensitive) 2. .impeccable/hook.json `enabled: false` 3. /impeccable hooks off slash command (writes the JSON) Inline ignores are language-aware. `// impeccable: ignore <rule>` for JS/TS, `<!-- impeccable: ignore <rule> -->` for HTML/Vue/Svelte/Astro, `{/* impeccable: ignore <rule> */}` for JSX/TSX, `/* impeccable: ignore <rule> */` for CSS. `*` matches any rule. Directive applies to the next non-blank line. Same shape as ESLint, Stylelint, Biome. Build pipeline - scripts/lib/transformers/hooks.js: per-provider hooks.json builders, plus the slim .codex-plugin/plugin.json manifest. - providers.js: emitHooks: 'claude' for claude-code, emitHooks: 'codex' for codex and agents. Codex also emits emitCodexPlugin. - factory.js: emits hooks/hooks.json next to the skills tree. - build.js: syncs hooks/ into harness roots and into the slim plugin/ subtree; writes .codex-plugin/plugin.json. Build is idempotent (verified: 98 staged files unchanged across two runs). Claude Code wiring uses exec form (command + args) and the ${CLAUDE_PLUGIN_ROOT} placeholder. Matcher: Edit|Write|MultiEdit. `if:` glob filters to UI extensions before spawning Node. PostToolUse timeout 5s, SessionStart timeout 3s. Codex wiring uses ${PLUGIN_ROOT} (Codex's native placeholder), matcher Edit|Write|apply_patch, no `if:` analog (the script does the extension filter). macOS and Linux only; hooks are disabled on Windows in current Codex builds. The trust ceremony and feature flag are documented in README.md. Routing - /impeccable hooks lives outside the 23-command router table on purpose: it is plumbing, not a design skill. The hidden routing slot is added to SKILL.md alongside pin/unpin so the LLM knows to dispatch it. The 23-command count and all stale-count validators remain happy. Tests - tests/hook.test.mjs: 38 unit tests covering env parsing, config load + defaults + malformed, cache round-trip + GC, ignoreRules/minSeverity/inline ignores (all four languages), globbing with **/*/{a,b}, render template with cap + clamp + 0-line prefix drop, audit log NDJSON, payload event-name parameterization, re-entrancy, kill switches, sensitive-path + generated-path + traversal skips, allowlist filter, config ignoreFiles, edit counter cycle including the 7th-edit notice, MultiEdit and apply_patch payload shapes, detector throw swallow, malformed stdin, missing file race. - tests/hook-build.test.mjs: 18 integration tests covering hook manifest shape (matcher, timeouts, exec form, if: glob, placeholders), Codex differences (${PLUGIN_ROOT}, no if:, no SessionStart), Codex plugin manifest (no inline hooks field to avoid the duplicate-file error), routing across the hooksJsonFor table, and presence of all three committed artifacts plus the bundled detector the runtime relative-import path depends on. Full suite: 175 bun tests + 186 node tests, all green. Docs - README.md: new "Design hook" section explaining default behavior, per-project / global / inline disable paths, the JSON schema knobs, the audit log debug flag, and the slop / a11y coverage split. - HARNESSES.md: flips the `hooks` row for Codex from No -> Yes (Claude was already Yes), adds a per-harness hook-surface table with the manifest location and matcher each provider uses. Open questions from the PRD intentionally deferred to v2: Bash-write blind spot, effort-aware suppression, Stop-hook session summary, per-rule severity, async hook mode. None block v1. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix Codex hook scanning: apply_patch paths and co-located stylesheets Parse file targets from Codex apply_patch command bodies, co-scan imported and sibling CSS when UI components are edited, drop the git-sweep PostToolUse group, and align Codex SessionStart manifest and trust docs with the official hooks spec. Co-authored-by: Cursor <cursoragent@cursor.com> * Gitignore hook session cache and drop local test HTML Hook dedup/throttle state in .impeccable/hook.cache.json is per-project runtime data like other .impeccable/ sidecars. Remove an untracked bad-nested-flexbox scratch page from site/public/. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix Claude Code hook: drop Edit-only if filter so Write/MultiEdit fire Claude's if permission rule binds to one tool name, so Edit(*.{…}) never spawned the hook on Write or MultiEdit despite the matcher listing them. Extension filtering now lives in hook-lib on both Claude and Codex. Co-authored-by: Cursor <cursoragent@cursor.com> * Surface Cursor design findings via stop-hook followup Replace dropped postToolUse additional_context with afterFileEdit recording and a one-shot stop followup_message so anti-pattern nudges reach the agent. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix design hook packaging and scans * Fix Cursor hook pending bucket fallback * Fix Sass hook scan coverage * Fix Cursor hook review findings * Fix session start dead hook normalization * Fix hook config and relative scan paths * Remove SessionStart design hook * Remove redundant afterFileEdit normalization * Fix Cursor suppression and module style scans * Fix sensitive path hook filter * Fix disabled Cursor stop hook emission * Refresh hook harness artifacts * Fix Cursor hook manifest install * Add hook ignore-value support * Ignore hook runtime files locally * Fix Codex plugin hook packaging * fix: address PR review bot findings Block numeric hook depth counters from re-entering. Avoid following stylesheet imports from traversal-looking hook targets. * fix: gate ignore-value suggestions by supported rules Only render exact ignore-value commands when the same finding can be suppressed by ignoreValues. * Package Codex plugin as hook-only * Remove Codex plugin packaging * Recover hook install probe plumbing * Remove Codex hook packaging follow-up doc * Remove extra hook docs and skill wording changes * Install real design hooks via skills CLI * Add provider hook smoke runner * Fix Cursor hook delivery with preToolUse gate * Simplify Cursor hook install to preToolUse * Clarify confirmed hook exceptions * Persist hook ignores in shared config * Guard font hook exceptions * Fix hook install after main rebase * Fix hook scan target handling * fix: address hook review findings * Address hook review feedback * Stabilize DeepSeek insert live fixture * Fix Cursor hook Python shell write bypass --------- Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
5fbe37c97c |
skill: remove Copy section from main design skill
Copy guidance (em-dash bans, buzzword bans, button-label / link-text phrasing, aphoristic-cadence) doesn't belong in the main design skill. It's not design-specific — the skill is trying to do too much. The six rules being dropped (every-word-earns, no-em-dashes, no-aphoristic-cadence, no-buzzwords, button-verb-object, link-standalone) are now better served by: - The impeccable engine's antipattern detectors (em-dash-overuse, marketing-buzzword, aphoristic-cadence, copy-slop) for linting at scan time. - The /clarify subcommand for surfacing the same checks when reviewing copy specifically. The em-dash ban for the SKILL prose itself still lives in STYLE.md and the build-time prose validator — that's separate from the skill's guidance to agents. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
d37be057ea |
skill: drop 4 redundant typography rules + add EMPIRICAL_VALIDATION
The v2.1 ablation sweep (n=10 × 4 brand niches × 3 providers, anchored to commit 54c3a502, ~544 cells) confirmed these four rules carry no weight in the skill: - skill-typo-no-all-caps-body — duplicate of brand-ban-all-caps-body; brand version is more specific (reserves caps for labels + headings) - skill-typo-codex-hero-ceiling-repeat — the codex-block restatement of skill-typo-hero-ceiling didn't add reinforcement on top of the universal rule - skill-typo-scale-ratio — duplicate of brand-typo-modular-scale; same signal, brand version carries the clamp() / fluid implementation detail - skill-typo-font-count — models don't reach for ≥4 font families in any niche we test, so the rule has no measurable effect Each deletion is the Agent A / B / C / D Phase-2 audit recommendation; none of the four ever validated under either prose state. Adds EMPIRICAL_VALIDATION.md naming the seven cross-provider winners as the trustworthy core, and documents the systemic findings (self-priming, detector saturation, vocabulary anchoring) so future skill edits can avoid the same traps. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
d0c934c03b |
chore(skill): rebuild harness SKILL.md outputs from source
Mirrors the 5 prose changes in skill/SKILL.src.md + skill/reference/brand.md out to every harness directory (`.claude`, `.gemini`, `.cursor`, `.codex`, `.agents`, etc.) so the staged skill that workers / agents read matches the source. Auto-generated by `bun run build:skills`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
b210dd71e7 |
skill: strip self-priming examples from 5 rules
Phase-2 ablation audit caught these rules causing the exact behavior they
ban via the literal examples in their own prose. Verified: OpenAI samples
under skill-on produced "fake theater", "vendor theater", "heatmap theater"
as verbatim copies of the 'X theater' example. Same pattern for the
restrained-on-cream example, the aphoristic-cadence template, and the
"reserve uppercase for…" enumeration.
- skill-ban-codex-x-theater: drop the 3 syntactic templates + 3 example
phrases ("Productivity theater" etc.)
- brand-imagery-required: drop the niche enumeration that cued
"imagery not required elsewhere"
- skill-typo-no-all-caps-body: drop the "Reserve uppercase for labels /
eyebrows / badges" enumeration that primed uppercase usage
- brand-color-no-converge: drop the "restrained-on-cream" example that
was priming cream-heavy palettes
- skill-copy-no-aphoristic-cadence: drop the literal cadence template
("serious statement, then punchy short negation") that named the
rhythm it bans
Ablation re-run pending in impeccable-evals to measure impact.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
96af55aa80 | add text-wrap: balance to the skill, which seems to be quite effective in ablation runs | ||
|
|
0047981a95 |
fix(skill): target local files for detect, never a URL (#159)
Rework context-signals' detect target after review: a URL meant a costly Puppeteer render (and a probed port might not even be this project), and the index.html-or-bail fallback failed most real apps (no root index.html). New priority: (1) the scannable markup/style files in the dirty git tree (what the user is working on, small and local); (2) a local source dir (src / app / components / pages / public — the detector walks these and skips node_modules / dist / build); (3) a root index.html, else the project root as a last resort when there's code. Emits `scan.targets` (a list) + `scan.via`. Never a URL. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
772aa73aa3 |
feat(skill): make bare /impeccable context-aware (re: #159)
Reshape of the "/impeccable suggest" proposal in #159. Instead of adding a 24th command (menu pollution + the command-add tax + its own discoverability problem), upgrade the path users already hit: bare `/impeccable` with no argument. - New skill/scripts/context-signals.mjs gathers cheap, deterministic signals (setup gaps, register, latest cached critique score, git change scope, a dev-server port probe, and a `scan.detectTarget` for the detector) and emits JSON. It does NOT score or rank, and it does NOT run the detector itself (the engine isn't importable in an installed skill, and shelling npx+jsdom would risk a hang) — the agent reasons over the raw signals. - SKILL.md routing rule 1 now leads with the 2-3 highest-value next commands, each with a reason from the signals, then the full menu. Never auto-runs; always confirms. Reuses init's "Recommend starting points" vocabulary. When a project has never been critiqued it offers critique; when scan.detectTarget is set it runs `npx impeccable detect --fast --json` and folds the hits in. - Export extractRegister from context.mjs for reuse. Stays 23 commands; no metadata/pin/site-data changes. Unit-tested, including a regression guard for porcelain leading-space path parsing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
5793e84292 |
feat(skill): make bare /impeccable context-aware (re: #159)
Reshape of the "/impeccable suggest" proposal in #159. Instead of adding a 24th command (menu pollution + the command-add tax + its own discoverability problem), upgrade the path users already hit: bare `/impeccable` with no argument. - New skill/scripts/context-signals.mjs gathers cheap, deterministic signals (setup gaps, register, latest cached critique score, git change scope, a dev-server port probe) and emits JSON. It does NOT score or rank — no brittle weights table — the agent reasons over the raw signals. - SKILL.md routing rule 1 now leads with the 2-3 highest-value next commands, each with a reason from the signals, then the full menu. Never auto-runs; always confirms. Reuses init's "Recommend starting points" vocabulary. - Export extractRegister from context.mjs for reuse. Stays 23 commands; no metadata/pin/site-data changes. Unit-tested, including a regression guard for porcelain leading-space path parsing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
9ffd3211d5 |
Neo Kinpaku design system + Live Mode v3 (#169)
* Add neo kinpaku design system page * skill: rip out baked-in category recipes and saturated-default motion tropes Programmatic bias mining (impeccable-evals) traced four major defects back to specific lines in this skill that contradicted SKILL.md's own first-order-reflex warning: - brand.md "Pairing and voice" prescribed four category→aesthetic recipes (editorial → serif+sans, tech/dev/fintech → tight tracking, consumer/food/travel → script/display serif, creative → rule-break). These directly drove OpenAI's 76% extreme-negative letter-spacing on tech briefs and Anthropic/Google's 28-34% italic-serif-display slop on editorial/food briefs. Replaced with one sentence: the shape depends on the brand, not on the brand's category. - brand.md "Brand permissions" had "Typographic risk. Enormous display type, unexpected italic cuts, mixed cases, hand-drawn headlines, a single oversize word as a hero." — a four-for-one slop driver behind 97% OpenAI comically-large H1, 42% bad-SVG illustration, and the editorial-italic slop. Deleted outright. - typeset.md and teach.md repeated the same category recipes; trimmed to the principle without the recipe. - SKILL.md Typography: added a hard hero-H1 ceiling (clamp() max ≤ 6rem ≈ 96px), with a <codex> block to make it explicit since OpenAI over-indexes here (97% ≥128px vs 24% for Anthropic). - animate.md, bolder.md, brand.md: removed "staggered reveals" and "scroll-triggered transitions" as the prescribed default ambitious motion. By 2026 that's the saturated AI tell, not a choreography. Reserved stagger for legitimate list-sibling rhythm. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * skill: anti-cream + codex-specific defect bans + universal slop bans Second pass after measuring more biases against the eval corpus. - SKILL.md Color: explicit "cream/sand/beige body bg is the saturated AI default of 2026" rule. Tone down the "tint every neutral" line so it doesn't read as "default to warm-tinted near-white" (which OpenAI hits at 74% and Anthropic at 31%-47%). - SKILL.md Absolute bans: add universal bans for two slop patterns detected at 55-95% across providers — tiny uppercase tracked eyebrow above every section (the 2023-era kicker that's now AI grammar) and numbered section markers (01/02/03). Also explicit "text that overflows its container is the universal defect on tablet/mobile." - SKILL.md Absolute bans → <codex> block: ban the GPT-specific defects Paul annotated repeatedly — `border:1px solid` + soft-wide-shadow (≥16px blur) "ghost cards", `border-radius:32px+` over-rounding, hand-drawn/sketchy SVG illustrations (loose-sketch / *-sketch classes, feTurbulence paper-grain filters), repeating-linear-gradient stripes, "X theater" AI-slop copy phrases. - SKILL.md Motion → <gemini> block: the image :hover transform tell (38% Google skill-on rate). Hover effects on images add no info; the image isn't an action target. Animate card chrome, not the image. - SKILL.md Typography: hard display letter-spacing floor ≥-0.04em (OpenAI defaults to -0.075em → cramped). Existing hero ceiling <codex> block extended with the letter-spacing rule. - codex.md Step A example: stop seeding "warm-grounded (deep oxblood + cream)" as the warm-palette template, which primes the cream default. - colorize.md Tinted backgrounds: stop printing the literal cream recipe `oklch(97% 0.01 60)`; replace with brand-anchored guidance. - document.md examples: warm-ash-cream → cool-paper so the example doesn't seed cream as the canonical neutral example. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * skill: universal anti-slop bans + contrast/font-count/all-caps-body rules Third pass after measuring the rest of the cross-provider matrix: - Color: explicit "Verify contrast" rule. Low-contrast text fires at 68% across all providers skill-on (90+% off). The most common failure is muted gray body on a tinted near-white; light-gray-for- elegance is named as the single biggest cause of unreadable AI pages. - Typography: max-3-font-families rule. Overused-fonts (>4 families) fires at 28% Anthropic / 36% Google / 0% OpenAI skill-on; >50% off. Also: universal "no all-caps body copy" (moved from brand-only ban to Shared design laws since product-register also overuses caps). - Copy: anti-aphoristic-cadence ban targets Anthropic's signature "X. No Y." / "X. Just Y." voice (63% skill-on copy-slop rate, 77% off — the worst rate in the matrix). Once-is-voice / three-or-more- is-tell framing per the runner's copy-slop detector. - Copy: anti-SaaS-buzzword-string ban with the literal phrase list the detector watches for (streamline/empower/supercharge, trusted- by-leading, best-in-class/enterprise-grade/cutting-edge, etc). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * skill: strengthen anti-cream rule across full warm-neutral band Smoke validation showed the cream fix worked for Google + OpenAI but Anthropic Sonnet italian-restaurant still shipped `--paper: oklch(90% .018 88)` — cream just outside the L≥95% band the rule cited. Broaden the rule: - Band: OKLCH L 0.84-0.97, C < 0.06, hue 40-100 (was 95-97% / 60-95). - Name the token-name tells explicitly (paper / cream / sand / bone / flour / linen / parchment / wheat / biscuit / ivory) — the model defaults to one of these regardless of what hex it lands on. - Call out the specific brief patterns ("warm, traditional, family- coastal-Italian" / "editorial-restraint") that the model translates into cream by reflex. Then provide three explicit non-cream options: saturated brand color, true off-white at C=0, or darker mid-tone. Warmth in the brand is carried by accent + typography + imagery, not by body bg. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * v3.2.0: skill bias-fix release Bumps version from 3.1.1 to mark the four-commit skill cleanup that rips out baked-in category recipes (brand.md), saturated-default motion tropes (staggered reveals everywhere), the cream/sand body-bg AI tell, codex-specific defects (1px+wide-shadow, over-rounding, hand-drawn SVGs, stripes, X-theater copy), the extreme-letter-spacing default, and universal slop bans (all-caps eyebrow on every section, numbered-section markers, all-caps body, font-family-count > 3, aphoristic copy cadence, SaaS buzzword strings). Plus a hard hero-H1 ceiling (clamp() ≤6rem) and a Gemini-specific image:hover transform block. Validated against ~190 post-fix samples — see impeccable-evals biases tab for per-provider deltas. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * drop "no pure black/white" rule entirely The rule was contested in the design world and causing more damage than good — pushing every page into the tinted-near-white default which is the cream/sand AI tell we already explicitly ban elsewhere. Vercel, SVKMS, Brutalist sites, et al. use pure black/white successfully; the skill shouldn't second-guess that. Skill markdown deletions: - SKILL.md Color: drop the "Never use #000 or #fff" bullet. - color-and-contrast.md: drop the "Never Use Pure Gray or Pure Black" subsection, the "Never pure black" table-row prescription, and the "Avoid: Using pure black for large areas" bullet. - colorize.md: drop the "NEVER use pure black or pure white for large areas" bullet. - polish.md: drop the "Tinted neutrals: No pure gray or pure black" half of the bullet (the gray-on-color bullet survives). Detector code (cli/engine): - registry/antipatterns.mjs: remove the `pure-black-white` entry. - rules/checks.mjs: remove the three `findings.push({ id: 'pure-black-white', ... })` emit points (inline #000 bg, Tailwind bg-black class, plain-HTML scan path). - engines/regex/detect-text.mjs: remove the two pure-black-white regex rules (CSS `background: #000…` + Tailwind `bg-black`). - detect-antipatterns-browser.js: regenerated via scripts/build-browser-detector.js. Tests: - detect-antipatterns-fixtures.test.mjs: invert the assertion that pure-black-white fires; expect it to NOT fire post-v3.2. Drop the Tailwind bg-black-opacity edge-case test (no longer relevant). - detect-antipatterns.test.js: drop the standalone "detects pure- black-white in styled-components" test and remove pure-black-white from the multi-detector assertions in PricingCard, globals.css, and GlobalStyle.tsx tests. 166 bun tests pass; 24 node fixture tests pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * skill: strip example patterns from copy rules, strengthen gemini block v3.2 rerun validation surfaced two issues: 1. Copy-slop detector fires more on Gemini under v3.2 (48% → 84%) than under no-skill baseline. Root cause: the anti-aphoristic-cadence rule printed the literal "X. No Y." / "X. Just Y." patterns as examples, and Gemini imitated them as the recommended voice. Same recipe-becomes- bias trap we hit with brand.md:116's "Enormous display type, unexpected italic cuts, mixed cases, hand-drawn headlines" enumeration. Fix: describe the cadence as a rhythm ("serious statement, then punchy short negation") without printing literal patterns. Buzzword list trimmed to a single inline phrase family rather than quoted strings. 2. Gemini image:hover transform Gemini-tell hadn't dropped (31% off → 32% v3.2). Strengthen the <gemini> block: explicit "Never animate <img> elements on hover", call out the Tailwind group-hover:scale / group-hover:rotate / group-hover:translate parent-hover patterns by name (Gemini was reaching for these via Tailwind even though the prior text talked about :hover on the image directly). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * skill: simplify context loading and inline register directive Replaces load-context.mjs's JSON output with a tight markdown block from the renamed context.mjs. The script now extracts PRODUCT.md's `## Register` field and appends a `NEXT STEP:` directive naming the matching reference (brand.md / product.md), which moved Gemini from skipping the register load entirely to honoring it. Drops the `.impeccable.md` auto-migration; makes IMPECCABLE_CONTEXT_DIR a lazy escape hatch consulted only when the default paths come up empty. Setup is now four bullets in one list. The DESIGN.md nudge is gone; in its place, a "familiarize with the existing design system" step that calls out CSS / tokens / running app as authoritative sources alongside DESIGN.md. The standalone `### Register` H3 stays for the cascade rules (task cue → surface → register field). New LLM-backed test suite at tests/skill-behavior/ runs five scenarios against claude-haiku-4-5, gpt-5.4-mini, and gemini-3.1-flash-lite via Vercel AI SDK. Captures real tool traces, asserts on context.mjs calls, brand.md loads, and teach.md fallback. Skips cleanly when API keys are unset. 13-14/15 pass; only stable failure is the v3.2.0-era gpt-mini S4 "don't re-run" regression. Adds @ai-sdk/google as devDep and the test:skill-behavior npm script. Touches em-dashes in skill/SKILL.md and four reference files so `bun run build:skills` passes its skill-prose validator. teach.md and document.md drop their "re-run the loader to refresh session cache" steps since the agent's own write is now the freshest source. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * skill: merge orphan reference files into command sub-skills + inline S-tier invariants Two related restructurings: 1. SKILL.md now carries the cross-domain invariants that catch defects in any project (contrast/placeholder/gray-on-color, similar-font pairing, text-wrap, tabular-nums, centered-stack default, Flex/Grid choice, auto-fit grids, semantic z-index, reduced motion, stagger vs section-fade, premium motion materials, focus-visible, placeholders-aren't-labels, dropdown overflow trap, button/link copy). Greenfield-only rules (theme picking, color strategy, tinted neutrals) live under "New projects only". 2. Reference files merged into their command counterparts: - spatial-design.md -> layout.md - motion-design.md -> animate.md - color-and-contrast.md -> colorize.md - responsive-design.md -> adapt.md - ux-writing.md -> clarify.md - typography.md -> typeset.md (bolder.md redirected) - cognitive-load.md + heuristics-scoring.md + personas.md -> critique.md craft.md and shape.md "load references" lists updated to new file homes. interaction-design.md stays standalone (no 1:1 command verb). Net: 36 -> 27 reference files. Same content, fewer files, no orphaned reference loaded only from craft.md. Also extends the routing rules: if the user's first word doesn't match a command but the intent clearly maps to one, load that command's reference and proceed as if invoked. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * skill: add sub-command + existing-project scenarios; move sub-command load to step 2 Adds three new LLM-backed scenarios to tests/skill-behavior: - S6: `/impeccable polish` → loads polish.md - S7: `/impeccable audit` → loads audit.md - S8: existing SvelteKit project (PRODUCT.md + DESIGN.md + src/app.css + src/lib/components/*.svelte + src/routes/+page.svelte) → agent reads at least one project code file to understand the existing design system S6/S7 surface a real model-floor: gpt-5.4-mini reads brand.md, reads the target index.html, and just does the polish/audit without ever loading the sub-command reference. Stronger SKILL.md wording didn't move it. Captured in the README baseline as a known weakness. Claude and Gemini honor the load reliably. To fix Gemini on S6/S7, sub-command reference loading is now Setup step 2 (right after context.mjs), not step 4 — placing it before the model gets focused on "doing the work". Step 3 (design-system familiarization) is tightened to require at least one project code read even when a sub-command reference loads in step 2, so Claude doesn't laser-focus on the sub-command flow and skip the broader exploration. Two new fixtures: MINIMAL_LANDING_HTML (a tiny static landing page for S6/S7) and SVELTE_PROJECT_FILES (a minimal SvelteKit scaffold with tokens, components, and a routes/+page.svelte for S8). Both designed to look real enough that agents treat them as production code. Suite is now 24 tests across three providers; baseline is 21-22/24, with the stable failures being gpt-5.4-mini scenarios 6 and 7. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * skill: add reveal-animation safety rule (must enhance, not gate visibility) Class-triggered visibility transitions pause on hidden tabs and headless renderers. The italian-restaurant smoke produced a build where 2 sections shipped opacity:0 because the CSS transition never advanced past currentTime=0 (timeline paused). Added one-liner under Motion to prevent the antipattern: reveals must enhance an already-visible default, never gate content visibility on a class-triggered transition. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * skill: restore prescriptive cream/sand/beige paragraph Bisection across 5 historical skill commits on Gemini 3.5 flash fast lane n=3 found that |