* Add neo kinpaku design system page * skill: rip out baked-in category recipes and saturated-default motion tropes Programmatic bias mining (impeccable-evals) traced four major defects back to specific lines in this skill that contradicted SKILL.md's own first-order-reflex warning: - brand.md "Pairing and voice" prescribed four category→aesthetic recipes (editorial → serif+sans, tech/dev/fintech → tight tracking, consumer/food/travel → script/display serif, creative → rule-break). These directly drove OpenAI's 76% extreme-negative letter-spacing on tech briefs and Anthropic/Google's 28-34% italic-serif-display slop on editorial/food briefs. Replaced with one sentence: the shape depends on the brand, not on the brand's category. - brand.md "Brand permissions" had "Typographic risk. Enormous display type, unexpected italic cuts, mixed cases, hand-drawn headlines, a single oversize word as a hero." — a four-for-one slop driver behind 97% OpenAI comically-large H1, 42% bad-SVG illustration, and the editorial-italic slop. Deleted outright. - typeset.md and teach.md repeated the same category recipes; trimmed to the principle without the recipe. - SKILL.md Typography: added a hard hero-H1 ceiling (clamp() max ≤ 6rem ≈ 96px), with a <codex> block to make it explicit since OpenAI over-indexes here (97% ≥128px vs 24% for Anthropic). - animate.md, bolder.md, brand.md: removed "staggered reveals" and "scroll-triggered transitions" as the prescribed default ambitious motion. By 2026 that's the saturated AI tell, not a choreography. Reserved stagger for legitimate list-sibling rhythm. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * skill: anti-cream + codex-specific defect bans + universal slop bans Second pass after measuring more biases against the eval corpus. - SKILL.md Color: explicit "cream/sand/beige body bg is the saturated AI default of 2026" rule. Tone down the "tint every neutral" line so it doesn't read as "default to warm-tinted near-white" (which OpenAI hits at 74% and Anthropic at 31%-47%). - SKILL.md Absolute bans: add universal bans for two slop patterns detected at 55-95% across providers — tiny uppercase tracked eyebrow above every section (the 2023-era kicker that's now AI grammar) and numbered section markers (01/02/03). Also explicit "text that overflows its container is the universal defect on tablet/mobile." - SKILL.md Absolute bans → <codex> block: ban the GPT-specific defects Paul annotated repeatedly — `border:1px solid` + soft-wide-shadow (≥16px blur) "ghost cards", `border-radius:32px+` over-rounding, hand-drawn/sketchy SVG illustrations (loose-sketch / *-sketch classes, feTurbulence paper-grain filters), repeating-linear-gradient stripes, "X theater" AI-slop copy phrases. - SKILL.md Motion → <gemini> block: the image :hover transform tell (38% Google skill-on rate). Hover effects on images add no info; the image isn't an action target. Animate card chrome, not the image. - SKILL.md Typography: hard display letter-spacing floor ≥-0.04em (OpenAI defaults to -0.075em → cramped). Existing hero ceiling <codex> block extended with the letter-spacing rule. - codex.md Step A example: stop seeding "warm-grounded (deep oxblood + cream)" as the warm-palette template, which primes the cream default. - colorize.md Tinted backgrounds: stop printing the literal cream recipe `oklch(97% 0.01 60)`; replace with brand-anchored guidance. - document.md examples: warm-ash-cream → cool-paper so the example doesn't seed cream as the canonical neutral example. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * skill: universal anti-slop bans + contrast/font-count/all-caps-body rules Third pass after measuring the rest of the cross-provider matrix: - Color: explicit "Verify contrast" rule. Low-contrast text fires at 68% across all providers skill-on (90+% off). The most common failure is muted gray body on a tinted near-white; light-gray-for- elegance is named as the single biggest cause of unreadable AI pages. - Typography: max-3-font-families rule. Overused-fonts (>4 families) fires at 28% Anthropic / 36% Google / 0% OpenAI skill-on; >50% off. Also: universal "no all-caps body copy" (moved from brand-only ban to Shared design laws since product-register also overuses caps). - Copy: anti-aphoristic-cadence ban targets Anthropic's signature "X. No Y." / "X. Just Y." voice (63% skill-on copy-slop rate, 77% off — the worst rate in the matrix). Once-is-voice / three-or-more- is-tell framing per the runner's copy-slop detector. - Copy: anti-SaaS-buzzword-string ban with the literal phrase list the detector watches for (streamline/empower/supercharge, trusted- by-leading, best-in-class/enterprise-grade/cutting-edge, etc). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * skill: strengthen anti-cream rule across full warm-neutral band Smoke validation showed the cream fix worked for Google + OpenAI but Anthropic Sonnet italian-restaurant still shipped `--paper: oklch(90% .018 88)` — cream just outside the L≥95% band the rule cited. Broaden the rule: - Band: OKLCH L 0.84-0.97, C < 0.06, hue 40-100 (was 95-97% / 60-95). - Name the token-name tells explicitly (paper / cream / sand / bone / flour / linen / parchment / wheat / biscuit / ivory) — the model defaults to one of these regardless of what hex it lands on. - Call out the specific brief patterns ("warm, traditional, family- coastal-Italian" / "editorial-restraint") that the model translates into cream by reflex. Then provide three explicit non-cream options: saturated brand color, true off-white at C=0, or darker mid-tone. Warmth in the brand is carried by accent + typography + imagery, not by body bg. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * v3.2.0: skill bias-fix release Bumps version from 3.1.1 to mark the four-commit skill cleanup that rips out baked-in category recipes (brand.md), saturated-default motion tropes (staggered reveals everywhere), the cream/sand body-bg AI tell, codex-specific defects (1px+wide-shadow, over-rounding, hand-drawn SVGs, stripes, X-theater copy), the extreme-letter-spacing default, and universal slop bans (all-caps eyebrow on every section, numbered-section markers, all-caps body, font-family-count > 3, aphoristic copy cadence, SaaS buzzword strings). Plus a hard hero-H1 ceiling (clamp() ≤6rem) and a Gemini-specific image:hover transform block. Validated against ~190 post-fix samples — see impeccable-evals biases tab for per-provider deltas. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * drop "no pure black/white" rule entirely The rule was contested in the design world and causing more damage than good — pushing every page into the tinted-near-white default which is the cream/sand AI tell we already explicitly ban elsewhere. Vercel, SVKMS, Brutalist sites, et al. use pure black/white successfully; the skill shouldn't second-guess that. Skill markdown deletions: - SKILL.md Color: drop the "Never use #000 or #fff" bullet. - color-and-contrast.md: drop the "Never Use Pure Gray or Pure Black" subsection, the "Never pure black" table-row prescription, and the "Avoid: Using pure black for large areas" bullet. - colorize.md: drop the "NEVER use pure black or pure white for large areas" bullet. - polish.md: drop the "Tinted neutrals: No pure gray or pure black" half of the bullet (the gray-on-color bullet survives). Detector code (cli/engine): - registry/antipatterns.mjs: remove the `pure-black-white` entry. - rules/checks.mjs: remove the three `findings.push({ id: 'pure-black-white', ... })` emit points (inline #000 bg, Tailwind bg-black class, plain-HTML scan path). - engines/regex/detect-text.mjs: remove the two pure-black-white regex rules (CSS `background: #000…` + Tailwind `bg-black`). - detect-antipatterns-browser.js: regenerated via scripts/build-browser-detector.js. Tests: - detect-antipatterns-fixtures.test.mjs: invert the assertion that pure-black-white fires; expect it to NOT fire post-v3.2. Drop the Tailwind bg-black-opacity edge-case test (no longer relevant). - detect-antipatterns.test.js: drop the standalone "detects pure- black-white in styled-components" test and remove pure-black-white from the multi-detector assertions in PricingCard, globals.css, and GlobalStyle.tsx tests. 166 bun tests pass; 24 node fixture tests pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * skill: strip example patterns from copy rules, strengthen gemini block v3.2 rerun validation surfaced two issues: 1. Copy-slop detector fires more on Gemini under v3.2 (48% → 84%) than under no-skill baseline. Root cause: the anti-aphoristic-cadence rule printed the literal "X. No Y." / "X. Just Y." patterns as examples, and Gemini imitated them as the recommended voice. Same recipe-becomes- bias trap we hit with brand.md:116's "Enormous display type, unexpected italic cuts, mixed cases, hand-drawn headlines" enumeration. Fix: describe the cadence as a rhythm ("serious statement, then punchy short negation") without printing literal patterns. Buzzword list trimmed to a single inline phrase family rather than quoted strings. 2. Gemini image:hover transform Gemini-tell hadn't dropped (31% off → 32% v3.2). Strengthen the <gemini> block: explicit "Never animate <img> elements on hover", call out the Tailwind group-hover:scale / group-hover:rotate / group-hover:translate parent-hover patterns by name (Gemini was reaching for these via Tailwind even though the prior text talked about :hover on the image directly). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * skill: simplify context loading and inline register directive Replaces load-context.mjs's JSON output with a tight markdown block from the renamed context.mjs. The script now extracts PRODUCT.md's `## Register` field and appends a `NEXT STEP:` directive naming the matching reference (brand.md / product.md), which moved Gemini from skipping the register load entirely to honoring it. Drops the `.impeccable.md` auto-migration; makes IMPECCABLE_CONTEXT_DIR a lazy escape hatch consulted only when the default paths come up empty. Setup is now four bullets in one list. The DESIGN.md nudge is gone; in its place, a "familiarize with the existing design system" step that calls out CSS / tokens / running app as authoritative sources alongside DESIGN.md. The standalone `### Register` H3 stays for the cascade rules (task cue → surface → register field). New LLM-backed test suite at tests/skill-behavior/ runs five scenarios against claude-haiku-4-5, gpt-5.4-mini, and gemini-3.1-flash-lite via Vercel AI SDK. Captures real tool traces, asserts on context.mjs calls, brand.md loads, and teach.md fallback. Skips cleanly when API keys are unset. 13-14/15 pass; only stable failure is the v3.2.0-era gpt-mini S4 "don't re-run" regression. Adds @ai-sdk/google as devDep and the test:skill-behavior npm script. Touches em-dashes in skill/SKILL.md and four reference files so `bun run build:skills` passes its skill-prose validator. teach.md and document.md drop their "re-run the loader to refresh session cache" steps since the agent's own write is now the freshest source. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * skill: merge orphan reference files into command sub-skills + inline S-tier invariants Two related restructurings: 1. SKILL.md now carries the cross-domain invariants that catch defects in any project (contrast/placeholder/gray-on-color, similar-font pairing, text-wrap, tabular-nums, centered-stack default, Flex/Grid choice, auto-fit grids, semantic z-index, reduced motion, stagger vs section-fade, premium motion materials, focus-visible, placeholders-aren't-labels, dropdown overflow trap, button/link copy). Greenfield-only rules (theme picking, color strategy, tinted neutrals) live under "New projects only". 2. Reference files merged into their command counterparts: - spatial-design.md -> layout.md - motion-design.md -> animate.md - color-and-contrast.md -> colorize.md - responsive-design.md -> adapt.md - ux-writing.md -> clarify.md - typography.md -> typeset.md (bolder.md redirected) - cognitive-load.md + heuristics-scoring.md + personas.md -> critique.md craft.md and shape.md "load references" lists updated to new file homes. interaction-design.md stays standalone (no 1:1 command verb). Net: 36 -> 27 reference files. Same content, fewer files, no orphaned reference loaded only from craft.md. Also extends the routing rules: if the user's first word doesn't match a command but the intent clearly maps to one, load that command's reference and proceed as if invoked. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * skill: add sub-command + existing-project scenarios; move sub-command load to step 2 Adds three new LLM-backed scenarios to tests/skill-behavior: - S6: `/impeccable polish` → loads polish.md - S7: `/impeccable audit` → loads audit.md - S8: existing SvelteKit project (PRODUCT.md + DESIGN.md + src/app.css + src/lib/components/*.svelte + src/routes/+page.svelte) → agent reads at least one project code file to understand the existing design system S6/S7 surface a real model-floor: gpt-5.4-mini reads brand.md, reads the target index.html, and just does the polish/audit without ever loading the sub-command reference. Stronger SKILL.md wording didn't move it. Captured in the README baseline as a known weakness. Claude and Gemini honor the load reliably. To fix Gemini on S6/S7, sub-command reference loading is now Setup step 2 (right after context.mjs), not step 4 — placing it before the model gets focused on "doing the work". Step 3 (design-system familiarization) is tightened to require at least one project code read even when a sub-command reference loads in step 2, so Claude doesn't laser-focus on the sub-command flow and skip the broader exploration. Two new fixtures: MINIMAL_LANDING_HTML (a tiny static landing page for S6/S7) and SVELTE_PROJECT_FILES (a minimal SvelteKit scaffold with tokens, components, and a routes/+page.svelte for S8). Both designed to look real enough that agents treat them as production code. Suite is now 24 tests across three providers; baseline is 21-22/24, with the stable failures being gpt-5.4-mini scenarios 6 and 7. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * skill: add reveal-animation safety rule (must enhance, not gate visibility) Class-triggered visibility transitions pause on hidden tabs and headless renderers. The italian-restaurant smoke produced a build where 2 sections shipped opacity:0 because the CSS transition never advanced past currentTime=0 (timeline paused). Added one-liner under Motion to prevent the antipattern: reveals must enhance an already-visible default, never gate content visibility on a class-triggered transition. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * skill: restore prescriptive cream/sand/beige paragraph Bisection across 5 historical skill commits on Gemini 3.5 flash fast lane n=3 found that0cf2debdwas the peak quality state. The regression between0cf2debdand HEAD came from simplifying the long anti-cream paragraph into a one-liner. Restoring the paragraph (with em-dashes replaced by parens to satisfy prose lint) recovers ~0.22pt average on Gemini vs HEAD, with the largest gains on: - 09-luxury-hotel: +0.50 (restores editorial drama in photo-led briefs) - 10-food-magazine: +0.67 - 03-italian-restaurant: +0.51 The paragraph's load-bearing parts are the (a)(b)(c) alternatives that give the model actionable replacements for cream-tinted body bg ("saturated brand color as body", "true off-white at chroma 0", "darker mid-tone tinted neutral"). Without them, the one-line warning left the model with no concrete alternative. Cross-provider validation showed the pattern matches historical behavior: Gemini benefits from prescriptive scaffold (+0.12 over off), Sonnet is roughly neutral (+0.01), GPT-5.5 slightly regresses (-0.11 matching the v3.1.0 pattern of -0.11). The skill has never been uniformly better than skill-off across providers; this is the closest achievable state without provider-specific rework. The structural improvements from the prior restructure stay (file merges, S-tier inlines, routing rule extension, reveal-animation safety rule). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs: teach CLAUDE.md / AGENTS.md / DEVELOP.md about the skill-behavior tests Adds the `bun run test:skill-behavior` script to the test commands lists in all three docs. CLAUDE.md gets a full `### Skill-behavior tests` subsection paralleling the existing Live-mode E2E one: how the suite works (inlines source SKILL.md, scoped tools, asserts on the trace), which providers it always runs (claude-haiku-4-5, gpt-5.4-mini, gemini-3.1-flash-lite — all three every run), the eight scenarios, the baseline (21-22/24 with stable gpt-mini sub-command-routing failures), auth via repo-root `.env`, and how to add a scenario. AGENTS.md gets the one-liner plus a paragraph in Testing Guidelines that points contributors at the suite for Setup-touching edits (SKILL.md Setup section, context.mjs, teach.md, document.md, register / sub-command refs). DEVELOP.md gets a short Testing section that didn't exist before, plus a nudge in the "Test across providers" bullet pointing at the new suite as the automated way to do that. No code changes. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * detector: add 5 new antipatterns (em-dash-overuse, broken-image, marketing-buzzword, numbered-section-markers, aphoristic-cadence) Consolidates eval-side detection logic into the canonical impeccable detector. Before this change, the eval harness had its own duplicate implementations of em-dash, copy-slop, and broken-image checks. They now live alongside the existing 28 antipatterns in the impeccable registry, available to the CLI, browser extension, critique skill, and eval (via the existing slop grader child-process call). New antipatterns: - em-dash-overuse: 5+ em-dashes in body text content (threshold permits legitimate prose use of em-dash; only triggers on AI cadence-level density) - broken-image: <img> with empty src, missing src, or src="#" - marketing-buzzword: SaaS phrase list (streamline / empower / supercharge / enterprise-grade / cutting-edge / etc) - numbered-section-markers: repeated 01 / 02 / 03 sequence as section labels — the AI editorial scaffold one tier deeper than tracked eyebrow chips - aphoristic-cadence: 3+ manufactured-contrast ("Not a X. A Y.") or short-rebuttal ("Sentence. No clause." / "Sentence. Just clause.") constructions in body text Engine wiring: - broken-image runs as a static-html element rule (selector: img) and a fallback regex matcher (for non-HTML files) - em-dash / buzzword / numbered / aphoristic run as regex page-analyzers, factored into a new runTextContentAnalyzers() helper that both detectText (non-HTML) and detectHtml (HTML) call, so .html files get the same coverage as .css/.tsx Tests: 166 detector + 12 browser + 24 fixture all pass. Browser detector rebuilt (162.7 KB). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * skill: drop unvalidated anti-centering rule; add image-led hero carve-out The anti-centering rule ("Don't default to centering everything") was added without empirical support. We have a detector for it (everything-centered, threshold ≥70%) that fires on 0 / 998 samples in the corpus — never validated, never useful. Meanwhile the rule was almost certainly responsible for collapsing Gemini 3.5 flash's luxury-hotel skill-on output from the canonical "full-bleed photo + centered overlay headline" cinematic hero (the shape skill-off Gemini chooses 67% of the time) to a 50/50 magazine grid (full-bleed rate drops to 18% under skill-on, -49pp). Changes: - skill/SKILL.md #### Layout: drop "Don't default to centering..." - skill/reference/brand.md ## Layout: drop the same rule; replace with a positive carve-out — image-led briefs (hotels, restaurants, magazines, photography) often want full-bleed hero with overlaid menu and centered headline; let the photograph be the design - skill/reference/layout.md: drop the assessment question and the "asymmetric breaks centered-content pattern" framing Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Apply neo-kinpaku design system and improve live picker UX Restyle the live picker to match the site kinpaku kit, persist pick mode in localStorage, fix DESIGN.md color swatches in the parser, and land the neo-kinpaku site refresh with new tokens, assets, palette script, and detector rules. Co-authored-by: Cursor <cursoragent@cursor.com> * Add live Steer end-to-end: poll protocol, browser UI, and E2E harness. Wire page-level Steer through the live server and agent poll loop with steer_done unlock semantics, extend live.md for agents, and add smoke tests with LLM handleSteer plus recovery for hidden heroes, HMR lag, and dev-tool overlays. Co-authored-by: Cursor <cursoragent@cursor.com> * Add experimental live-poll --stream mode; keep one-shot default for Cursor. Stream keeps one process alive with ack-aware resume, but live.md documents that Cursor should stay on one-shot background notify after testing showed ~5s pickup vs sub-second on exit-based notify. Co-authored-by: Cursor <cursoragent@cursor.com> * Sync harness output and fix build validators for poll stream release. Regenerate provider skills after live-poll --stream work, update homepage detection counts to 41, and replace em dashes in site/skill copy so bun run build passes prose and count checks. Co-authored-by: Cursor <cursoragent@cursor.com> * homepage: add testimonials marquee section A two-row testimonial marquee on a tinted graphite plinth, sitting between the hero and the slop teaser. 29 testimonials sourced via api.fxtwitter.com (lightly cleaned: leading @-mention reply targets stripped, trailing self-links removed). Avatars downloaded into site/public/assets/testimonials/ so they're served locally. Quote order curated for impact — both rows lead with the punchiest quotes (Ben Davis spotlight, "Impeccable > Claude design", "THIS. This shit works.", "Uninstall whatever frontend skill you're using.") so the first viewport is loaded with the most memorable testimonials. Engineering notes: - Section uses width:100vw + margin-left:calc(50% - 50vw) to escape main.site-content's max-width + side padding (cards now clip cleanly at the actual viewport edges). - Marquee runs at 110s linear infinite. Both rows share the same duration so on-screen speeds match; track is doubled so the loop back to 0 reads as continuous. - Hero min-height reduced from 100svh to calc(100svh - 115px) so the dotted divider and top of row A peek above the fold on landing, signalling the section is there. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * homepage: keep the hero demo clear of the fixed header on short viewports The hero centers its content in the full viewport (the site header is a fixed overlay), so on shorter screens the tall Live Mode demo tucked under the nav. Raise the hero's top padding above the 97px header (113px wide, 108/92px when stacked) so content always pins below the header while still centering on tall viewports, and cap the demo frame to the viewport so the whole demo stays on screen. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Add steer voice input and refine processing animation. Wire Web Speech API on the Steer mic with auto-submit, block Cursor's preview browser with a clear message, and replace truncated "Working" text with a dots-only processing state. Co-authored-by: Cursor <cursoragent@cursor.com> * Add agent poll connectivity indicator and tighten global bar spacing. Surface poller state on the Impeccable mark via SSE and /status, with an instant disconnected tooltip, steer timeout failsafe, and matched brand/chat section gaps. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix steer focus to allow page text selection without losing type-to-steer. Blur the hidden steer input on page interaction, pause refocus during selection gestures, and reschedule focus recovery after clicks and cleared selections. Co-authored-by: Cursor <cursoragent@cursor.com> * site: rework "Design in production" section glyphs and audience band Put the three how-it-works steps back into thin-line cards and drop the overused browser-chrome bars from each glyph. Redraw the step 2 and 3 visuals to mirror the real Live Mode UI: step 2 shows the on-canvas pick outline with an attached comment bubble, step 3 shows the floating contextual accept bar plus the source-write confirmation. Re-treat the audience tiles as verdigris-lined text (no card box) under a "Who it's for" eyebrow, so each role reads as distinct from the gold step band. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Add live insert mode with HMR-safe placeholder recovery. Ships insert picking, scaffold helpers, variant cycling fixes for hidden variants, and placeholder snapshot/recreation so Astro HMR does not drop the wait-state box or re-anchor to the hero container. Co-authored-by: Cursor <cursoragent@cursor.com> * site: mobile pass — hamburger nav + designing hero overflow fix The header was rendering inline nav links + GitHub button that overflowed narrow viewports (~363px). Pre-existing display:none hacks hid Designing and Live to make the row fit, but those items still belonged in the menu. Header.astro: added a hamburger toggle button + inline script. The right cluster (nav + GitHub) becomes a collapsible drawer below the header on mobile, with data-nav-open driving the open/closed state and animating the two-line glyph into an X. kinpaku-kit.css: hamburger button (kinpaku-bordered glyph), mobile drawer panel (solid lacquer-deep bg, hairline separators between rows, full-width tappable rows), and overrides for the older sub-pages.css mobile rules (horizontal-scroll mask on the nav, hidden [data-nav="home"] item, hidden GitHub star label) — all redundant now that the drawer surfaces everything. home-kinpaku.css: dropped the @media (max-width: 560px) block that hid Designing / Live / GitHub. The drawer pattern shows them all. designing-kinpaku.css: hero h1 "Designing with Impeccable" was overflowing at narrow viewports. Three fixes: - grid-template-columns 1fr → minmax(0, 1fr) so the column shrinks to fit container instead of growing to "Impeccable"'s 472px intrinsic min-content width. - mobile h1 size override (clamp(2.2rem, 11vw, 3rem) at <=480px) since the display token's 3.4rem minimum is sized for desktop hero impact. - hide the decorative loop-wheel SVG below 600px (was overflowing 22px past the right edge). Verified clean at both 363px and 403px viewports across /, /docs, /docs/animate, /slop, /designing, /live-mode. scrollWidth matches viewport width on every page (no horizontal scroll). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * detector: refine new rules + run provider tells in browser env Follow-up to the detector port (rules landed in7648af00): - oversized-h1: flag long headlines set at display size, not punchy one/two-word heroes (length, not size alone, is the tell) - provider tells (--gpt/--gemini) now always run in a real browser env (detector page, live overlay, extension); gating is a CLI-output concern only, applied in the Node engine return paths - move theater-slop-phrase into checkHtmlPatterns so it runs in the bundled browser path, not just CLI/static (browser bundle excludes detect-text.mjs) - hero-eyebrow-chip overlay highlights the eyebrow, not the heading - gemini-tells fixture: data-URI images so the hover-zoom renders - rebuild browser bundle Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * site: migrate /detector lab to neo-kinpaku design system Rebuild the detector lab tool shell on --ks-* tokens (lacquer ground, gold hairlines, champagne/mono type) instead of the legacy warm-paper palette. Swap the "/" placeholder for the real carved-tile brand lockup, restyle the toolbar actions as kinpaku primary/secondary buttons, and recolor the finding overlay from off-brand magenta to vermilion. Update the global theme-color from #fafafa to #010101 (the sRGB render of the lacquer ground) so the browser chrome matches the dark site. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Homepage: hero finalist, compact live demo, real picker bar. Switch the hero to m-01-v2-01, tighten the in-hero demo layout, and replace the marketing gbar with a shared LiveDemoGbar that mirrors live-browser.js. Size the bar with max-content so controls are not clipped inside the capsule. Co-authored-by: Cursor <cursoragent@cursor.com> * site: migrate /cases/neo-mirai to neo-kinpaku design system Rebuild the Neo Mirai case-study page on --ks-* tokens: lacquer ground (drops the off-brand magenta radial spotlight), Alumni Sans Pinstripe display headings instead of the banned italic serif, gold eyebrow/labels, gold hairline image frames, kinpaku primary/secondary buttons, and a lacquer-deep command panel with a gold-bordered code block. Opt .neon-case-page into the shared kinpaku site-header/footer chrome in kinpaku-kit.css (per the "add new kinpaku pages to the selector list" note) so the global header and footer go dark to match the page. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * site: consolidate kinpaku header+footer into one reusable .kinpaku-chrome class The dark header/footer were not a reusable unit: the header was scoped to a per-page selector list, the github star pill was home-only, and the default footer was copy-pasted into four page stylesheets. Pages not on the lists (like /cases/neo-mirai) fell back to the legacy light chrome. Collapse all of it into one `.kinpaku-chrome` block in kinpaku-kit.css — header, github pill, and default footer — and opt every kinpaku page in via a single body class. Delete the four duplicated per-page footer blocks and the home-only github pill. The home page keeps its textured verdigris footer as a deliberate override, raised to body.home-kinpaku specificity so it wins regardless of import order. Genuinely light pages (privacy, tutorials) just omit the class. Fixes on /cases/neo-mirai: footer and github star now render dark/kinpaku (were legacy-light), and the content sections are wrapped in the .neon-case container so they sit in header-aligned gutters instead of bleeding to the viewport edge. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * site: migrate privacy + tutorials to kinpaku via a reusable surface class These were the last two light pages. Rather than rewrite their per-rule styling, add a reusable .kinpaku-surface class that remaps the legacy --color-* / --font-* tokens to kinpaku values at the body scope, so the existing legacy-token CSS (sub-pages.css prose, the pages' inline styles) renders dark for free. Same trick docs-kinpaku/slop-kinpaku use per page, lifted into one shared class. Pair it with .kinpaku-chrome for header + footer. privacy + both tutorials pages now carry both classes. Also force the sub-1.2rem headings (tutorial card titles, prose h1/h2) back to the upright body face: the legacy display face was italic serif, and the kinpaku Pinstripe face reads wrong synthesized-italic at small sizes. No light pages remain. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * site: re-add Tutorials to the /docs sidebar Tutorials lost its docs placement across two refactors: the Astro docs rebuild never carried over the sidebar tutorials list the old generated pages had, and the kinpaku homepage redesign dropped the "Full walkthrough" link. It survived only via /designing and /live-mode. Add a "Tutorials" group at the top of the docs sidebar (matching the command-category styling) linking the index plus all four tutorials, restoring the old information architecture. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * site: make kinpaku the default — flip legacy :root tokens to dark (phase 1) Repoint the legacy design tokens in tokens.css from light-mode to kinpaku: --font-* now reference the --ks-* brand faces (retiring Cormorant/Instrument/ Space Grotesk), surfaces carry dark-lacquer oklch, and --color-accent is gold instead of magenta. Values mirror the per-page kinpaku remaps. Every live page already overrides these at its body-class scope, so this changes the fallback (any classless/new page now renders kinpaku) without altering existing pages — verified home, designing, slop, live-mode, docs unchanged, and the deliberate-light demos (slop specimens, home's Aurelia mock) still render light via their own colors. First step toward removing the per-page remaps; those become redundant next. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * detector + slop: cream-palette rule, drop everything-centered, polish catalog - new deterministic cream-palette rule ("claude beige"): flags warm lightly-tinted off-white page backgrounds; wired into static + browser engines, with fixture + test - remove everything-centered rule entirely (no longer in the skill) from registry, regex analyzer (+ index-offset fix), checkPageLayout, and tests - catch Instrument Serif in overused-font (regex + OVERUSED_FONTS) - /slop: reconcile catalog (cream card in, everything-centered out; counts), and fix demo visuals — visible hairline border, gigantic clipped hero, more extreme crushed tracking, padded gray-on-color card, uniform-rhythm monotonous-spacing, long line-length line, elastic-overshoot dialog for bounce easing, real zooming image for image-hover; flip the demo surface off warm beige to a cool neutral Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * detector page: add cream-palette fixture to the catalog Surfaces the new cream/beige palette rule on /detector alongside the other Color specimens. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * site: shared docs sidebar + tutorial pages join the layout Extract the /docs section sidebar into a reusable DocsSidebar component and wire it into all three entry points so the navigation is consistent across docs index, command pages, and tutorial pages. site/components/DocsSidebar.astro (new): one source of truth. Loads the tutorials + skills collections, renders Tutorials → Commands grouped by category, and highlights the active entry via activeCommand / activeTutorial props. site/pages/docs/index.astro: swap the inline sidebar markup for the component. Drop the "All tutorials" link — the dedicated tutorials listing page wasn't earning its slot in the rail. site/layouts/Doc.astro: same swap. Command pages now also see the Tutorials section above Commands, matching /docs. site/pages/tutorials/[...slug].astro: rewrite from a standalone page (custom .tutorial-page wrapper, ad-hoc breadcrumb) to the full skills-layout shell with DocsSidebar in the left rail. Tutorial content now reads in the same layout as command reference pages. site/content/tutorials/brand-vs-product.md (deleted): the skill picks the register automatically from PRODUCT.md, so a tutorial telling users to pick it themselves was misleading. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * detector: catch Tailwind warm-light bg utilities in cream-palette The static engine can't resolve Tailwind classes to computed CSS, so a `bg-amber-50` on <body> slipped past the cream-palette rule. Add a class-list fallback that scans body/html for arbitrary `bg-[...]` values and named warm-light utilities (amber/orange/yellow/stone), each run through the same isCreamColor test so neutrals and over-saturated shades drop out. Fixture + test for the class-only case. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * site: drop redundant per-page token remaps (phase 2) With kinpaku now the :root default, the --color-* / --font-* remap blocks in docs/slop/designing/live-mode-kinpaku.css re-declared values identical to :root. Removed them, keeping only the --ks-muted alias (still read by name in those files) and each page's shell (gradient bg, color, min-height). home-kinpaku.css keeps its remap: it uses home-specific values (e.g. --color-charcoal: var(--ks-text), --color-cream: var(--ks-lacquer-raised)) plus the --cat-* gradient overrides, so it is not redundant. Verified designing (PRODUCT.md viz), slop (specimens stay light), docs, live-mode unchanged. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * site: drop italic from 15 dead editorial-serif heading rules Audited every font-style: italic in sub-pages.css and main.css against the live markup. Removed italic from the 15 rules whose selectors don't appear in any page/component/content/script: sub-pages.css: docs-home-card-title, docs-category-title, tutorial-embed-caption, skill-demo-caption, skill-source-card-subtitle, skill-references-heading, skill-reference-title main.css: hero-title-combined, hero-tagline-combined, impeccable-title, loading-state, install-primary-howto .install-path-desc em, install-howto-steps > li::before, install-step-status, consulting-title These were dormant remnants of the retired Cormorant italic-serif look — the kinpaku Pinstripe face renders them as bad synthesized-italic, but no markup matches the selectors so nothing rendered. Removed only the font-style declaration; the rest of each rule stays (whole-rule cleanup is out of scope). Kept the 5 live selectors (slop-section-heading, tutorial-card-title, visual-mode-demo-caption, visual-mode-method-name, gallery-card-title) per the "if they're not used anywhere" condition, plus .prose em (real emphasis) and .prose blockquote (conventional blockquote italic). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * skill: brand-seed palette.mjs + Setup step to run it New-brand color now starts from a curated seed color (129 OKLCH seeds) instead of the model guessing or defaulting to warm-cream. The script returns one seed + composition guidance (pure-bg architecture, perceptual text-on-fill, anti-cliché moods, jewel-tone range), with inverse-frequency hue weighting for fair rainbow exposure and deterministic --from picking. SKILL.md Setup step 5 makes it run for greenfield projects. Curation tooling lives in the impeccable-evals repo (tools/palette/). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Remove accidental live mode inject from Base.astro. The localhost live.js tag was left in the site layout after a dev session and should never ship in the Astro template. Co-authored-by: Cursor <cursoragent@cursor.com> * site: dedicated /changelog + /faq, epic v3.5.0 notes, Live Mode → Beta Split changelog and FAQ out of the homepage into two standalone kinpaku pages, linked from the footer (and a quiet hint under the Get-started CTA). /changelog: every release inline (no collapsible), newest first. The v3.5.0 entry leads with a one-line summary, a real before/after pair from the GPT-5.5 eval corpus (luxury-hotel brief, skill off vs on), and a stat row (74% cream-bg, 76% extreme tracking, 90%+ low-contrast — measured across ~190 samples). Then five scannable bold-led bullets, biggest takeaway first: per-provider skill compilation, the bias-fix, Live Mode, the 7 new detector rules, the tighter skill. Before/after JPGs optimized to ~470KB total (down from ~2.5MB PNGs). /faq: the six support questions, each deep-linkable. Live Mode is now Beta everywhere it surfaces: the /live-mode eyebrow badge and note, the homepage bento tile badge, and the changelog entry. The historical v3.0 changelog entry stays "Alpha" — accurate to what shipped then. Footer trimmed to the four links not already in the top nav (Changelog, FAQ, Privacy, GitHub). Version bumped 3.2.0 → 3.5.0 across the three plugin manifests; the 3.2 bias-fix work folds into this release rather than shipping separately. astro.config.mjs: disable the dev toolbar. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * site: point /design-system hero at the m-01-v2-01 finalist design-system.css referenced kintsugi-hero-v2.png, an untracked orphan that was never committed. Repoint it at the committed m-01-v2-01 finalist so /design-system and the homepage hero share one image, and the page no longer depends on a file outside the repo. The v2 orphan moved to tmp/. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * build: sync harness mirrors + green the prose gate Rebuild propagates the committed skill source (palette.mjs Setup step, detector rule updates, brand.md) into the 13 harness output dirs and the plugin subtree, which had drifted from source. Also fixes the prose validator, which had been red on six pre-existing hits across committed files: - Four em dashes in code comments (Testimonials.astro, LiveDemoGbar.astro, index.astro) and one in skill/reference/live.md — reworded to colons/commas. - Two in the slop catalog (an em-dash-overuse specimen and the marketing-buzzword rule naming "empower"). Those are intentional: the slop page documents every antipattern by example, so it must contain them. Exempted site/pages/slop from validateProse rather than neutering the specimens. `bun run build` is now green end to end: counts validate, prose passes, site builds. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * skill: rewrite no-section-fade rule to fix Gemini zero-motion overcorrection The old rule ("whole-section fade-on-scroll is the saturated AI motion reflex") drove Gemini to overcorrect into shipping pages with no motion at all: motion-variety 39% / zero-motion 12% with the skill on, vs ~74-78% variety and ~3% zero-motion without it. Rewrite keeps the legitimate-stagger carve-out, names the defect at shape level (one identical entrance on every section) without enumerating motion primitives, and adds an explicit clause that suppressing the reflex is never grounds for a static page. Validated on Gemini 3.5-flash (n=10, luxury-hotel + infra-platform): motion-variety 39% -> 70%, zero-motion 12% -> 0%, staggered-reveal stays 0% (reflex not re-inflated). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * release: bump CLI to 2.2.0 and extension to 1.1.0 Both ship the expanded detector: the 7 new rules (cream-palette, em-dash-overuse, marketing-buzzword, numbered-section-markers, aphoristic-cadence, broken-image, italic-serif-display) plus hero-eyebrow-chip, with everything-centered removed. 41 rules total. The extension settings page already supports toggling them: the rule list renders from detector/antipatterns.json, grouped by category, and disabledRules flows through chrome.storage.sync into the scan config, which detect.js honors by rule id. New rules are toggleable with no UI change. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * release: fix release.mjs for the moved changelog + add CLI/ext entries The changelog moved from site/pages/index.astro to its own site/pages/changelog.astro with new markup (cf-version / cf-entry / cf-items), which left release.mjs reading the wrong file with the old selectors. All three release commands would have failed at note extraction. Point it at changelog.astro, match cf-version, and scope notes to the <ul class="cf-items"> bullet list — that also skips the lead paragraph, before/after figure, and stat row on the v3.5.0 entry, keeping release notes to clean bullets. Add CLI v2.2.0 and Extension v1.1.0 changelog entries (the shared detector update: 7 new rules, everything-centered removed, 41 total; plus the extension's per-rule toggles) so release:cli and release:ext have notes to extract. Verified extraction for all three labels: v3.5.0 (5 bullets), CLI v2.2.0 (3), Extension v1.1.0 (2). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix: correct dev server port to 4321 and drop stale pnpm-lock Astro serves on 4321, not 3000 as the docs claimed; update CLAUDE.md, AGENTS.md, and screenshot-antipatterns.js. Remove the leftover pnpm-lock.yaml from the Astro migration so Cloudflare's frozen install uses the maintained, in-sync bun.lock instead of a drifted pnpm lockfile. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * site: rework /designing flow, rhythm, Live Mode mock, and CTA Restructure the page so iteration reads as the core value, not net-new. The four loop phases are wrapped in a track with a sticky scroll-spy nav (Start/Iterate/Polish/Maintain) that pins under the header and highlights the active phase; the surfaces section (skill/CLI/extension) moves out of the loop into the post-loop context group so the loop runs uninterrupted. Fix the iterate split: shared subgrid row tracks so the terminal and the Live Mode mock align on the same baseline regardless of paragraph length, wider intro measure (52ch, was a crammed 36ch), and a deeper picker stage so the context and global bars breathe instead of stacking on the card. Rebuild the Live Mode mock to mirror the real picker: carved-tile mark plus Pick / Insert / Detect / DESIGN.md controls on lacquer-deep with the gold border, and a /impeccable live entry line so the reader knows how to start. Reframe Start as the hard mode, move h3 subheads off the thin display face onto Albert Sans, and trim Start so it no longer dominates the loop. Rework the closing CTA into two standalone raised cards (the bento plinth made them read as boxes nested in a box), and fix the tutorials copy: there are three walkthroughs now, and the brand-vs-product tutorial is gone, so drop it from the CTA and remove the dead lane link to it. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * site: reorder Get Started so usage follows setup, link out to more Move the /impeccable usage examples below the Chrome extension, CLI, and Stay-updated block. Running a command is the logical next step once the skill, extension, CLI, and subscriptions are all in place, so the section now reads install -> set up the extras -> use it. Add a closing "Go deeper" line linking to the Designing with Impeccable workflow page and the docs. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix: install compiled per-provider skill variants, not uncompiled source `npx skills add` (and `impeccable skills install`, which wrapped it) installed the uncompiled skill/ source verbatim: the skills CLI dedupes discovery by name and picks skill/SKILL.md first, so installs shipped unresolved {{placeholders}} and no vendored detector (#168). - Rename skill/SKILL.md -> skill/SKILL.src.md so the skills CLI's discovery skips the source and falls through to a compiled .agents variant; update the build reader, skill-behavior harness, and docs to match. - Refactor `impeccable skills install` to copy each harness's compiled variant from the universal bundle (real dirs, no npx skills, no symlink), with project/global harness detection and a --providers override. - Fix stale unit tests (replacePlaceholders, readPatterns, transformer prefix/summary) that asserted removed pre-v3.0 behavior, and wire the three orphaned test files into `bun run test` so the drift can't recur. - Split skills-cli.test.js: pure blocks run by default, network blocks move behind a new `bun run test:cli-e2e`; fix its stale update assertions. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * site: default to `npx impeccable skills install`, restore install-method panel Get Started recommended `npx skills add`, which installs a single shared build across harnesses. Make our CLI the default (it installs the build compiled for each harness) and bring back the "Other install methods" disclosure the neo-kinpaku redesign dropped. - Homepage: primary command is now `npx impeccable skills install`; a native <details> panel offers the Claude Code plugin and `npx skills` (caveated as installing one shared build rather than the per-harness one). - FAQ: recommend `npx impeccable skills install` to install, `--force` to reinstall, and note the npx skills shared-build caveat. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * site: reword craft tagline so it doesn't lead with "Shape" The craft card's tagline began with the word "Shape", which reads like the name of the sibling /shape command and made the two cards look swapped (#166). Reword to "Design it, then build it, all in one flow." No data was actually swapped; this is a copy collision fix. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * skill: rename teach -> init and expand its setup flow Rename the `/impeccable teach` command to `/impeccable init` across the skill, site, CLI, and tests. `teach` stays as a deprecated router alias and /docs/teach + /skills/teach redirect to /docs/init. Expand the command beyond writing PRODUCT.md/DESIGN.md: the same codebase crawl now also pre-configures `.impeccable/live/config.json` (Step 6, with CSP consent) so live mode boots with no first-time detour, and the flow ends by recommending the best commands to run next from what the scan surfaced (Step 7). Fold two items into the unreleased v3.5.0 changelog entry: the init rename and the brand-seed palette picker. No version bump. Regenerates all harness skill output dirs and the _redirects file. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: lead README install + usage with the CLI installer Add `npx impeccable skills install` as the recommended install option and update the Usage section to the `/impeccable <command>` form, dropping the nonexistent `/normalize` example. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(skill-behavior): swap to production-tier models (sonnet + gpt-5.5) Replace the cheap-tier default lineup (claude-haiku-4-5, gpt-5.4-mini) with production-tier models (claude-sonnet-4-6, gpt-5.5) so the skill-behavior suite reflects what users actually run. gemini stays on flash-lite. Sync the docs (CLAUDE.md, AGENTS.md, tests/skill-behavior/README.md): new model names, cost estimate raised to ~$0.50-1.50/sweep, and the old 21-22/24 baseline reframed as previous-cheap-tier history pending re-measurement on the new lineup. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: self-updating skill via boot-time version check context.mjs now polls a new lightweight /api/version endpoint at most once per day (cached globally in ~/.impeccable) and appends an UPDATE_AVAILABLE directive when a newer skill version has shipped, prompting the agent to offer `npx impeccable skills update`. Best-effort and silent on any failure; asks before updating; suppresses re-prompts for a declined version for a week. Opt out with IMPECCABLE_NO_UPDATE_CHECK=1. - skill/scripts/context.mjs: version read, throttle + anti-nag cache, directive - scripts/build.js + _redirects: /api/version endpoint (from plugin.json version) - skill/SKILL.src.md: document the UPDATE_AVAILABLE boot branch - tests/context.test.mjs: coverage for cached/newer/suppressed/opt-out paths - changelog: v3.5.0 entry - synced harness skill dirs via bun run build Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test: cover the self-update path (network + LLM behavior) context.test.mjs: add a localhost stub-server integration test for the live fetch path (poll /api/version, cache a newer version, stay silent on same-or-older, fail silent + stamp lastCheck when unreachable). Runs against 127.0.0.1 only, never the real site; uses async spawn so the in-process stub isn't deadlocked by spawnSync blocking the event loop. skill-behavior: add scenario 9 asserting the agent surfaces UPDATE_AVAILABLE but never auto-runs `npx impeccable skills update` without asking. New prepareWorkspace `skillVersion` copy-mode (so context.mjs has a SKILL.md to version-check), env threading through runTurn -> execBash, and bash-output capture to prove the agent actually received the directive. Passed on claude-sonnet-4-6, gpt-5.5, and gemini-3.1-flash-lite. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> Co-authored-by: Cursor <cursoragent@cursor.com>
38 KiB
Purpose
Resolve one stable target, run two independent assessments, synthesize a design critique, persist a snapshot, and ask the user what to improve next. The chat response is the primary deliverable; the snapshot is an archive/backlog for future commands.
Hard Invariants
- Assessment A (design review) and Assessment B (detector/browser evidence) are both required.
- Assessment A must finish before detector findings enter the parent synthesis context. Detector output is deterministic, but it still anchors judgment.
- If sub-agents are unavailable, fall back sequentially: finish and record Assessment A first, then run Assessment B, then synthesize.
- A skipped detector is a failed critique run unless
detect.mjsis missing or crashes after a real attempt. - Viewable targets require browser inspection when available.
- Any local server started only for critique visualization must run in the background, have a recorded stop method, and be stopped before final reporting unless the user asks to keep it.
- Do not claim a user-visible overlay exists unless script injection succeeded and the detector ran in the page.
Setup
- Resolve the target to a concrete file path or URL. Prefer a source path over a dev-server URL when both identify the same surface; ports drift, paths do not.
- "the homepage" ->
site/pages/index.astroorindex.html - "the settings modal" -> the primary component file
- "this page" -> the current URL or source file
- "the homepage" ->
- Compute the slug:
Keep it. If the command exits non-zero, skip persistence and trend for this run, but continue the critique.
node {{scripts_path}}/critique-storage.mjs slug "<resolved-path-or-url>" - Read
.impeccable/critique/ignore.mdif it exists. Drop matching findings silently; it is the only prior-run input critique consumes.
Assessment Orchestration
Delegate Assessment A and Assessment B to separate sub-agents when possible. They must not see each other's output. Do not show findings to the user until synthesis.
Codex sub-agent gate: - If `spawn_agent` is exposed and the user explicitly allowed sub-agents, delegation, or parallel agent work, spawn A and B immediately. - If `spawn_agent` is exposed but the user did not explicitly allow sub-agents, ask exactly once: "Impeccable critique is designed to run two independent sub-agents for an unanchored assessment. May I use sub-agents for this critique?" Then stop until the user answers. - If allowed, spawn A and B. If declined, run sequentially and report `Assessment independence: degraded (sub-agents declined by user)`. - If `spawn_agent` is not exposed, do not ask; run sequentially and report `Assessment independence: degraded (spawn_agent unavailable in this session)`. - If spawning fails after permission, run sequentially and report `Assessment independence: degraded (sub-agent spawn failed: )`. Prefer `fork_context: false` with self-contained prompts containing cwd, target, live URL, references, product context, and output contract. If using `fork_context: true`, omit `agent_type`, `model`, and `reasoning_effort`.If browser automation is available, each assessment creates its own new tab. Never reuse an existing tab, even if it is already at the right URL.
Assessment A: Design Review
Read relevant source files and visually inspect the live page when browser automation is available. Think like a design director.
Evaluate:
- AI slop: Would someone believe "AI made this" immediately? Check all DON'T guidance from the parent Impeccable skill.
- Holistic design: hierarchy, IA, emotional fit, discoverability, composition, typography, color, accessibility, states, copy, and edge cases.
- Cognitive load: consult the Cognitive Load Assessment section below; report checklist failures and decision points with >4 visible options.
- Emotional journey: peak-end rule, emotional valleys, reassurance at high-stakes moments.
- Nielsen heuristics: consult the Heuristics Scoring Guide section below; score all 10 heuristics 0-4.
Return: AI slop verdict, heuristic scores, cognitive load, emotional journey, 2-3 strengths, 3-5 priority issues, persona red flags, minor observations, and provocative questions.
Assessment B: Detector + Browser Evidence
Run the bundled detector and browser visualization evidence. Assessment B is mandatory and must remain isolated from Assessment A until both are complete.
CLI scan:
node {{scripts_path}}/detect.mjs --json [--fast] [target]
- Pass markup files/directories as
[target]; do not pass CSS-only files. - For URLs, skip CLI scan and use browser visualization.
- For 200+ scannable files, use
--fast; for 500+, narrow scope or ask. - Exit code 0 = clean; 2 = findings.
- If the detector entrypoint is missing or fails to load, report deterministic scan unavailable and continue with browser/manual review.
Browser visualization is required for a viewable target when browser automation is available. Use a localhost dev/static URL for local files; avoid file:// unless the available browser explicitly supports this workflow. Overlay flow:
- Create a fresh tab and navigate.
- Preflight mutable injection by setting
document.titleand appending a<script>tag. Read-only evaluate APIs do not count. - If mutation is unavailable, skip live server, browser presentation, and injection; report fallback signal.
- If mutation is available, start
node {{scripts_path}}/live-server.mjs --background, present the browser if supported, label[Human], scroll top, injecthttp://localhost:PORT/detect.js, wait 2-3 seconds, readimpeccableconsole messages, then stop the live server. - For multi-view targets, inject on 3-5 representative pages.
Return: CLI findings JSON/counts, browser console findings if applicable, false positives, and skipped/failed browser steps with concrete reasons.
After Assessment B returns usable CLI findings, reuse them. Do not rerun detect.mjs in the parent unless Assessment B failed, was truncated, or omitted count, rule names, or file locations.
Generate Combined Critique Report
Synthesize both assessments into a single report. Do NOT simply concatenate. Weave the findings together, noting where the LLM review and detector agree, where the detector caught issues the LLM missed, and where detector findings are false positives.
The chat response is the primary user-facing deliverable. Present the full structured critique below in chat; do not replace it with a summary and a link. The persisted snapshot is only an archive/backlog for later commands.
Codex final-answer note: `$impeccable critique` produces a report artifact, so the final chat response should intentionally exceed the usual concise close-out style. Do not title the final response "Critique Summary" unless the user explicitly asked for a summary.Structure your feedback as a design director would:
Design Health Score
Consult the Heuristics Scoring Guide section below.
Present the Nielsen's 10 heuristics scores as a table:
| # | Heuristic | Score | Key Issue |
|---|---|---|---|
| 1 | Visibility of System Status | ? | [specific finding or "n/a" if solid] |
| 2 | Match System / Real World | ? | |
| 3 | User Control and Freedom | ? | |
| 4 | Consistency and Standards | ? | |
| 5 | Error Prevention | ? | |
| 6 | Recognition Rather Than Recall | ? | |
| 7 | Flexibility and Efficiency | ? | |
| 8 | Aesthetic and Minimalist Design | ? | |
| 9 | Error Recovery | ? | |
| 10 | Help and Documentation | ? | |
| Total | ??/40 | [Rating band] |
Be honest with scores. A 4 means genuinely excellent. Most real interfaces score 20-32.
Anti-Patterns Verdict
Start here. Does this look AI-generated?
LLM assessment: Your own evaluation of AI slop tells. Cover overall aesthetic feel, layout sameness, generic composition, missed opportunities for personality.
Deterministic scan: Summarize what the automated detector found, with counts and file locations. Note any additional issues the detector caught that you missed, and flag any false positives.
Visual overlays (if injection succeeded): Tell the user that overlays are now visible in the [Human] tab in their browser, highlighting the detected issues. Summarize what the console output reported. If browser visualization was attempted but injection failed, say that no reliable user-visible overlay is available and report the fallback signal instead.
Overall Impression
A brief gut reaction: what works, what doesn't, and the single biggest opportunity.
What's Working
Highlight 2-3 things done well. Be specific about why they work.
Priority Issues
The 3-5 most impactful design problems, ordered by importance.
For each issue, tag with P0-P3 severity (see Issue Severity below for definitions):
- [P?] What: Name the problem clearly
- Why it matters: How this hurts users or undermines goals
- Fix: What to do about it (be concrete)
- Suggested command: Which command could address this (from: {{available_commands}})
Persona Red Flags
Consult the Personas reference below.
Auto-select 2-3 personas most relevant to this interface type (use the selection table in the reference). If {{config_file}} contains a ## Design Context section from impeccable init, also generate 1-2 project-specific personas from the audience/brand info.
For each selected persona, walk through the primary user action and list specific red flags found:
Alex (Power User): No keyboard shortcuts detected. Form requires 8 clicks for primary action. Forced modal onboarding. High abandonment risk.
Jordan (First-Timer): Icon-only nav in sidebar. Technical jargon in error messages ("404 Not Found"). No visible help. Will abandon at step 2.
Be specific. Name the exact elements and interactions that fail each persona. Don't write generic persona descriptions; write what broke for them.
Minor Observations
Quick notes on smaller issues worth addressing.
Questions to Consider
Provocative questions that might unlock better solutions:
- "What if the primary action were more prominent?"
- "Does this need to feel this complex?"
- "What would a confident version of this look like?"
Codex Run Notes are final-chat only. Do not include this section in the persisted snapshot body, because persistence, trend read, and temp cleanup happen after the snapshot write and would otherwise archive stale status such as "pending after persistence."
Remember:
- Be direct. Vague feedback wastes everyone's time.
- Be specific. "The submit button," not "some elements."
- Say what's wrong AND why it matters to users.
- Give concrete suggestions. Cut "consider exploring..." entirely.
- Prioritize ruthlessly. If everything is important, nothing is.
- Don't soften criticism. Developers need honest feedback to ship great design.
Persist the Snapshot
Once the report above is finalized, write it to .impeccable/critique/ so the user can refer back, and so {{command_prefix}}impeccable polish can pick up the priority issues without a copy-paste.
Skip this step if the Setup slug was null (vague or root-level target).
-
Write the body to a temp file so you can pipe it to the helper. Use the full critique report (heuristic table, anti-patterns verdict, priority issues, persona red flags, minor observations, and questions), but stop before the "Ask the User" / "Recommended Actions" sections that come later.
Codex: exclude Run Notes from the temp body file; Run Notes are final-chat only because persistence, trend read, and temp cleanup happen after the snapshot write. -
Pass the structured metadata through
IMPECCABLE_CRITIQUE_META(JSON), then run the write command:IMPECCABLE_CRITIQUE_META='{"target":"<user phrasing>","total_score":<n>,"p0_count":<n>,"p1_count":<n>}' \ node {{scripts_path}}/critique-storage.mjs write <slug> <body-file>The helper prints the absolute path it wrote.
-
Delete the temp body file after the write attempt completes, whether the write succeeded or failed. If deletion fails, mention
temp-file cleanup failed: <reason>briefly in the final output, but do not block the critique. -
Read the trend for context:
node {{scripts_path}}/critique-storage.mjs trend <slug> 5This returns a JSON array of the last 5 frontmatter entries (including the one you just wrote).
-
Append a single line to the user-visible output, after the report and before the questions:
Trend for
<slug>(last 5 runs): 24 → 28 → 32 → 29 → 32 Wrote.impeccable/critique/<filename>.If this is the first run for the slug, the trend is just one score; say so: "First run for this target, no trend yet."
This is fire-and-forget. Do not show the user the helper's JSON output; only the human-readable trend line and the written path. Failures here should not block the rest of the flow; print the error and move on.
Ask the User
After presenting findings, use targeted questions based on what was actually found. {{ask_instruction}} These answers will shape the action plan.
Ask questions along these lines (adapt to the specific findings; do NOT ask generic questions):
-
Priority direction: Based on the issues found, ask which category matters most to the user right now. For example: "I found problems with visual hierarchy, color usage, and information overload. Which area should we tackle first?" Offer the top 2-3 issue categories as options.
-
Design intent: If the critique found a tonal mismatch, ask whether it was intentional. For example: "The interface feels clinical and corporate. Is that the intended tone, or should it feel warmer/bolder/more playful?" Offer 2-3 tonal directions as options based on what would fix the issues found.
-
Scope: Ask how much the user wants to take on. For example: "I found N issues. Want to address everything, or focus on the top 3?" Offer scope options like "Top 3 only", "All issues", "Critical issues only".
-
Constraints (optional; only ask if relevant): If the findings touch many areas, ask if anything is off-limits. For example: "Should any sections stay as-is?" This prevents the plan from touching things the user considers done.
Rules for questions:
- Every question must reference specific findings from the report. Never ask generic "who is your audience?" questions.
- Keep it to 2-4 questions maximum. Respect the user's time.
- Offer concrete options, not open-ended prompts.
- If findings are straightforward (e.g., only 1-2 clear issues), skip questions and go directly to Recommended Actions.
Recommended Actions
After receiving the user's answers, present a prioritized action summary reflecting the user's priorities and scope from Ask the User.
Action Summary
List recommended commands in priority order, based on the user's answers:
{{command_prefix}}command-name: Brief description of what to fix (specific context from critique findings){{command_prefix}}command-name: Brief description (specific context) ...
Rules for recommendations:
- Only recommend commands from: {{available_commands}}
- Order by the user's stated priorities first, then by impact
- Each item's description should carry enough context that the command knows what to focus on
- Map each Priority Issue to the appropriate command
- Skip commands that would address zero issues
- If the user chose a limited scope, only include items within that scope
- If the user marked areas as off-limits, exclude commands that would touch those areas
- End with
{{command_prefix}}impeccable polishas the final step if any fixes were recommended
After presenting the summary, tell the user:
You can ask me to run these one at a time, all at once, or in any order you prefer.
Re-run
{{command_prefix}}impeccable critiqueafter fixes to see your score improve.
Reference Material
The sections below were previously separate reference files (cognitive-load.md, heuristics-scoring.md, personas.md). They live inline now so the critique flow has all its deep context in one place.
Cognitive Load Assessment
Cognitive load is the total mental effort required to use an interface. Overloaded users make mistakes, get frustrated, and leave. This reference helps identify and fix cognitive overload.
Three Types of Cognitive Load
Intrinsic Load: The Task Itself
Complexity inherent to what the user is trying to do. You can't eliminate this, but you can structure it.
Manage it by:
- Breaking complex tasks into discrete steps
- Providing scaffolding (templates, defaults, examples)
- Progressive disclosure: show what's needed now, hide the rest
- Grouping related decisions together
Extraneous Load: Bad Design
Mental effort caused by poor design choices. Eliminate this ruthlessly. It's pure waste.
Common sources:
- Confusing navigation that requires mental mapping
- Unclear labels that force users to guess meaning
- Visual clutter competing for attention
- Inconsistent patterns that prevent learning
- Unnecessary steps between user intent and result
Germane Load: Learning Effort
Mental effort spent building understanding. This is good cognitive load; it leads to mastery.
Support it by:
- Progressive disclosure that reveals complexity gradually
- Consistent patterns that reward learning
- Feedback that confirms correct understanding
- Onboarding that teaches through action, not walls of text
Cognitive Load Checklist
Evaluate the interface against these 8 items:
- Single focus: Can the user complete their primary task without distraction from competing elements?
- Chunking: Is information presented in digestible groups (≤4 items per group)?
- Grouping: Are related items visually grouped together (proximity, borders, shared background)?
- Visual hierarchy: Is it immediately clear what's most important on the screen?
- One thing at a time: Can the user focus on a single decision before moving to the next?
- Minimal choices: Are decisions simplified (≤4 visible options at any decision point)?
- Working memory: Does the user need to remember information from a previous screen to act on the current one?
- Progressive disclosure: Is complexity revealed only when the user needs it?
Scoring: Count the failed items. 0–1 failures = low cognitive load (good). 2–3 = moderate (address soon). 4+ = high cognitive load (critical fix needed).
The Working Memory Rule
Humans can hold ≤4 items in working memory at once (Miller's Law revised by Cowan, 2001).
At any decision point, count the number of distinct options, actions, or pieces of information a user must simultaneously consider:
- ≤4 items: Within working memory limits, manageable
- 5–7 items: Pushing the boundary; consider grouping or progressive disclosure
- 8+ items: Overloaded; users will skip, misclick, or abandon
Practical applications:
- Navigation menus: ≤5 top-level items (group the rest under clear categories)
- Form sections: ≤4 fields visible per group before a visual break
- Action buttons: 1 primary, 1–2 secondary, group the rest in a menu
- Dashboard widgets: ≤4 key metrics visible without scrolling
- Pricing tiers: ≤3 options (more causes analysis paralysis)
Common Cognitive Load Violations
1. The Wall of Options
Problem: Presenting 10+ choices at once with no hierarchy. Fix: Group into categories, highlight recommended, use progressive disclosure.
2. The Memory Bridge
Problem: User must remember info from step 1 to complete step 3. Fix: Keep relevant context visible, or repeat it where it's needed.
3. The Hidden Navigation
Problem: User must build a mental map of where things are. Fix: Always show current location (breadcrumbs, active states, progress indicators).
4. The Jargon Barrier
Problem: Technical or domain language forces translation effort. Fix: Use plain language. If domain terms are unavoidable, define them inline.
5. The Visual Noise Floor
Problem: Every element has the same visual weight; nothing stands out. Fix: Establish clear hierarchy: one primary element, 2–3 secondary, everything else muted.
6. The Inconsistent Pattern
Problem: Similar actions work differently in different places. Fix: Standardize interaction patterns. Same type of action = same type of UI.
7. The Multi-Task Demand
Problem: Interface requires processing multiple simultaneous inputs (reading + deciding + navigating). Fix: Sequence the steps. Let the user do one thing at a time.
8. The Context Switch
Problem: User must jump between screens/tabs/modals to gather info for a single decision. Fix: Co-locate the information needed for each decision. Reduce back-and-forth.
Heuristics Scoring Guide
Score each of Nielsen's 10 Usability Heuristics on a 0–4 scale. Be honest: a 4 means genuinely excellent, not "good enough."
Nielsen's 10 Heuristics
1. Visibility of System Status
Keep users informed about what's happening through timely, appropriate feedback.
Check for:
- Loading indicators during async operations
- Confirmation of user actions (save, submit, delete)
- Progress indicators for multi-step processes
- Current location in navigation (breadcrumbs, active states)
- Form validation feedback (inline, not just on submit)
Scoring:
| Score | Criteria |
|---|---|
| 0 | No feedback; user is guessing what happened |
| 1 | Rare feedback; most actions produce no visible response |
| 2 | Partial; some states communicated, major gaps remain |
| 3 | Good; most operations give clear feedback, minor gaps |
| 4 | Excellent; every action confirms, progress is always visible |
2. Match Between System and Real World
Speak the user's language. Follow real-world conventions. Information appears in natural, logical order.
Check for:
- Familiar terminology (no unexplained jargon)
- Logical information order matching user expectations
- Recognizable icons and metaphors
- Domain-appropriate language for the target audience
- Natural reading flow (left-to-right, top-to-bottom priority)
Scoring:
| Score | Criteria |
|---|---|
| 0 | Pure tech jargon, alien to users |
| 1 | Mostly confusing; requires domain expertise to navigate |
| 2 | Mixed; some plain language, some jargon leaks through |
| 3 | Mostly natural; occasional term needs context |
| 4 | Speaks the user's language fluently throughout |
3. User Control and Freedom
Users need a clear "emergency exit" from unwanted states without extended dialogue.
Check for:
- Undo/redo functionality
- Cancel buttons on forms and modals
- Clear navigation back to safety (home, previous)
- Easy way to clear filters, search, selections
- Escape from long or multi-step processes
Scoring:
| Score | Criteria |
|---|---|
| 0 | Users get trapped; no way out without refreshing |
| 1 | Difficult exits; must find obscure paths to escape |
| 2 | Some exits; main flows have escape, edge cases don't |
| 3 | Good control; users can exit and undo most actions |
| 4 | Full control; undo, cancel, back, and escape everywhere |
4. Consistency and Standards
Users shouldn't wonder whether different words, situations, or actions mean the same thing.
Check for:
- Consistent terminology throughout the interface
- Same actions produce same results everywhere
- Platform conventions followed (standard UI patterns)
- Visual consistency (colors, typography, spacing, components)
- Consistent interaction patterns (same gesture = same behavior)
Scoring:
| Score | Criteria |
|---|---|
| 0 | Inconsistent everywhere; feels like different products stitched together |
| 1 | Many inconsistencies; similar things look/behave differently |
| 2 | Partially consistent; main flows match, details diverge |
| 3 | Mostly consistent; occasional deviation, nothing confusing |
| 4 | Fully consistent; cohesive system, predictable behavior |
5. Error Prevention
Better than good error messages is a design that prevents problems in the first place.
Check for:
- Confirmation before destructive actions (delete, overwrite)
- Constraints preventing invalid input (date pickers, dropdowns)
- Smart defaults that reduce errors
- Clear labels that prevent misunderstanding
- Autosave and draft recovery
Scoring:
| Score | Criteria |
|---|---|
| 0 | Errors easy to make; no guardrails anywhere |
| 1 | Few safeguards; some inputs validated, most aren't |
| 2 | Partial prevention; common errors caught, edge cases slip |
| 3 | Good prevention; most error paths blocked proactively |
| 4 | Excellent; errors nearly impossible through smart constraints |
6. Recognition Rather Than Recall
Minimize memory load. Make objects, actions, and options visible or easily retrievable.
Check for:
- Visible options (not buried in hidden menus)
- Contextual help when needed (tooltips, inline hints)
- Recent items and history
- Autocomplete and suggestions
- Labels on icons (not icon-only navigation)
Scoring:
| Score | Criteria |
|---|---|
| 0 | Heavy memorization; users must remember paths and commands |
| 1 | Mostly recall; many hidden features, few visible cues |
| 2 | Some aids; main actions visible, secondary features hidden |
| 3 | Good recognition; most things discoverable, few memory demands |
| 4 | Everything discoverable; users never need to memorize |
7. Flexibility and Efficiency of Use
Accelerators, invisible to novices, speed up expert interaction.
Check for:
- Keyboard shortcuts for common actions
- Customizable interface elements
- Recent items and favorites
- Bulk/batch actions
- Power user features that don't complicate the basics
Scoring:
| Score | Criteria |
|---|---|
| 0 | One rigid path; no shortcuts or alternatives |
| 1 | Limited flexibility; few alternatives to the main path |
| 2 | Some shortcuts; basic keyboard support, limited bulk actions |
| 3 | Good accelerators; keyboard nav, some customization |
| 4 | Highly flexible; multiple paths, power features, customizable |
8. Aesthetic and Minimalist Design
Interfaces should not contain irrelevant or rarely needed information. Every element should serve a purpose.
Check for:
- Only necessary information visible at each step
- Clear visual hierarchy directing attention
- Purposeful use of color and emphasis
- No decorative clutter competing for attention
- Focused, uncluttered layouts
Scoring:
| Score | Criteria |
|---|---|
| 0 | Overwhelming; everything competes for attention equally |
| 1 | Cluttered; too much noise, hard to find what matters |
| 2 | Some clutter; main content clear, periphery noisy |
| 3 | Mostly clean; focused design, minor visual noise |
| 4 | Perfectly minimal; every element earns its pixel |
9. Help Users Recognize, Diagnose, and Recover from Errors
Error messages should use plain language, precisely indicate the problem, and constructively suggest a solution.
Check for:
- Plain language error messages (no error codes for users)
- Specific problem identification ("Email is missing @" not "Invalid input")
- Actionable recovery suggestions
- Errors displayed near the source of the problem
- Non-blocking error handling (don't wipe the form)
Scoring:
| Score | Criteria |
|---|---|
| 0 | Cryptic errors; codes, jargon, or no message at all |
| 1 | Vague errors; "Something went wrong" with no guidance |
| 2 | Clear but unhelpful; names the problem but not the fix |
| 3 | Clear with suggestions; identifies problem and offers next steps |
| 4 | Perfect recovery; pinpoints issue, suggests fix, preserves user work |
10. Help and Documentation
Even if the system is usable without docs, help should be easy to find, task-focused, and concise.
Check for:
- Searchable help or documentation
- Contextual help (tooltips, inline hints, guided tours)
- Task-focused organization (not feature-organized)
- Concise, scannable content
- Easy access without leaving current context
Scoring:
| Score | Criteria |
|---|---|
| 0 | No help available anywhere |
| 1 | Help exists but hard to find or irrelevant |
| 2 | Basic help; FAQ or docs exist, not contextual |
| 3 | Good documentation; searchable, mostly task-focused |
| 4 | Excellent contextual help; right info at the right moment |
Score Summary
Total possible: 40 points (10 heuristics × 4 max)
| Score Range | Rating | What It Means |
|---|---|---|
| 36–40 | Excellent | Minor polish only; ship it |
| 28–35 | Good | Address weak areas, solid foundation |
| 20–27 | Acceptable | Significant improvements needed before users are happy |
| 12–19 | Poor | Major UX overhaul required; core experience broken |
| 0–11 | Critical | Redesign needed; unusable in current state |
Issue Severity (P0–P3)
Tag each individual issue found during scoring with a priority level:
| Priority | Name | Description | Action |
|---|---|---|---|
| P0 | Blocking | Prevents task completion entirely | Fix immediately; this is a showstopper |
| P1 | Major | Causes significant difficulty or confusion | Fix before release |
| P2 | Minor | Annoyance, but workaround exists | Fix in next pass |
| P3 | Polish | Nice-to-fix, no real user impact | Fix if time permits |
Tip: If you're unsure between two levels, ask: "Would a user contact support about this?" If yes, it's at least P1.
Persona-Based Design Testing
Test the interface through the eyes of 5 distinct user archetypes. Each persona exposes different failure modes that a single "design director" perspective would miss.
How to use: Select 2–3 personas most relevant to the interface being critiqued. Walk through the primary user action as each persona. Report specific red flags, not generic concerns.
1. Impatient Power User: "Alex"
Profile: Expert with similar products. Expects efficiency, hates hand-holding. Will find shortcuts or leave.
Behaviors:
- Skips all onboarding and instructions
- Looks for keyboard shortcuts immediately
- Tries to bulk-select, batch-edit, and automate
- Gets frustrated by required steps that feel unnecessary
- Abandons if anything feels slow or patronizing
Test Questions:
- Can Alex complete the core task in under 60 seconds?
- Are there keyboard shortcuts for common actions?
- Can onboarding be skipped entirely?
- Do modals have keyboard dismiss (Esc)?
- Is there a "power user" path (shortcuts, bulk actions)?
Red Flags (report these specifically):
- Forced tutorials or unskippable onboarding
- No keyboard navigation for primary actions
- Slow animations that can't be skipped
- One-item-at-a-time workflows where batch would be natural
- Redundant confirmation steps for low-risk actions
2. Confused First-Timer: "Jordan"
Profile: Never used this type of product. Needs guidance at every step. Will abandon rather than figure it out.
Behaviors:
- Reads all instructions carefully
- Hesitates before clicking anything unfamiliar
- Looks for help or support constantly
- Misunderstands jargon and abbreviations
- Takes the most literal interpretation of any label
Test Questions:
- Is the first action obviously clear within 5 seconds?
- Are all icons labeled with text?
- Is there contextual help at decision points?
- Does terminology assume prior knowledge?
- Is there a clear "back" or "undo" at every step?
Red Flags (report these specifically):
- Icon-only navigation with no labels
- Technical jargon without explanation
- No visible help option or guidance
- Ambiguous next steps after completing an action
- No confirmation that an action succeeded
3. Accessibility-Dependent User: "Sam"
Profile: Uses screen reader (VoiceOver/NVDA), keyboard-only navigation. May have low vision, motor impairment, or cognitive differences.
Behaviors:
- Tabs through the interface linearly
- Relies on ARIA labels and heading structure
- Cannot see hover states or visual-only indicators
- Needs adequate color contrast (4.5:1 minimum)
- May use browser zoom up to 200%
Test Questions:
- Can the entire primary flow be completed keyboard-only?
- Are all interactive elements focusable with visible focus indicators?
- Do images have meaningful alt text?
- Is color contrast WCAG AA compliant (4.5:1 for text)?
- Does the screen reader announce state changes (loading, success, errors)?
Red Flags (report these specifically):
- Click-only interactions with no keyboard alternative
- Missing or invisible focus indicators
- Meaning conveyed by color alone (red = error, green = success)
- Unlabeled form fields or buttons
- Time-limited actions without extension option
- Custom components that break screen reader flow
4. Deliberate Stress Tester: "Riley"
Profile: Methodical user who pushes interfaces beyond the happy path. Tests edge cases, tries unexpected inputs, and probes for gaps in the experience.
Behaviors:
- Tests edge cases intentionally (empty states, long strings, special characters)
- Submits forms with unexpected data (emoji, RTL text, very long values)
- Tries to break workflows by navigating backwards, refreshing mid-flow, or opening in multiple tabs
- Looks for inconsistencies between what the UI promises and what actually happens
- Documents problems methodically
Test Questions:
- What happens at the edges (0 items, 1000 items, very long text)?
- Do error states recover gracefully or leave the UI in a broken state?
- What happens on refresh mid-workflow? Is state preserved?
- Are there features that appear to work but produce broken results?
- How does the UI handle unexpected input (emoji, special chars, paste from Excel)?
Red Flags (report these specifically):
- Features that appear to work but silently fail or produce wrong results
- Error handling that exposes technical details or leaves UI in a broken state
- Empty states that show nothing useful ("No results" with no guidance)
- Workflows that lose user data on refresh or navigation
- Inconsistent behavior between similar interactions in different parts of the UI
5. Distracted Mobile User: "Casey"
Profile: Using phone one-handed on the go. Frequently interrupted. Possibly on a slow connection.
Behaviors:
- Uses thumb only; prefers bottom-of-screen actions
- Gets interrupted mid-flow and returns later
- Switches between apps frequently
- Has limited attention span and low patience
- Types as little as possible, prefers taps and selections
Test Questions:
- Are primary actions in the thumb zone (bottom half of screen)?
- Is state preserved if the user leaves and returns?
- Does it work on slow connections (3G)?
- Can forms use autocomplete and smart defaults?
- Are touch targets at least 44×44pt?
Red Flags (report these specifically):
- Important actions positioned at the top of the screen (unreachable by thumb)
- No state persistence; progress lost on tab switch or interruption
- Large text inputs required where selection would work
- Heavy assets loading on every page (no lazy loading)
- Tiny tap targets or targets too close together
Selecting Personas
Choose personas based on the interface type:
| Interface Type | Primary Personas | Why |
|---|---|---|
| Landing page / marketing | Jordan, Riley, Casey | First impressions, trust, mobile |
| Dashboard / admin | Alex, Sam | Power users, accessibility |
| E-commerce / checkout | Casey, Riley, Jordan | Mobile, edge cases, clarity |
| Onboarding flow | Jordan, Casey | Confusion, interruption |
| Data-heavy / analytics | Alex, Sam | Efficiency, keyboard nav |
| Form-heavy / wizard | Jordan, Sam, Casey | Clarity, accessibility, mobile |
Project-Specific Personas
If {{config_file}} contains a ## Design Context section (generated by impeccable init), derive 1–2 additional personas from the audience and brand information:
- Read the target audience description
- Identify the primary user archetype not covered by the 5 predefined personas
- Create a persona following this template:
##### [Role]: "[Name]"
**Profile**: [2-3 key characteristics derived from Design Context]
**Behaviors**: [3-4 specific behaviors based on the described audience]
**Red Flags**: [3-4 things that would alienate this specific user type]
Only generate project-specific personas when real Design Context data is available. Don't invent audience details; use the 5 predefined personas when no context exists.