diff --git a/.agents/skills/audit/SKILL.md b/.agents/skills/audit/SKILL.md index 74bb05abc..1debe043e 100644 --- a/.agents/skills/audit/SKILL.md +++ b/.agents/skills/audit/SKILL.md @@ -1,17 +1,23 @@ --- name: audit -description: Perform a comprehensive audit of interface quality across accessibility, performance, theming, and responsive design. Generates a scored report with severity ratings and actionable plan. Use when the user wants a design review, accessibility check, quality audit, or a full list of UI issues to fix. +description: Run technical quality checks across accessibility, performance, theming, responsive design, and anti-patterns. Generates a scored report with P0-P3 severity ratings and actionable plan. Use when the user wants an accessibility check, performance audit, or technical quality review. user-invocable: true argument-hint: "[area (feature, page, component...)]" --- -Run systematic quality checks and generate a comprehensive audit report with quantitative scoring, prioritized issues, and an actionable plan. Don't fix issues — document them for other commands to address. +## MANDATORY PREPARATION -**First**: Invoke /frontend-design for design principles and anti-patterns. +Invoke /frontend-design — it contains design principles, anti-patterns, and the **Context Gathering Protocol**. Follow the protocol before proceeding — if no design context exists yet, you MUST run /teach-impeccable first. + +--- + +Run systematic **technical** quality checks and generate a comprehensive report. Don't fix issues — document them for other commands to address. + +This is a code-level audit, not a design critique. Check what's measurable and verifiable in the implementation. ## Diagnostic Scan -Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using the criteria below. +Run comprehensive checks across 5 dimensions. Score each dimension 0-4 using the criteria below. ### 1. Accessibility (A11y) @@ -23,7 +29,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Alt text**: Missing or poor image descriptions - **Form issues**: Inputs without labels, poor error messaging, missing required indicators -**Score 0–4**: 0=Inaccessible (fails WCAG A), 1=Major gaps (few ARIA labels, no keyboard nav), 2=Partial (some a11y effort, significant gaps), 3=Good (WCAG AA mostly met, minor gaps), 4=Excellent (WCAG AA fully met, approaches AAA) +**Score 0-4**: 0=Inaccessible (fails WCAG A), 1=Major gaps (few ARIA labels, no keyboard nav), 2=Partial (some a11y effort, significant gaps), 3=Good (WCAG AA mostly met, minor gaps), 4=Excellent (WCAG AA fully met, approaches AAA) ### 2. Performance @@ -34,7 +40,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Bundle size**: Unnecessary imports, unused dependencies - **Render performance**: Unnecessary re-renders, missing memoization -**Score 0–4**: 0=Severe issues (layout thrash, unoptimized everything), 1=Major problems (no lazy loading, expensive animations), 2=Partial (some optimization, gaps remain), 3=Good (mostly optimized, minor improvements possible), 4=Excellent (fast, lean, well-optimized) +**Score 0-4**: 0=Severe issues (layout thrash, unoptimized everything), 1=Major problems (no lazy loading, expensive animations), 2=Partial (some optimization, gaps remain), 3=Good (mostly optimized, minor improvements possible), 4=Excellent (fast, lean, well-optimized) ### 3. Theming @@ -44,7 +50,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Inconsistent tokens**: Using wrong tokens, mixing token types - **Theme switching issues**: Values that don't update on theme change -**Score 0–4**: 0=No theming (hard-coded everything), 1=Minimal tokens (mostly hard-coded), 2=Partial (tokens exist but inconsistently used), 3=Good (tokens used, minor hard-coded values), 4=Excellent (full token system, dark mode works perfectly) +**Score 0-4**: 0=No theming (hard-coded everything), 1=Minimal tokens (mostly hard-coded), 2=Partial (tokens exist but inconsistently used), 3=Good (tokens used, minor hard-coded values), 4=Excellent (full token system, dark mode works perfectly) ### 4. Responsive Design @@ -55,113 +61,73 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Text scaling**: Layouts that break when text size increases - **Missing breakpoints**: No mobile/tablet variants -**Score 0–4**: 0=Desktop-only (breaks on mobile), 1=Major issues (some breakpoints, many failures), 2=Partial (works on mobile, rough edges), 3=Good (responsive, minor touch target or overflow issues), 4=Excellent (fluid, all viewports, proper touch targets) +**Score 0-4**: 0=Desktop-only (breaks on mobile), 1=Major issues (some breakpoints, many failures), 2=Partial (works on mobile, rough edges), 3=Good (responsive, minor touch target or overflow issues), 4=Excellent (fluid, all viewports, proper touch targets) ### 5. Anti-Patterns (CRITICAL) Check against ALL the **DON'T** guidelines in the frontend-design skill. Look for AI slop tells (AI color palette, gradient text, glassmorphism, hero metrics, card grids, generic fonts) and general design anti-patterns (gray on color, nested cards, bounce easing, redundant copy). -**Score 0–4**: 0=AI slop gallery (5+ tells), 1=Heavy AI aesthetic (3-4 tells), 2=Some tells (1-2 noticeable), 3=Mostly clean (subtle issues only), 4=No AI tells (distinctive, intentional design) +**Score 0-4**: 0=AI slop gallery (5+ tells), 1=Heavy AI aesthetic (3-4 tells), 2=Some tells (1-2 noticeable), 3=Mostly clean (subtle issues only), 4=No AI tells (distinctive, intentional design) -**CRITICAL**: This is an audit, not a fix. Document issues thoroughly with clear explanations of impact. Use other commands to fix issues after audit. - -## Generate Comprehensive Report +## Generate Report ### Audit Health Score -Present the dimension scores as a table: - | # | Dimension | Score | Key Finding | |---|-----------|-------|-------------| -| 1 | Accessibility | ? | [most critical a11y issue or "—"] | +| 1 | Accessibility | ? | [most critical a11y issue or "--"] | | 2 | Performance | ? | | | 3 | Responsive Design | ? | | | 4 | Theming | ? | | | 5 | Anti-Patterns | ? | | | **Total** | | **??/20** | **[Rating band]** | -**Rating bands**: -| Score | Rating | Action | -|-------|--------|--------| -| 18–20 | Excellent | Minor polish only | -| 14–17 | Good | Address weak dimensions | -| 10–13 | Acceptable | Significant work needed | -| 6–9 | Poor | Major quality overhaul | -| 0–5 | Critical | Fundamental issues across the board | +**Rating bands**: 18-20 Excellent (minor polish), 14-17 Good (address weak dimensions), 10-13 Acceptable (significant work needed), 6-9 Poor (major overhaul), 0-5 Critical (fundamental issues) ### Anti-Patterns Verdict -**Start here.** Pass/fail: Does this look AI-generated? List specific tells from the skill's Anti-Patterns section. Be brutally honest. +**Start here.** Pass/fail: Does this look AI-generated? List specific tells. Be brutally honest. ### Executive Summary - Audit Health Score: **??/20** ([rating band]) - Total issues found (count by severity: P0/P1/P2/P3) -- Most critical issues (top 3-5) +- Top 3-5 critical issues - Recommended next steps ### Detailed Findings by Severity -Tag every issue with **P0–P3 severity**: -| Priority | Name | Description | -|----------|------|-------------| -| **P0** | Blocking | Prevents task completion — fix immediately | -| **P1** | Major | Significant difficulty or WCAG AA violation — fix before release | -| **P2** | Minor | Annoyance, workaround exists — fix in next pass | -| **P3** | Polish | Nice-to-fix, no real user impact — fix if time permits | +Tag every issue with **P0-P3 severity**: +- **P0 Blocking**: Prevents task completion — fix immediately +- **P1 Major**: Significant difficulty or WCAG AA violation — fix before release +- **P2 Minor**: Annoyance, workaround exists — fix in next pass +- **P3 Polish**: Nice-to-fix, no real user impact — fix if time permits For each issue, document: - **[P?] Issue name** -- **Location**: Where it occurs (component, file, line) +- **Location**: Component, file, line - **Category**: Accessibility / Performance / Theming / Responsive / Anti-Pattern -- **Description**: What the issue is - **Impact**: How it affects users - **WCAG/Standard**: Which standard it violates (if applicable) - **Recommendation**: How to fix it -- **Suggested command**: Which command to use (prefer: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive — or other installed skills you're sure exist) - -#### P0 — Blocking Issues -[Issues that prevent task completion or violate WCAG A] - -#### P1 — Major Issues -[Significant usability/accessibility impact, WCAG AA violations] - -#### P2 — Minor Issues -[Quality issues, WCAG AAA violations, performance concerns] - -#### P3 — Polish Issues -[Minor inconsistencies, optimization opportunities] +- **Suggested command**: Which command to use (prefer: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive) ### Patterns & Systemic Issues -Identify recurring problems: +Identify recurring problems that indicate systemic gaps rather than one-off mistakes: - "Hard-coded colors appear in 15+ components, should use design tokens" - "Touch targets consistently too small (<44px) throughout mobile experience" -- "Missing focus indicators on all custom interactive components" ### Positive Findings -Note what's working well: -- Good practices to maintain -- Exemplary implementations to replicate elsewhere +Note what's working well — good practices to maintain and replicate. ## Recommended Actions -Present a prioritized action summary. Order is determined by severity automatically (P0 first, then P1, then P2). - -### Action Summary - -List recommended commands in priority order: +List recommended commands in priority order (P0 first, then P1, then P2): 1. **[P?] `/command-name`** — Brief description (specific context from audit findings) 2. **[P?] `/command-name`** — Brief description (specific context) -... -**Rules for recommendations**: -- Only recommend commands from: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive -- Order by severity: P0 issues first, then P1, then P2 (skip P3 unless user has few issues) -- Each item's description should carry enough context that the command knows what to focus on -- Map findings to the most appropriate command -- Skip commands that would address zero issues -- End with `/polish` as the final step if any fixes were recommended +**Rules**: Only recommend commands from: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive. Map findings to the most appropriate command. End with `/polish` as the final step if any fixes were recommended. After presenting the summary, tell the user: @@ -173,10 +139,9 @@ After presenting the summary, tell the user: **NEVER**: - Report issues without explaining impact (why does this matter?) -- Mix severity levels inconsistently -- Skip positive findings (celebrate what works) - Provide generic recommendations (be specific and actionable) +- Skip positive findings (celebrate what works) - Forget to prioritize (everything can't be P0) - Report false positives without verification -Remember: You're a quality auditor with exceptional attention to detail. Document systematically, prioritize ruthlessly, and provide clear paths to improvement. A good audit makes fixing easy. \ No newline at end of file +Remember: You're a technical quality auditor. Document systematically, prioritize ruthlessly, cite specific code locations, and provide clear paths to improvement. \ No newline at end of file diff --git a/.agents/skills/critique/SKILL.md b/.agents/skills/critique/SKILL.md index 9e1322a33..70ac82f1d 100644 --- a/.agents/skills/critique/SKILL.md +++ b/.agents/skills/critique/SKILL.md @@ -1,6 +1,6 @@ --- name: critique -description: Evaluate design effectiveness from a UX perspective. Assesses visual hierarchy, information architecture, emotional resonance, cognitive load, and overall design quality with quantitative scoring, persona-based testing, and actionable feedback. Use when the user asks to review, critique, evaluate, or give feedback on a design, UI, or component. +description: Evaluate design from a UX perspective, assessing visual hierarchy, information architecture, emotional resonance, cognitive load, and overall quality with quantitative scoring, persona-based testing, and actionable feedback. Use when the user asks to review, critique, evaluate, or give feedback on a design or component. user-invocable: true argument-hint: "[area (feature, page, component...)]" --- @@ -11,7 +11,7 @@ Invoke /frontend-design — it contains design principles, anti-patterns, and th --- -Conduct a holistic design critique, evaluating whether the interface actually works—not just technically, but as a designed experience. Think like a design director giving feedback. +Conduct a holistic design critique, evaluating whether the interface actually works — not just technically, but as a designed experience. Think like a design director giving feedback. ## Phase 1: Design Critique @@ -21,7 +21,7 @@ Evaluate the interface across these dimensions: **This is the most important check.** Does this look like every other AI-generated interface from 2024-2025? -Review the design against ALL the **DON'T** guidelines in the frontend-design skill—they are the fingerprints of AI-generated work. Check for the AI color palette, gradient text, dark mode with glowing accents, glassmorphism, hero metric layouts, identical card grids, generic fonts, and all other tells. +Review the design against ALL the **DON'T** guidelines in the frontend-design skill — they are the fingerprints of AI-generated work. Check for the AI color palette, gradient text, dark mode with glowing accents, glassmorphism, hero metric layouts, identical card grids, generic fonts, and all other tells. **The test**: If you showed this to someone and said "AI made this," would they believe you immediately? If yes, that's the problem. @@ -31,20 +31,19 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Do size, color, and position communicate importance correctly? - Is there visual competition between elements that should have different weights? -### 3. Information Architecture -→ *Consult [cognitive-load](reference/cognitive-load.md) for the working memory rule and checklist* +### 3. Information Architecture & Cognitive Load +> *Consult [cognitive-load](reference/cognitive-load.md) for the working memory rule and 8-item checklist* - Is the structure intuitive? Would a new user understand the organization? - Is related content grouped logically? - Are there too many choices at once? Count visible options at each decision point — if >4, flag it - Is the navigation clear and predictable? - **Progressive disclosure**: Is complexity revealed only when needed, or dumped on the user upfront? -- **Cognitive load sub-check**: Run the 8-item cognitive load checklist from the reference. Report the number of failures. +- **Run the 8-item cognitive load checklist** from the reference. Report failure count: 0–1 = low (good), 2–3 = moderate, 4+ = critical. ### 4. Emotional Journey -→ *Consult [cognitive-load](reference/cognitive-load.md) for emotional intervention patterns* - What emotion does this interface evoke? Is that intentional? - Does it match the brand personality? -- Does it feel trustworthy, approachable, premium, playful—whatever it should feel? +- Does it feel trustworthy, approachable, premium, playful — whatever it should feel? - Would the target user feel "this is for me"? - **Peak-end rule**: Is the most intense moment positive? Does the experience end well (confirmation, celebration, clear next step)? - **Emotional valleys**: Check for onboarding frustration, error cliffs, feature discovery gaps, or anxiety spikes at high-stakes moments (payment, delete, commit) @@ -72,7 +71,7 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Is color used to communicate, not just decorate? - Does the palette feel cohesive? - Are accent colors drawing attention to the right things? -- Does it work for colorblind users? (not just technically—does meaning still come through?) +- Does it work for colorblind users? (not just technically — does meaning still come through?) ### 9. States & Edge Cases - Empty states: Do they guide users toward action, or just say "nothing here"? @@ -86,27 +85,19 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Are labels and buttons unambiguous? - Does error copy help users fix the problem? -### 11. Cognitive Load -→ *Consult [cognitive-load](reference/cognitive-load.md)* -- **Intrinsic vs. extraneous**: Is the mental effort coming from the task itself (acceptable) or from poor design choices (eliminate)? -- **Decision points**: Count visible choices at key moments. More than 4 simultaneous options = overload. -- **Working memory burden**: Does the user need to remember information from a previous screen to act on the current one? -- **Information chunking**: Is content broken into digestible groups, or presented as undifferentiated walls? -- Run the 8-item cognitive load checklist. Report failures count: 0–1 = low (good), 2–3 = moderate, 4+ = critical. - ## Phase 2: Present Findings Structure your feedback as a design director would: ### Design Health Score -→ *Consult [heuristics-scoring](reference/heuristics-scoring.md)* +> *Consult [heuristics-scoring](reference/heuristics-scoring.md)* Score each of Nielsen's 10 heuristics 0–4. Present as a table: | # | Heuristic | Score | Key Issue | |---|-----------|-------|-----------| | 1 | Visibility of System Status | ? | [specific finding or "—" if solid] | -| 2 | Match System ↔ Real World | ? | | +| 2 | Match System / Real World | ? | | | 3 | User Control and Freedom | ? | | | 4 | Consistency and Standards | ? | | | 5 | Error Prevention | ? | | @@ -123,13 +114,13 @@ Be honest with scores. A 4 means genuinely excellent. Most real interfaces score **Start here.** Pass/fail: Does this look AI-generated? List specific tells from the skill's Anti-Patterns section. Be brutally honest. ### Overall Impression -A brief gut reaction—what works, what doesn't, and the single biggest opportunity. +A brief gut reaction — what works, what doesn't, and the single biggest opportunity. ### What's Working -Highlight 2-3 things done well. Be specific about why they work. +Highlight 2–3 things done well. Be specific about why they work. ### Priority Issues -The 3-5 most impactful design problems, ordered by importance. +The 3–5 most impactful design problems, ordered by importance. For each issue, tag with **P0–P3 severity** (consult [heuristics-scoring](reference/heuristics-scoring.md) for severity definitions): - **[P?] What**: Name the problem clearly @@ -138,7 +129,7 @@ For each issue, tag with **P0–P3 severity** (consult [heuristics-scoring](refe - **Suggested command**: Which command could address this (from: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive) ### Persona Red Flags -→ *Consult [personas](reference/personas.md)* +> *Consult [personas](reference/personas.md)* Auto-select 2–3 personas most relevant to this interface type (use the selection table in the reference). If `.github/copilot-instructions.md` contains a `## Design Context` section from `teach-impeccable`, also generate 1–2 project-specific personas from the audience/brand info. @@ -154,12 +145,12 @@ Be specific — name the exact elements and interactions that fail each persona. Quick notes on smaller issues worth addressing. **Remember**: -- Be direct—vague feedback wastes everyone's time -- Be specific—"the submit button" not "some elements" +- Be direct — vague feedback wastes everyone's time +- Be specific — "the submit button" not "some elements" - Say what's wrong AND why it matters to users - Give concrete suggestions, not just "consider exploring..." -- Prioritize ruthlessly—if everything is important, nothing is -- Don't soften criticism—developers need honest feedback to ship great design +- Prioritize ruthlessly — if everything is important, nothing is +- Don't soften criticism — developers need honest feedback to ship great design ## Phase 3: Ask the User @@ -167,9 +158,9 @@ Quick notes on smaller issues worth addressing. Ask questions along these lines (adapt to the specific findings — do NOT ask generic questions): -1. **Priority direction**: Based on the issues found, ask which category matters most to the user right now. For example: "I found problems with visual hierarchy, color usage, and information overload. Which area should we tackle first?" Offer the top 2-3 issue categories as options. +1. **Priority direction**: Based on the issues found, ask which category matters most to the user right now. For example: "I found problems with visual hierarchy, color usage, and information overload. Which area should we tackle first?" Offer the top 2–3 issue categories as options. -2. **Design intent**: If the critique found a tonal mismatch, ask whether it was intentional. For example: "The interface feels clinical and corporate. Is that the intended tone, or should it feel warmer/bolder/more playful?" Offer 2-3 tonal directions as options based on what would fix the issues found. +2. **Design intent**: If the critique found a tonal mismatch, ask whether it was intentional. For example: "The interface feels clinical and corporate. Is that the intended tone, or should it feel warmer/bolder/more playful?" Offer 2–3 tonal directions as options based on what would fix the issues found. 3. **Scope**: Ask how much the user wants to take on. For example: "I found N issues. Want to address everything, or focus on the top 3?" Offer scope options like "Top 3 only", "All issues", "Critical issues only". @@ -177,9 +168,9 @@ Ask questions along these lines (adapt to the specific findings — do NOT ask g **Rules for questions**: - Every question must reference specific findings from Phase 2 — never ask generic "who is your audience?" questions -- Keep it to 2-4 questions maximum — respect the user's time +- Keep it to 2–4 questions maximum — respect the user's time - Offer concrete options, not open-ended prompts -- If findings are straightforward (e.g., only 1-2 clear issues), skip questions and go directly to Phase 4 +- If findings are straightforward (e.g., only 1–2 clear issues), skip questions and go directly to Phase 4 ## Phase 4: Recommended Actions diff --git a/.agents/skills/critique/reference/personas.md b/.agents/skills/critique/reference/personas.md index c91c78f74..eb2f2b6d1 100644 --- a/.agents/skills/critique/reference/personas.md +++ b/.agents/skills/critique/reference/personas.md @@ -88,30 +88,30 @@ Test the interface through the eyes of 5 distinct user archetypes. Each persona --- -## 4. Skeptical Evaluator — "Riley" +## 4. Deliberate Stress Tester — "Riley" -**Profile**: Evaluating the product for their team or company. Looking for reasons to reject. Comparing against competitors. +**Profile**: Methodical user who pushes interfaces beyond the happy path. Tests edge cases, tries unexpected inputs, and probes for gaps in the experience. **Behaviors**: - Tests edge cases intentionally (empty states, long strings, special characters) -- Looks for pricing catches and hidden limitations -- Reads fine print and terms of service -- Tries to break things deliberately +- Submits forms with unexpected data (emoji, RTL text, very long values) +- Tries to break workflows by navigating backwards, refreshing mid-flow, or opening in multiple tabs +- Looks for inconsistencies between what the UI promises and what actually happens - Documents problems methodically **Test Questions**: - What happens at the edges (0 items, 1000 items, very long text)? -- Is pricing and value proposition transparent? -- Are there hidden limitations or gotchas? -- How polished is error handling? -- What data is collected and why? +- Do error states recover gracefully or leave the UI in a broken state? +- What happens on refresh mid-workflow? Is state preserved? +- Are there features that appear to work but produce broken results? +- How does the UI handle unexpected input (emoji, special chars, paste from Excel)? **Red Flags** (report these specifically): -- Hidden pricing or "contact sales" for basic information -- Features that appear to work but produce broken results -- Poor error handling that exposes technical details -- Unclear data practices or missing privacy information +- Features that appear to work but silently fail or produce wrong results +- Error handling that exposes technical details or leaves UI in a broken state - Empty states that show nothing useful ("No results" with no guidance) +- Workflows that lose user data on refresh or navigation +- Inconsistent behavior between similar interactions in different parts of the UI --- @@ -150,7 +150,7 @@ Choose personas based on the interface type: |---------------|-----------------|-----| | Landing page / marketing | Jordan, Riley, Casey | First impressions, trust, mobile | | Dashboard / admin | Alex, Sam | Power users, accessibility | -| E-commerce / checkout | Casey, Riley, Jordan | Mobile, trust, clarity | +| E-commerce / checkout | Casey, Riley, Jordan | Mobile, edge cases, clarity | | Onboarding flow | Jordan, Casey | Confusion, interruption | | Data-heavy / analytics | Alex, Sam | Efficiency, keyboard nav | | Form-heavy / wizard | Jordan, Sam, Casey | Clarity, accessibility, mobile | diff --git a/.claude/skills/audit/SKILL.md b/.claude/skills/audit/SKILL.md index 74bb05abc..1debe043e 100644 --- a/.claude/skills/audit/SKILL.md +++ b/.claude/skills/audit/SKILL.md @@ -1,17 +1,23 @@ --- name: audit -description: Perform a comprehensive audit of interface quality across accessibility, performance, theming, and responsive design. Generates a scored report with severity ratings and actionable plan. Use when the user wants a design review, accessibility check, quality audit, or a full list of UI issues to fix. +description: Run technical quality checks across accessibility, performance, theming, responsive design, and anti-patterns. Generates a scored report with P0-P3 severity ratings and actionable plan. Use when the user wants an accessibility check, performance audit, or technical quality review. user-invocable: true argument-hint: "[area (feature, page, component...)]" --- -Run systematic quality checks and generate a comprehensive audit report with quantitative scoring, prioritized issues, and an actionable plan. Don't fix issues — document them for other commands to address. +## MANDATORY PREPARATION -**First**: Invoke /frontend-design for design principles and anti-patterns. +Invoke /frontend-design — it contains design principles, anti-patterns, and the **Context Gathering Protocol**. Follow the protocol before proceeding — if no design context exists yet, you MUST run /teach-impeccable first. + +--- + +Run systematic **technical** quality checks and generate a comprehensive report. Don't fix issues — document them for other commands to address. + +This is a code-level audit, not a design critique. Check what's measurable and verifiable in the implementation. ## Diagnostic Scan -Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using the criteria below. +Run comprehensive checks across 5 dimensions. Score each dimension 0-4 using the criteria below. ### 1. Accessibility (A11y) @@ -23,7 +29,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Alt text**: Missing or poor image descriptions - **Form issues**: Inputs without labels, poor error messaging, missing required indicators -**Score 0–4**: 0=Inaccessible (fails WCAG A), 1=Major gaps (few ARIA labels, no keyboard nav), 2=Partial (some a11y effort, significant gaps), 3=Good (WCAG AA mostly met, minor gaps), 4=Excellent (WCAG AA fully met, approaches AAA) +**Score 0-4**: 0=Inaccessible (fails WCAG A), 1=Major gaps (few ARIA labels, no keyboard nav), 2=Partial (some a11y effort, significant gaps), 3=Good (WCAG AA mostly met, minor gaps), 4=Excellent (WCAG AA fully met, approaches AAA) ### 2. Performance @@ -34,7 +40,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Bundle size**: Unnecessary imports, unused dependencies - **Render performance**: Unnecessary re-renders, missing memoization -**Score 0–4**: 0=Severe issues (layout thrash, unoptimized everything), 1=Major problems (no lazy loading, expensive animations), 2=Partial (some optimization, gaps remain), 3=Good (mostly optimized, minor improvements possible), 4=Excellent (fast, lean, well-optimized) +**Score 0-4**: 0=Severe issues (layout thrash, unoptimized everything), 1=Major problems (no lazy loading, expensive animations), 2=Partial (some optimization, gaps remain), 3=Good (mostly optimized, minor improvements possible), 4=Excellent (fast, lean, well-optimized) ### 3. Theming @@ -44,7 +50,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Inconsistent tokens**: Using wrong tokens, mixing token types - **Theme switching issues**: Values that don't update on theme change -**Score 0–4**: 0=No theming (hard-coded everything), 1=Minimal tokens (mostly hard-coded), 2=Partial (tokens exist but inconsistently used), 3=Good (tokens used, minor hard-coded values), 4=Excellent (full token system, dark mode works perfectly) +**Score 0-4**: 0=No theming (hard-coded everything), 1=Minimal tokens (mostly hard-coded), 2=Partial (tokens exist but inconsistently used), 3=Good (tokens used, minor hard-coded values), 4=Excellent (full token system, dark mode works perfectly) ### 4. Responsive Design @@ -55,113 +61,73 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Text scaling**: Layouts that break when text size increases - **Missing breakpoints**: No mobile/tablet variants -**Score 0–4**: 0=Desktop-only (breaks on mobile), 1=Major issues (some breakpoints, many failures), 2=Partial (works on mobile, rough edges), 3=Good (responsive, minor touch target or overflow issues), 4=Excellent (fluid, all viewports, proper touch targets) +**Score 0-4**: 0=Desktop-only (breaks on mobile), 1=Major issues (some breakpoints, many failures), 2=Partial (works on mobile, rough edges), 3=Good (responsive, minor touch target or overflow issues), 4=Excellent (fluid, all viewports, proper touch targets) ### 5. Anti-Patterns (CRITICAL) Check against ALL the **DON'T** guidelines in the frontend-design skill. Look for AI slop tells (AI color palette, gradient text, glassmorphism, hero metrics, card grids, generic fonts) and general design anti-patterns (gray on color, nested cards, bounce easing, redundant copy). -**Score 0–4**: 0=AI slop gallery (5+ tells), 1=Heavy AI aesthetic (3-4 tells), 2=Some tells (1-2 noticeable), 3=Mostly clean (subtle issues only), 4=No AI tells (distinctive, intentional design) +**Score 0-4**: 0=AI slop gallery (5+ tells), 1=Heavy AI aesthetic (3-4 tells), 2=Some tells (1-2 noticeable), 3=Mostly clean (subtle issues only), 4=No AI tells (distinctive, intentional design) -**CRITICAL**: This is an audit, not a fix. Document issues thoroughly with clear explanations of impact. Use other commands to fix issues after audit. - -## Generate Comprehensive Report +## Generate Report ### Audit Health Score -Present the dimension scores as a table: - | # | Dimension | Score | Key Finding | |---|-----------|-------|-------------| -| 1 | Accessibility | ? | [most critical a11y issue or "—"] | +| 1 | Accessibility | ? | [most critical a11y issue or "--"] | | 2 | Performance | ? | | | 3 | Responsive Design | ? | | | 4 | Theming | ? | | | 5 | Anti-Patterns | ? | | | **Total** | | **??/20** | **[Rating band]** | -**Rating bands**: -| Score | Rating | Action | -|-------|--------|--------| -| 18–20 | Excellent | Minor polish only | -| 14–17 | Good | Address weak dimensions | -| 10–13 | Acceptable | Significant work needed | -| 6–9 | Poor | Major quality overhaul | -| 0–5 | Critical | Fundamental issues across the board | +**Rating bands**: 18-20 Excellent (minor polish), 14-17 Good (address weak dimensions), 10-13 Acceptable (significant work needed), 6-9 Poor (major overhaul), 0-5 Critical (fundamental issues) ### Anti-Patterns Verdict -**Start here.** Pass/fail: Does this look AI-generated? List specific tells from the skill's Anti-Patterns section. Be brutally honest. +**Start here.** Pass/fail: Does this look AI-generated? List specific tells. Be brutally honest. ### Executive Summary - Audit Health Score: **??/20** ([rating band]) - Total issues found (count by severity: P0/P1/P2/P3) -- Most critical issues (top 3-5) +- Top 3-5 critical issues - Recommended next steps ### Detailed Findings by Severity -Tag every issue with **P0–P3 severity**: -| Priority | Name | Description | -|----------|------|-------------| -| **P0** | Blocking | Prevents task completion — fix immediately | -| **P1** | Major | Significant difficulty or WCAG AA violation — fix before release | -| **P2** | Minor | Annoyance, workaround exists — fix in next pass | -| **P3** | Polish | Nice-to-fix, no real user impact — fix if time permits | +Tag every issue with **P0-P3 severity**: +- **P0 Blocking**: Prevents task completion — fix immediately +- **P1 Major**: Significant difficulty or WCAG AA violation — fix before release +- **P2 Minor**: Annoyance, workaround exists — fix in next pass +- **P3 Polish**: Nice-to-fix, no real user impact — fix if time permits For each issue, document: - **[P?] Issue name** -- **Location**: Where it occurs (component, file, line) +- **Location**: Component, file, line - **Category**: Accessibility / Performance / Theming / Responsive / Anti-Pattern -- **Description**: What the issue is - **Impact**: How it affects users - **WCAG/Standard**: Which standard it violates (if applicable) - **Recommendation**: How to fix it -- **Suggested command**: Which command to use (prefer: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive — or other installed skills you're sure exist) - -#### P0 — Blocking Issues -[Issues that prevent task completion or violate WCAG A] - -#### P1 — Major Issues -[Significant usability/accessibility impact, WCAG AA violations] - -#### P2 — Minor Issues -[Quality issues, WCAG AAA violations, performance concerns] - -#### P3 — Polish Issues -[Minor inconsistencies, optimization opportunities] +- **Suggested command**: Which command to use (prefer: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive) ### Patterns & Systemic Issues -Identify recurring problems: +Identify recurring problems that indicate systemic gaps rather than one-off mistakes: - "Hard-coded colors appear in 15+ components, should use design tokens" - "Touch targets consistently too small (<44px) throughout mobile experience" -- "Missing focus indicators on all custom interactive components" ### Positive Findings -Note what's working well: -- Good practices to maintain -- Exemplary implementations to replicate elsewhere +Note what's working well — good practices to maintain and replicate. ## Recommended Actions -Present a prioritized action summary. Order is determined by severity automatically (P0 first, then P1, then P2). - -### Action Summary - -List recommended commands in priority order: +List recommended commands in priority order (P0 first, then P1, then P2): 1. **[P?] `/command-name`** — Brief description (specific context from audit findings) 2. **[P?] `/command-name`** — Brief description (specific context) -... -**Rules for recommendations**: -- Only recommend commands from: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive -- Order by severity: P0 issues first, then P1, then P2 (skip P3 unless user has few issues) -- Each item's description should carry enough context that the command knows what to focus on -- Map findings to the most appropriate command -- Skip commands that would address zero issues -- End with `/polish` as the final step if any fixes were recommended +**Rules**: Only recommend commands from: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive. Map findings to the most appropriate command. End with `/polish` as the final step if any fixes were recommended. After presenting the summary, tell the user: @@ -173,10 +139,9 @@ After presenting the summary, tell the user: **NEVER**: - Report issues without explaining impact (why does this matter?) -- Mix severity levels inconsistently -- Skip positive findings (celebrate what works) - Provide generic recommendations (be specific and actionable) +- Skip positive findings (celebrate what works) - Forget to prioritize (everything can't be P0) - Report false positives without verification -Remember: You're a quality auditor with exceptional attention to detail. Document systematically, prioritize ruthlessly, and provide clear paths to improvement. A good audit makes fixing easy. \ No newline at end of file +Remember: You're a technical quality auditor. Document systematically, prioritize ruthlessly, cite specific code locations, and provide clear paths to improvement. \ No newline at end of file diff --git a/.claude/skills/critique/SKILL.md b/.claude/skills/critique/SKILL.md index 7acdc2b45..2c8e073ed 100644 --- a/.claude/skills/critique/SKILL.md +++ b/.claude/skills/critique/SKILL.md @@ -1,6 +1,6 @@ --- name: critique -description: Evaluate design effectiveness from a UX perspective. Assesses visual hierarchy, information architecture, emotional resonance, cognitive load, and overall design quality with quantitative scoring, persona-based testing, and actionable feedback. Use when the user asks to review, critique, evaluate, or give feedback on a design, UI, or component. +description: Evaluate design from a UX perspective, assessing visual hierarchy, information architecture, emotional resonance, cognitive load, and overall quality with quantitative scoring, persona-based testing, and actionable feedback. Use when the user asks to review, critique, evaluate, or give feedback on a design or component. user-invocable: true argument-hint: "[area (feature, page, component...)]" --- @@ -11,7 +11,7 @@ Invoke /frontend-design — it contains design principles, anti-patterns, and th --- -Conduct a holistic design critique, evaluating whether the interface actually works—not just technically, but as a designed experience. Think like a design director giving feedback. +Conduct a holistic design critique, evaluating whether the interface actually works — not just technically, but as a designed experience. Think like a design director giving feedback. ## Phase 1: Design Critique @@ -21,7 +21,7 @@ Evaluate the interface across these dimensions: **This is the most important check.** Does this look like every other AI-generated interface from 2024-2025? -Review the design against ALL the **DON'T** guidelines in the frontend-design skill—they are the fingerprints of AI-generated work. Check for the AI color palette, gradient text, dark mode with glowing accents, glassmorphism, hero metric layouts, identical card grids, generic fonts, and all other tells. +Review the design against ALL the **DON'T** guidelines in the frontend-design skill — they are the fingerprints of AI-generated work. Check for the AI color palette, gradient text, dark mode with glowing accents, glassmorphism, hero metric layouts, identical card grids, generic fonts, and all other tells. **The test**: If you showed this to someone and said "AI made this," would they believe you immediately? If yes, that's the problem. @@ -31,20 +31,19 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Do size, color, and position communicate importance correctly? - Is there visual competition between elements that should have different weights? -### 3. Information Architecture -→ *Consult [cognitive-load](reference/cognitive-load.md) for the working memory rule and checklist* +### 3. Information Architecture & Cognitive Load +> *Consult [cognitive-load](reference/cognitive-load.md) for the working memory rule and 8-item checklist* - Is the structure intuitive? Would a new user understand the organization? - Is related content grouped logically? - Are there too many choices at once? Count visible options at each decision point — if >4, flag it - Is the navigation clear and predictable? - **Progressive disclosure**: Is complexity revealed only when needed, or dumped on the user upfront? -- **Cognitive load sub-check**: Run the 8-item cognitive load checklist from the reference. Report the number of failures. +- **Run the 8-item cognitive load checklist** from the reference. Report failure count: 0–1 = low (good), 2–3 = moderate, 4+ = critical. ### 4. Emotional Journey -→ *Consult [cognitive-load](reference/cognitive-load.md) for emotional intervention patterns* - What emotion does this interface evoke? Is that intentional? - Does it match the brand personality? -- Does it feel trustworthy, approachable, premium, playful—whatever it should feel? +- Does it feel trustworthy, approachable, premium, playful — whatever it should feel? - Would the target user feel "this is for me"? - **Peak-end rule**: Is the most intense moment positive? Does the experience end well (confirmation, celebration, clear next step)? - **Emotional valleys**: Check for onboarding frustration, error cliffs, feature discovery gaps, or anxiety spikes at high-stakes moments (payment, delete, commit) @@ -72,7 +71,7 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Is color used to communicate, not just decorate? - Does the palette feel cohesive? - Are accent colors drawing attention to the right things? -- Does it work for colorblind users? (not just technically—does meaning still come through?) +- Does it work for colorblind users? (not just technically — does meaning still come through?) ### 9. States & Edge Cases - Empty states: Do they guide users toward action, or just say "nothing here"? @@ -86,27 +85,19 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Are labels and buttons unambiguous? - Does error copy help users fix the problem? -### 11. Cognitive Load -→ *Consult [cognitive-load](reference/cognitive-load.md)* -- **Intrinsic vs. extraneous**: Is the mental effort coming from the task itself (acceptable) or from poor design choices (eliminate)? -- **Decision points**: Count visible choices at key moments. More than 4 simultaneous options = overload. -- **Working memory burden**: Does the user need to remember information from a previous screen to act on the current one? -- **Information chunking**: Is content broken into digestible groups, or presented as undifferentiated walls? -- Run the 8-item cognitive load checklist. Report failures count: 0–1 = low (good), 2–3 = moderate, 4+ = critical. - ## Phase 2: Present Findings Structure your feedback as a design director would: ### Design Health Score -→ *Consult [heuristics-scoring](reference/heuristics-scoring.md)* +> *Consult [heuristics-scoring](reference/heuristics-scoring.md)* Score each of Nielsen's 10 heuristics 0–4. Present as a table: | # | Heuristic | Score | Key Issue | |---|-----------|-------|-----------| | 1 | Visibility of System Status | ? | [specific finding or "—" if solid] | -| 2 | Match System ↔ Real World | ? | | +| 2 | Match System / Real World | ? | | | 3 | User Control and Freedom | ? | | | 4 | Consistency and Standards | ? | | | 5 | Error Prevention | ? | | @@ -123,13 +114,13 @@ Be honest with scores. A 4 means genuinely excellent. Most real interfaces score **Start here.** Pass/fail: Does this look AI-generated? List specific tells from the skill's Anti-Patterns section. Be brutally honest. ### Overall Impression -A brief gut reaction—what works, what doesn't, and the single biggest opportunity. +A brief gut reaction — what works, what doesn't, and the single biggest opportunity. ### What's Working -Highlight 2-3 things done well. Be specific about why they work. +Highlight 2–3 things done well. Be specific about why they work. ### Priority Issues -The 3-5 most impactful design problems, ordered by importance. +The 3–5 most impactful design problems, ordered by importance. For each issue, tag with **P0–P3 severity** (consult [heuristics-scoring](reference/heuristics-scoring.md) for severity definitions): - **[P?] What**: Name the problem clearly @@ -138,7 +129,7 @@ For each issue, tag with **P0–P3 severity** (consult [heuristics-scoring](refe - **Suggested command**: Which command could address this (from: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive) ### Persona Red Flags -→ *Consult [personas](reference/personas.md)* +> *Consult [personas](reference/personas.md)* Auto-select 2–3 personas most relevant to this interface type (use the selection table in the reference). If `CLAUDE.md` contains a `## Design Context` section from `teach-impeccable`, also generate 1–2 project-specific personas from the audience/brand info. @@ -154,12 +145,12 @@ Be specific — name the exact elements and interactions that fail each persona. Quick notes on smaller issues worth addressing. **Remember**: -- Be direct—vague feedback wastes everyone's time -- Be specific—"the submit button" not "some elements" +- Be direct — vague feedback wastes everyone's time +- Be specific — "the submit button" not "some elements" - Say what's wrong AND why it matters to users - Give concrete suggestions, not just "consider exploring..." -- Prioritize ruthlessly—if everything is important, nothing is -- Don't soften criticism—developers need honest feedback to ship great design +- Prioritize ruthlessly — if everything is important, nothing is +- Don't soften criticism — developers need honest feedback to ship great design ## Phase 3: Ask the User @@ -167,9 +158,9 @@ Quick notes on smaller issues worth addressing. Ask questions along these lines (adapt to the specific findings — do NOT ask generic questions): -1. **Priority direction**: Based on the issues found, ask which category matters most to the user right now. For example: "I found problems with visual hierarchy, color usage, and information overload. Which area should we tackle first?" Offer the top 2-3 issue categories as options. +1. **Priority direction**: Based on the issues found, ask which category matters most to the user right now. For example: "I found problems with visual hierarchy, color usage, and information overload. Which area should we tackle first?" Offer the top 2–3 issue categories as options. -2. **Design intent**: If the critique found a tonal mismatch, ask whether it was intentional. For example: "The interface feels clinical and corporate. Is that the intended tone, or should it feel warmer/bolder/more playful?" Offer 2-3 tonal directions as options based on what would fix the issues found. +2. **Design intent**: If the critique found a tonal mismatch, ask whether it was intentional. For example: "The interface feels clinical and corporate. Is that the intended tone, or should it feel warmer/bolder/more playful?" Offer 2–3 tonal directions as options based on what would fix the issues found. 3. **Scope**: Ask how much the user wants to take on. For example: "I found N issues. Want to address everything, or focus on the top 3?" Offer scope options like "Top 3 only", "All issues", "Critical issues only". @@ -177,9 +168,9 @@ Ask questions along these lines (adapt to the specific findings — do NOT ask g **Rules for questions**: - Every question must reference specific findings from Phase 2 — never ask generic "who is your audience?" questions -- Keep it to 2-4 questions maximum — respect the user's time +- Keep it to 2–4 questions maximum — respect the user's time - Offer concrete options, not open-ended prompts -- If findings are straightforward (e.g., only 1-2 clear issues), skip questions and go directly to Phase 4 +- If findings are straightforward (e.g., only 1–2 clear issues), skip questions and go directly to Phase 4 ## Phase 4: Recommended Actions diff --git a/.claude/skills/critique/reference/personas.md b/.claude/skills/critique/reference/personas.md index 67cf47d6b..1960220aa 100644 --- a/.claude/skills/critique/reference/personas.md +++ b/.claude/skills/critique/reference/personas.md @@ -88,30 +88,30 @@ Test the interface through the eyes of 5 distinct user archetypes. Each persona --- -## 4. Skeptical Evaluator — "Riley" +## 4. Deliberate Stress Tester — "Riley" -**Profile**: Evaluating the product for their team or company. Looking for reasons to reject. Comparing against competitors. +**Profile**: Methodical user who pushes interfaces beyond the happy path. Tests edge cases, tries unexpected inputs, and probes for gaps in the experience. **Behaviors**: - Tests edge cases intentionally (empty states, long strings, special characters) -- Looks for pricing catches and hidden limitations -- Reads fine print and terms of service -- Tries to break things deliberately +- Submits forms with unexpected data (emoji, RTL text, very long values) +- Tries to break workflows by navigating backwards, refreshing mid-flow, or opening in multiple tabs +- Looks for inconsistencies between what the UI promises and what actually happens - Documents problems methodically **Test Questions**: - What happens at the edges (0 items, 1000 items, very long text)? -- Is pricing and value proposition transparent? -- Are there hidden limitations or gotchas? -- How polished is error handling? -- What data is collected and why? +- Do error states recover gracefully or leave the UI in a broken state? +- What happens on refresh mid-workflow? Is state preserved? +- Are there features that appear to work but produce broken results? +- How does the UI handle unexpected input (emoji, special chars, paste from Excel)? **Red Flags** (report these specifically): -- Hidden pricing or "contact sales" for basic information -- Features that appear to work but produce broken results -- Poor error handling that exposes technical details -- Unclear data practices or missing privacy information +- Features that appear to work but silently fail or produce wrong results +- Error handling that exposes technical details or leaves UI in a broken state - Empty states that show nothing useful ("No results" with no guidance) +- Workflows that lose user data on refresh or navigation +- Inconsistent behavior between similar interactions in different parts of the UI --- @@ -150,7 +150,7 @@ Choose personas based on the interface type: |---------------|-----------------|-----| | Landing page / marketing | Jordan, Riley, Casey | First impressions, trust, mobile | | Dashboard / admin | Alex, Sam | Power users, accessibility | -| E-commerce / checkout | Casey, Riley, Jordan | Mobile, trust, clarity | +| E-commerce / checkout | Casey, Riley, Jordan | Mobile, edge cases, clarity | | Onboarding flow | Jordan, Casey | Confusion, interruption | | Data-heavy / analytics | Alex, Sam | Efficiency, keyboard nav | | Form-heavy / wizard | Jordan, Sam, Casey | Clarity, accessibility, mobile | diff --git a/.codex/skills/audit/SKILL.md b/.codex/skills/audit/SKILL.md index 5bb698aa3..def1967dc 100644 --- a/.codex/skills/audit/SKILL.md +++ b/.codex/skills/audit/SKILL.md @@ -1,16 +1,22 @@ --- name: audit -description: Perform a comprehensive audit of interface quality across accessibility, performance, theming, and responsive design. Generates a scored report with severity ratings and actionable plan. Use when the user wants a design review, accessibility check, quality audit, or a full list of UI issues to fix. +description: Run technical quality checks across accessibility, performance, theming, responsive design, and anti-patterns. Generates a scored report with P0-P3 severity ratings and actionable plan. Use when the user wants an accessibility check, performance audit, or technical quality review. argument-hint: "[area (feature, page, component...)]" --- -Run systematic quality checks and generate a comprehensive audit report with quantitative scoring, prioritized issues, and an actionable plan. Don't fix issues — document them for other commands to address. +## MANDATORY PREPARATION -**First**: Invoke $frontend-design for design principles and anti-patterns. +Invoke $frontend-design — it contains design principles, anti-patterns, and the **Context Gathering Protocol**. Follow the protocol before proceeding — if no design context exists yet, you MUST run $teach-impeccable first. + +--- + +Run systematic **technical** quality checks and generate a comprehensive report. Don't fix issues — document them for other commands to address. + +This is a code-level audit, not a design critique. Check what's measurable and verifiable in the implementation. ## Diagnostic Scan -Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using the criteria below. +Run comprehensive checks across 5 dimensions. Score each dimension 0-4 using the criteria below. ### 1. Accessibility (A11y) @@ -22,7 +28,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Alt text**: Missing or poor image descriptions - **Form issues**: Inputs without labels, poor error messaging, missing required indicators -**Score 0–4**: 0=Inaccessible (fails WCAG A), 1=Major gaps (few ARIA labels, no keyboard nav), 2=Partial (some a11y effort, significant gaps), 3=Good (WCAG AA mostly met, minor gaps), 4=Excellent (WCAG AA fully met, approaches AAA) +**Score 0-4**: 0=Inaccessible (fails WCAG A), 1=Major gaps (few ARIA labels, no keyboard nav), 2=Partial (some a11y effort, significant gaps), 3=Good (WCAG AA mostly met, minor gaps), 4=Excellent (WCAG AA fully met, approaches AAA) ### 2. Performance @@ -33,7 +39,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Bundle size**: Unnecessary imports, unused dependencies - **Render performance**: Unnecessary re-renders, missing memoization -**Score 0–4**: 0=Severe issues (layout thrash, unoptimized everything), 1=Major problems (no lazy loading, expensive animations), 2=Partial (some optimization, gaps remain), 3=Good (mostly optimized, minor improvements possible), 4=Excellent (fast, lean, well-optimized) +**Score 0-4**: 0=Severe issues (layout thrash, unoptimized everything), 1=Major problems (no lazy loading, expensive animations), 2=Partial (some optimization, gaps remain), 3=Good (mostly optimized, minor improvements possible), 4=Excellent (fast, lean, well-optimized) ### 3. Theming @@ -43,7 +49,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Inconsistent tokens**: Using wrong tokens, mixing token types - **Theme switching issues**: Values that don't update on theme change -**Score 0–4**: 0=No theming (hard-coded everything), 1=Minimal tokens (mostly hard-coded), 2=Partial (tokens exist but inconsistently used), 3=Good (tokens used, minor hard-coded values), 4=Excellent (full token system, dark mode works perfectly) +**Score 0-4**: 0=No theming (hard-coded everything), 1=Minimal tokens (mostly hard-coded), 2=Partial (tokens exist but inconsistently used), 3=Good (tokens used, minor hard-coded values), 4=Excellent (full token system, dark mode works perfectly) ### 4. Responsive Design @@ -54,113 +60,73 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Text scaling**: Layouts that break when text size increases - **Missing breakpoints**: No mobile/tablet variants -**Score 0–4**: 0=Desktop-only (breaks on mobile), 1=Major issues (some breakpoints, many failures), 2=Partial (works on mobile, rough edges), 3=Good (responsive, minor touch target or overflow issues), 4=Excellent (fluid, all viewports, proper touch targets) +**Score 0-4**: 0=Desktop-only (breaks on mobile), 1=Major issues (some breakpoints, many failures), 2=Partial (works on mobile, rough edges), 3=Good (responsive, minor touch target or overflow issues), 4=Excellent (fluid, all viewports, proper touch targets) ### 5. Anti-Patterns (CRITICAL) Check against ALL the **DON'T** guidelines in the frontend-design skill. Look for AI slop tells (AI color palette, gradient text, glassmorphism, hero metrics, card grids, generic fonts) and general design anti-patterns (gray on color, nested cards, bounce easing, redundant copy). -**Score 0–4**: 0=AI slop gallery (5+ tells), 1=Heavy AI aesthetic (3-4 tells), 2=Some tells (1-2 noticeable), 3=Mostly clean (subtle issues only), 4=No AI tells (distinctive, intentional design) +**Score 0-4**: 0=AI slop gallery (5+ tells), 1=Heavy AI aesthetic (3-4 tells), 2=Some tells (1-2 noticeable), 3=Mostly clean (subtle issues only), 4=No AI tells (distinctive, intentional design) -**CRITICAL**: This is an audit, not a fix. Document issues thoroughly with clear explanations of impact. Use other commands to fix issues after audit. - -## Generate Comprehensive Report +## Generate Report ### Audit Health Score -Present the dimension scores as a table: - | # | Dimension | Score | Key Finding | |---|-----------|-------|-------------| -| 1 | Accessibility | ? | [most critical a11y issue or "—"] | +| 1 | Accessibility | ? | [most critical a11y issue or "--"] | | 2 | Performance | ? | | | 3 | Responsive Design | ? | | | 4 | Theming | ? | | | 5 | Anti-Patterns | ? | | | **Total** | | **??/20** | **[Rating band]** | -**Rating bands**: -| Score | Rating | Action | -|-------|--------|--------| -| 18–20 | Excellent | Minor polish only | -| 14–17 | Good | Address weak dimensions | -| 10–13 | Acceptable | Significant work needed | -| 6–9 | Poor | Major quality overhaul | -| 0–5 | Critical | Fundamental issues across the board | +**Rating bands**: 18-20 Excellent (minor polish), 14-17 Good (address weak dimensions), 10-13 Acceptable (significant work needed), 6-9 Poor (major overhaul), 0-5 Critical (fundamental issues) ### Anti-Patterns Verdict -**Start here.** Pass/fail: Does this look AI-generated? List specific tells from the skill's Anti-Patterns section. Be brutally honest. +**Start here.** Pass/fail: Does this look AI-generated? List specific tells. Be brutally honest. ### Executive Summary - Audit Health Score: **??/20** ([rating band]) - Total issues found (count by severity: P0/P1/P2/P3) -- Most critical issues (top 3-5) +- Top 3-5 critical issues - Recommended next steps ### Detailed Findings by Severity -Tag every issue with **P0–P3 severity**: -| Priority | Name | Description | -|----------|------|-------------| -| **P0** | Blocking | Prevents task completion — fix immediately | -| **P1** | Major | Significant difficulty or WCAG AA violation — fix before release | -| **P2** | Minor | Annoyance, workaround exists — fix in next pass | -| **P3** | Polish | Nice-to-fix, no real user impact — fix if time permits | +Tag every issue with **P0-P3 severity**: +- **P0 Blocking**: Prevents task completion — fix immediately +- **P1 Major**: Significant difficulty or WCAG AA violation — fix before release +- **P2 Minor**: Annoyance, workaround exists — fix in next pass +- **P3 Polish**: Nice-to-fix, no real user impact — fix if time permits For each issue, document: - **[P?] Issue name** -- **Location**: Where it occurs (component, file, line) +- **Location**: Component, file, line - **Category**: Accessibility / Performance / Theming / Responsive / Anti-Pattern -- **Description**: What the issue is - **Impact**: How it affects users - **WCAG/Standard**: Which standard it violates (if applicable) - **Recommendation**: How to fix it -- **Suggested command**: Which command to use (prefer: $animate, $quieter, $optimize, $adapt, $clarify, $distill, $delight, $onboard, $normalize, $audit, $harden, $polish, $extract, $bolder, $arrange, $typeset, $critique, $colorize, $overdrive — or other installed skills you're sure exist) - -#### P0 — Blocking Issues -[Issues that prevent task completion or violate WCAG A] - -#### P1 — Major Issues -[Significant usability/accessibility impact, WCAG AA violations] - -#### P2 — Minor Issues -[Quality issues, WCAG AAA violations, performance concerns] - -#### P3 — Polish Issues -[Minor inconsistencies, optimization opportunities] +- **Suggested command**: Which command to use (prefer: $animate, $quieter, $optimize, $adapt, $clarify, $distill, $delight, $onboard, $normalize, $audit, $harden, $polish, $extract, $bolder, $arrange, $typeset, $critique, $colorize, $overdrive) ### Patterns & Systemic Issues -Identify recurring problems: +Identify recurring problems that indicate systemic gaps rather than one-off mistakes: - "Hard-coded colors appear in 15+ components, should use design tokens" - "Touch targets consistently too small (<44px) throughout mobile experience" -- "Missing focus indicators on all custom interactive components" ### Positive Findings -Note what's working well: -- Good practices to maintain -- Exemplary implementations to replicate elsewhere +Note what's working well — good practices to maintain and replicate. ## Recommended Actions -Present a prioritized action summary. Order is determined by severity automatically (P0 first, then P1, then P2). - -### Action Summary - -List recommended commands in priority order: +List recommended commands in priority order (P0 first, then P1, then P2): 1. **[P?] `$command-name`** — Brief description (specific context from audit findings) 2. **[P?] `$command-name`** — Brief description (specific context) -... -**Rules for recommendations**: -- Only recommend commands from: $animate, $quieter, $optimize, $adapt, $clarify, $distill, $delight, $onboard, $normalize, $audit, $harden, $polish, $extract, $bolder, $arrange, $typeset, $critique, $colorize, $overdrive -- Order by severity: P0 issues first, then P1, then P2 (skip P3 unless user has few issues) -- Each item's description should carry enough context that the command knows what to focus on -- Map findings to the most appropriate command -- Skip commands that would address zero issues -- End with `$polish` as the final step if any fixes were recommended +**Rules**: Only recommend commands from: $animate, $quieter, $optimize, $adapt, $clarify, $distill, $delight, $onboard, $normalize, $audit, $harden, $polish, $extract, $bolder, $arrange, $typeset, $critique, $colorize, $overdrive. Map findings to the most appropriate command. End with `$polish` as the final step if any fixes were recommended. After presenting the summary, tell the user: @@ -172,10 +138,9 @@ After presenting the summary, tell the user: **NEVER**: - Report issues without explaining impact (why does this matter?) -- Mix severity levels inconsistently -- Skip positive findings (celebrate what works) - Provide generic recommendations (be specific and actionable) +- Skip positive findings (celebrate what works) - Forget to prioritize (everything can't be P0) - Report false positives without verification -Remember: You're a quality auditor with exceptional attention to detail. Document systematically, prioritize ruthlessly, and provide clear paths to improvement. A good audit makes fixing easy. \ No newline at end of file +Remember: You're a technical quality auditor. Document systematically, prioritize ruthlessly, cite specific code locations, and provide clear paths to improvement. \ No newline at end of file diff --git a/.codex/skills/critique/SKILL.md b/.codex/skills/critique/SKILL.md index 7f2ed0450..5cb4af114 100644 --- a/.codex/skills/critique/SKILL.md +++ b/.codex/skills/critique/SKILL.md @@ -1,6 +1,6 @@ --- name: critique -description: Evaluate design effectiveness from a UX perspective. Assesses visual hierarchy, information architecture, emotional resonance, cognitive load, and overall design quality with quantitative scoring, persona-based testing, and actionable feedback. Use when the user asks to review, critique, evaluate, or give feedback on a design, UI, or component. +description: Evaluate design from a UX perspective, assessing visual hierarchy, information architecture, emotional resonance, cognitive load, and overall quality with quantitative scoring, persona-based testing, and actionable feedback. Use when the user asks to review, critique, evaluate, or give feedback on a design or component. argument-hint: "[area (feature, page, component...)]" --- @@ -10,7 +10,7 @@ Invoke $frontend-design — it contains design principles, anti-patterns, and th --- -Conduct a holistic design critique, evaluating whether the interface actually works—not just technically, but as a designed experience. Think like a design director giving feedback. +Conduct a holistic design critique, evaluating whether the interface actually works — not just technically, but as a designed experience. Think like a design director giving feedback. ## Phase 1: Design Critique @@ -20,7 +20,7 @@ Evaluate the interface across these dimensions: **This is the most important check.** Does this look like every other AI-generated interface from 2024-2025? -Review the design against ALL the **DON'T** guidelines in the frontend-design skill—they are the fingerprints of AI-generated work. Check for the AI color palette, gradient text, dark mode with glowing accents, glassmorphism, hero metric layouts, identical card grids, generic fonts, and all other tells. +Review the design against ALL the **DON'T** guidelines in the frontend-design skill — they are the fingerprints of AI-generated work. Check for the AI color palette, gradient text, dark mode with glowing accents, glassmorphism, hero metric layouts, identical card grids, generic fonts, and all other tells. **The test**: If you showed this to someone and said "AI made this," would they believe you immediately? If yes, that's the problem. @@ -30,20 +30,19 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Do size, color, and position communicate importance correctly? - Is there visual competition between elements that should have different weights? -### 3. Information Architecture -→ *Consult [cognitive-load](reference/cognitive-load.md) for the working memory rule and checklist* +### 3. Information Architecture & Cognitive Load +> *Consult [cognitive-load](reference/cognitive-load.md) for the working memory rule and 8-item checklist* - Is the structure intuitive? Would a new user understand the organization? - Is related content grouped logically? - Are there too many choices at once? Count visible options at each decision point — if >4, flag it - Is the navigation clear and predictable? - **Progressive disclosure**: Is complexity revealed only when needed, or dumped on the user upfront? -- **Cognitive load sub-check**: Run the 8-item cognitive load checklist from the reference. Report the number of failures. +- **Run the 8-item cognitive load checklist** from the reference. Report failure count: 0–1 = low (good), 2–3 = moderate, 4+ = critical. ### 4. Emotional Journey -→ *Consult [cognitive-load](reference/cognitive-load.md) for emotional intervention patterns* - What emotion does this interface evoke? Is that intentional? - Does it match the brand personality? -- Does it feel trustworthy, approachable, premium, playful—whatever it should feel? +- Does it feel trustworthy, approachable, premium, playful — whatever it should feel? - Would the target user feel "this is for me"? - **Peak-end rule**: Is the most intense moment positive? Does the experience end well (confirmation, celebration, clear next step)? - **Emotional valleys**: Check for onboarding frustration, error cliffs, feature discovery gaps, or anxiety spikes at high-stakes moments (payment, delete, commit) @@ -71,7 +70,7 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Is color used to communicate, not just decorate? - Does the palette feel cohesive? - Are accent colors drawing attention to the right things? -- Does it work for colorblind users? (not just technically—does meaning still come through?) +- Does it work for colorblind users? (not just technically — does meaning still come through?) ### 9. States & Edge Cases - Empty states: Do they guide users toward action, or just say "nothing here"? @@ -85,27 +84,19 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Are labels and buttons unambiguous? - Does error copy help users fix the problem? -### 11. Cognitive Load -→ *Consult [cognitive-load](reference/cognitive-load.md)* -- **Intrinsic vs. extraneous**: Is the mental effort coming from the task itself (acceptable) or from poor design choices (eliminate)? -- **Decision points**: Count visible choices at key moments. More than 4 simultaneous options = overload. -- **Working memory burden**: Does the user need to remember information from a previous screen to act on the current one? -- **Information chunking**: Is content broken into digestible groups, or presented as undifferentiated walls? -- Run the 8-item cognitive load checklist. Report failures count: 0–1 = low (good), 2–3 = moderate, 4+ = critical. - ## Phase 2: Present Findings Structure your feedback as a design director would: ### Design Health Score -→ *Consult [heuristics-scoring](reference/heuristics-scoring.md)* +> *Consult [heuristics-scoring](reference/heuristics-scoring.md)* Score each of Nielsen's 10 heuristics 0–4. Present as a table: | # | Heuristic | Score | Key Issue | |---|-----------|-------|-----------| | 1 | Visibility of System Status | ? | [specific finding or "—" if solid] | -| 2 | Match System ↔ Real World | ? | | +| 2 | Match System / Real World | ? | | | 3 | User Control and Freedom | ? | | | 4 | Consistency and Standards | ? | | | 5 | Error Prevention | ? | | @@ -122,13 +113,13 @@ Be honest with scores. A 4 means genuinely excellent. Most real interfaces score **Start here.** Pass/fail: Does this look AI-generated? List specific tells from the skill's Anti-Patterns section. Be brutally honest. ### Overall Impression -A brief gut reaction—what works, what doesn't, and the single biggest opportunity. +A brief gut reaction — what works, what doesn't, and the single biggest opportunity. ### What's Working -Highlight 2-3 things done well. Be specific about why they work. +Highlight 2–3 things done well. Be specific about why they work. ### Priority Issues -The 3-5 most impactful design problems, ordered by importance. +The 3–5 most impactful design problems, ordered by importance. For each issue, tag with **P0–P3 severity** (consult [heuristics-scoring](reference/heuristics-scoring.md) for severity definitions): - **[P?] What**: Name the problem clearly @@ -137,7 +128,7 @@ For each issue, tag with **P0–P3 severity** (consult [heuristics-scoring](refe - **Suggested command**: Which command could address this (from: $animate, $quieter, $optimize, $adapt, $clarify, $distill, $delight, $onboard, $normalize, $audit, $harden, $polish, $extract, $bolder, $arrange, $typeset, $critique, $colorize, $overdrive) ### Persona Red Flags -→ *Consult [personas](reference/personas.md)* +> *Consult [personas](reference/personas.md)* Auto-select 2–3 personas most relevant to this interface type (use the selection table in the reference). If `AGENTS.md` contains a `## Design Context` section from `teach-impeccable`, also generate 1–2 project-specific personas from the audience/brand info. @@ -153,12 +144,12 @@ Be specific — name the exact elements and interactions that fail each persona. Quick notes on smaller issues worth addressing. **Remember**: -- Be direct—vague feedback wastes everyone's time -- Be specific—"the submit button" not "some elements" +- Be direct — vague feedback wastes everyone's time +- Be specific — "the submit button" not "some elements" - Say what's wrong AND why it matters to users - Give concrete suggestions, not just "consider exploring..." -- Prioritize ruthlessly—if everything is important, nothing is -- Don't soften criticism—developers need honest feedback to ship great design +- Prioritize ruthlessly — if everything is important, nothing is +- Don't soften criticism — developers need honest feedback to ship great design ## Phase 3: Ask the User @@ -166,9 +157,9 @@ Quick notes on smaller issues worth addressing. Ask questions along these lines (adapt to the specific findings — do NOT ask generic questions): -1. **Priority direction**: Based on the issues found, ask which category matters most to the user right now. For example: "I found problems with visual hierarchy, color usage, and information overload. Which area should we tackle first?" Offer the top 2-3 issue categories as options. +1. **Priority direction**: Based on the issues found, ask which category matters most to the user right now. For example: "I found problems with visual hierarchy, color usage, and information overload. Which area should we tackle first?" Offer the top 2–3 issue categories as options. -2. **Design intent**: If the critique found a tonal mismatch, ask whether it was intentional. For example: "The interface feels clinical and corporate. Is that the intended tone, or should it feel warmer$bolder/more playful?" Offer 2-3 tonal directions as options based on what would fix the issues found. +2. **Design intent**: If the critique found a tonal mismatch, ask whether it was intentional. For example: "The interface feels clinical and corporate. Is that the intended tone, or should it feel warmer$bolder/more playful?" Offer 2–3 tonal directions as options based on what would fix the issues found. 3. **Scope**: Ask how much the user wants to take on. For example: "I found N issues. Want to address everything, or focus on the top 3?" Offer scope options like "Top 3 only", "All issues", "Critical issues only". @@ -176,9 +167,9 @@ Ask questions along these lines (adapt to the specific findings — do NOT ask g **Rules for questions**: - Every question must reference specific findings from Phase 2 — never ask generic "who is your audience?" questions -- Keep it to 2-4 questions maximum — respect the user's time +- Keep it to 2–4 questions maximum — respect the user's time - Offer concrete options, not open-ended prompts -- If findings are straightforward (e.g., only 1-2 clear issues), skip questions and go directly to Phase 4 +- If findings are straightforward (e.g., only 1–2 clear issues), skip questions and go directly to Phase 4 ## Phase 4: Recommended Actions diff --git a/.codex/skills/critique/reference/personas.md b/.codex/skills/critique/reference/personas.md index fdc88e20b..2d0f9cbf3 100644 --- a/.codex/skills/critique/reference/personas.md +++ b/.codex/skills/critique/reference/personas.md @@ -88,30 +88,30 @@ Test the interface through the eyes of 5 distinct user archetypes. Each persona --- -## 4. Skeptical Evaluator — "Riley" +## 4. Deliberate Stress Tester — "Riley" -**Profile**: Evaluating the product for their team or company. Looking for reasons to reject. Comparing against competitors. +**Profile**: Methodical user who pushes interfaces beyond the happy path. Tests edge cases, tries unexpected inputs, and probes for gaps in the experience. **Behaviors**: - Tests edge cases intentionally (empty states, long strings, special characters) -- Looks for pricing catches and hidden limitations -- Reads fine print and terms of service -- Tries to break things deliberately +- Submits forms with unexpected data (emoji, RTL text, very long values) +- Tries to break workflows by navigating backwards, refreshing mid-flow, or opening in multiple tabs +- Looks for inconsistencies between what the UI promises and what actually happens - Documents problems methodically **Test Questions**: - What happens at the edges (0 items, 1000 items, very long text)? -- Is pricing and value proposition transparent? -- Are there hidden limitations or gotchas? -- How polished is error handling? -- What data is collected and why? +- Do error states recover gracefully or leave the UI in a broken state? +- What happens on refresh mid-workflow? Is state preserved? +- Are there features that appear to work but produce broken results? +- How does the UI handle unexpected input (emoji, special chars, paste from Excel)? **Red Flags** (report these specifically): -- Hidden pricing or "contact sales" for basic information -- Features that appear to work but produce broken results -- Poor error handling that exposes technical details -- Unclear data practices or missing privacy information +- Features that appear to work but silently fail or produce wrong results +- Error handling that exposes technical details or leaves UI in a broken state - Empty states that show nothing useful ("No results" with no guidance) +- Workflows that lose user data on refresh or navigation +- Inconsistent behavior between similar interactions in different parts of the UI --- @@ -150,7 +150,7 @@ Choose personas based on the interface type: |---------------|-----------------|-----| | Landing page / marketing | Jordan, Riley, Casey | First impressions, trust, mobile | | Dashboard / admin | Alex, Sam | Power users, accessibility | -| E-commerce / checkout | Casey, Riley, Jordan | Mobile, trust, clarity | +| E-commerce / checkout | Casey, Riley, Jordan | Mobile, edge cases, clarity | | Onboarding flow | Jordan, Casey | Confusion, interruption | | Data-heavy / analytics | Alex, Sam | Efficiency, keyboard nav | | Form-heavy / wizard | Jordan, Sam, Casey | Clarity, accessibility, mobile | diff --git a/.cursor/skills/audit/SKILL.md b/.cursor/skills/audit/SKILL.md index d84428320..6dc747d60 100644 --- a/.cursor/skills/audit/SKILL.md +++ b/.cursor/skills/audit/SKILL.md @@ -1,15 +1,21 @@ --- name: audit -description: Perform a comprehensive audit of interface quality across accessibility, performance, theming, and responsive design. Generates a scored report with severity ratings and actionable plan. Use when the user wants a design review, accessibility check, quality audit, or a full list of UI issues to fix. +description: Run technical quality checks across accessibility, performance, theming, responsive design, and anti-patterns. Generates a scored report with P0-P3 severity ratings and actionable plan. Use when the user wants an accessibility check, performance audit, or technical quality review. --- -Run systematic quality checks and generate a comprehensive audit report with quantitative scoring, prioritized issues, and an actionable plan. Don't fix issues — document them for other commands to address. +## MANDATORY PREPARATION -**First**: Invoke /frontend-design for design principles and anti-patterns. +Invoke /frontend-design — it contains design principles, anti-patterns, and the **Context Gathering Protocol**. Follow the protocol before proceeding — if no design context exists yet, you MUST run /teach-impeccable first. + +--- + +Run systematic **technical** quality checks and generate a comprehensive report. Don't fix issues — document them for other commands to address. + +This is a code-level audit, not a design critique. Check what's measurable and verifiable in the implementation. ## Diagnostic Scan -Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using the criteria below. +Run comprehensive checks across 5 dimensions. Score each dimension 0-4 using the criteria below. ### 1. Accessibility (A11y) @@ -21,7 +27,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Alt text**: Missing or poor image descriptions - **Form issues**: Inputs without labels, poor error messaging, missing required indicators -**Score 0–4**: 0=Inaccessible (fails WCAG A), 1=Major gaps (few ARIA labels, no keyboard nav), 2=Partial (some a11y effort, significant gaps), 3=Good (WCAG AA mostly met, minor gaps), 4=Excellent (WCAG AA fully met, approaches AAA) +**Score 0-4**: 0=Inaccessible (fails WCAG A), 1=Major gaps (few ARIA labels, no keyboard nav), 2=Partial (some a11y effort, significant gaps), 3=Good (WCAG AA mostly met, minor gaps), 4=Excellent (WCAG AA fully met, approaches AAA) ### 2. Performance @@ -32,7 +38,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Bundle size**: Unnecessary imports, unused dependencies - **Render performance**: Unnecessary re-renders, missing memoization -**Score 0–4**: 0=Severe issues (layout thrash, unoptimized everything), 1=Major problems (no lazy loading, expensive animations), 2=Partial (some optimization, gaps remain), 3=Good (mostly optimized, minor improvements possible), 4=Excellent (fast, lean, well-optimized) +**Score 0-4**: 0=Severe issues (layout thrash, unoptimized everything), 1=Major problems (no lazy loading, expensive animations), 2=Partial (some optimization, gaps remain), 3=Good (mostly optimized, minor improvements possible), 4=Excellent (fast, lean, well-optimized) ### 3. Theming @@ -42,7 +48,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Inconsistent tokens**: Using wrong tokens, mixing token types - **Theme switching issues**: Values that don't update on theme change -**Score 0–4**: 0=No theming (hard-coded everything), 1=Minimal tokens (mostly hard-coded), 2=Partial (tokens exist but inconsistently used), 3=Good (tokens used, minor hard-coded values), 4=Excellent (full token system, dark mode works perfectly) +**Score 0-4**: 0=No theming (hard-coded everything), 1=Minimal tokens (mostly hard-coded), 2=Partial (tokens exist but inconsistently used), 3=Good (tokens used, minor hard-coded values), 4=Excellent (full token system, dark mode works perfectly) ### 4. Responsive Design @@ -53,113 +59,73 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Text scaling**: Layouts that break when text size increases - **Missing breakpoints**: No mobile/tablet variants -**Score 0–4**: 0=Desktop-only (breaks on mobile), 1=Major issues (some breakpoints, many failures), 2=Partial (works on mobile, rough edges), 3=Good (responsive, minor touch target or overflow issues), 4=Excellent (fluid, all viewports, proper touch targets) +**Score 0-4**: 0=Desktop-only (breaks on mobile), 1=Major issues (some breakpoints, many failures), 2=Partial (works on mobile, rough edges), 3=Good (responsive, minor touch target or overflow issues), 4=Excellent (fluid, all viewports, proper touch targets) ### 5. Anti-Patterns (CRITICAL) Check against ALL the **DON'T** guidelines in the frontend-design skill. Look for AI slop tells (AI color palette, gradient text, glassmorphism, hero metrics, card grids, generic fonts) and general design anti-patterns (gray on color, nested cards, bounce easing, redundant copy). -**Score 0–4**: 0=AI slop gallery (5+ tells), 1=Heavy AI aesthetic (3-4 tells), 2=Some tells (1-2 noticeable), 3=Mostly clean (subtle issues only), 4=No AI tells (distinctive, intentional design) +**Score 0-4**: 0=AI slop gallery (5+ tells), 1=Heavy AI aesthetic (3-4 tells), 2=Some tells (1-2 noticeable), 3=Mostly clean (subtle issues only), 4=No AI tells (distinctive, intentional design) -**CRITICAL**: This is an audit, not a fix. Document issues thoroughly with clear explanations of impact. Use other commands to fix issues after audit. - -## Generate Comprehensive Report +## Generate Report ### Audit Health Score -Present the dimension scores as a table: - | # | Dimension | Score | Key Finding | |---|-----------|-------|-------------| -| 1 | Accessibility | ? | [most critical a11y issue or "—"] | +| 1 | Accessibility | ? | [most critical a11y issue or "--"] | | 2 | Performance | ? | | | 3 | Responsive Design | ? | | | 4 | Theming | ? | | | 5 | Anti-Patterns | ? | | | **Total** | | **??/20** | **[Rating band]** | -**Rating bands**: -| Score | Rating | Action | -|-------|--------|--------| -| 18–20 | Excellent | Minor polish only | -| 14–17 | Good | Address weak dimensions | -| 10–13 | Acceptable | Significant work needed | -| 6–9 | Poor | Major quality overhaul | -| 0–5 | Critical | Fundamental issues across the board | +**Rating bands**: 18-20 Excellent (minor polish), 14-17 Good (address weak dimensions), 10-13 Acceptable (significant work needed), 6-9 Poor (major overhaul), 0-5 Critical (fundamental issues) ### Anti-Patterns Verdict -**Start here.** Pass/fail: Does this look AI-generated? List specific tells from the skill's Anti-Patterns section. Be brutally honest. +**Start here.** Pass/fail: Does this look AI-generated? List specific tells. Be brutally honest. ### Executive Summary - Audit Health Score: **??/20** ([rating band]) - Total issues found (count by severity: P0/P1/P2/P3) -- Most critical issues (top 3-5) +- Top 3-5 critical issues - Recommended next steps ### Detailed Findings by Severity -Tag every issue with **P0–P3 severity**: -| Priority | Name | Description | -|----------|------|-------------| -| **P0** | Blocking | Prevents task completion — fix immediately | -| **P1** | Major | Significant difficulty or WCAG AA violation — fix before release | -| **P2** | Minor | Annoyance, workaround exists — fix in next pass | -| **P3** | Polish | Nice-to-fix, no real user impact — fix if time permits | +Tag every issue with **P0-P3 severity**: +- **P0 Blocking**: Prevents task completion — fix immediately +- **P1 Major**: Significant difficulty or WCAG AA violation — fix before release +- **P2 Minor**: Annoyance, workaround exists — fix in next pass +- **P3 Polish**: Nice-to-fix, no real user impact — fix if time permits For each issue, document: - **[P?] Issue name** -- **Location**: Where it occurs (component, file, line) +- **Location**: Component, file, line - **Category**: Accessibility / Performance / Theming / Responsive / Anti-Pattern -- **Description**: What the issue is - **Impact**: How it affects users - **WCAG/Standard**: Which standard it violates (if applicable) - **Recommendation**: How to fix it -- **Suggested command**: Which command to use (prefer: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive — or other installed skills you're sure exist) - -#### P0 — Blocking Issues -[Issues that prevent task completion or violate WCAG A] - -#### P1 — Major Issues -[Significant usability/accessibility impact, WCAG AA violations] - -#### P2 — Minor Issues -[Quality issues, WCAG AAA violations, performance concerns] - -#### P3 — Polish Issues -[Minor inconsistencies, optimization opportunities] +- **Suggested command**: Which command to use (prefer: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive) ### Patterns & Systemic Issues -Identify recurring problems: +Identify recurring problems that indicate systemic gaps rather than one-off mistakes: - "Hard-coded colors appear in 15+ components, should use design tokens" - "Touch targets consistently too small (<44px) throughout mobile experience" -- "Missing focus indicators on all custom interactive components" ### Positive Findings -Note what's working well: -- Good practices to maintain -- Exemplary implementations to replicate elsewhere +Note what's working well — good practices to maintain and replicate. ## Recommended Actions -Present a prioritized action summary. Order is determined by severity automatically (P0 first, then P1, then P2). - -### Action Summary - -List recommended commands in priority order: +List recommended commands in priority order (P0 first, then P1, then P2): 1. **[P?] `/command-name`** — Brief description (specific context from audit findings) 2. **[P?] `/command-name`** — Brief description (specific context) -... -**Rules for recommendations**: -- Only recommend commands from: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive -- Order by severity: P0 issues first, then P1, then P2 (skip P3 unless user has few issues) -- Each item's description should carry enough context that the command knows what to focus on -- Map findings to the most appropriate command -- Skip commands that would address zero issues -- End with `/polish` as the final step if any fixes were recommended +**Rules**: Only recommend commands from: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive. Map findings to the most appropriate command. End with `/polish` as the final step if any fixes were recommended. After presenting the summary, tell the user: @@ -171,10 +137,9 @@ After presenting the summary, tell the user: **NEVER**: - Report issues without explaining impact (why does this matter?) -- Mix severity levels inconsistently -- Skip positive findings (celebrate what works) - Provide generic recommendations (be specific and actionable) +- Skip positive findings (celebrate what works) - Forget to prioritize (everything can't be P0) - Report false positives without verification -Remember: You're a quality auditor with exceptional attention to detail. Document systematically, prioritize ruthlessly, and provide clear paths to improvement. A good audit makes fixing easy. \ No newline at end of file +Remember: You're a technical quality auditor. Document systematically, prioritize ruthlessly, cite specific code locations, and provide clear paths to improvement. \ No newline at end of file diff --git a/.cursor/skills/critique/SKILL.md b/.cursor/skills/critique/SKILL.md index 7ae547bf5..3d2746f69 100644 --- a/.cursor/skills/critique/SKILL.md +++ b/.cursor/skills/critique/SKILL.md @@ -1,6 +1,6 @@ --- name: critique -description: Evaluate design effectiveness from a UX perspective. Assesses visual hierarchy, information architecture, emotional resonance, cognitive load, and overall design quality with quantitative scoring, persona-based testing, and actionable feedback. Use when the user asks to review, critique, evaluate, or give feedback on a design, UI, or component. +description: Evaluate design from a UX perspective, assessing visual hierarchy, information architecture, emotional resonance, cognitive load, and overall quality with quantitative scoring, persona-based testing, and actionable feedback. Use when the user asks to review, critique, evaluate, or give feedback on a design or component. --- ## MANDATORY PREPARATION @@ -9,7 +9,7 @@ Invoke /frontend-design — it contains design principles, anti-patterns, and th --- -Conduct a holistic design critique, evaluating whether the interface actually works—not just technically, but as a designed experience. Think like a design director giving feedback. +Conduct a holistic design critique, evaluating whether the interface actually works — not just technically, but as a designed experience. Think like a design director giving feedback. ## Phase 1: Design Critique @@ -19,7 +19,7 @@ Evaluate the interface across these dimensions: **This is the most important check.** Does this look like every other AI-generated interface from 2024-2025? -Review the design against ALL the **DON'T** guidelines in the frontend-design skill—they are the fingerprints of AI-generated work. Check for the AI color palette, gradient text, dark mode with glowing accents, glassmorphism, hero metric layouts, identical card grids, generic fonts, and all other tells. +Review the design against ALL the **DON'T** guidelines in the frontend-design skill — they are the fingerprints of AI-generated work. Check for the AI color palette, gradient text, dark mode with glowing accents, glassmorphism, hero metric layouts, identical card grids, generic fonts, and all other tells. **The test**: If you showed this to someone and said "AI made this," would they believe you immediately? If yes, that's the problem. @@ -29,20 +29,19 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Do size, color, and position communicate importance correctly? - Is there visual competition between elements that should have different weights? -### 3. Information Architecture -→ *Consult [cognitive-load](reference/cognitive-load.md) for the working memory rule and checklist* +### 3. Information Architecture & Cognitive Load +> *Consult [cognitive-load](reference/cognitive-load.md) for the working memory rule and 8-item checklist* - Is the structure intuitive? Would a new user understand the organization? - Is related content grouped logically? - Are there too many choices at once? Count visible options at each decision point — if >4, flag it - Is the navigation clear and predictable? - **Progressive disclosure**: Is complexity revealed only when needed, or dumped on the user upfront? -- **Cognitive load sub-check**: Run the 8-item cognitive load checklist from the reference. Report the number of failures. +- **Run the 8-item cognitive load checklist** from the reference. Report failure count: 0–1 = low (good), 2–3 = moderate, 4+ = critical. ### 4. Emotional Journey -→ *Consult [cognitive-load](reference/cognitive-load.md) for emotional intervention patterns* - What emotion does this interface evoke? Is that intentional? - Does it match the brand personality? -- Does it feel trustworthy, approachable, premium, playful—whatever it should feel? +- Does it feel trustworthy, approachable, premium, playful — whatever it should feel? - Would the target user feel "this is for me"? - **Peak-end rule**: Is the most intense moment positive? Does the experience end well (confirmation, celebration, clear next step)? - **Emotional valleys**: Check for onboarding frustration, error cliffs, feature discovery gaps, or anxiety spikes at high-stakes moments (payment, delete, commit) @@ -70,7 +69,7 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Is color used to communicate, not just decorate? - Does the palette feel cohesive? - Are accent colors drawing attention to the right things? -- Does it work for colorblind users? (not just technically—does meaning still come through?) +- Does it work for colorblind users? (not just technically — does meaning still come through?) ### 9. States & Edge Cases - Empty states: Do they guide users toward action, or just say "nothing here"? @@ -84,27 +83,19 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Are labels and buttons unambiguous? - Does error copy help users fix the problem? -### 11. Cognitive Load -→ *Consult [cognitive-load](reference/cognitive-load.md)* -- **Intrinsic vs. extraneous**: Is the mental effort coming from the task itself (acceptable) or from poor design choices (eliminate)? -- **Decision points**: Count visible choices at key moments. More than 4 simultaneous options = overload. -- **Working memory burden**: Does the user need to remember information from a previous screen to act on the current one? -- **Information chunking**: Is content broken into digestible groups, or presented as undifferentiated walls? -- Run the 8-item cognitive load checklist. Report failures count: 0–1 = low (good), 2–3 = moderate, 4+ = critical. - ## Phase 2: Present Findings Structure your feedback as a design director would: ### Design Health Score -→ *Consult [heuristics-scoring](reference/heuristics-scoring.md)* +> *Consult [heuristics-scoring](reference/heuristics-scoring.md)* Score each of Nielsen's 10 heuristics 0–4. Present as a table: | # | Heuristic | Score | Key Issue | |---|-----------|-------|-----------| | 1 | Visibility of System Status | ? | [specific finding or "—" if solid] | -| 2 | Match System ↔ Real World | ? | | +| 2 | Match System / Real World | ? | | | 3 | User Control and Freedom | ? | | | 4 | Consistency and Standards | ? | | | 5 | Error Prevention | ? | | @@ -121,13 +112,13 @@ Be honest with scores. A 4 means genuinely excellent. Most real interfaces score **Start here.** Pass/fail: Does this look AI-generated? List specific tells from the skill's Anti-Patterns section. Be brutally honest. ### Overall Impression -A brief gut reaction—what works, what doesn't, and the single biggest opportunity. +A brief gut reaction — what works, what doesn't, and the single biggest opportunity. ### What's Working -Highlight 2-3 things done well. Be specific about why they work. +Highlight 2–3 things done well. Be specific about why they work. ### Priority Issues -The 3-5 most impactful design problems, ordered by importance. +The 3–5 most impactful design problems, ordered by importance. For each issue, tag with **P0–P3 severity** (consult [heuristics-scoring](reference/heuristics-scoring.md) for severity definitions): - **[P?] What**: Name the problem clearly @@ -136,7 +127,7 @@ For each issue, tag with **P0–P3 severity** (consult [heuristics-scoring](refe - **Suggested command**: Which command could address this (from: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive) ### Persona Red Flags -→ *Consult [personas](reference/personas.md)* +> *Consult [personas](reference/personas.md)* Auto-select 2–3 personas most relevant to this interface type (use the selection table in the reference). If `.cursorrules` contains a `## Design Context` section from `teach-impeccable`, also generate 1–2 project-specific personas from the audience/brand info. @@ -152,12 +143,12 @@ Be specific — name the exact elements and interactions that fail each persona. Quick notes on smaller issues worth addressing. **Remember**: -- Be direct—vague feedback wastes everyone's time -- Be specific—"the submit button" not "some elements" +- Be direct — vague feedback wastes everyone's time +- Be specific — "the submit button" not "some elements" - Say what's wrong AND why it matters to users - Give concrete suggestions, not just "consider exploring..." -- Prioritize ruthlessly—if everything is important, nothing is -- Don't soften criticism—developers need honest feedback to ship great design +- Prioritize ruthlessly — if everything is important, nothing is +- Don't soften criticism — developers need honest feedback to ship great design ## Phase 3: Ask the User @@ -165,9 +156,9 @@ Quick notes on smaller issues worth addressing. Ask questions along these lines (adapt to the specific findings — do NOT ask generic questions): -1. **Priority direction**: Based on the issues found, ask which category matters most to the user right now. For example: "I found problems with visual hierarchy, color usage, and information overload. Which area should we tackle first?" Offer the top 2-3 issue categories as options. +1. **Priority direction**: Based on the issues found, ask which category matters most to the user right now. For example: "I found problems with visual hierarchy, color usage, and information overload. Which area should we tackle first?" Offer the top 2–3 issue categories as options. -2. **Design intent**: If the critique found a tonal mismatch, ask whether it was intentional. For example: "The interface feels clinical and corporate. Is that the intended tone, or should it feel warmer/bolder/more playful?" Offer 2-3 tonal directions as options based on what would fix the issues found. +2. **Design intent**: If the critique found a tonal mismatch, ask whether it was intentional. For example: "The interface feels clinical and corporate. Is that the intended tone, or should it feel warmer/bolder/more playful?" Offer 2–3 tonal directions as options based on what would fix the issues found. 3. **Scope**: Ask how much the user wants to take on. For example: "I found N issues. Want to address everything, or focus on the top 3?" Offer scope options like "Top 3 only", "All issues", "Critical issues only". @@ -175,9 +166,9 @@ Ask questions along these lines (adapt to the specific findings — do NOT ask g **Rules for questions**: - Every question must reference specific findings from Phase 2 — never ask generic "who is your audience?" questions -- Keep it to 2-4 questions maximum — respect the user's time +- Keep it to 2–4 questions maximum — respect the user's time - Offer concrete options, not open-ended prompts -- If findings are straightforward (e.g., only 1-2 clear issues), skip questions and go directly to Phase 4 +- If findings are straightforward (e.g., only 1–2 clear issues), skip questions and go directly to Phase 4 ## Phase 4: Recommended Actions diff --git a/.cursor/skills/critique/reference/personas.md b/.cursor/skills/critique/reference/personas.md index 773021f95..84689b598 100644 --- a/.cursor/skills/critique/reference/personas.md +++ b/.cursor/skills/critique/reference/personas.md @@ -88,30 +88,30 @@ Test the interface through the eyes of 5 distinct user archetypes. Each persona --- -## 4. Skeptical Evaluator — "Riley" +## 4. Deliberate Stress Tester — "Riley" -**Profile**: Evaluating the product for their team or company. Looking for reasons to reject. Comparing against competitors. +**Profile**: Methodical user who pushes interfaces beyond the happy path. Tests edge cases, tries unexpected inputs, and probes for gaps in the experience. **Behaviors**: - Tests edge cases intentionally (empty states, long strings, special characters) -- Looks for pricing catches and hidden limitations -- Reads fine print and terms of service -- Tries to break things deliberately +- Submits forms with unexpected data (emoji, RTL text, very long values) +- Tries to break workflows by navigating backwards, refreshing mid-flow, or opening in multiple tabs +- Looks for inconsistencies between what the UI promises and what actually happens - Documents problems methodically **Test Questions**: - What happens at the edges (0 items, 1000 items, very long text)? -- Is pricing and value proposition transparent? -- Are there hidden limitations or gotchas? -- How polished is error handling? -- What data is collected and why? +- Do error states recover gracefully or leave the UI in a broken state? +- What happens on refresh mid-workflow? Is state preserved? +- Are there features that appear to work but produce broken results? +- How does the UI handle unexpected input (emoji, special chars, paste from Excel)? **Red Flags** (report these specifically): -- Hidden pricing or "contact sales" for basic information -- Features that appear to work but produce broken results -- Poor error handling that exposes technical details -- Unclear data practices or missing privacy information +- Features that appear to work but silently fail or produce wrong results +- Error handling that exposes technical details or leaves UI in a broken state - Empty states that show nothing useful ("No results" with no guidance) +- Workflows that lose user data on refresh or navigation +- Inconsistent behavior between similar interactions in different parts of the UI --- @@ -150,7 +150,7 @@ Choose personas based on the interface type: |---------------|-----------------|-----| | Landing page / marketing | Jordan, Riley, Casey | First impressions, trust, mobile | | Dashboard / admin | Alex, Sam | Power users, accessibility | -| E-commerce / checkout | Casey, Riley, Jordan | Mobile, trust, clarity | +| E-commerce / checkout | Casey, Riley, Jordan | Mobile, edge cases, clarity | | Onboarding flow | Jordan, Casey | Confusion, interruption | | Data-heavy / analytics | Alex, Sam | Efficiency, keyboard nav | | Form-heavy / wizard | Jordan, Sam, Casey | Clarity, accessibility, mobile | diff --git a/.gemini/skills/audit/SKILL.md b/.gemini/skills/audit/SKILL.md index d84428320..6dc747d60 100644 --- a/.gemini/skills/audit/SKILL.md +++ b/.gemini/skills/audit/SKILL.md @@ -1,15 +1,21 @@ --- name: audit -description: Perform a comprehensive audit of interface quality across accessibility, performance, theming, and responsive design. Generates a scored report with severity ratings and actionable plan. Use when the user wants a design review, accessibility check, quality audit, or a full list of UI issues to fix. +description: Run technical quality checks across accessibility, performance, theming, responsive design, and anti-patterns. Generates a scored report with P0-P3 severity ratings and actionable plan. Use when the user wants an accessibility check, performance audit, or technical quality review. --- -Run systematic quality checks and generate a comprehensive audit report with quantitative scoring, prioritized issues, and an actionable plan. Don't fix issues — document them for other commands to address. +## MANDATORY PREPARATION -**First**: Invoke /frontend-design for design principles and anti-patterns. +Invoke /frontend-design — it contains design principles, anti-patterns, and the **Context Gathering Protocol**. Follow the protocol before proceeding — if no design context exists yet, you MUST run /teach-impeccable first. + +--- + +Run systematic **technical** quality checks and generate a comprehensive report. Don't fix issues — document them for other commands to address. + +This is a code-level audit, not a design critique. Check what's measurable and verifiable in the implementation. ## Diagnostic Scan -Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using the criteria below. +Run comprehensive checks across 5 dimensions. Score each dimension 0-4 using the criteria below. ### 1. Accessibility (A11y) @@ -21,7 +27,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Alt text**: Missing or poor image descriptions - **Form issues**: Inputs without labels, poor error messaging, missing required indicators -**Score 0–4**: 0=Inaccessible (fails WCAG A), 1=Major gaps (few ARIA labels, no keyboard nav), 2=Partial (some a11y effort, significant gaps), 3=Good (WCAG AA mostly met, minor gaps), 4=Excellent (WCAG AA fully met, approaches AAA) +**Score 0-4**: 0=Inaccessible (fails WCAG A), 1=Major gaps (few ARIA labels, no keyboard nav), 2=Partial (some a11y effort, significant gaps), 3=Good (WCAG AA mostly met, minor gaps), 4=Excellent (WCAG AA fully met, approaches AAA) ### 2. Performance @@ -32,7 +38,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Bundle size**: Unnecessary imports, unused dependencies - **Render performance**: Unnecessary re-renders, missing memoization -**Score 0–4**: 0=Severe issues (layout thrash, unoptimized everything), 1=Major problems (no lazy loading, expensive animations), 2=Partial (some optimization, gaps remain), 3=Good (mostly optimized, minor improvements possible), 4=Excellent (fast, lean, well-optimized) +**Score 0-4**: 0=Severe issues (layout thrash, unoptimized everything), 1=Major problems (no lazy loading, expensive animations), 2=Partial (some optimization, gaps remain), 3=Good (mostly optimized, minor improvements possible), 4=Excellent (fast, lean, well-optimized) ### 3. Theming @@ -42,7 +48,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Inconsistent tokens**: Using wrong tokens, mixing token types - **Theme switching issues**: Values that don't update on theme change -**Score 0–4**: 0=No theming (hard-coded everything), 1=Minimal tokens (mostly hard-coded), 2=Partial (tokens exist but inconsistently used), 3=Good (tokens used, minor hard-coded values), 4=Excellent (full token system, dark mode works perfectly) +**Score 0-4**: 0=No theming (hard-coded everything), 1=Minimal tokens (mostly hard-coded), 2=Partial (tokens exist but inconsistently used), 3=Good (tokens used, minor hard-coded values), 4=Excellent (full token system, dark mode works perfectly) ### 4. Responsive Design @@ -53,113 +59,73 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Text scaling**: Layouts that break when text size increases - **Missing breakpoints**: No mobile/tablet variants -**Score 0–4**: 0=Desktop-only (breaks on mobile), 1=Major issues (some breakpoints, many failures), 2=Partial (works on mobile, rough edges), 3=Good (responsive, minor touch target or overflow issues), 4=Excellent (fluid, all viewports, proper touch targets) +**Score 0-4**: 0=Desktop-only (breaks on mobile), 1=Major issues (some breakpoints, many failures), 2=Partial (works on mobile, rough edges), 3=Good (responsive, minor touch target or overflow issues), 4=Excellent (fluid, all viewports, proper touch targets) ### 5. Anti-Patterns (CRITICAL) Check against ALL the **DON'T** guidelines in the frontend-design skill. Look for AI slop tells (AI color palette, gradient text, glassmorphism, hero metrics, card grids, generic fonts) and general design anti-patterns (gray on color, nested cards, bounce easing, redundant copy). -**Score 0–4**: 0=AI slop gallery (5+ tells), 1=Heavy AI aesthetic (3-4 tells), 2=Some tells (1-2 noticeable), 3=Mostly clean (subtle issues only), 4=No AI tells (distinctive, intentional design) +**Score 0-4**: 0=AI slop gallery (5+ tells), 1=Heavy AI aesthetic (3-4 tells), 2=Some tells (1-2 noticeable), 3=Mostly clean (subtle issues only), 4=No AI tells (distinctive, intentional design) -**CRITICAL**: This is an audit, not a fix. Document issues thoroughly with clear explanations of impact. Use other commands to fix issues after audit. - -## Generate Comprehensive Report +## Generate Report ### Audit Health Score -Present the dimension scores as a table: - | # | Dimension | Score | Key Finding | |---|-----------|-------|-------------| -| 1 | Accessibility | ? | [most critical a11y issue or "—"] | +| 1 | Accessibility | ? | [most critical a11y issue or "--"] | | 2 | Performance | ? | | | 3 | Responsive Design | ? | | | 4 | Theming | ? | | | 5 | Anti-Patterns | ? | | | **Total** | | **??/20** | **[Rating band]** | -**Rating bands**: -| Score | Rating | Action | -|-------|--------|--------| -| 18–20 | Excellent | Minor polish only | -| 14–17 | Good | Address weak dimensions | -| 10–13 | Acceptable | Significant work needed | -| 6–9 | Poor | Major quality overhaul | -| 0–5 | Critical | Fundamental issues across the board | +**Rating bands**: 18-20 Excellent (minor polish), 14-17 Good (address weak dimensions), 10-13 Acceptable (significant work needed), 6-9 Poor (major overhaul), 0-5 Critical (fundamental issues) ### Anti-Patterns Verdict -**Start here.** Pass/fail: Does this look AI-generated? List specific tells from the skill's Anti-Patterns section. Be brutally honest. +**Start here.** Pass/fail: Does this look AI-generated? List specific tells. Be brutally honest. ### Executive Summary - Audit Health Score: **??/20** ([rating band]) - Total issues found (count by severity: P0/P1/P2/P3) -- Most critical issues (top 3-5) +- Top 3-5 critical issues - Recommended next steps ### Detailed Findings by Severity -Tag every issue with **P0–P3 severity**: -| Priority | Name | Description | -|----------|------|-------------| -| **P0** | Blocking | Prevents task completion — fix immediately | -| **P1** | Major | Significant difficulty or WCAG AA violation — fix before release | -| **P2** | Minor | Annoyance, workaround exists — fix in next pass | -| **P3** | Polish | Nice-to-fix, no real user impact — fix if time permits | +Tag every issue with **P0-P3 severity**: +- **P0 Blocking**: Prevents task completion — fix immediately +- **P1 Major**: Significant difficulty or WCAG AA violation — fix before release +- **P2 Minor**: Annoyance, workaround exists — fix in next pass +- **P3 Polish**: Nice-to-fix, no real user impact — fix if time permits For each issue, document: - **[P?] Issue name** -- **Location**: Where it occurs (component, file, line) +- **Location**: Component, file, line - **Category**: Accessibility / Performance / Theming / Responsive / Anti-Pattern -- **Description**: What the issue is - **Impact**: How it affects users - **WCAG/Standard**: Which standard it violates (if applicable) - **Recommendation**: How to fix it -- **Suggested command**: Which command to use (prefer: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive — or other installed skills you're sure exist) - -#### P0 — Blocking Issues -[Issues that prevent task completion or violate WCAG A] - -#### P1 — Major Issues -[Significant usability/accessibility impact, WCAG AA violations] - -#### P2 — Minor Issues -[Quality issues, WCAG AAA violations, performance concerns] - -#### P3 — Polish Issues -[Minor inconsistencies, optimization opportunities] +- **Suggested command**: Which command to use (prefer: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive) ### Patterns & Systemic Issues -Identify recurring problems: +Identify recurring problems that indicate systemic gaps rather than one-off mistakes: - "Hard-coded colors appear in 15+ components, should use design tokens" - "Touch targets consistently too small (<44px) throughout mobile experience" -- "Missing focus indicators on all custom interactive components" ### Positive Findings -Note what's working well: -- Good practices to maintain -- Exemplary implementations to replicate elsewhere +Note what's working well — good practices to maintain and replicate. ## Recommended Actions -Present a prioritized action summary. Order is determined by severity automatically (P0 first, then P1, then P2). - -### Action Summary - -List recommended commands in priority order: +List recommended commands in priority order (P0 first, then P1, then P2): 1. **[P?] `/command-name`** — Brief description (specific context from audit findings) 2. **[P?] `/command-name`** — Brief description (specific context) -... -**Rules for recommendations**: -- Only recommend commands from: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive -- Order by severity: P0 issues first, then P1, then P2 (skip P3 unless user has few issues) -- Each item's description should carry enough context that the command knows what to focus on -- Map findings to the most appropriate command -- Skip commands that would address zero issues -- End with `/polish` as the final step if any fixes were recommended +**Rules**: Only recommend commands from: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive. Map findings to the most appropriate command. End with `/polish` as the final step if any fixes were recommended. After presenting the summary, tell the user: @@ -171,10 +137,9 @@ After presenting the summary, tell the user: **NEVER**: - Report issues without explaining impact (why does this matter?) -- Mix severity levels inconsistently -- Skip positive findings (celebrate what works) - Provide generic recommendations (be specific and actionable) +- Skip positive findings (celebrate what works) - Forget to prioritize (everything can't be P0) - Report false positives without verification -Remember: You're a quality auditor with exceptional attention to detail. Document systematically, prioritize ruthlessly, and provide clear paths to improvement. A good audit makes fixing easy. \ No newline at end of file +Remember: You're a technical quality auditor. Document systematically, prioritize ruthlessly, cite specific code locations, and provide clear paths to improvement. \ No newline at end of file diff --git a/.gemini/skills/critique/SKILL.md b/.gemini/skills/critique/SKILL.md index 359b222e8..12180ce73 100644 --- a/.gemini/skills/critique/SKILL.md +++ b/.gemini/skills/critique/SKILL.md @@ -1,6 +1,6 @@ --- name: critique -description: Evaluate design effectiveness from a UX perspective. Assesses visual hierarchy, information architecture, emotional resonance, cognitive load, and overall design quality with quantitative scoring, persona-based testing, and actionable feedback. Use when the user asks to review, critique, evaluate, or give feedback on a design, UI, or component. +description: Evaluate design from a UX perspective, assessing visual hierarchy, information architecture, emotional resonance, cognitive load, and overall quality with quantitative scoring, persona-based testing, and actionable feedback. Use when the user asks to review, critique, evaluate, or give feedback on a design or component. --- ## MANDATORY PREPARATION @@ -9,7 +9,7 @@ Invoke /frontend-design — it contains design principles, anti-patterns, and th --- -Conduct a holistic design critique, evaluating whether the interface actually works—not just technically, but as a designed experience. Think like a design director giving feedback. +Conduct a holistic design critique, evaluating whether the interface actually works — not just technically, but as a designed experience. Think like a design director giving feedback. ## Phase 1: Design Critique @@ -19,7 +19,7 @@ Evaluate the interface across these dimensions: **This is the most important check.** Does this look like every other AI-generated interface from 2024-2025? -Review the design against ALL the **DON'T** guidelines in the frontend-design skill—they are the fingerprints of AI-generated work. Check for the AI color palette, gradient text, dark mode with glowing accents, glassmorphism, hero metric layouts, identical card grids, generic fonts, and all other tells. +Review the design against ALL the **DON'T** guidelines in the frontend-design skill — they are the fingerprints of AI-generated work. Check for the AI color palette, gradient text, dark mode with glowing accents, glassmorphism, hero metric layouts, identical card grids, generic fonts, and all other tells. **The test**: If you showed this to someone and said "AI made this," would they believe you immediately? If yes, that's the problem. @@ -29,20 +29,19 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Do size, color, and position communicate importance correctly? - Is there visual competition between elements that should have different weights? -### 3. Information Architecture -→ *Consult [cognitive-load](reference/cognitive-load.md) for the working memory rule and checklist* +### 3. Information Architecture & Cognitive Load +> *Consult [cognitive-load](reference/cognitive-load.md) for the working memory rule and 8-item checklist* - Is the structure intuitive? Would a new user understand the organization? - Is related content grouped logically? - Are there too many choices at once? Count visible options at each decision point — if >4, flag it - Is the navigation clear and predictable? - **Progressive disclosure**: Is complexity revealed only when needed, or dumped on the user upfront? -- **Cognitive load sub-check**: Run the 8-item cognitive load checklist from the reference. Report the number of failures. +- **Run the 8-item cognitive load checklist** from the reference. Report failure count: 0–1 = low (good), 2–3 = moderate, 4+ = critical. ### 4. Emotional Journey -→ *Consult [cognitive-load](reference/cognitive-load.md) for emotional intervention patterns* - What emotion does this interface evoke? Is that intentional? - Does it match the brand personality? -- Does it feel trustworthy, approachable, premium, playful—whatever it should feel? +- Does it feel trustworthy, approachable, premium, playful — whatever it should feel? - Would the target user feel "this is for me"? - **Peak-end rule**: Is the most intense moment positive? Does the experience end well (confirmation, celebration, clear next step)? - **Emotional valleys**: Check for onboarding frustration, error cliffs, feature discovery gaps, or anxiety spikes at high-stakes moments (payment, delete, commit) @@ -70,7 +69,7 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Is color used to communicate, not just decorate? - Does the palette feel cohesive? - Are accent colors drawing attention to the right things? -- Does it work for colorblind users? (not just technically—does meaning still come through?) +- Does it work for colorblind users? (not just technically — does meaning still come through?) ### 9. States & Edge Cases - Empty states: Do they guide users toward action, or just say "nothing here"? @@ -84,27 +83,19 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Are labels and buttons unambiguous? - Does error copy help users fix the problem? -### 11. Cognitive Load -→ *Consult [cognitive-load](reference/cognitive-load.md)* -- **Intrinsic vs. extraneous**: Is the mental effort coming from the task itself (acceptable) or from poor design choices (eliminate)? -- **Decision points**: Count visible choices at key moments. More than 4 simultaneous options = overload. -- **Working memory burden**: Does the user need to remember information from a previous screen to act on the current one? -- **Information chunking**: Is content broken into digestible groups, or presented as undifferentiated walls? -- Run the 8-item cognitive load checklist. Report failures count: 0–1 = low (good), 2–3 = moderate, 4+ = critical. - ## Phase 2: Present Findings Structure your feedback as a design director would: ### Design Health Score -→ *Consult [heuristics-scoring](reference/heuristics-scoring.md)* +> *Consult [heuristics-scoring](reference/heuristics-scoring.md)* Score each of Nielsen's 10 heuristics 0–4. Present as a table: | # | Heuristic | Score | Key Issue | |---|-----------|-------|-----------| | 1 | Visibility of System Status | ? | [specific finding or "—" if solid] | -| 2 | Match System ↔ Real World | ? | | +| 2 | Match System / Real World | ? | | | 3 | User Control and Freedom | ? | | | 4 | Consistency and Standards | ? | | | 5 | Error Prevention | ? | | @@ -121,13 +112,13 @@ Be honest with scores. A 4 means genuinely excellent. Most real interfaces score **Start here.** Pass/fail: Does this look AI-generated? List specific tells from the skill's Anti-Patterns section. Be brutally honest. ### Overall Impression -A brief gut reaction—what works, what doesn't, and the single biggest opportunity. +A brief gut reaction — what works, what doesn't, and the single biggest opportunity. ### What's Working -Highlight 2-3 things done well. Be specific about why they work. +Highlight 2–3 things done well. Be specific about why they work. ### Priority Issues -The 3-5 most impactful design problems, ordered by importance. +The 3–5 most impactful design problems, ordered by importance. For each issue, tag with **P0–P3 severity** (consult [heuristics-scoring](reference/heuristics-scoring.md) for severity definitions): - **[P?] What**: Name the problem clearly @@ -136,7 +127,7 @@ For each issue, tag with **P0–P3 severity** (consult [heuristics-scoring](refe - **Suggested command**: Which command could address this (from: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive) ### Persona Red Flags -→ *Consult [personas](reference/personas.md)* +> *Consult [personas](reference/personas.md)* Auto-select 2–3 personas most relevant to this interface type (use the selection table in the reference). If `GEMINI.md` contains a `## Design Context` section from `teach-impeccable`, also generate 1–2 project-specific personas from the audience/brand info. @@ -152,12 +143,12 @@ Be specific — name the exact elements and interactions that fail each persona. Quick notes on smaller issues worth addressing. **Remember**: -- Be direct—vague feedback wastes everyone's time -- Be specific—"the submit button" not "some elements" +- Be direct — vague feedback wastes everyone's time +- Be specific — "the submit button" not "some elements" - Say what's wrong AND why it matters to users - Give concrete suggestions, not just "consider exploring..." -- Prioritize ruthlessly—if everything is important, nothing is -- Don't soften criticism—developers need honest feedback to ship great design +- Prioritize ruthlessly — if everything is important, nothing is +- Don't soften criticism — developers need honest feedback to ship great design ## Phase 3: Ask the User @@ -165,9 +156,9 @@ Quick notes on smaller issues worth addressing. Ask questions along these lines (adapt to the specific findings — do NOT ask generic questions): -1. **Priority direction**: Based on the issues found, ask which category matters most to the user right now. For example: "I found problems with visual hierarchy, color usage, and information overload. Which area should we tackle first?" Offer the top 2-3 issue categories as options. +1. **Priority direction**: Based on the issues found, ask which category matters most to the user right now. For example: "I found problems with visual hierarchy, color usage, and information overload. Which area should we tackle first?" Offer the top 2–3 issue categories as options. -2. **Design intent**: If the critique found a tonal mismatch, ask whether it was intentional. For example: "The interface feels clinical and corporate. Is that the intended tone, or should it feel warmer/bolder/more playful?" Offer 2-3 tonal directions as options based on what would fix the issues found. +2. **Design intent**: If the critique found a tonal mismatch, ask whether it was intentional. For example: "The interface feels clinical and corporate. Is that the intended tone, or should it feel warmer/bolder/more playful?" Offer 2–3 tonal directions as options based on what would fix the issues found. 3. **Scope**: Ask how much the user wants to take on. For example: "I found N issues. Want to address everything, or focus on the top 3?" Offer scope options like "Top 3 only", "All issues", "Critical issues only". @@ -175,9 +166,9 @@ Ask questions along these lines (adapt to the specific findings — do NOT ask g **Rules for questions**: - Every question must reference specific findings from Phase 2 — never ask generic "who is your audience?" questions -- Keep it to 2-4 questions maximum — respect the user's time +- Keep it to 2–4 questions maximum — respect the user's time - Offer concrete options, not open-ended prompts -- If findings are straightforward (e.g., only 1-2 clear issues), skip questions and go directly to Phase 4 +- If findings are straightforward (e.g., only 1–2 clear issues), skip questions and go directly to Phase 4 ## Phase 4: Recommended Actions diff --git a/.gemini/skills/critique/reference/personas.md b/.gemini/skills/critique/reference/personas.md index 009244180..80b38a939 100644 --- a/.gemini/skills/critique/reference/personas.md +++ b/.gemini/skills/critique/reference/personas.md @@ -88,30 +88,30 @@ Test the interface through the eyes of 5 distinct user archetypes. Each persona --- -## 4. Skeptical Evaluator — "Riley" +## 4. Deliberate Stress Tester — "Riley" -**Profile**: Evaluating the product for their team or company. Looking for reasons to reject. Comparing against competitors. +**Profile**: Methodical user who pushes interfaces beyond the happy path. Tests edge cases, tries unexpected inputs, and probes for gaps in the experience. **Behaviors**: - Tests edge cases intentionally (empty states, long strings, special characters) -- Looks for pricing catches and hidden limitations -- Reads fine print and terms of service -- Tries to break things deliberately +- Submits forms with unexpected data (emoji, RTL text, very long values) +- Tries to break workflows by navigating backwards, refreshing mid-flow, or opening in multiple tabs +- Looks for inconsistencies between what the UI promises and what actually happens - Documents problems methodically **Test Questions**: - What happens at the edges (0 items, 1000 items, very long text)? -- Is pricing and value proposition transparent? -- Are there hidden limitations or gotchas? -- How polished is error handling? -- What data is collected and why? +- Do error states recover gracefully or leave the UI in a broken state? +- What happens on refresh mid-workflow? Is state preserved? +- Are there features that appear to work but produce broken results? +- How does the UI handle unexpected input (emoji, special chars, paste from Excel)? **Red Flags** (report these specifically): -- Hidden pricing or "contact sales" for basic information -- Features that appear to work but produce broken results -- Poor error handling that exposes technical details -- Unclear data practices or missing privacy information +- Features that appear to work but silently fail or produce wrong results +- Error handling that exposes technical details or leaves UI in a broken state - Empty states that show nothing useful ("No results" with no guidance) +- Workflows that lose user data on refresh or navigation +- Inconsistent behavior between similar interactions in different parts of the UI --- @@ -150,7 +150,7 @@ Choose personas based on the interface type: |---------------|-----------------|-----| | Landing page / marketing | Jordan, Riley, Casey | First impressions, trust, mobile | | Dashboard / admin | Alex, Sam | Power users, accessibility | -| E-commerce / checkout | Casey, Riley, Jordan | Mobile, trust, clarity | +| E-commerce / checkout | Casey, Riley, Jordan | Mobile, edge cases, clarity | | Onboarding flow | Jordan, Casey | Confusion, interruption | | Data-heavy / analytics | Alex, Sam | Efficiency, keyboard nav | | Form-heavy / wizard | Jordan, Sam, Casey | Clarity, accessibility, mobile | diff --git a/.kiro/skills/audit/SKILL.md b/.kiro/skills/audit/SKILL.md index d84428320..6dc747d60 100644 --- a/.kiro/skills/audit/SKILL.md +++ b/.kiro/skills/audit/SKILL.md @@ -1,15 +1,21 @@ --- name: audit -description: Perform a comprehensive audit of interface quality across accessibility, performance, theming, and responsive design. Generates a scored report with severity ratings and actionable plan. Use when the user wants a design review, accessibility check, quality audit, or a full list of UI issues to fix. +description: Run technical quality checks across accessibility, performance, theming, responsive design, and anti-patterns. Generates a scored report with P0-P3 severity ratings and actionable plan. Use when the user wants an accessibility check, performance audit, or technical quality review. --- -Run systematic quality checks and generate a comprehensive audit report with quantitative scoring, prioritized issues, and an actionable plan. Don't fix issues — document them for other commands to address. +## MANDATORY PREPARATION -**First**: Invoke /frontend-design for design principles and anti-patterns. +Invoke /frontend-design — it contains design principles, anti-patterns, and the **Context Gathering Protocol**. Follow the protocol before proceeding — if no design context exists yet, you MUST run /teach-impeccable first. + +--- + +Run systematic **technical** quality checks and generate a comprehensive report. Don't fix issues — document them for other commands to address. + +This is a code-level audit, not a design critique. Check what's measurable and verifiable in the implementation. ## Diagnostic Scan -Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using the criteria below. +Run comprehensive checks across 5 dimensions. Score each dimension 0-4 using the criteria below. ### 1. Accessibility (A11y) @@ -21,7 +27,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Alt text**: Missing or poor image descriptions - **Form issues**: Inputs without labels, poor error messaging, missing required indicators -**Score 0–4**: 0=Inaccessible (fails WCAG A), 1=Major gaps (few ARIA labels, no keyboard nav), 2=Partial (some a11y effort, significant gaps), 3=Good (WCAG AA mostly met, minor gaps), 4=Excellent (WCAG AA fully met, approaches AAA) +**Score 0-4**: 0=Inaccessible (fails WCAG A), 1=Major gaps (few ARIA labels, no keyboard nav), 2=Partial (some a11y effort, significant gaps), 3=Good (WCAG AA mostly met, minor gaps), 4=Excellent (WCAG AA fully met, approaches AAA) ### 2. Performance @@ -32,7 +38,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Bundle size**: Unnecessary imports, unused dependencies - **Render performance**: Unnecessary re-renders, missing memoization -**Score 0–4**: 0=Severe issues (layout thrash, unoptimized everything), 1=Major problems (no lazy loading, expensive animations), 2=Partial (some optimization, gaps remain), 3=Good (mostly optimized, minor improvements possible), 4=Excellent (fast, lean, well-optimized) +**Score 0-4**: 0=Severe issues (layout thrash, unoptimized everything), 1=Major problems (no lazy loading, expensive animations), 2=Partial (some optimization, gaps remain), 3=Good (mostly optimized, minor improvements possible), 4=Excellent (fast, lean, well-optimized) ### 3. Theming @@ -42,7 +48,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Inconsistent tokens**: Using wrong tokens, mixing token types - **Theme switching issues**: Values that don't update on theme change -**Score 0–4**: 0=No theming (hard-coded everything), 1=Minimal tokens (mostly hard-coded), 2=Partial (tokens exist but inconsistently used), 3=Good (tokens used, minor hard-coded values), 4=Excellent (full token system, dark mode works perfectly) +**Score 0-4**: 0=No theming (hard-coded everything), 1=Minimal tokens (mostly hard-coded), 2=Partial (tokens exist but inconsistently used), 3=Good (tokens used, minor hard-coded values), 4=Excellent (full token system, dark mode works perfectly) ### 4. Responsive Design @@ -53,113 +59,73 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Text scaling**: Layouts that break when text size increases - **Missing breakpoints**: No mobile/tablet variants -**Score 0–4**: 0=Desktop-only (breaks on mobile), 1=Major issues (some breakpoints, many failures), 2=Partial (works on mobile, rough edges), 3=Good (responsive, minor touch target or overflow issues), 4=Excellent (fluid, all viewports, proper touch targets) +**Score 0-4**: 0=Desktop-only (breaks on mobile), 1=Major issues (some breakpoints, many failures), 2=Partial (works on mobile, rough edges), 3=Good (responsive, minor touch target or overflow issues), 4=Excellent (fluid, all viewports, proper touch targets) ### 5. Anti-Patterns (CRITICAL) Check against ALL the **DON'T** guidelines in the frontend-design skill. Look for AI slop tells (AI color palette, gradient text, glassmorphism, hero metrics, card grids, generic fonts) and general design anti-patterns (gray on color, nested cards, bounce easing, redundant copy). -**Score 0–4**: 0=AI slop gallery (5+ tells), 1=Heavy AI aesthetic (3-4 tells), 2=Some tells (1-2 noticeable), 3=Mostly clean (subtle issues only), 4=No AI tells (distinctive, intentional design) +**Score 0-4**: 0=AI slop gallery (5+ tells), 1=Heavy AI aesthetic (3-4 tells), 2=Some tells (1-2 noticeable), 3=Mostly clean (subtle issues only), 4=No AI tells (distinctive, intentional design) -**CRITICAL**: This is an audit, not a fix. Document issues thoroughly with clear explanations of impact. Use other commands to fix issues after audit. - -## Generate Comprehensive Report +## Generate Report ### Audit Health Score -Present the dimension scores as a table: - | # | Dimension | Score | Key Finding | |---|-----------|-------|-------------| -| 1 | Accessibility | ? | [most critical a11y issue or "—"] | +| 1 | Accessibility | ? | [most critical a11y issue or "--"] | | 2 | Performance | ? | | | 3 | Responsive Design | ? | | | 4 | Theming | ? | | | 5 | Anti-Patterns | ? | | | **Total** | | **??/20** | **[Rating band]** | -**Rating bands**: -| Score | Rating | Action | -|-------|--------|--------| -| 18–20 | Excellent | Minor polish only | -| 14–17 | Good | Address weak dimensions | -| 10–13 | Acceptable | Significant work needed | -| 6–9 | Poor | Major quality overhaul | -| 0–5 | Critical | Fundamental issues across the board | +**Rating bands**: 18-20 Excellent (minor polish), 14-17 Good (address weak dimensions), 10-13 Acceptable (significant work needed), 6-9 Poor (major overhaul), 0-5 Critical (fundamental issues) ### Anti-Patterns Verdict -**Start here.** Pass/fail: Does this look AI-generated? List specific tells from the skill's Anti-Patterns section. Be brutally honest. +**Start here.** Pass/fail: Does this look AI-generated? List specific tells. Be brutally honest. ### Executive Summary - Audit Health Score: **??/20** ([rating band]) - Total issues found (count by severity: P0/P1/P2/P3) -- Most critical issues (top 3-5) +- Top 3-5 critical issues - Recommended next steps ### Detailed Findings by Severity -Tag every issue with **P0–P3 severity**: -| Priority | Name | Description | -|----------|------|-------------| -| **P0** | Blocking | Prevents task completion — fix immediately | -| **P1** | Major | Significant difficulty or WCAG AA violation — fix before release | -| **P2** | Minor | Annoyance, workaround exists — fix in next pass | -| **P3** | Polish | Nice-to-fix, no real user impact — fix if time permits | +Tag every issue with **P0-P3 severity**: +- **P0 Blocking**: Prevents task completion — fix immediately +- **P1 Major**: Significant difficulty or WCAG AA violation — fix before release +- **P2 Minor**: Annoyance, workaround exists — fix in next pass +- **P3 Polish**: Nice-to-fix, no real user impact — fix if time permits For each issue, document: - **[P?] Issue name** -- **Location**: Where it occurs (component, file, line) +- **Location**: Component, file, line - **Category**: Accessibility / Performance / Theming / Responsive / Anti-Pattern -- **Description**: What the issue is - **Impact**: How it affects users - **WCAG/Standard**: Which standard it violates (if applicable) - **Recommendation**: How to fix it -- **Suggested command**: Which command to use (prefer: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive — or other installed skills you're sure exist) - -#### P0 — Blocking Issues -[Issues that prevent task completion or violate WCAG A] - -#### P1 — Major Issues -[Significant usability/accessibility impact, WCAG AA violations] - -#### P2 — Minor Issues -[Quality issues, WCAG AAA violations, performance concerns] - -#### P3 — Polish Issues -[Minor inconsistencies, optimization opportunities] +- **Suggested command**: Which command to use (prefer: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive) ### Patterns & Systemic Issues -Identify recurring problems: +Identify recurring problems that indicate systemic gaps rather than one-off mistakes: - "Hard-coded colors appear in 15+ components, should use design tokens" - "Touch targets consistently too small (<44px) throughout mobile experience" -- "Missing focus indicators on all custom interactive components" ### Positive Findings -Note what's working well: -- Good practices to maintain -- Exemplary implementations to replicate elsewhere +Note what's working well — good practices to maintain and replicate. ## Recommended Actions -Present a prioritized action summary. Order is determined by severity automatically (P0 first, then P1, then P2). - -### Action Summary - -List recommended commands in priority order: +List recommended commands in priority order (P0 first, then P1, then P2): 1. **[P?] `/command-name`** — Brief description (specific context from audit findings) 2. **[P?] `/command-name`** — Brief description (specific context) -... -**Rules for recommendations**: -- Only recommend commands from: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive -- Order by severity: P0 issues first, then P1, then P2 (skip P3 unless user has few issues) -- Each item's description should carry enough context that the command knows what to focus on -- Map findings to the most appropriate command -- Skip commands that would address zero issues -- End with `/polish` as the final step if any fixes were recommended +**Rules**: Only recommend commands from: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive. Map findings to the most appropriate command. End with `/polish` as the final step if any fixes were recommended. After presenting the summary, tell the user: @@ -171,10 +137,9 @@ After presenting the summary, tell the user: **NEVER**: - Report issues without explaining impact (why does this matter?) -- Mix severity levels inconsistently -- Skip positive findings (celebrate what works) - Provide generic recommendations (be specific and actionable) +- Skip positive findings (celebrate what works) - Forget to prioritize (everything can't be P0) - Report false positives without verification -Remember: You're a quality auditor with exceptional attention to detail. Document systematically, prioritize ruthlessly, and provide clear paths to improvement. A good audit makes fixing easy. \ No newline at end of file +Remember: You're a technical quality auditor. Document systematically, prioritize ruthlessly, cite specific code locations, and provide clear paths to improvement. \ No newline at end of file diff --git a/.kiro/skills/critique/SKILL.md b/.kiro/skills/critique/SKILL.md index 28590d3f7..5d9660aae 100644 --- a/.kiro/skills/critique/SKILL.md +++ b/.kiro/skills/critique/SKILL.md @@ -1,6 +1,6 @@ --- name: critique -description: Evaluate design effectiveness from a UX perspective. Assesses visual hierarchy, information architecture, emotional resonance, cognitive load, and overall design quality with quantitative scoring, persona-based testing, and actionable feedback. Use when the user asks to review, critique, evaluate, or give feedback on a design, UI, or component. +description: Evaluate design from a UX perspective, assessing visual hierarchy, information architecture, emotional resonance, cognitive load, and overall quality with quantitative scoring, persona-based testing, and actionable feedback. Use when the user asks to review, critique, evaluate, or give feedback on a design or component. --- ## MANDATORY PREPARATION @@ -9,7 +9,7 @@ Invoke /frontend-design — it contains design principles, anti-patterns, and th --- -Conduct a holistic design critique, evaluating whether the interface actually works—not just technically, but as a designed experience. Think like a design director giving feedback. +Conduct a holistic design critique, evaluating whether the interface actually works — not just technically, but as a designed experience. Think like a design director giving feedback. ## Phase 1: Design Critique @@ -19,7 +19,7 @@ Evaluate the interface across these dimensions: **This is the most important check.** Does this look like every other AI-generated interface from 2024-2025? -Review the design against ALL the **DON'T** guidelines in the frontend-design skill—they are the fingerprints of AI-generated work. Check for the AI color palette, gradient text, dark mode with glowing accents, glassmorphism, hero metric layouts, identical card grids, generic fonts, and all other tells. +Review the design against ALL the **DON'T** guidelines in the frontend-design skill — they are the fingerprints of AI-generated work. Check for the AI color palette, gradient text, dark mode with glowing accents, glassmorphism, hero metric layouts, identical card grids, generic fonts, and all other tells. **The test**: If you showed this to someone and said "AI made this," would they believe you immediately? If yes, that's the problem. @@ -29,20 +29,19 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Do size, color, and position communicate importance correctly? - Is there visual competition between elements that should have different weights? -### 3. Information Architecture -→ *Consult [cognitive-load](reference/cognitive-load.md) for the working memory rule and checklist* +### 3. Information Architecture & Cognitive Load +> *Consult [cognitive-load](reference/cognitive-load.md) for the working memory rule and 8-item checklist* - Is the structure intuitive? Would a new user understand the organization? - Is related content grouped logically? - Are there too many choices at once? Count visible options at each decision point — if >4, flag it - Is the navigation clear and predictable? - **Progressive disclosure**: Is complexity revealed only when needed, or dumped on the user upfront? -- **Cognitive load sub-check**: Run the 8-item cognitive load checklist from the reference. Report the number of failures. +- **Run the 8-item cognitive load checklist** from the reference. Report failure count: 0–1 = low (good), 2–3 = moderate, 4+ = critical. ### 4. Emotional Journey -→ *Consult [cognitive-load](reference/cognitive-load.md) for emotional intervention patterns* - What emotion does this interface evoke? Is that intentional? - Does it match the brand personality? -- Does it feel trustworthy, approachable, premium, playful—whatever it should feel? +- Does it feel trustworthy, approachable, premium, playful — whatever it should feel? - Would the target user feel "this is for me"? - **Peak-end rule**: Is the most intense moment positive? Does the experience end well (confirmation, celebration, clear next step)? - **Emotional valleys**: Check for onboarding frustration, error cliffs, feature discovery gaps, or anxiety spikes at high-stakes moments (payment, delete, commit) @@ -70,7 +69,7 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Is color used to communicate, not just decorate? - Does the palette feel cohesive? - Are accent colors drawing attention to the right things? -- Does it work for colorblind users? (not just technically—does meaning still come through?) +- Does it work for colorblind users? (not just technically — does meaning still come through?) ### 9. States & Edge Cases - Empty states: Do they guide users toward action, or just say "nothing here"? @@ -84,27 +83,19 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Are labels and buttons unambiguous? - Does error copy help users fix the problem? -### 11. Cognitive Load -→ *Consult [cognitive-load](reference/cognitive-load.md)* -- **Intrinsic vs. extraneous**: Is the mental effort coming from the task itself (acceptable) or from poor design choices (eliminate)? -- **Decision points**: Count visible choices at key moments. More than 4 simultaneous options = overload. -- **Working memory burden**: Does the user need to remember information from a previous screen to act on the current one? -- **Information chunking**: Is content broken into digestible groups, or presented as undifferentiated walls? -- Run the 8-item cognitive load checklist. Report failures count: 0–1 = low (good), 2–3 = moderate, 4+ = critical. - ## Phase 2: Present Findings Structure your feedback as a design director would: ### Design Health Score -→ *Consult [heuristics-scoring](reference/heuristics-scoring.md)* +> *Consult [heuristics-scoring](reference/heuristics-scoring.md)* Score each of Nielsen's 10 heuristics 0–4. Present as a table: | # | Heuristic | Score | Key Issue | |---|-----------|-------|-----------| | 1 | Visibility of System Status | ? | [specific finding or "—" if solid] | -| 2 | Match System ↔ Real World | ? | | +| 2 | Match System / Real World | ? | | | 3 | User Control and Freedom | ? | | | 4 | Consistency and Standards | ? | | | 5 | Error Prevention | ? | | @@ -121,13 +112,13 @@ Be honest with scores. A 4 means genuinely excellent. Most real interfaces score **Start here.** Pass/fail: Does this look AI-generated? List specific tells from the skill's Anti-Patterns section. Be brutally honest. ### Overall Impression -A brief gut reaction—what works, what doesn't, and the single biggest opportunity. +A brief gut reaction — what works, what doesn't, and the single biggest opportunity. ### What's Working -Highlight 2-3 things done well. Be specific about why they work. +Highlight 2–3 things done well. Be specific about why they work. ### Priority Issues -The 3-5 most impactful design problems, ordered by importance. +The 3–5 most impactful design problems, ordered by importance. For each issue, tag with **P0–P3 severity** (consult [heuristics-scoring](reference/heuristics-scoring.md) for severity definitions): - **[P?] What**: Name the problem clearly @@ -136,7 +127,7 @@ For each issue, tag with **P0–P3 severity** (consult [heuristics-scoring](refe - **Suggested command**: Which command could address this (from: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive) ### Persona Red Flags -→ *Consult [personas](reference/personas.md)* +> *Consult [personas](reference/personas.md)* Auto-select 2–3 personas most relevant to this interface type (use the selection table in the reference). If `.kiro/settings.json` contains a `## Design Context` section from `teach-impeccable`, also generate 1–2 project-specific personas from the audience/brand info. @@ -152,12 +143,12 @@ Be specific — name the exact elements and interactions that fail each persona. Quick notes on smaller issues worth addressing. **Remember**: -- Be direct—vague feedback wastes everyone's time -- Be specific—"the submit button" not "some elements" +- Be direct — vague feedback wastes everyone's time +- Be specific — "the submit button" not "some elements" - Say what's wrong AND why it matters to users - Give concrete suggestions, not just "consider exploring..." -- Prioritize ruthlessly—if everything is important, nothing is -- Don't soften criticism—developers need honest feedback to ship great design +- Prioritize ruthlessly — if everything is important, nothing is +- Don't soften criticism — developers need honest feedback to ship great design ## Phase 3: Ask the User @@ -165,9 +156,9 @@ Quick notes on smaller issues worth addressing. Ask questions along these lines (adapt to the specific findings — do NOT ask generic questions): -1. **Priority direction**: Based on the issues found, ask which category matters most to the user right now. For example: "I found problems with visual hierarchy, color usage, and information overload. Which area should we tackle first?" Offer the top 2-3 issue categories as options. +1. **Priority direction**: Based on the issues found, ask which category matters most to the user right now. For example: "I found problems with visual hierarchy, color usage, and information overload. Which area should we tackle first?" Offer the top 2–3 issue categories as options. -2. **Design intent**: If the critique found a tonal mismatch, ask whether it was intentional. For example: "The interface feels clinical and corporate. Is that the intended tone, or should it feel warmer/bolder/more playful?" Offer 2-3 tonal directions as options based on what would fix the issues found. +2. **Design intent**: If the critique found a tonal mismatch, ask whether it was intentional. For example: "The interface feels clinical and corporate. Is that the intended tone, or should it feel warmer/bolder/more playful?" Offer 2–3 tonal directions as options based on what would fix the issues found. 3. **Scope**: Ask how much the user wants to take on. For example: "I found N issues. Want to address everything, or focus on the top 3?" Offer scope options like "Top 3 only", "All issues", "Critical issues only". @@ -175,9 +166,9 @@ Ask questions along these lines (adapt to the specific findings — do NOT ask g **Rules for questions**: - Every question must reference specific findings from Phase 2 — never ask generic "who is your audience?" questions -- Keep it to 2-4 questions maximum — respect the user's time +- Keep it to 2–4 questions maximum — respect the user's time - Offer concrete options, not open-ended prompts -- If findings are straightforward (e.g., only 1-2 clear issues), skip questions and go directly to Phase 4 +- If findings are straightforward (e.g., only 1–2 clear issues), skip questions and go directly to Phase 4 ## Phase 4: Recommended Actions diff --git a/.kiro/skills/critique/reference/personas.md b/.kiro/skills/critique/reference/personas.md index bf9296c16..615d30d74 100644 --- a/.kiro/skills/critique/reference/personas.md +++ b/.kiro/skills/critique/reference/personas.md @@ -88,30 +88,30 @@ Test the interface through the eyes of 5 distinct user archetypes. Each persona --- -## 4. Skeptical Evaluator — "Riley" +## 4. Deliberate Stress Tester — "Riley" -**Profile**: Evaluating the product for their team or company. Looking for reasons to reject. Comparing against competitors. +**Profile**: Methodical user who pushes interfaces beyond the happy path. Tests edge cases, tries unexpected inputs, and probes for gaps in the experience. **Behaviors**: - Tests edge cases intentionally (empty states, long strings, special characters) -- Looks for pricing catches and hidden limitations -- Reads fine print and terms of service -- Tries to break things deliberately +- Submits forms with unexpected data (emoji, RTL text, very long values) +- Tries to break workflows by navigating backwards, refreshing mid-flow, or opening in multiple tabs +- Looks for inconsistencies between what the UI promises and what actually happens - Documents problems methodically **Test Questions**: - What happens at the edges (0 items, 1000 items, very long text)? -- Is pricing and value proposition transparent? -- Are there hidden limitations or gotchas? -- How polished is error handling? -- What data is collected and why? +- Do error states recover gracefully or leave the UI in a broken state? +- What happens on refresh mid-workflow? Is state preserved? +- Are there features that appear to work but produce broken results? +- How does the UI handle unexpected input (emoji, special chars, paste from Excel)? **Red Flags** (report these specifically): -- Hidden pricing or "contact sales" for basic information -- Features that appear to work but produce broken results -- Poor error handling that exposes technical details -- Unclear data practices or missing privacy information +- Features that appear to work but silently fail or produce wrong results +- Error handling that exposes technical details or leaves UI in a broken state - Empty states that show nothing useful ("No results" with no guidance) +- Workflows that lose user data on refresh or navigation +- Inconsistent behavior between similar interactions in different parts of the UI --- @@ -150,7 +150,7 @@ Choose personas based on the interface type: |---------------|-----------------|-----| | Landing page / marketing | Jordan, Riley, Casey | First impressions, trust, mobile | | Dashboard / admin | Alex, Sam | Power users, accessibility | -| E-commerce / checkout | Casey, Riley, Jordan | Mobile, trust, clarity | +| E-commerce / checkout | Casey, Riley, Jordan | Mobile, edge cases, clarity | | Onboarding flow | Jordan, Casey | Confusion, interruption | | Data-heavy / analytics | Alex, Sam | Efficiency, keyboard nav | | Form-heavy / wizard | Jordan, Sam, Casey | Clarity, accessibility, mobile | diff --git a/.opencode/skills/audit/SKILL.md b/.opencode/skills/audit/SKILL.md index 74bb05abc..1debe043e 100644 --- a/.opencode/skills/audit/SKILL.md +++ b/.opencode/skills/audit/SKILL.md @@ -1,17 +1,23 @@ --- name: audit -description: Perform a comprehensive audit of interface quality across accessibility, performance, theming, and responsive design. Generates a scored report with severity ratings and actionable plan. Use when the user wants a design review, accessibility check, quality audit, or a full list of UI issues to fix. +description: Run technical quality checks across accessibility, performance, theming, responsive design, and anti-patterns. Generates a scored report with P0-P3 severity ratings and actionable plan. Use when the user wants an accessibility check, performance audit, or technical quality review. user-invocable: true argument-hint: "[area (feature, page, component...)]" --- -Run systematic quality checks and generate a comprehensive audit report with quantitative scoring, prioritized issues, and an actionable plan. Don't fix issues — document them for other commands to address. +## MANDATORY PREPARATION -**First**: Invoke /frontend-design for design principles and anti-patterns. +Invoke /frontend-design — it contains design principles, anti-patterns, and the **Context Gathering Protocol**. Follow the protocol before proceeding — if no design context exists yet, you MUST run /teach-impeccable first. + +--- + +Run systematic **technical** quality checks and generate a comprehensive report. Don't fix issues — document them for other commands to address. + +This is a code-level audit, not a design critique. Check what's measurable and verifiable in the implementation. ## Diagnostic Scan -Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using the criteria below. +Run comprehensive checks across 5 dimensions. Score each dimension 0-4 using the criteria below. ### 1. Accessibility (A11y) @@ -23,7 +29,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Alt text**: Missing or poor image descriptions - **Form issues**: Inputs without labels, poor error messaging, missing required indicators -**Score 0–4**: 0=Inaccessible (fails WCAG A), 1=Major gaps (few ARIA labels, no keyboard nav), 2=Partial (some a11y effort, significant gaps), 3=Good (WCAG AA mostly met, minor gaps), 4=Excellent (WCAG AA fully met, approaches AAA) +**Score 0-4**: 0=Inaccessible (fails WCAG A), 1=Major gaps (few ARIA labels, no keyboard nav), 2=Partial (some a11y effort, significant gaps), 3=Good (WCAG AA mostly met, minor gaps), 4=Excellent (WCAG AA fully met, approaches AAA) ### 2. Performance @@ -34,7 +40,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Bundle size**: Unnecessary imports, unused dependencies - **Render performance**: Unnecessary re-renders, missing memoization -**Score 0–4**: 0=Severe issues (layout thrash, unoptimized everything), 1=Major problems (no lazy loading, expensive animations), 2=Partial (some optimization, gaps remain), 3=Good (mostly optimized, minor improvements possible), 4=Excellent (fast, lean, well-optimized) +**Score 0-4**: 0=Severe issues (layout thrash, unoptimized everything), 1=Major problems (no lazy loading, expensive animations), 2=Partial (some optimization, gaps remain), 3=Good (mostly optimized, minor improvements possible), 4=Excellent (fast, lean, well-optimized) ### 3. Theming @@ -44,7 +50,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Inconsistent tokens**: Using wrong tokens, mixing token types - **Theme switching issues**: Values that don't update on theme change -**Score 0–4**: 0=No theming (hard-coded everything), 1=Minimal tokens (mostly hard-coded), 2=Partial (tokens exist but inconsistently used), 3=Good (tokens used, minor hard-coded values), 4=Excellent (full token system, dark mode works perfectly) +**Score 0-4**: 0=No theming (hard-coded everything), 1=Minimal tokens (mostly hard-coded), 2=Partial (tokens exist but inconsistently used), 3=Good (tokens used, minor hard-coded values), 4=Excellent (full token system, dark mode works perfectly) ### 4. Responsive Design @@ -55,113 +61,73 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Text scaling**: Layouts that break when text size increases - **Missing breakpoints**: No mobile/tablet variants -**Score 0–4**: 0=Desktop-only (breaks on mobile), 1=Major issues (some breakpoints, many failures), 2=Partial (works on mobile, rough edges), 3=Good (responsive, minor touch target or overflow issues), 4=Excellent (fluid, all viewports, proper touch targets) +**Score 0-4**: 0=Desktop-only (breaks on mobile), 1=Major issues (some breakpoints, many failures), 2=Partial (works on mobile, rough edges), 3=Good (responsive, minor touch target or overflow issues), 4=Excellent (fluid, all viewports, proper touch targets) ### 5. Anti-Patterns (CRITICAL) Check against ALL the **DON'T** guidelines in the frontend-design skill. Look for AI slop tells (AI color palette, gradient text, glassmorphism, hero metrics, card grids, generic fonts) and general design anti-patterns (gray on color, nested cards, bounce easing, redundant copy). -**Score 0–4**: 0=AI slop gallery (5+ tells), 1=Heavy AI aesthetic (3-4 tells), 2=Some tells (1-2 noticeable), 3=Mostly clean (subtle issues only), 4=No AI tells (distinctive, intentional design) +**Score 0-4**: 0=AI slop gallery (5+ tells), 1=Heavy AI aesthetic (3-4 tells), 2=Some tells (1-2 noticeable), 3=Mostly clean (subtle issues only), 4=No AI tells (distinctive, intentional design) -**CRITICAL**: This is an audit, not a fix. Document issues thoroughly with clear explanations of impact. Use other commands to fix issues after audit. - -## Generate Comprehensive Report +## Generate Report ### Audit Health Score -Present the dimension scores as a table: - | # | Dimension | Score | Key Finding | |---|-----------|-------|-------------| -| 1 | Accessibility | ? | [most critical a11y issue or "—"] | +| 1 | Accessibility | ? | [most critical a11y issue or "--"] | | 2 | Performance | ? | | | 3 | Responsive Design | ? | | | 4 | Theming | ? | | | 5 | Anti-Patterns | ? | | | **Total** | | **??/20** | **[Rating band]** | -**Rating bands**: -| Score | Rating | Action | -|-------|--------|--------| -| 18–20 | Excellent | Minor polish only | -| 14–17 | Good | Address weak dimensions | -| 10–13 | Acceptable | Significant work needed | -| 6–9 | Poor | Major quality overhaul | -| 0–5 | Critical | Fundamental issues across the board | +**Rating bands**: 18-20 Excellent (minor polish), 14-17 Good (address weak dimensions), 10-13 Acceptable (significant work needed), 6-9 Poor (major overhaul), 0-5 Critical (fundamental issues) ### Anti-Patterns Verdict -**Start here.** Pass/fail: Does this look AI-generated? List specific tells from the skill's Anti-Patterns section. Be brutally honest. +**Start here.** Pass/fail: Does this look AI-generated? List specific tells. Be brutally honest. ### Executive Summary - Audit Health Score: **??/20** ([rating band]) - Total issues found (count by severity: P0/P1/P2/P3) -- Most critical issues (top 3-5) +- Top 3-5 critical issues - Recommended next steps ### Detailed Findings by Severity -Tag every issue with **P0–P3 severity**: -| Priority | Name | Description | -|----------|------|-------------| -| **P0** | Blocking | Prevents task completion — fix immediately | -| **P1** | Major | Significant difficulty or WCAG AA violation — fix before release | -| **P2** | Minor | Annoyance, workaround exists — fix in next pass | -| **P3** | Polish | Nice-to-fix, no real user impact — fix if time permits | +Tag every issue with **P0-P3 severity**: +- **P0 Blocking**: Prevents task completion — fix immediately +- **P1 Major**: Significant difficulty or WCAG AA violation — fix before release +- **P2 Minor**: Annoyance, workaround exists — fix in next pass +- **P3 Polish**: Nice-to-fix, no real user impact — fix if time permits For each issue, document: - **[P?] Issue name** -- **Location**: Where it occurs (component, file, line) +- **Location**: Component, file, line - **Category**: Accessibility / Performance / Theming / Responsive / Anti-Pattern -- **Description**: What the issue is - **Impact**: How it affects users - **WCAG/Standard**: Which standard it violates (if applicable) - **Recommendation**: How to fix it -- **Suggested command**: Which command to use (prefer: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive — or other installed skills you're sure exist) - -#### P0 — Blocking Issues -[Issues that prevent task completion or violate WCAG A] - -#### P1 — Major Issues -[Significant usability/accessibility impact, WCAG AA violations] - -#### P2 — Minor Issues -[Quality issues, WCAG AAA violations, performance concerns] - -#### P3 — Polish Issues -[Minor inconsistencies, optimization opportunities] +- **Suggested command**: Which command to use (prefer: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive) ### Patterns & Systemic Issues -Identify recurring problems: +Identify recurring problems that indicate systemic gaps rather than one-off mistakes: - "Hard-coded colors appear in 15+ components, should use design tokens" - "Touch targets consistently too small (<44px) throughout mobile experience" -- "Missing focus indicators on all custom interactive components" ### Positive Findings -Note what's working well: -- Good practices to maintain -- Exemplary implementations to replicate elsewhere +Note what's working well — good practices to maintain and replicate. ## Recommended Actions -Present a prioritized action summary. Order is determined by severity automatically (P0 first, then P1, then P2). - -### Action Summary - -List recommended commands in priority order: +List recommended commands in priority order (P0 first, then P1, then P2): 1. **[P?] `/command-name`** — Brief description (specific context from audit findings) 2. **[P?] `/command-name`** — Brief description (specific context) -... -**Rules for recommendations**: -- Only recommend commands from: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive -- Order by severity: P0 issues first, then P1, then P2 (skip P3 unless user has few issues) -- Each item's description should carry enough context that the command knows what to focus on -- Map findings to the most appropriate command -- Skip commands that would address zero issues -- End with `/polish` as the final step if any fixes were recommended +**Rules**: Only recommend commands from: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive. Map findings to the most appropriate command. End with `/polish` as the final step if any fixes were recommended. After presenting the summary, tell the user: @@ -173,10 +139,9 @@ After presenting the summary, tell the user: **NEVER**: - Report issues without explaining impact (why does this matter?) -- Mix severity levels inconsistently -- Skip positive findings (celebrate what works) - Provide generic recommendations (be specific and actionable) +- Skip positive findings (celebrate what works) - Forget to prioritize (everything can't be P0) - Report false positives without verification -Remember: You're a quality auditor with exceptional attention to detail. Document systematically, prioritize ruthlessly, and provide clear paths to improvement. A good audit makes fixing easy. \ No newline at end of file +Remember: You're a technical quality auditor. Document systematically, prioritize ruthlessly, cite specific code locations, and provide clear paths to improvement. \ No newline at end of file diff --git a/.opencode/skills/critique/SKILL.md b/.opencode/skills/critique/SKILL.md index acc09f8c0..74f39ab4e 100644 --- a/.opencode/skills/critique/SKILL.md +++ b/.opencode/skills/critique/SKILL.md @@ -1,6 +1,6 @@ --- name: critique -description: Evaluate design effectiveness from a UX perspective. Assesses visual hierarchy, information architecture, emotional resonance, cognitive load, and overall design quality with quantitative scoring, persona-based testing, and actionable feedback. Use when the user asks to review, critique, evaluate, or give feedback on a design, UI, or component. +description: Evaluate design from a UX perspective, assessing visual hierarchy, information architecture, emotional resonance, cognitive load, and overall quality with quantitative scoring, persona-based testing, and actionable feedback. Use when the user asks to review, critique, evaluate, or give feedback on a design or component. user-invocable: true argument-hint: "[area (feature, page, component...)]" --- @@ -11,7 +11,7 @@ Invoke /frontend-design — it contains design principles, anti-patterns, and th --- -Conduct a holistic design critique, evaluating whether the interface actually works—not just technically, but as a designed experience. Think like a design director giving feedback. +Conduct a holistic design critique, evaluating whether the interface actually works — not just technically, but as a designed experience. Think like a design director giving feedback. ## Phase 1: Design Critique @@ -21,7 +21,7 @@ Evaluate the interface across these dimensions: **This is the most important check.** Does this look like every other AI-generated interface from 2024-2025? -Review the design against ALL the **DON'T** guidelines in the frontend-design skill—they are the fingerprints of AI-generated work. Check for the AI color palette, gradient text, dark mode with glowing accents, glassmorphism, hero metric layouts, identical card grids, generic fonts, and all other tells. +Review the design against ALL the **DON'T** guidelines in the frontend-design skill — they are the fingerprints of AI-generated work. Check for the AI color palette, gradient text, dark mode with glowing accents, glassmorphism, hero metric layouts, identical card grids, generic fonts, and all other tells. **The test**: If you showed this to someone and said "AI made this," would they believe you immediately? If yes, that's the problem. @@ -31,20 +31,19 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Do size, color, and position communicate importance correctly? - Is there visual competition between elements that should have different weights? -### 3. Information Architecture -→ *Consult [cognitive-load](reference/cognitive-load.md) for the working memory rule and checklist* +### 3. Information Architecture & Cognitive Load +> *Consult [cognitive-load](reference/cognitive-load.md) for the working memory rule and 8-item checklist* - Is the structure intuitive? Would a new user understand the organization? - Is related content grouped logically? - Are there too many choices at once? Count visible options at each decision point — if >4, flag it - Is the navigation clear and predictable? - **Progressive disclosure**: Is complexity revealed only when needed, or dumped on the user upfront? -- **Cognitive load sub-check**: Run the 8-item cognitive load checklist from the reference. Report the number of failures. +- **Run the 8-item cognitive load checklist** from the reference. Report failure count: 0–1 = low (good), 2–3 = moderate, 4+ = critical. ### 4. Emotional Journey -→ *Consult [cognitive-load](reference/cognitive-load.md) for emotional intervention patterns* - What emotion does this interface evoke? Is that intentional? - Does it match the brand personality? -- Does it feel trustworthy, approachable, premium, playful—whatever it should feel? +- Does it feel trustworthy, approachable, premium, playful — whatever it should feel? - Would the target user feel "this is for me"? - **Peak-end rule**: Is the most intense moment positive? Does the experience end well (confirmation, celebration, clear next step)? - **Emotional valleys**: Check for onboarding frustration, error cliffs, feature discovery gaps, or anxiety spikes at high-stakes moments (payment, delete, commit) @@ -72,7 +71,7 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Is color used to communicate, not just decorate? - Does the palette feel cohesive? - Are accent colors drawing attention to the right things? -- Does it work for colorblind users? (not just technically—does meaning still come through?) +- Does it work for colorblind users? (not just technically — does meaning still come through?) ### 9. States & Edge Cases - Empty states: Do they guide users toward action, or just say "nothing here"? @@ -86,27 +85,19 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Are labels and buttons unambiguous? - Does error copy help users fix the problem? -### 11. Cognitive Load -→ *Consult [cognitive-load](reference/cognitive-load.md)* -- **Intrinsic vs. extraneous**: Is the mental effort coming from the task itself (acceptable) or from poor design choices (eliminate)? -- **Decision points**: Count visible choices at key moments. More than 4 simultaneous options = overload. -- **Working memory burden**: Does the user need to remember information from a previous screen to act on the current one? -- **Information chunking**: Is content broken into digestible groups, or presented as undifferentiated walls? -- Run the 8-item cognitive load checklist. Report failures count: 0–1 = low (good), 2–3 = moderate, 4+ = critical. - ## Phase 2: Present Findings Structure your feedback as a design director would: ### Design Health Score -→ *Consult [heuristics-scoring](reference/heuristics-scoring.md)* +> *Consult [heuristics-scoring](reference/heuristics-scoring.md)* Score each of Nielsen's 10 heuristics 0–4. Present as a table: | # | Heuristic | Score | Key Issue | |---|-----------|-------|-----------| | 1 | Visibility of System Status | ? | [specific finding or "—" if solid] | -| 2 | Match System ↔ Real World | ? | | +| 2 | Match System / Real World | ? | | | 3 | User Control and Freedom | ? | | | 4 | Consistency and Standards | ? | | | 5 | Error Prevention | ? | | @@ -123,13 +114,13 @@ Be honest with scores. A 4 means genuinely excellent. Most real interfaces score **Start here.** Pass/fail: Does this look AI-generated? List specific tells from the skill's Anti-Patterns section. Be brutally honest. ### Overall Impression -A brief gut reaction—what works, what doesn't, and the single biggest opportunity. +A brief gut reaction — what works, what doesn't, and the single biggest opportunity. ### What's Working -Highlight 2-3 things done well. Be specific about why they work. +Highlight 2–3 things done well. Be specific about why they work. ### Priority Issues -The 3-5 most impactful design problems, ordered by importance. +The 3–5 most impactful design problems, ordered by importance. For each issue, tag with **P0–P3 severity** (consult [heuristics-scoring](reference/heuristics-scoring.md) for severity definitions): - **[P?] What**: Name the problem clearly @@ -138,7 +129,7 @@ For each issue, tag with **P0–P3 severity** (consult [heuristics-scoring](refe - **Suggested command**: Which command could address this (from: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive) ### Persona Red Flags -→ *Consult [personas](reference/personas.md)* +> *Consult [personas](reference/personas.md)* Auto-select 2–3 personas most relevant to this interface type (use the selection table in the reference). If `AGENTS.md` contains a `## Design Context` section from `teach-impeccable`, also generate 1–2 project-specific personas from the audience/brand info. @@ -154,12 +145,12 @@ Be specific — name the exact elements and interactions that fail each persona. Quick notes on smaller issues worth addressing. **Remember**: -- Be direct—vague feedback wastes everyone's time -- Be specific—"the submit button" not "some elements" +- Be direct — vague feedback wastes everyone's time +- Be specific — "the submit button" not "some elements" - Say what's wrong AND why it matters to users - Give concrete suggestions, not just "consider exploring..." -- Prioritize ruthlessly—if everything is important, nothing is -- Don't soften criticism—developers need honest feedback to ship great design +- Prioritize ruthlessly — if everything is important, nothing is +- Don't soften criticism — developers need honest feedback to ship great design ## Phase 3: Ask the User @@ -167,9 +158,9 @@ Quick notes on smaller issues worth addressing. Ask questions along these lines (adapt to the specific findings — do NOT ask generic questions): -1. **Priority direction**: Based on the issues found, ask which category matters most to the user right now. For example: "I found problems with visual hierarchy, color usage, and information overload. Which area should we tackle first?" Offer the top 2-3 issue categories as options. +1. **Priority direction**: Based on the issues found, ask which category matters most to the user right now. For example: "I found problems with visual hierarchy, color usage, and information overload. Which area should we tackle first?" Offer the top 2–3 issue categories as options. -2. **Design intent**: If the critique found a tonal mismatch, ask whether it was intentional. For example: "The interface feels clinical and corporate. Is that the intended tone, or should it feel warmer/bolder/more playful?" Offer 2-3 tonal directions as options based on what would fix the issues found. +2. **Design intent**: If the critique found a tonal mismatch, ask whether it was intentional. For example: "The interface feels clinical and corporate. Is that the intended tone, or should it feel warmer/bolder/more playful?" Offer 2–3 tonal directions as options based on what would fix the issues found. 3. **Scope**: Ask how much the user wants to take on. For example: "I found N issues. Want to address everything, or focus on the top 3?" Offer scope options like "Top 3 only", "All issues", "Critical issues only". @@ -177,9 +168,9 @@ Ask questions along these lines (adapt to the specific findings — do NOT ask g **Rules for questions**: - Every question must reference specific findings from Phase 2 — never ask generic "who is your audience?" questions -- Keep it to 2-4 questions maximum — respect the user's time +- Keep it to 2–4 questions maximum — respect the user's time - Offer concrete options, not open-ended prompts -- If findings are straightforward (e.g., only 1-2 clear issues), skip questions and go directly to Phase 4 +- If findings are straightforward (e.g., only 1–2 clear issues), skip questions and go directly to Phase 4 ## Phase 4: Recommended Actions diff --git a/.opencode/skills/critique/reference/personas.md b/.opencode/skills/critique/reference/personas.md index fdc88e20b..2d0f9cbf3 100644 --- a/.opencode/skills/critique/reference/personas.md +++ b/.opencode/skills/critique/reference/personas.md @@ -88,30 +88,30 @@ Test the interface through the eyes of 5 distinct user archetypes. Each persona --- -## 4. Skeptical Evaluator — "Riley" +## 4. Deliberate Stress Tester — "Riley" -**Profile**: Evaluating the product for their team or company. Looking for reasons to reject. Comparing against competitors. +**Profile**: Methodical user who pushes interfaces beyond the happy path. Tests edge cases, tries unexpected inputs, and probes for gaps in the experience. **Behaviors**: - Tests edge cases intentionally (empty states, long strings, special characters) -- Looks for pricing catches and hidden limitations -- Reads fine print and terms of service -- Tries to break things deliberately +- Submits forms with unexpected data (emoji, RTL text, very long values) +- Tries to break workflows by navigating backwards, refreshing mid-flow, or opening in multiple tabs +- Looks for inconsistencies between what the UI promises and what actually happens - Documents problems methodically **Test Questions**: - What happens at the edges (0 items, 1000 items, very long text)? -- Is pricing and value proposition transparent? -- Are there hidden limitations or gotchas? -- How polished is error handling? -- What data is collected and why? +- Do error states recover gracefully or leave the UI in a broken state? +- What happens on refresh mid-workflow? Is state preserved? +- Are there features that appear to work but produce broken results? +- How does the UI handle unexpected input (emoji, special chars, paste from Excel)? **Red Flags** (report these specifically): -- Hidden pricing or "contact sales" for basic information -- Features that appear to work but produce broken results -- Poor error handling that exposes technical details -- Unclear data practices or missing privacy information +- Features that appear to work but silently fail or produce wrong results +- Error handling that exposes technical details or leaves UI in a broken state - Empty states that show nothing useful ("No results" with no guidance) +- Workflows that lose user data on refresh or navigation +- Inconsistent behavior between similar interactions in different parts of the UI --- @@ -150,7 +150,7 @@ Choose personas based on the interface type: |---------------|-----------------|-----| | Landing page / marketing | Jordan, Riley, Casey | First impressions, trust, mobile | | Dashboard / admin | Alex, Sam | Power users, accessibility | -| E-commerce / checkout | Casey, Riley, Jordan | Mobile, trust, clarity | +| E-commerce / checkout | Casey, Riley, Jordan | Mobile, edge cases, clarity | | Onboarding flow | Jordan, Casey | Confusion, interruption | | Data-heavy / analytics | Alex, Sam | Efficiency, keyboard nav | | Form-heavy / wizard | Jordan, Sam, Casey | Clarity, accessibility, mobile | diff --git a/.pi/skills/audit/SKILL.md b/.pi/skills/audit/SKILL.md index d84428320..6dc747d60 100644 --- a/.pi/skills/audit/SKILL.md +++ b/.pi/skills/audit/SKILL.md @@ -1,15 +1,21 @@ --- name: audit -description: Perform a comprehensive audit of interface quality across accessibility, performance, theming, and responsive design. Generates a scored report with severity ratings and actionable plan. Use when the user wants a design review, accessibility check, quality audit, or a full list of UI issues to fix. +description: Run technical quality checks across accessibility, performance, theming, responsive design, and anti-patterns. Generates a scored report with P0-P3 severity ratings and actionable plan. Use when the user wants an accessibility check, performance audit, or technical quality review. --- -Run systematic quality checks and generate a comprehensive audit report with quantitative scoring, prioritized issues, and an actionable plan. Don't fix issues — document them for other commands to address. +## MANDATORY PREPARATION -**First**: Invoke /frontend-design for design principles and anti-patterns. +Invoke /frontend-design — it contains design principles, anti-patterns, and the **Context Gathering Protocol**. Follow the protocol before proceeding — if no design context exists yet, you MUST run /teach-impeccable first. + +--- + +Run systematic **technical** quality checks and generate a comprehensive report. Don't fix issues — document them for other commands to address. + +This is a code-level audit, not a design critique. Check what's measurable and verifiable in the implementation. ## Diagnostic Scan -Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using the criteria below. +Run comprehensive checks across 5 dimensions. Score each dimension 0-4 using the criteria below. ### 1. Accessibility (A11y) @@ -21,7 +27,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Alt text**: Missing or poor image descriptions - **Form issues**: Inputs without labels, poor error messaging, missing required indicators -**Score 0–4**: 0=Inaccessible (fails WCAG A), 1=Major gaps (few ARIA labels, no keyboard nav), 2=Partial (some a11y effort, significant gaps), 3=Good (WCAG AA mostly met, minor gaps), 4=Excellent (WCAG AA fully met, approaches AAA) +**Score 0-4**: 0=Inaccessible (fails WCAG A), 1=Major gaps (few ARIA labels, no keyboard nav), 2=Partial (some a11y effort, significant gaps), 3=Good (WCAG AA mostly met, minor gaps), 4=Excellent (WCAG AA fully met, approaches AAA) ### 2. Performance @@ -32,7 +38,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Bundle size**: Unnecessary imports, unused dependencies - **Render performance**: Unnecessary re-renders, missing memoization -**Score 0–4**: 0=Severe issues (layout thrash, unoptimized everything), 1=Major problems (no lazy loading, expensive animations), 2=Partial (some optimization, gaps remain), 3=Good (mostly optimized, minor improvements possible), 4=Excellent (fast, lean, well-optimized) +**Score 0-4**: 0=Severe issues (layout thrash, unoptimized everything), 1=Major problems (no lazy loading, expensive animations), 2=Partial (some optimization, gaps remain), 3=Good (mostly optimized, minor improvements possible), 4=Excellent (fast, lean, well-optimized) ### 3. Theming @@ -42,7 +48,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Inconsistent tokens**: Using wrong tokens, mixing token types - **Theme switching issues**: Values that don't update on theme change -**Score 0–4**: 0=No theming (hard-coded everything), 1=Minimal tokens (mostly hard-coded), 2=Partial (tokens exist but inconsistently used), 3=Good (tokens used, minor hard-coded values), 4=Excellent (full token system, dark mode works perfectly) +**Score 0-4**: 0=No theming (hard-coded everything), 1=Minimal tokens (mostly hard-coded), 2=Partial (tokens exist but inconsistently used), 3=Good (tokens used, minor hard-coded values), 4=Excellent (full token system, dark mode works perfectly) ### 4. Responsive Design @@ -53,113 +59,73 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Text scaling**: Layouts that break when text size increases - **Missing breakpoints**: No mobile/tablet variants -**Score 0–4**: 0=Desktop-only (breaks on mobile), 1=Major issues (some breakpoints, many failures), 2=Partial (works on mobile, rough edges), 3=Good (responsive, minor touch target or overflow issues), 4=Excellent (fluid, all viewports, proper touch targets) +**Score 0-4**: 0=Desktop-only (breaks on mobile), 1=Major issues (some breakpoints, many failures), 2=Partial (works on mobile, rough edges), 3=Good (responsive, minor touch target or overflow issues), 4=Excellent (fluid, all viewports, proper touch targets) ### 5. Anti-Patterns (CRITICAL) Check against ALL the **DON'T** guidelines in the frontend-design skill. Look for AI slop tells (AI color palette, gradient text, glassmorphism, hero metrics, card grids, generic fonts) and general design anti-patterns (gray on color, nested cards, bounce easing, redundant copy). -**Score 0–4**: 0=AI slop gallery (5+ tells), 1=Heavy AI aesthetic (3-4 tells), 2=Some tells (1-2 noticeable), 3=Mostly clean (subtle issues only), 4=No AI tells (distinctive, intentional design) +**Score 0-4**: 0=AI slop gallery (5+ tells), 1=Heavy AI aesthetic (3-4 tells), 2=Some tells (1-2 noticeable), 3=Mostly clean (subtle issues only), 4=No AI tells (distinctive, intentional design) -**CRITICAL**: This is an audit, not a fix. Document issues thoroughly with clear explanations of impact. Use other commands to fix issues after audit. - -## Generate Comprehensive Report +## Generate Report ### Audit Health Score -Present the dimension scores as a table: - | # | Dimension | Score | Key Finding | |---|-----------|-------|-------------| -| 1 | Accessibility | ? | [most critical a11y issue or "—"] | +| 1 | Accessibility | ? | [most critical a11y issue or "--"] | | 2 | Performance | ? | | | 3 | Responsive Design | ? | | | 4 | Theming | ? | | | 5 | Anti-Patterns | ? | | | **Total** | | **??/20** | **[Rating band]** | -**Rating bands**: -| Score | Rating | Action | -|-------|--------|--------| -| 18–20 | Excellent | Minor polish only | -| 14–17 | Good | Address weak dimensions | -| 10–13 | Acceptable | Significant work needed | -| 6–9 | Poor | Major quality overhaul | -| 0–5 | Critical | Fundamental issues across the board | +**Rating bands**: 18-20 Excellent (minor polish), 14-17 Good (address weak dimensions), 10-13 Acceptable (significant work needed), 6-9 Poor (major overhaul), 0-5 Critical (fundamental issues) ### Anti-Patterns Verdict -**Start here.** Pass/fail: Does this look AI-generated? List specific tells from the skill's Anti-Patterns section. Be brutally honest. +**Start here.** Pass/fail: Does this look AI-generated? List specific tells. Be brutally honest. ### Executive Summary - Audit Health Score: **??/20** ([rating band]) - Total issues found (count by severity: P0/P1/P2/P3) -- Most critical issues (top 3-5) +- Top 3-5 critical issues - Recommended next steps ### Detailed Findings by Severity -Tag every issue with **P0–P3 severity**: -| Priority | Name | Description | -|----------|------|-------------| -| **P0** | Blocking | Prevents task completion — fix immediately | -| **P1** | Major | Significant difficulty or WCAG AA violation — fix before release | -| **P2** | Minor | Annoyance, workaround exists — fix in next pass | -| **P3** | Polish | Nice-to-fix, no real user impact — fix if time permits | +Tag every issue with **P0-P3 severity**: +- **P0 Blocking**: Prevents task completion — fix immediately +- **P1 Major**: Significant difficulty or WCAG AA violation — fix before release +- **P2 Minor**: Annoyance, workaround exists — fix in next pass +- **P3 Polish**: Nice-to-fix, no real user impact — fix if time permits For each issue, document: - **[P?] Issue name** -- **Location**: Where it occurs (component, file, line) +- **Location**: Component, file, line - **Category**: Accessibility / Performance / Theming / Responsive / Anti-Pattern -- **Description**: What the issue is - **Impact**: How it affects users - **WCAG/Standard**: Which standard it violates (if applicable) - **Recommendation**: How to fix it -- **Suggested command**: Which command to use (prefer: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive — or other installed skills you're sure exist) - -#### P0 — Blocking Issues -[Issues that prevent task completion or violate WCAG A] - -#### P1 — Major Issues -[Significant usability/accessibility impact, WCAG AA violations] - -#### P2 — Minor Issues -[Quality issues, WCAG AAA violations, performance concerns] - -#### P3 — Polish Issues -[Minor inconsistencies, optimization opportunities] +- **Suggested command**: Which command to use (prefer: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive) ### Patterns & Systemic Issues -Identify recurring problems: +Identify recurring problems that indicate systemic gaps rather than one-off mistakes: - "Hard-coded colors appear in 15+ components, should use design tokens" - "Touch targets consistently too small (<44px) throughout mobile experience" -- "Missing focus indicators on all custom interactive components" ### Positive Findings -Note what's working well: -- Good practices to maintain -- Exemplary implementations to replicate elsewhere +Note what's working well — good practices to maintain and replicate. ## Recommended Actions -Present a prioritized action summary. Order is determined by severity automatically (P0 first, then P1, then P2). - -### Action Summary - -List recommended commands in priority order: +List recommended commands in priority order (P0 first, then P1, then P2): 1. **[P?] `/command-name`** — Brief description (specific context from audit findings) 2. **[P?] `/command-name`** — Brief description (specific context) -... -**Rules for recommendations**: -- Only recommend commands from: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive -- Order by severity: P0 issues first, then P1, then P2 (skip P3 unless user has few issues) -- Each item's description should carry enough context that the command knows what to focus on -- Map findings to the most appropriate command -- Skip commands that would address zero issues -- End with `/polish` as the final step if any fixes were recommended +**Rules**: Only recommend commands from: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive. Map findings to the most appropriate command. End with `/polish` as the final step if any fixes were recommended. After presenting the summary, tell the user: @@ -171,10 +137,9 @@ After presenting the summary, tell the user: **NEVER**: - Report issues without explaining impact (why does this matter?) -- Mix severity levels inconsistently -- Skip positive findings (celebrate what works) - Provide generic recommendations (be specific and actionable) +- Skip positive findings (celebrate what works) - Forget to prioritize (everything can't be P0) - Report false positives without verification -Remember: You're a quality auditor with exceptional attention to detail. Document systematically, prioritize ruthlessly, and provide clear paths to improvement. A good audit makes fixing easy. \ No newline at end of file +Remember: You're a technical quality auditor. Document systematically, prioritize ruthlessly, cite specific code locations, and provide clear paths to improvement. \ No newline at end of file diff --git a/.pi/skills/critique/SKILL.md b/.pi/skills/critique/SKILL.md index f05987cbb..dba9b20aa 100644 --- a/.pi/skills/critique/SKILL.md +++ b/.pi/skills/critique/SKILL.md @@ -1,6 +1,6 @@ --- name: critique -description: Evaluate design effectiveness from a UX perspective. Assesses visual hierarchy, information architecture, emotional resonance, cognitive load, and overall design quality with quantitative scoring, persona-based testing, and actionable feedback. Use when the user asks to review, critique, evaluate, or give feedback on a design, UI, or component. +description: Evaluate design from a UX perspective, assessing visual hierarchy, information architecture, emotional resonance, cognitive load, and overall quality with quantitative scoring, persona-based testing, and actionable feedback. Use when the user asks to review, critique, evaluate, or give feedback on a design or component. --- ## MANDATORY PREPARATION @@ -9,7 +9,7 @@ Invoke /frontend-design — it contains design principles, anti-patterns, and th --- -Conduct a holistic design critique, evaluating whether the interface actually works—not just technically, but as a designed experience. Think like a design director giving feedback. +Conduct a holistic design critique, evaluating whether the interface actually works — not just technically, but as a designed experience. Think like a design director giving feedback. ## Phase 1: Design Critique @@ -19,7 +19,7 @@ Evaluate the interface across these dimensions: **This is the most important check.** Does this look like every other AI-generated interface from 2024-2025? -Review the design against ALL the **DON'T** guidelines in the frontend-design skill—they are the fingerprints of AI-generated work. Check for the AI color palette, gradient text, dark mode with glowing accents, glassmorphism, hero metric layouts, identical card grids, generic fonts, and all other tells. +Review the design against ALL the **DON'T** guidelines in the frontend-design skill — they are the fingerprints of AI-generated work. Check for the AI color palette, gradient text, dark mode with glowing accents, glassmorphism, hero metric layouts, identical card grids, generic fonts, and all other tells. **The test**: If you showed this to someone and said "AI made this," would they believe you immediately? If yes, that's the problem. @@ -29,20 +29,19 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Do size, color, and position communicate importance correctly? - Is there visual competition between elements that should have different weights? -### 3. Information Architecture -→ *Consult [cognitive-load](reference/cognitive-load.md) for the working memory rule and checklist* +### 3. Information Architecture & Cognitive Load +> *Consult [cognitive-load](reference/cognitive-load.md) for the working memory rule and 8-item checklist* - Is the structure intuitive? Would a new user understand the organization? - Is related content grouped logically? - Are there too many choices at once? Count visible options at each decision point — if >4, flag it - Is the navigation clear and predictable? - **Progressive disclosure**: Is complexity revealed only when needed, or dumped on the user upfront? -- **Cognitive load sub-check**: Run the 8-item cognitive load checklist from the reference. Report the number of failures. +- **Run the 8-item cognitive load checklist** from the reference. Report failure count: 0–1 = low (good), 2–3 = moderate, 4+ = critical. ### 4. Emotional Journey -→ *Consult [cognitive-load](reference/cognitive-load.md) for emotional intervention patterns* - What emotion does this interface evoke? Is that intentional? - Does it match the brand personality? -- Does it feel trustworthy, approachable, premium, playful—whatever it should feel? +- Does it feel trustworthy, approachable, premium, playful — whatever it should feel? - Would the target user feel "this is for me"? - **Peak-end rule**: Is the most intense moment positive? Does the experience end well (confirmation, celebration, clear next step)? - **Emotional valleys**: Check for onboarding frustration, error cliffs, feature discovery gaps, or anxiety spikes at high-stakes moments (payment, delete, commit) @@ -70,7 +69,7 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Is color used to communicate, not just decorate? - Does the palette feel cohesive? - Are accent colors drawing attention to the right things? -- Does it work for colorblind users? (not just technically—does meaning still come through?) +- Does it work for colorblind users? (not just technically — does meaning still come through?) ### 9. States & Edge Cases - Empty states: Do they guide users toward action, or just say "nothing here"? @@ -84,27 +83,19 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Are labels and buttons unambiguous? - Does error copy help users fix the problem? -### 11. Cognitive Load -→ *Consult [cognitive-load](reference/cognitive-load.md)* -- **Intrinsic vs. extraneous**: Is the mental effort coming from the task itself (acceptable) or from poor design choices (eliminate)? -- **Decision points**: Count visible choices at key moments. More than 4 simultaneous options = overload. -- **Working memory burden**: Does the user need to remember information from a previous screen to act on the current one? -- **Information chunking**: Is content broken into digestible groups, or presented as undifferentiated walls? -- Run the 8-item cognitive load checklist. Report failures count: 0–1 = low (good), 2–3 = moderate, 4+ = critical. - ## Phase 2: Present Findings Structure your feedback as a design director would: ### Design Health Score -→ *Consult [heuristics-scoring](reference/heuristics-scoring.md)* +> *Consult [heuristics-scoring](reference/heuristics-scoring.md)* Score each of Nielsen's 10 heuristics 0–4. Present as a table: | # | Heuristic | Score | Key Issue | |---|-----------|-------|-----------| | 1 | Visibility of System Status | ? | [specific finding or "—" if solid] | -| 2 | Match System ↔ Real World | ? | | +| 2 | Match System / Real World | ? | | | 3 | User Control and Freedom | ? | | | 4 | Consistency and Standards | ? | | | 5 | Error Prevention | ? | | @@ -121,13 +112,13 @@ Be honest with scores. A 4 means genuinely excellent. Most real interfaces score **Start here.** Pass/fail: Does this look AI-generated? List specific tells from the skill's Anti-Patterns section. Be brutally honest. ### Overall Impression -A brief gut reaction—what works, what doesn't, and the single biggest opportunity. +A brief gut reaction — what works, what doesn't, and the single biggest opportunity. ### What's Working -Highlight 2-3 things done well. Be specific about why they work. +Highlight 2–3 things done well. Be specific about why they work. ### Priority Issues -The 3-5 most impactful design problems, ordered by importance. +The 3–5 most impactful design problems, ordered by importance. For each issue, tag with **P0–P3 severity** (consult [heuristics-scoring](reference/heuristics-scoring.md) for severity definitions): - **[P?] What**: Name the problem clearly @@ -136,7 +127,7 @@ For each issue, tag with **P0–P3 severity** (consult [heuristics-scoring](refe - **Suggested command**: Which command could address this (from: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive) ### Persona Red Flags -→ *Consult [personas](reference/personas.md)* +> *Consult [personas](reference/personas.md)* Auto-select 2–3 personas most relevant to this interface type (use the selection table in the reference). If `AGENTS.md` contains a `## Design Context` section from `teach-impeccable`, also generate 1–2 project-specific personas from the audience/brand info. @@ -152,12 +143,12 @@ Be specific — name the exact elements and interactions that fail each persona. Quick notes on smaller issues worth addressing. **Remember**: -- Be direct—vague feedback wastes everyone's time -- Be specific—"the submit button" not "some elements" +- Be direct — vague feedback wastes everyone's time +- Be specific — "the submit button" not "some elements" - Say what's wrong AND why it matters to users - Give concrete suggestions, not just "consider exploring..." -- Prioritize ruthlessly—if everything is important, nothing is -- Don't soften criticism—developers need honest feedback to ship great design +- Prioritize ruthlessly — if everything is important, nothing is +- Don't soften criticism — developers need honest feedback to ship great design ## Phase 3: Ask the User @@ -165,9 +156,9 @@ Quick notes on smaller issues worth addressing. Ask questions along these lines (adapt to the specific findings — do NOT ask generic questions): -1. **Priority direction**: Based on the issues found, ask which category matters most to the user right now. For example: "I found problems with visual hierarchy, color usage, and information overload. Which area should we tackle first?" Offer the top 2-3 issue categories as options. +1. **Priority direction**: Based on the issues found, ask which category matters most to the user right now. For example: "I found problems with visual hierarchy, color usage, and information overload. Which area should we tackle first?" Offer the top 2–3 issue categories as options. -2. **Design intent**: If the critique found a tonal mismatch, ask whether it was intentional. For example: "The interface feels clinical and corporate. Is that the intended tone, or should it feel warmer/bolder/more playful?" Offer 2-3 tonal directions as options based on what would fix the issues found. +2. **Design intent**: If the critique found a tonal mismatch, ask whether it was intentional. For example: "The interface feels clinical and corporate. Is that the intended tone, or should it feel warmer/bolder/more playful?" Offer 2–3 tonal directions as options based on what would fix the issues found. 3. **Scope**: Ask how much the user wants to take on. For example: "I found N issues. Want to address everything, or focus on the top 3?" Offer scope options like "Top 3 only", "All issues", "Critical issues only". @@ -175,9 +166,9 @@ Ask questions along these lines (adapt to the specific findings — do NOT ask g **Rules for questions**: - Every question must reference specific findings from Phase 2 — never ask generic "who is your audience?" questions -- Keep it to 2-4 questions maximum — respect the user's time +- Keep it to 2–4 questions maximum — respect the user's time - Offer concrete options, not open-ended prompts -- If findings are straightforward (e.g., only 1-2 clear issues), skip questions and go directly to Phase 4 +- If findings are straightforward (e.g., only 1–2 clear issues), skip questions and go directly to Phase 4 ## Phase 4: Recommended Actions diff --git a/.pi/skills/critique/reference/personas.md b/.pi/skills/critique/reference/personas.md index fdc88e20b..2d0f9cbf3 100644 --- a/.pi/skills/critique/reference/personas.md +++ b/.pi/skills/critique/reference/personas.md @@ -88,30 +88,30 @@ Test the interface through the eyes of 5 distinct user archetypes. Each persona --- -## 4. Skeptical Evaluator — "Riley" +## 4. Deliberate Stress Tester — "Riley" -**Profile**: Evaluating the product for their team or company. Looking for reasons to reject. Comparing against competitors. +**Profile**: Methodical user who pushes interfaces beyond the happy path. Tests edge cases, tries unexpected inputs, and probes for gaps in the experience. **Behaviors**: - Tests edge cases intentionally (empty states, long strings, special characters) -- Looks for pricing catches and hidden limitations -- Reads fine print and terms of service -- Tries to break things deliberately +- Submits forms with unexpected data (emoji, RTL text, very long values) +- Tries to break workflows by navigating backwards, refreshing mid-flow, or opening in multiple tabs +- Looks for inconsistencies between what the UI promises and what actually happens - Documents problems methodically **Test Questions**: - What happens at the edges (0 items, 1000 items, very long text)? -- Is pricing and value proposition transparent? -- Are there hidden limitations or gotchas? -- How polished is error handling? -- What data is collected and why? +- Do error states recover gracefully or leave the UI in a broken state? +- What happens on refresh mid-workflow? Is state preserved? +- Are there features that appear to work but produce broken results? +- How does the UI handle unexpected input (emoji, special chars, paste from Excel)? **Red Flags** (report these specifically): -- Hidden pricing or "contact sales" for basic information -- Features that appear to work but produce broken results -- Poor error handling that exposes technical details -- Unclear data practices or missing privacy information +- Features that appear to work but silently fail or produce wrong results +- Error handling that exposes technical details or leaves UI in a broken state - Empty states that show nothing useful ("No results" with no guidance) +- Workflows that lose user data on refresh or navigation +- Inconsistent behavior between similar interactions in different parts of the UI --- @@ -150,7 +150,7 @@ Choose personas based on the interface type: |---------------|-----------------|-----| | Landing page / marketing | Jordan, Riley, Casey | First impressions, trust, mobile | | Dashboard / admin | Alex, Sam | Power users, accessibility | -| E-commerce / checkout | Casey, Riley, Jordan | Mobile, trust, clarity | +| E-commerce / checkout | Casey, Riley, Jordan | Mobile, edge cases, clarity | | Onboarding flow | Jordan, Casey | Confusion, interruption | | Data-heavy / analytics | Alex, Sam | Efficiency, keyboard nav | | Form-heavy / wizard | Jordan, Sam, Casey | Clarity, accessibility, mobile | diff --git a/.trae-cn/skills/audit/SKILL.md b/.trae-cn/skills/audit/SKILL.md index 74bb05abc..1debe043e 100644 --- a/.trae-cn/skills/audit/SKILL.md +++ b/.trae-cn/skills/audit/SKILL.md @@ -1,17 +1,23 @@ --- name: audit -description: Perform a comprehensive audit of interface quality across accessibility, performance, theming, and responsive design. Generates a scored report with severity ratings and actionable plan. Use when the user wants a design review, accessibility check, quality audit, or a full list of UI issues to fix. +description: Run technical quality checks across accessibility, performance, theming, responsive design, and anti-patterns. Generates a scored report with P0-P3 severity ratings and actionable plan. Use when the user wants an accessibility check, performance audit, or technical quality review. user-invocable: true argument-hint: "[area (feature, page, component...)]" --- -Run systematic quality checks and generate a comprehensive audit report with quantitative scoring, prioritized issues, and an actionable plan. Don't fix issues — document them for other commands to address. +## MANDATORY PREPARATION -**First**: Invoke /frontend-design for design principles and anti-patterns. +Invoke /frontend-design — it contains design principles, anti-patterns, and the **Context Gathering Protocol**. Follow the protocol before proceeding — if no design context exists yet, you MUST run /teach-impeccable first. + +--- + +Run systematic **technical** quality checks and generate a comprehensive report. Don't fix issues — document them for other commands to address. + +This is a code-level audit, not a design critique. Check what's measurable and verifiable in the implementation. ## Diagnostic Scan -Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using the criteria below. +Run comprehensive checks across 5 dimensions. Score each dimension 0-4 using the criteria below. ### 1. Accessibility (A11y) @@ -23,7 +29,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Alt text**: Missing or poor image descriptions - **Form issues**: Inputs without labels, poor error messaging, missing required indicators -**Score 0–4**: 0=Inaccessible (fails WCAG A), 1=Major gaps (few ARIA labels, no keyboard nav), 2=Partial (some a11y effort, significant gaps), 3=Good (WCAG AA mostly met, minor gaps), 4=Excellent (WCAG AA fully met, approaches AAA) +**Score 0-4**: 0=Inaccessible (fails WCAG A), 1=Major gaps (few ARIA labels, no keyboard nav), 2=Partial (some a11y effort, significant gaps), 3=Good (WCAG AA mostly met, minor gaps), 4=Excellent (WCAG AA fully met, approaches AAA) ### 2. Performance @@ -34,7 +40,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Bundle size**: Unnecessary imports, unused dependencies - **Render performance**: Unnecessary re-renders, missing memoization -**Score 0–4**: 0=Severe issues (layout thrash, unoptimized everything), 1=Major problems (no lazy loading, expensive animations), 2=Partial (some optimization, gaps remain), 3=Good (mostly optimized, minor improvements possible), 4=Excellent (fast, lean, well-optimized) +**Score 0-4**: 0=Severe issues (layout thrash, unoptimized everything), 1=Major problems (no lazy loading, expensive animations), 2=Partial (some optimization, gaps remain), 3=Good (mostly optimized, minor improvements possible), 4=Excellent (fast, lean, well-optimized) ### 3. Theming @@ -44,7 +50,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Inconsistent tokens**: Using wrong tokens, mixing token types - **Theme switching issues**: Values that don't update on theme change -**Score 0–4**: 0=No theming (hard-coded everything), 1=Minimal tokens (mostly hard-coded), 2=Partial (tokens exist but inconsistently used), 3=Good (tokens used, minor hard-coded values), 4=Excellent (full token system, dark mode works perfectly) +**Score 0-4**: 0=No theming (hard-coded everything), 1=Minimal tokens (mostly hard-coded), 2=Partial (tokens exist but inconsistently used), 3=Good (tokens used, minor hard-coded values), 4=Excellent (full token system, dark mode works perfectly) ### 4. Responsive Design @@ -55,113 +61,73 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Text scaling**: Layouts that break when text size increases - **Missing breakpoints**: No mobile/tablet variants -**Score 0–4**: 0=Desktop-only (breaks on mobile), 1=Major issues (some breakpoints, many failures), 2=Partial (works on mobile, rough edges), 3=Good (responsive, minor touch target or overflow issues), 4=Excellent (fluid, all viewports, proper touch targets) +**Score 0-4**: 0=Desktop-only (breaks on mobile), 1=Major issues (some breakpoints, many failures), 2=Partial (works on mobile, rough edges), 3=Good (responsive, minor touch target or overflow issues), 4=Excellent (fluid, all viewports, proper touch targets) ### 5. Anti-Patterns (CRITICAL) Check against ALL the **DON'T** guidelines in the frontend-design skill. Look for AI slop tells (AI color palette, gradient text, glassmorphism, hero metrics, card grids, generic fonts) and general design anti-patterns (gray on color, nested cards, bounce easing, redundant copy). -**Score 0–4**: 0=AI slop gallery (5+ tells), 1=Heavy AI aesthetic (3-4 tells), 2=Some tells (1-2 noticeable), 3=Mostly clean (subtle issues only), 4=No AI tells (distinctive, intentional design) +**Score 0-4**: 0=AI slop gallery (5+ tells), 1=Heavy AI aesthetic (3-4 tells), 2=Some tells (1-2 noticeable), 3=Mostly clean (subtle issues only), 4=No AI tells (distinctive, intentional design) -**CRITICAL**: This is an audit, not a fix. Document issues thoroughly with clear explanations of impact. Use other commands to fix issues after audit. - -## Generate Comprehensive Report +## Generate Report ### Audit Health Score -Present the dimension scores as a table: - | # | Dimension | Score | Key Finding | |---|-----------|-------|-------------| -| 1 | Accessibility | ? | [most critical a11y issue or "—"] | +| 1 | Accessibility | ? | [most critical a11y issue or "--"] | | 2 | Performance | ? | | | 3 | Responsive Design | ? | | | 4 | Theming | ? | | | 5 | Anti-Patterns | ? | | | **Total** | | **??/20** | **[Rating band]** | -**Rating bands**: -| Score | Rating | Action | -|-------|--------|--------| -| 18–20 | Excellent | Minor polish only | -| 14–17 | Good | Address weak dimensions | -| 10–13 | Acceptable | Significant work needed | -| 6–9 | Poor | Major quality overhaul | -| 0–5 | Critical | Fundamental issues across the board | +**Rating bands**: 18-20 Excellent (minor polish), 14-17 Good (address weak dimensions), 10-13 Acceptable (significant work needed), 6-9 Poor (major overhaul), 0-5 Critical (fundamental issues) ### Anti-Patterns Verdict -**Start here.** Pass/fail: Does this look AI-generated? List specific tells from the skill's Anti-Patterns section. Be brutally honest. +**Start here.** Pass/fail: Does this look AI-generated? List specific tells. Be brutally honest. ### Executive Summary - Audit Health Score: **??/20** ([rating band]) - Total issues found (count by severity: P0/P1/P2/P3) -- Most critical issues (top 3-5) +- Top 3-5 critical issues - Recommended next steps ### Detailed Findings by Severity -Tag every issue with **P0–P3 severity**: -| Priority | Name | Description | -|----------|------|-------------| -| **P0** | Blocking | Prevents task completion — fix immediately | -| **P1** | Major | Significant difficulty or WCAG AA violation — fix before release | -| **P2** | Minor | Annoyance, workaround exists — fix in next pass | -| **P3** | Polish | Nice-to-fix, no real user impact — fix if time permits | +Tag every issue with **P0-P3 severity**: +- **P0 Blocking**: Prevents task completion — fix immediately +- **P1 Major**: Significant difficulty or WCAG AA violation — fix before release +- **P2 Minor**: Annoyance, workaround exists — fix in next pass +- **P3 Polish**: Nice-to-fix, no real user impact — fix if time permits For each issue, document: - **[P?] Issue name** -- **Location**: Where it occurs (component, file, line) +- **Location**: Component, file, line - **Category**: Accessibility / Performance / Theming / Responsive / Anti-Pattern -- **Description**: What the issue is - **Impact**: How it affects users - **WCAG/Standard**: Which standard it violates (if applicable) - **Recommendation**: How to fix it -- **Suggested command**: Which command to use (prefer: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive — or other installed skills you're sure exist) - -#### P0 — Blocking Issues -[Issues that prevent task completion or violate WCAG A] - -#### P1 — Major Issues -[Significant usability/accessibility impact, WCAG AA violations] - -#### P2 — Minor Issues -[Quality issues, WCAG AAA violations, performance concerns] - -#### P3 — Polish Issues -[Minor inconsistencies, optimization opportunities] +- **Suggested command**: Which command to use (prefer: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive) ### Patterns & Systemic Issues -Identify recurring problems: +Identify recurring problems that indicate systemic gaps rather than one-off mistakes: - "Hard-coded colors appear in 15+ components, should use design tokens" - "Touch targets consistently too small (<44px) throughout mobile experience" -- "Missing focus indicators on all custom interactive components" ### Positive Findings -Note what's working well: -- Good practices to maintain -- Exemplary implementations to replicate elsewhere +Note what's working well — good practices to maintain and replicate. ## Recommended Actions -Present a prioritized action summary. Order is determined by severity automatically (P0 first, then P1, then P2). - -### Action Summary - -List recommended commands in priority order: +List recommended commands in priority order (P0 first, then P1, then P2): 1. **[P?] `/command-name`** — Brief description (specific context from audit findings) 2. **[P?] `/command-name`** — Brief description (specific context) -... -**Rules for recommendations**: -- Only recommend commands from: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive -- Order by severity: P0 issues first, then P1, then P2 (skip P3 unless user has few issues) -- Each item's description should carry enough context that the command knows what to focus on -- Map findings to the most appropriate command -- Skip commands that would address zero issues -- End with `/polish` as the final step if any fixes were recommended +**Rules**: Only recommend commands from: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive. Map findings to the most appropriate command. End with `/polish` as the final step if any fixes were recommended. After presenting the summary, tell the user: @@ -173,10 +139,9 @@ After presenting the summary, tell the user: **NEVER**: - Report issues without explaining impact (why does this matter?) -- Mix severity levels inconsistently -- Skip positive findings (celebrate what works) - Provide generic recommendations (be specific and actionable) +- Skip positive findings (celebrate what works) - Forget to prioritize (everything can't be P0) - Report false positives without verification -Remember: You're a quality auditor with exceptional attention to detail. Document systematically, prioritize ruthlessly, and provide clear paths to improvement. A good audit makes fixing easy. \ No newline at end of file +Remember: You're a technical quality auditor. Document systematically, prioritize ruthlessly, cite specific code locations, and provide clear paths to improvement. \ No newline at end of file diff --git a/.trae-cn/skills/critique/SKILL.md b/.trae-cn/skills/critique/SKILL.md index 1b1193ed2..8a409115a 100644 --- a/.trae-cn/skills/critique/SKILL.md +++ b/.trae-cn/skills/critique/SKILL.md @@ -1,6 +1,6 @@ --- name: critique -description: Evaluate design effectiveness from a UX perspective. Assesses visual hierarchy, information architecture, emotional resonance, cognitive load, and overall design quality with quantitative scoring, persona-based testing, and actionable feedback. Use when the user asks to review, critique, evaluate, or give feedback on a design, UI, or component. +description: Evaluate design from a UX perspective, assessing visual hierarchy, information architecture, emotional resonance, cognitive load, and overall quality with quantitative scoring, persona-based testing, and actionable feedback. Use when the user asks to review, critique, evaluate, or give feedback on a design or component. user-invocable: true argument-hint: "[area (feature, page, component...)]" --- @@ -11,7 +11,7 @@ Invoke /frontend-design — it contains design principles, anti-patterns, and th --- -Conduct a holistic design critique, evaluating whether the interface actually works—not just technically, but as a designed experience. Think like a design director giving feedback. +Conduct a holistic design critique, evaluating whether the interface actually works — not just technically, but as a designed experience. Think like a design director giving feedback. ## Phase 1: Design Critique @@ -21,7 +21,7 @@ Evaluate the interface across these dimensions: **This is the most important check.** Does this look like every other AI-generated interface from 2024-2025? -Review the design against ALL the **DON'T** guidelines in the frontend-design skill—they are the fingerprints of AI-generated work. Check for the AI color palette, gradient text, dark mode with glowing accents, glassmorphism, hero metric layouts, identical card grids, generic fonts, and all other tells. +Review the design against ALL the **DON'T** guidelines in the frontend-design skill — they are the fingerprints of AI-generated work. Check for the AI color palette, gradient text, dark mode with glowing accents, glassmorphism, hero metric layouts, identical card grids, generic fonts, and all other tells. **The test**: If you showed this to someone and said "AI made this," would they believe you immediately? If yes, that's the problem. @@ -31,20 +31,19 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Do size, color, and position communicate importance correctly? - Is there visual competition between elements that should have different weights? -### 3. Information Architecture -→ *Consult [cognitive-load](reference/cognitive-load.md) for the working memory rule and checklist* +### 3. Information Architecture & Cognitive Load +> *Consult [cognitive-load](reference/cognitive-load.md) for the working memory rule and 8-item checklist* - Is the structure intuitive? Would a new user understand the organization? - Is related content grouped logically? - Are there too many choices at once? Count visible options at each decision point — if >4, flag it - Is the navigation clear and predictable? - **Progressive disclosure**: Is complexity revealed only when needed, or dumped on the user upfront? -- **Cognitive load sub-check**: Run the 8-item cognitive load checklist from the reference. Report the number of failures. +- **Run the 8-item cognitive load checklist** from the reference. Report failure count: 0–1 = low (good), 2–3 = moderate, 4+ = critical. ### 4. Emotional Journey -→ *Consult [cognitive-load](reference/cognitive-load.md) for emotional intervention patterns* - What emotion does this interface evoke? Is that intentional? - Does it match the brand personality? -- Does it feel trustworthy, approachable, premium, playful—whatever it should feel? +- Does it feel trustworthy, approachable, premium, playful — whatever it should feel? - Would the target user feel "this is for me"? - **Peak-end rule**: Is the most intense moment positive? Does the experience end well (confirmation, celebration, clear next step)? - **Emotional valleys**: Check for onboarding frustration, error cliffs, feature discovery gaps, or anxiety spikes at high-stakes moments (payment, delete, commit) @@ -72,7 +71,7 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Is color used to communicate, not just decorate? - Does the palette feel cohesive? - Are accent colors drawing attention to the right things? -- Does it work for colorblind users? (not just technically—does meaning still come through?) +- Does it work for colorblind users? (not just technically — does meaning still come through?) ### 9. States & Edge Cases - Empty states: Do they guide users toward action, or just say "nothing here"? @@ -86,27 +85,19 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Are labels and buttons unambiguous? - Does error copy help users fix the problem? -### 11. Cognitive Load -→ *Consult [cognitive-load](reference/cognitive-load.md)* -- **Intrinsic vs. extraneous**: Is the mental effort coming from the task itself (acceptable) or from poor design choices (eliminate)? -- **Decision points**: Count visible choices at key moments. More than 4 simultaneous options = overload. -- **Working memory burden**: Does the user need to remember information from a previous screen to act on the current one? -- **Information chunking**: Is content broken into digestible groups, or presented as undifferentiated walls? -- Run the 8-item cognitive load checklist. Report failures count: 0–1 = low (good), 2–3 = moderate, 4+ = critical. - ## Phase 2: Present Findings Structure your feedback as a design director would: ### Design Health Score -→ *Consult [heuristics-scoring](reference/heuristics-scoring.md)* +> *Consult [heuristics-scoring](reference/heuristics-scoring.md)* Score each of Nielsen's 10 heuristics 0–4. Present as a table: | # | Heuristic | Score | Key Issue | |---|-----------|-------|-----------| | 1 | Visibility of System Status | ? | [specific finding or "—" if solid] | -| 2 | Match System ↔ Real World | ? | | +| 2 | Match System / Real World | ? | | | 3 | User Control and Freedom | ? | | | 4 | Consistency and Standards | ? | | | 5 | Error Prevention | ? | | @@ -123,13 +114,13 @@ Be honest with scores. A 4 means genuinely excellent. Most real interfaces score **Start here.** Pass/fail: Does this look AI-generated? List specific tells from the skill's Anti-Patterns section. Be brutally honest. ### Overall Impression -A brief gut reaction—what works, what doesn't, and the single biggest opportunity. +A brief gut reaction — what works, what doesn't, and the single biggest opportunity. ### What's Working -Highlight 2-3 things done well. Be specific about why they work. +Highlight 2–3 things done well. Be specific about why they work. ### Priority Issues -The 3-5 most impactful design problems, ordered by importance. +The 3–5 most impactful design problems, ordered by importance. For each issue, tag with **P0–P3 severity** (consult [heuristics-scoring](reference/heuristics-scoring.md) for severity definitions): - **[P?] What**: Name the problem clearly @@ -138,7 +129,7 @@ For each issue, tag with **P0–P3 severity** (consult [heuristics-scoring](refe - **Suggested command**: Which command could address this (from: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive) ### Persona Red Flags -→ *Consult [personas](reference/personas.md)* +> *Consult [personas](reference/personas.md)* Auto-select 2–3 personas most relevant to this interface type (use the selection table in the reference). If `RULES.md` contains a `## Design Context` section from `teach-impeccable`, also generate 1–2 project-specific personas from the audience/brand info. @@ -154,12 +145,12 @@ Be specific — name the exact elements and interactions that fail each persona. Quick notes on smaller issues worth addressing. **Remember**: -- Be direct—vague feedback wastes everyone's time -- Be specific—"the submit button" not "some elements" +- Be direct — vague feedback wastes everyone's time +- Be specific — "the submit button" not "some elements" - Say what's wrong AND why it matters to users - Give concrete suggestions, not just "consider exploring..." -- Prioritize ruthlessly—if everything is important, nothing is -- Don't soften criticism—developers need honest feedback to ship great design +- Prioritize ruthlessly — if everything is important, nothing is +- Don't soften criticism — developers need honest feedback to ship great design ## Phase 3: Ask the User @@ -167,9 +158,9 @@ Quick notes on smaller issues worth addressing. Ask questions along these lines (adapt to the specific findings — do NOT ask generic questions): -1. **Priority direction**: Based on the issues found, ask which category matters most to the user right now. For example: "I found problems with visual hierarchy, color usage, and information overload. Which area should we tackle first?" Offer the top 2-3 issue categories as options. +1. **Priority direction**: Based on the issues found, ask which category matters most to the user right now. For example: "I found problems with visual hierarchy, color usage, and information overload. Which area should we tackle first?" Offer the top 2–3 issue categories as options. -2. **Design intent**: If the critique found a tonal mismatch, ask whether it was intentional. For example: "The interface feels clinical and corporate. Is that the intended tone, or should it feel warmer/bolder/more playful?" Offer 2-3 tonal directions as options based on what would fix the issues found. +2. **Design intent**: If the critique found a tonal mismatch, ask whether it was intentional. For example: "The interface feels clinical and corporate. Is that the intended tone, or should it feel warmer/bolder/more playful?" Offer 2–3 tonal directions as options based on what would fix the issues found. 3. **Scope**: Ask how much the user wants to take on. For example: "I found N issues. Want to address everything, or focus on the top 3?" Offer scope options like "Top 3 only", "All issues", "Critical issues only". @@ -177,9 +168,9 @@ Ask questions along these lines (adapt to the specific findings — do NOT ask g **Rules for questions**: - Every question must reference specific findings from Phase 2 — never ask generic "who is your audience?" questions -- Keep it to 2-4 questions maximum — respect the user's time +- Keep it to 2–4 questions maximum — respect the user's time - Offer concrete options, not open-ended prompts -- If findings are straightforward (e.g., only 1-2 clear issues), skip questions and go directly to Phase 4 +- If findings are straightforward (e.g., only 1–2 clear issues), skip questions and go directly to Phase 4 ## Phase 4: Recommended Actions diff --git a/.trae-cn/skills/critique/reference/personas.md b/.trae-cn/skills/critique/reference/personas.md index 95daae85e..6325e2839 100644 --- a/.trae-cn/skills/critique/reference/personas.md +++ b/.trae-cn/skills/critique/reference/personas.md @@ -88,30 +88,30 @@ Test the interface through the eyes of 5 distinct user archetypes. Each persona --- -## 4. Skeptical Evaluator — "Riley" +## 4. Deliberate Stress Tester — "Riley" -**Profile**: Evaluating the product for their team or company. Looking for reasons to reject. Comparing against competitors. +**Profile**: Methodical user who pushes interfaces beyond the happy path. Tests edge cases, tries unexpected inputs, and probes for gaps in the experience. **Behaviors**: - Tests edge cases intentionally (empty states, long strings, special characters) -- Looks for pricing catches and hidden limitations -- Reads fine print and terms of service -- Tries to break things deliberately +- Submits forms with unexpected data (emoji, RTL text, very long values) +- Tries to break workflows by navigating backwards, refreshing mid-flow, or opening in multiple tabs +- Looks for inconsistencies between what the UI promises and what actually happens - Documents problems methodically **Test Questions**: - What happens at the edges (0 items, 1000 items, very long text)? -- Is pricing and value proposition transparent? -- Are there hidden limitations or gotchas? -- How polished is error handling? -- What data is collected and why? +- Do error states recover gracefully or leave the UI in a broken state? +- What happens on refresh mid-workflow? Is state preserved? +- Are there features that appear to work but produce broken results? +- How does the UI handle unexpected input (emoji, special chars, paste from Excel)? **Red Flags** (report these specifically): -- Hidden pricing or "contact sales" for basic information -- Features that appear to work but produce broken results -- Poor error handling that exposes technical details -- Unclear data practices or missing privacy information +- Features that appear to work but silently fail or produce wrong results +- Error handling that exposes technical details or leaves UI in a broken state - Empty states that show nothing useful ("No results" with no guidance) +- Workflows that lose user data on refresh or navigation +- Inconsistent behavior between similar interactions in different parts of the UI --- @@ -150,7 +150,7 @@ Choose personas based on the interface type: |---------------|-----------------|-----| | Landing page / marketing | Jordan, Riley, Casey | First impressions, trust, mobile | | Dashboard / admin | Alex, Sam | Power users, accessibility | -| E-commerce / checkout | Casey, Riley, Jordan | Mobile, trust, clarity | +| E-commerce / checkout | Casey, Riley, Jordan | Mobile, edge cases, clarity | | Onboarding flow | Jordan, Casey | Confusion, interruption | | Data-heavy / analytics | Alex, Sam | Efficiency, keyboard nav | | Form-heavy / wizard | Jordan, Sam, Casey | Clarity, accessibility, mobile | diff --git a/.trae/skills/audit/SKILL.md b/.trae/skills/audit/SKILL.md index 74bb05abc..1debe043e 100644 --- a/.trae/skills/audit/SKILL.md +++ b/.trae/skills/audit/SKILL.md @@ -1,17 +1,23 @@ --- name: audit -description: Perform a comprehensive audit of interface quality across accessibility, performance, theming, and responsive design. Generates a scored report with severity ratings and actionable plan. Use when the user wants a design review, accessibility check, quality audit, or a full list of UI issues to fix. +description: Run technical quality checks across accessibility, performance, theming, responsive design, and anti-patterns. Generates a scored report with P0-P3 severity ratings and actionable plan. Use when the user wants an accessibility check, performance audit, or technical quality review. user-invocable: true argument-hint: "[area (feature, page, component...)]" --- -Run systematic quality checks and generate a comprehensive audit report with quantitative scoring, prioritized issues, and an actionable plan. Don't fix issues — document them for other commands to address. +## MANDATORY PREPARATION -**First**: Invoke /frontend-design for design principles and anti-patterns. +Invoke /frontend-design — it contains design principles, anti-patterns, and the **Context Gathering Protocol**. Follow the protocol before proceeding — if no design context exists yet, you MUST run /teach-impeccable first. + +--- + +Run systematic **technical** quality checks and generate a comprehensive report. Don't fix issues — document them for other commands to address. + +This is a code-level audit, not a design critique. Check what's measurable and verifiable in the implementation. ## Diagnostic Scan -Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using the criteria below. +Run comprehensive checks across 5 dimensions. Score each dimension 0-4 using the criteria below. ### 1. Accessibility (A11y) @@ -23,7 +29,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Alt text**: Missing or poor image descriptions - **Form issues**: Inputs without labels, poor error messaging, missing required indicators -**Score 0–4**: 0=Inaccessible (fails WCAG A), 1=Major gaps (few ARIA labels, no keyboard nav), 2=Partial (some a11y effort, significant gaps), 3=Good (WCAG AA mostly met, minor gaps), 4=Excellent (WCAG AA fully met, approaches AAA) +**Score 0-4**: 0=Inaccessible (fails WCAG A), 1=Major gaps (few ARIA labels, no keyboard nav), 2=Partial (some a11y effort, significant gaps), 3=Good (WCAG AA mostly met, minor gaps), 4=Excellent (WCAG AA fully met, approaches AAA) ### 2. Performance @@ -34,7 +40,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Bundle size**: Unnecessary imports, unused dependencies - **Render performance**: Unnecessary re-renders, missing memoization -**Score 0–4**: 0=Severe issues (layout thrash, unoptimized everything), 1=Major problems (no lazy loading, expensive animations), 2=Partial (some optimization, gaps remain), 3=Good (mostly optimized, minor improvements possible), 4=Excellent (fast, lean, well-optimized) +**Score 0-4**: 0=Severe issues (layout thrash, unoptimized everything), 1=Major problems (no lazy loading, expensive animations), 2=Partial (some optimization, gaps remain), 3=Good (mostly optimized, minor improvements possible), 4=Excellent (fast, lean, well-optimized) ### 3. Theming @@ -44,7 +50,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Inconsistent tokens**: Using wrong tokens, mixing token types - **Theme switching issues**: Values that don't update on theme change -**Score 0–4**: 0=No theming (hard-coded everything), 1=Minimal tokens (mostly hard-coded), 2=Partial (tokens exist but inconsistently used), 3=Good (tokens used, minor hard-coded values), 4=Excellent (full token system, dark mode works perfectly) +**Score 0-4**: 0=No theming (hard-coded everything), 1=Minimal tokens (mostly hard-coded), 2=Partial (tokens exist but inconsistently used), 3=Good (tokens used, minor hard-coded values), 4=Excellent (full token system, dark mode works perfectly) ### 4. Responsive Design @@ -55,113 +61,73 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Text scaling**: Layouts that break when text size increases - **Missing breakpoints**: No mobile/tablet variants -**Score 0–4**: 0=Desktop-only (breaks on mobile), 1=Major issues (some breakpoints, many failures), 2=Partial (works on mobile, rough edges), 3=Good (responsive, minor touch target or overflow issues), 4=Excellent (fluid, all viewports, proper touch targets) +**Score 0-4**: 0=Desktop-only (breaks on mobile), 1=Major issues (some breakpoints, many failures), 2=Partial (works on mobile, rough edges), 3=Good (responsive, minor touch target or overflow issues), 4=Excellent (fluid, all viewports, proper touch targets) ### 5. Anti-Patterns (CRITICAL) Check against ALL the **DON'T** guidelines in the frontend-design skill. Look for AI slop tells (AI color palette, gradient text, glassmorphism, hero metrics, card grids, generic fonts) and general design anti-patterns (gray on color, nested cards, bounce easing, redundant copy). -**Score 0–4**: 0=AI slop gallery (5+ tells), 1=Heavy AI aesthetic (3-4 tells), 2=Some tells (1-2 noticeable), 3=Mostly clean (subtle issues only), 4=No AI tells (distinctive, intentional design) +**Score 0-4**: 0=AI slop gallery (5+ tells), 1=Heavy AI aesthetic (3-4 tells), 2=Some tells (1-2 noticeable), 3=Mostly clean (subtle issues only), 4=No AI tells (distinctive, intentional design) -**CRITICAL**: This is an audit, not a fix. Document issues thoroughly with clear explanations of impact. Use other commands to fix issues after audit. - -## Generate Comprehensive Report +## Generate Report ### Audit Health Score -Present the dimension scores as a table: - | # | Dimension | Score | Key Finding | |---|-----------|-------|-------------| -| 1 | Accessibility | ? | [most critical a11y issue or "—"] | +| 1 | Accessibility | ? | [most critical a11y issue or "--"] | | 2 | Performance | ? | | | 3 | Responsive Design | ? | | | 4 | Theming | ? | | | 5 | Anti-Patterns | ? | | | **Total** | | **??/20** | **[Rating band]** | -**Rating bands**: -| Score | Rating | Action | -|-------|--------|--------| -| 18–20 | Excellent | Minor polish only | -| 14–17 | Good | Address weak dimensions | -| 10–13 | Acceptable | Significant work needed | -| 6–9 | Poor | Major quality overhaul | -| 0–5 | Critical | Fundamental issues across the board | +**Rating bands**: 18-20 Excellent (minor polish), 14-17 Good (address weak dimensions), 10-13 Acceptable (significant work needed), 6-9 Poor (major overhaul), 0-5 Critical (fundamental issues) ### Anti-Patterns Verdict -**Start here.** Pass/fail: Does this look AI-generated? List specific tells from the skill's Anti-Patterns section. Be brutally honest. +**Start here.** Pass/fail: Does this look AI-generated? List specific tells. Be brutally honest. ### Executive Summary - Audit Health Score: **??/20** ([rating band]) - Total issues found (count by severity: P0/P1/P2/P3) -- Most critical issues (top 3-5) +- Top 3-5 critical issues - Recommended next steps ### Detailed Findings by Severity -Tag every issue with **P0–P3 severity**: -| Priority | Name | Description | -|----------|------|-------------| -| **P0** | Blocking | Prevents task completion — fix immediately | -| **P1** | Major | Significant difficulty or WCAG AA violation — fix before release | -| **P2** | Minor | Annoyance, workaround exists — fix in next pass | -| **P3** | Polish | Nice-to-fix, no real user impact — fix if time permits | +Tag every issue with **P0-P3 severity**: +- **P0 Blocking**: Prevents task completion — fix immediately +- **P1 Major**: Significant difficulty or WCAG AA violation — fix before release +- **P2 Minor**: Annoyance, workaround exists — fix in next pass +- **P3 Polish**: Nice-to-fix, no real user impact — fix if time permits For each issue, document: - **[P?] Issue name** -- **Location**: Where it occurs (component, file, line) +- **Location**: Component, file, line - **Category**: Accessibility / Performance / Theming / Responsive / Anti-Pattern -- **Description**: What the issue is - **Impact**: How it affects users - **WCAG/Standard**: Which standard it violates (if applicable) - **Recommendation**: How to fix it -- **Suggested command**: Which command to use (prefer: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive — or other installed skills you're sure exist) - -#### P0 — Blocking Issues -[Issues that prevent task completion or violate WCAG A] - -#### P1 — Major Issues -[Significant usability/accessibility impact, WCAG AA violations] - -#### P2 — Minor Issues -[Quality issues, WCAG AAA violations, performance concerns] - -#### P3 — Polish Issues -[Minor inconsistencies, optimization opportunities] +- **Suggested command**: Which command to use (prefer: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive) ### Patterns & Systemic Issues -Identify recurring problems: +Identify recurring problems that indicate systemic gaps rather than one-off mistakes: - "Hard-coded colors appear in 15+ components, should use design tokens" - "Touch targets consistently too small (<44px) throughout mobile experience" -- "Missing focus indicators on all custom interactive components" ### Positive Findings -Note what's working well: -- Good practices to maintain -- Exemplary implementations to replicate elsewhere +Note what's working well — good practices to maintain and replicate. ## Recommended Actions -Present a prioritized action summary. Order is determined by severity automatically (P0 first, then P1, then P2). - -### Action Summary - -List recommended commands in priority order: +List recommended commands in priority order (P0 first, then P1, then P2): 1. **[P?] `/command-name`** — Brief description (specific context from audit findings) 2. **[P?] `/command-name`** — Brief description (specific context) -... -**Rules for recommendations**: -- Only recommend commands from: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive -- Order by severity: P0 issues first, then P1, then P2 (skip P3 unless user has few issues) -- Each item's description should carry enough context that the command knows what to focus on -- Map findings to the most appropriate command -- Skip commands that would address zero issues -- End with `/polish` as the final step if any fixes were recommended +**Rules**: Only recommend commands from: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive. Map findings to the most appropriate command. End with `/polish` as the final step if any fixes were recommended. After presenting the summary, tell the user: @@ -173,10 +139,9 @@ After presenting the summary, tell the user: **NEVER**: - Report issues without explaining impact (why does this matter?) -- Mix severity levels inconsistently -- Skip positive findings (celebrate what works) - Provide generic recommendations (be specific and actionable) +- Skip positive findings (celebrate what works) - Forget to prioritize (everything can't be P0) - Report false positives without verification -Remember: You're a quality auditor with exceptional attention to detail. Document systematically, prioritize ruthlessly, and provide clear paths to improvement. A good audit makes fixing easy. \ No newline at end of file +Remember: You're a technical quality auditor. Document systematically, prioritize ruthlessly, cite specific code locations, and provide clear paths to improvement. \ No newline at end of file diff --git a/.trae/skills/critique/SKILL.md b/.trae/skills/critique/SKILL.md index 1b1193ed2..8a409115a 100644 --- a/.trae/skills/critique/SKILL.md +++ b/.trae/skills/critique/SKILL.md @@ -1,6 +1,6 @@ --- name: critique -description: Evaluate design effectiveness from a UX perspective. Assesses visual hierarchy, information architecture, emotional resonance, cognitive load, and overall design quality with quantitative scoring, persona-based testing, and actionable feedback. Use when the user asks to review, critique, evaluate, or give feedback on a design, UI, or component. +description: Evaluate design from a UX perspective, assessing visual hierarchy, information architecture, emotional resonance, cognitive load, and overall quality with quantitative scoring, persona-based testing, and actionable feedback. Use when the user asks to review, critique, evaluate, or give feedback on a design or component. user-invocable: true argument-hint: "[area (feature, page, component...)]" --- @@ -11,7 +11,7 @@ Invoke /frontend-design — it contains design principles, anti-patterns, and th --- -Conduct a holistic design critique, evaluating whether the interface actually works—not just technically, but as a designed experience. Think like a design director giving feedback. +Conduct a holistic design critique, evaluating whether the interface actually works — not just technically, but as a designed experience. Think like a design director giving feedback. ## Phase 1: Design Critique @@ -21,7 +21,7 @@ Evaluate the interface across these dimensions: **This is the most important check.** Does this look like every other AI-generated interface from 2024-2025? -Review the design against ALL the **DON'T** guidelines in the frontend-design skill—they are the fingerprints of AI-generated work. Check for the AI color palette, gradient text, dark mode with glowing accents, glassmorphism, hero metric layouts, identical card grids, generic fonts, and all other tells. +Review the design against ALL the **DON'T** guidelines in the frontend-design skill — they are the fingerprints of AI-generated work. Check for the AI color palette, gradient text, dark mode with glowing accents, glassmorphism, hero metric layouts, identical card grids, generic fonts, and all other tells. **The test**: If you showed this to someone and said "AI made this," would they believe you immediately? If yes, that's the problem. @@ -31,20 +31,19 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Do size, color, and position communicate importance correctly? - Is there visual competition between elements that should have different weights? -### 3. Information Architecture -→ *Consult [cognitive-load](reference/cognitive-load.md) for the working memory rule and checklist* +### 3. Information Architecture & Cognitive Load +> *Consult [cognitive-load](reference/cognitive-load.md) for the working memory rule and 8-item checklist* - Is the structure intuitive? Would a new user understand the organization? - Is related content grouped logically? - Are there too many choices at once? Count visible options at each decision point — if >4, flag it - Is the navigation clear and predictable? - **Progressive disclosure**: Is complexity revealed only when needed, or dumped on the user upfront? -- **Cognitive load sub-check**: Run the 8-item cognitive load checklist from the reference. Report the number of failures. +- **Run the 8-item cognitive load checklist** from the reference. Report failure count: 0–1 = low (good), 2–3 = moderate, 4+ = critical. ### 4. Emotional Journey -→ *Consult [cognitive-load](reference/cognitive-load.md) for emotional intervention patterns* - What emotion does this interface evoke? Is that intentional? - Does it match the brand personality? -- Does it feel trustworthy, approachable, premium, playful—whatever it should feel? +- Does it feel trustworthy, approachable, premium, playful — whatever it should feel? - Would the target user feel "this is for me"? - **Peak-end rule**: Is the most intense moment positive? Does the experience end well (confirmation, celebration, clear next step)? - **Emotional valleys**: Check for onboarding frustration, error cliffs, feature discovery gaps, or anxiety spikes at high-stakes moments (payment, delete, commit) @@ -72,7 +71,7 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Is color used to communicate, not just decorate? - Does the palette feel cohesive? - Are accent colors drawing attention to the right things? -- Does it work for colorblind users? (not just technically—does meaning still come through?) +- Does it work for colorblind users? (not just technically — does meaning still come through?) ### 9. States & Edge Cases - Empty states: Do they guide users toward action, or just say "nothing here"? @@ -86,27 +85,19 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Are labels and buttons unambiguous? - Does error copy help users fix the problem? -### 11. Cognitive Load -→ *Consult [cognitive-load](reference/cognitive-load.md)* -- **Intrinsic vs. extraneous**: Is the mental effort coming from the task itself (acceptable) or from poor design choices (eliminate)? -- **Decision points**: Count visible choices at key moments. More than 4 simultaneous options = overload. -- **Working memory burden**: Does the user need to remember information from a previous screen to act on the current one? -- **Information chunking**: Is content broken into digestible groups, or presented as undifferentiated walls? -- Run the 8-item cognitive load checklist. Report failures count: 0–1 = low (good), 2–3 = moderate, 4+ = critical. - ## Phase 2: Present Findings Structure your feedback as a design director would: ### Design Health Score -→ *Consult [heuristics-scoring](reference/heuristics-scoring.md)* +> *Consult [heuristics-scoring](reference/heuristics-scoring.md)* Score each of Nielsen's 10 heuristics 0–4. Present as a table: | # | Heuristic | Score | Key Issue | |---|-----------|-------|-----------| | 1 | Visibility of System Status | ? | [specific finding or "—" if solid] | -| 2 | Match System ↔ Real World | ? | | +| 2 | Match System / Real World | ? | | | 3 | User Control and Freedom | ? | | | 4 | Consistency and Standards | ? | | | 5 | Error Prevention | ? | | @@ -123,13 +114,13 @@ Be honest with scores. A 4 means genuinely excellent. Most real interfaces score **Start here.** Pass/fail: Does this look AI-generated? List specific tells from the skill's Anti-Patterns section. Be brutally honest. ### Overall Impression -A brief gut reaction—what works, what doesn't, and the single biggest opportunity. +A brief gut reaction — what works, what doesn't, and the single biggest opportunity. ### What's Working -Highlight 2-3 things done well. Be specific about why they work. +Highlight 2–3 things done well. Be specific about why they work. ### Priority Issues -The 3-5 most impactful design problems, ordered by importance. +The 3–5 most impactful design problems, ordered by importance. For each issue, tag with **P0–P3 severity** (consult [heuristics-scoring](reference/heuristics-scoring.md) for severity definitions): - **[P?] What**: Name the problem clearly @@ -138,7 +129,7 @@ For each issue, tag with **P0–P3 severity** (consult [heuristics-scoring](refe - **Suggested command**: Which command could address this (from: /animate, /quieter, /optimize, /adapt, /clarify, /distill, /delight, /onboard, /normalize, /audit, /harden, /polish, /extract, /bolder, /arrange, /typeset, /critique, /colorize, /overdrive) ### Persona Red Flags -→ *Consult [personas](reference/personas.md)* +> *Consult [personas](reference/personas.md)* Auto-select 2–3 personas most relevant to this interface type (use the selection table in the reference). If `RULES.md` contains a `## Design Context` section from `teach-impeccable`, also generate 1–2 project-specific personas from the audience/brand info. @@ -154,12 +145,12 @@ Be specific — name the exact elements and interactions that fail each persona. Quick notes on smaller issues worth addressing. **Remember**: -- Be direct—vague feedback wastes everyone's time -- Be specific—"the submit button" not "some elements" +- Be direct — vague feedback wastes everyone's time +- Be specific — "the submit button" not "some elements" - Say what's wrong AND why it matters to users - Give concrete suggestions, not just "consider exploring..." -- Prioritize ruthlessly—if everything is important, nothing is -- Don't soften criticism—developers need honest feedback to ship great design +- Prioritize ruthlessly — if everything is important, nothing is +- Don't soften criticism — developers need honest feedback to ship great design ## Phase 3: Ask the User @@ -167,9 +158,9 @@ Quick notes on smaller issues worth addressing. Ask questions along these lines (adapt to the specific findings — do NOT ask generic questions): -1. **Priority direction**: Based on the issues found, ask which category matters most to the user right now. For example: "I found problems with visual hierarchy, color usage, and information overload. Which area should we tackle first?" Offer the top 2-3 issue categories as options. +1. **Priority direction**: Based on the issues found, ask which category matters most to the user right now. For example: "I found problems with visual hierarchy, color usage, and information overload. Which area should we tackle first?" Offer the top 2–3 issue categories as options. -2. **Design intent**: If the critique found a tonal mismatch, ask whether it was intentional. For example: "The interface feels clinical and corporate. Is that the intended tone, or should it feel warmer/bolder/more playful?" Offer 2-3 tonal directions as options based on what would fix the issues found. +2. **Design intent**: If the critique found a tonal mismatch, ask whether it was intentional. For example: "The interface feels clinical and corporate. Is that the intended tone, or should it feel warmer/bolder/more playful?" Offer 2–3 tonal directions as options based on what would fix the issues found. 3. **Scope**: Ask how much the user wants to take on. For example: "I found N issues. Want to address everything, or focus on the top 3?" Offer scope options like "Top 3 only", "All issues", "Critical issues only". @@ -177,9 +168,9 @@ Ask questions along these lines (adapt to the specific findings — do NOT ask g **Rules for questions**: - Every question must reference specific findings from Phase 2 — never ask generic "who is your audience?" questions -- Keep it to 2-4 questions maximum — respect the user's time +- Keep it to 2–4 questions maximum — respect the user's time - Offer concrete options, not open-ended prompts -- If findings are straightforward (e.g., only 1-2 clear issues), skip questions and go directly to Phase 4 +- If findings are straightforward (e.g., only 1–2 clear issues), skip questions and go directly to Phase 4 ## Phase 4: Recommended Actions diff --git a/.trae/skills/critique/reference/personas.md b/.trae/skills/critique/reference/personas.md index 95daae85e..6325e2839 100644 --- a/.trae/skills/critique/reference/personas.md +++ b/.trae/skills/critique/reference/personas.md @@ -88,30 +88,30 @@ Test the interface through the eyes of 5 distinct user archetypes. Each persona --- -## 4. Skeptical Evaluator — "Riley" +## 4. Deliberate Stress Tester — "Riley" -**Profile**: Evaluating the product for their team or company. Looking for reasons to reject. Comparing against competitors. +**Profile**: Methodical user who pushes interfaces beyond the happy path. Tests edge cases, tries unexpected inputs, and probes for gaps in the experience. **Behaviors**: - Tests edge cases intentionally (empty states, long strings, special characters) -- Looks for pricing catches and hidden limitations -- Reads fine print and terms of service -- Tries to break things deliberately +- Submits forms with unexpected data (emoji, RTL text, very long values) +- Tries to break workflows by navigating backwards, refreshing mid-flow, or opening in multiple tabs +- Looks for inconsistencies between what the UI promises and what actually happens - Documents problems methodically **Test Questions**: - What happens at the edges (0 items, 1000 items, very long text)? -- Is pricing and value proposition transparent? -- Are there hidden limitations or gotchas? -- How polished is error handling? -- What data is collected and why? +- Do error states recover gracefully or leave the UI in a broken state? +- What happens on refresh mid-workflow? Is state preserved? +- Are there features that appear to work but produce broken results? +- How does the UI handle unexpected input (emoji, special chars, paste from Excel)? **Red Flags** (report these specifically): -- Hidden pricing or "contact sales" for basic information -- Features that appear to work but produce broken results -- Poor error handling that exposes technical details -- Unclear data practices or missing privacy information +- Features that appear to work but silently fail or produce wrong results +- Error handling that exposes technical details or leaves UI in a broken state - Empty states that show nothing useful ("No results" with no guidance) +- Workflows that lose user data on refresh or navigation +- Inconsistent behavior between similar interactions in different parts of the UI --- @@ -150,7 +150,7 @@ Choose personas based on the interface type: |---------------|-----------------|-----| | Landing page / marketing | Jordan, Riley, Casey | First impressions, trust, mobile | | Dashboard / admin | Alex, Sam | Power users, accessibility | -| E-commerce / checkout | Casey, Riley, Jordan | Mobile, trust, clarity | +| E-commerce / checkout | Casey, Riley, Jordan | Mobile, edge cases, clarity | | Onboarding flow | Jordan, Casey | Confusion, interruption | | Data-heavy / analytics | Alex, Sam | Efficiency, keyboard nav | | Form-heavy / wizard | Jordan, Sam, Casey | Clarity, accessibility, mobile | diff --git a/source/skills/audit/SKILL.md b/source/skills/audit/SKILL.md index 519c29e9d..2cbf118ff 100644 --- a/source/skills/audit/SKILL.md +++ b/source/skills/audit/SKILL.md @@ -1,17 +1,23 @@ --- name: audit -description: "Perform a comprehensive audit of interface quality across accessibility, performance, theming, and responsive design. Generates a scored report with severity ratings and actionable plan. Use when the user wants a design review, accessibility check, quality audit, or a full list of UI issues to fix." +description: "Run technical quality checks across accessibility, performance, theming, responsive design, and anti-patterns. Generates a scored report with P0-P3 severity ratings and actionable plan. Use when the user wants an accessibility check, performance audit, or technical quality review." argument-hint: "[area (feature, page, component...)]" user-invocable: true --- -Run systematic quality checks and generate a comprehensive audit report with quantitative scoring, prioritized issues, and an actionable plan. Don't fix issues — document them for other commands to address. +## MANDATORY PREPARATION -**First**: Invoke {{command_prefix}}frontend-design for design principles and anti-patterns. +Invoke {{command_prefix}}frontend-design — it contains design principles, anti-patterns, and the **Context Gathering Protocol**. Follow the protocol before proceeding — if no design context exists yet, you MUST run {{command_prefix}}teach-impeccable first. + +--- + +Run systematic **technical** quality checks and generate a comprehensive report. Don't fix issues — document them for other commands to address. + +This is a code-level audit, not a design critique. Check what's measurable and verifiable in the implementation. ## Diagnostic Scan -Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using the criteria below. +Run comprehensive checks across 5 dimensions. Score each dimension 0-4 using the criteria below. ### 1. Accessibility (A11y) @@ -23,7 +29,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Alt text**: Missing or poor image descriptions - **Form issues**: Inputs without labels, poor error messaging, missing required indicators -**Score 0–4**: 0=Inaccessible (fails WCAG A), 1=Major gaps (few ARIA labels, no keyboard nav), 2=Partial (some a11y effort, significant gaps), 3=Good (WCAG AA mostly met, minor gaps), 4=Excellent (WCAG AA fully met, approaches AAA) +**Score 0-4**: 0=Inaccessible (fails WCAG A), 1=Major gaps (few ARIA labels, no keyboard nav), 2=Partial (some a11y effort, significant gaps), 3=Good (WCAG AA mostly met, minor gaps), 4=Excellent (WCAG AA fully met, approaches AAA) ### 2. Performance @@ -34,7 +40,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Bundle size**: Unnecessary imports, unused dependencies - **Render performance**: Unnecessary re-renders, missing memoization -**Score 0–4**: 0=Severe issues (layout thrash, unoptimized everything), 1=Major problems (no lazy loading, expensive animations), 2=Partial (some optimization, gaps remain), 3=Good (mostly optimized, minor improvements possible), 4=Excellent (fast, lean, well-optimized) +**Score 0-4**: 0=Severe issues (layout thrash, unoptimized everything), 1=Major problems (no lazy loading, expensive animations), 2=Partial (some optimization, gaps remain), 3=Good (mostly optimized, minor improvements possible), 4=Excellent (fast, lean, well-optimized) ### 3. Theming @@ -44,7 +50,7 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Inconsistent tokens**: Using wrong tokens, mixing token types - **Theme switching issues**: Values that don't update on theme change -**Score 0–4**: 0=No theming (hard-coded everything), 1=Minimal tokens (mostly hard-coded), 2=Partial (tokens exist but inconsistently used), 3=Good (tokens used, minor hard-coded values), 4=Excellent (full token system, dark mode works perfectly) +**Score 0-4**: 0=No theming (hard-coded everything), 1=Minimal tokens (mostly hard-coded), 2=Partial (tokens exist but inconsistently used), 3=Good (tokens used, minor hard-coded values), 4=Excellent (full token system, dark mode works perfectly) ### 4. Responsive Design @@ -55,113 +61,73 @@ Run comprehensive checks across 5 dimensions. Score each dimension 0–4 using t - **Text scaling**: Layouts that break when text size increases - **Missing breakpoints**: No mobile/tablet variants -**Score 0–4**: 0=Desktop-only (breaks on mobile), 1=Major issues (some breakpoints, many failures), 2=Partial (works on mobile, rough edges), 3=Good (responsive, minor touch target or overflow issues), 4=Excellent (fluid, all viewports, proper touch targets) +**Score 0-4**: 0=Desktop-only (breaks on mobile), 1=Major issues (some breakpoints, many failures), 2=Partial (works on mobile, rough edges), 3=Good (responsive, minor touch target or overflow issues), 4=Excellent (fluid, all viewports, proper touch targets) ### 5. Anti-Patterns (CRITICAL) Check against ALL the **DON'T** guidelines in the frontend-design skill. Look for AI slop tells (AI color palette, gradient text, glassmorphism, hero metrics, card grids, generic fonts) and general design anti-patterns (gray on color, nested cards, bounce easing, redundant copy). -**Score 0–4**: 0=AI slop gallery (5+ tells), 1=Heavy AI aesthetic (3-4 tells), 2=Some tells (1-2 noticeable), 3=Mostly clean (subtle issues only), 4=No AI tells (distinctive, intentional design) +**Score 0-4**: 0=AI slop gallery (5+ tells), 1=Heavy AI aesthetic (3-4 tells), 2=Some tells (1-2 noticeable), 3=Mostly clean (subtle issues only), 4=No AI tells (distinctive, intentional design) -**CRITICAL**: This is an audit, not a fix. Document issues thoroughly with clear explanations of impact. Use other commands to fix issues after audit. - -## Generate Comprehensive Report +## Generate Report ### Audit Health Score -Present the dimension scores as a table: - | # | Dimension | Score | Key Finding | |---|-----------|-------|-------------| -| 1 | Accessibility | ? | [most critical a11y issue or "—"] | +| 1 | Accessibility | ? | [most critical a11y issue or "--"] | | 2 | Performance | ? | | | 3 | Responsive Design | ? | | | 4 | Theming | ? | | | 5 | Anti-Patterns | ? | | | **Total** | | **??/20** | **[Rating band]** | -**Rating bands**: -| Score | Rating | Action | -|-------|--------|--------| -| 18–20 | Excellent | Minor polish only | -| 14–17 | Good | Address weak dimensions | -| 10–13 | Acceptable | Significant work needed | -| 6–9 | Poor | Major quality overhaul | -| 0–5 | Critical | Fundamental issues across the board | +**Rating bands**: 18-20 Excellent (minor polish), 14-17 Good (address weak dimensions), 10-13 Acceptable (significant work needed), 6-9 Poor (major overhaul), 0-5 Critical (fundamental issues) ### Anti-Patterns Verdict -**Start here.** Pass/fail: Does this look AI-generated? List specific tells from the skill's Anti-Patterns section. Be brutally honest. +**Start here.** Pass/fail: Does this look AI-generated? List specific tells. Be brutally honest. ### Executive Summary - Audit Health Score: **??/20** ([rating band]) - Total issues found (count by severity: P0/P1/P2/P3) -- Most critical issues (top 3-5) +- Top 3-5 critical issues - Recommended next steps ### Detailed Findings by Severity -Tag every issue with **P0–P3 severity**: -| Priority | Name | Description | -|----------|------|-------------| -| **P0** | Blocking | Prevents task completion — fix immediately | -| **P1** | Major | Significant difficulty or WCAG AA violation — fix before release | -| **P2** | Minor | Annoyance, workaround exists — fix in next pass | -| **P3** | Polish | Nice-to-fix, no real user impact — fix if time permits | +Tag every issue with **P0-P3 severity**: +- **P0 Blocking**: Prevents task completion — fix immediately +- **P1 Major**: Significant difficulty or WCAG AA violation — fix before release +- **P2 Minor**: Annoyance, workaround exists — fix in next pass +- **P3 Polish**: Nice-to-fix, no real user impact — fix if time permits For each issue, document: - **[P?] Issue name** -- **Location**: Where it occurs (component, file, line) +- **Location**: Component, file, line - **Category**: Accessibility / Performance / Theming / Responsive / Anti-Pattern -- **Description**: What the issue is - **Impact**: How it affects users - **WCAG/Standard**: Which standard it violates (if applicable) - **Recommendation**: How to fix it -- **Suggested command**: Which command to use (prefer: {{available_commands}} — or other installed skills you're sure exist) - -#### P0 — Blocking Issues -[Issues that prevent task completion or violate WCAG A] - -#### P1 — Major Issues -[Significant usability/accessibility impact, WCAG AA violations] - -#### P2 — Minor Issues -[Quality issues, WCAG AAA violations, performance concerns] - -#### P3 — Polish Issues -[Minor inconsistencies, optimization opportunities] +- **Suggested command**: Which command to use (prefer: {{available_commands}}) ### Patterns & Systemic Issues -Identify recurring problems: +Identify recurring problems that indicate systemic gaps rather than one-off mistakes: - "Hard-coded colors appear in 15+ components, should use design tokens" - "Touch targets consistently too small (<44px) throughout mobile experience" -- "Missing focus indicators on all custom interactive components" ### Positive Findings -Note what's working well: -- Good practices to maintain -- Exemplary implementations to replicate elsewhere +Note what's working well — good practices to maintain and replicate. ## Recommended Actions -Present a prioritized action summary. Order is determined by severity automatically (P0 first, then P1, then P2). - -### Action Summary - -List recommended commands in priority order: +List recommended commands in priority order (P0 first, then P1, then P2): 1. **[P?] `{{command_prefix}}command-name`** — Brief description (specific context from audit findings) 2. **[P?] `{{command_prefix}}command-name`** — Brief description (specific context) -... -**Rules for recommendations**: -- Only recommend commands from: {{available_commands}} -- Order by severity: P0 issues first, then P1, then P2 (skip P3 unless user has few issues) -- Each item's description should carry enough context that the command knows what to focus on -- Map findings to the most appropriate command -- Skip commands that would address zero issues -- End with `{{command_prefix}}polish` as the final step if any fixes were recommended +**Rules**: Only recommend commands from: {{available_commands}}. Map findings to the most appropriate command. End with `{{command_prefix}}polish` as the final step if any fixes were recommended. After presenting the summary, tell the user: @@ -173,10 +139,9 @@ After presenting the summary, tell the user: **NEVER**: - Report issues without explaining impact (why does this matter?) -- Mix severity levels inconsistently -- Skip positive findings (celebrate what works) - Provide generic recommendations (be specific and actionable) +- Skip positive findings (celebrate what works) - Forget to prioritize (everything can't be P0) - Report false positives without verification -Remember: You're a quality auditor with exceptional attention to detail. Document systematically, prioritize ruthlessly, and provide clear paths to improvement. A good audit makes fixing easy. +Remember: You're a technical quality auditor. Document systematically, prioritize ruthlessly, cite specific code locations, and provide clear paths to improvement. diff --git a/source/skills/critique/SKILL.md b/source/skills/critique/SKILL.md index dbcfdfda2..f8ae5c428 100644 --- a/source/skills/critique/SKILL.md +++ b/source/skills/critique/SKILL.md @@ -1,6 +1,6 @@ --- name: critique -description: "Evaluate design effectiveness from a UX perspective. Assesses visual hierarchy, information architecture, emotional resonance, cognitive load, and overall design quality with quantitative scoring, persona-based testing, and actionable feedback. Use when the user asks to review, critique, evaluate, or give feedback on a design, UI, or component." +description: "Evaluate design from a UX perspective, assessing visual hierarchy, information architecture, emotional resonance, cognitive load, and overall quality with quantitative scoring, persona-based testing, and actionable feedback. Use when the user asks to review, critique, evaluate, or give feedback on a design or component." argument-hint: "[area (feature, page, component...)]" user-invocable: true --- @@ -11,7 +11,7 @@ Invoke {{command_prefix}}frontend-design — it contains design principles, anti --- -Conduct a holistic design critique, evaluating whether the interface actually works—not just technically, but as a designed experience. Think like a design director giving feedback. +Conduct a holistic design critique, evaluating whether the interface actually works — not just technically, but as a designed experience. Think like a design director giving feedback. ## Phase 1: Design Critique @@ -21,7 +21,7 @@ Evaluate the interface across these dimensions: **This is the most important check.** Does this look like every other AI-generated interface from 2024-2025? -Review the design against ALL the **DON'T** guidelines in the frontend-design skill—they are the fingerprints of AI-generated work. Check for the AI color palette, gradient text, dark mode with glowing accents, glassmorphism, hero metric layouts, identical card grids, generic fonts, and all other tells. +Review the design against ALL the **DON'T** guidelines in the frontend-design skill — they are the fingerprints of AI-generated work. Check for the AI color palette, gradient text, dark mode with glowing accents, glassmorphism, hero metric layouts, identical card grids, generic fonts, and all other tells. **The test**: If you showed this to someone and said "AI made this," would they believe you immediately? If yes, that's the problem. @@ -31,20 +31,19 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Do size, color, and position communicate importance correctly? - Is there visual competition between elements that should have different weights? -### 3. Information Architecture -→ *Consult [cognitive-load](reference/cognitive-load.md) for the working memory rule and checklist* +### 3. Information Architecture & Cognitive Load +> *Consult [cognitive-load](reference/cognitive-load.md) for the working memory rule and 8-item checklist* - Is the structure intuitive? Would a new user understand the organization? - Is related content grouped logically? - Are there too many choices at once? Count visible options at each decision point — if >4, flag it - Is the navigation clear and predictable? - **Progressive disclosure**: Is complexity revealed only when needed, or dumped on the user upfront? -- **Cognitive load sub-check**: Run the 8-item cognitive load checklist from the reference. Report the number of failures. +- **Run the 8-item cognitive load checklist** from the reference. Report failure count: 0–1 = low (good), 2–3 = moderate, 4+ = critical. ### 4. Emotional Journey -→ *Consult [cognitive-load](reference/cognitive-load.md) for emotional intervention patterns* - What emotion does this interface evoke? Is that intentional? - Does it match the brand personality? -- Does it feel trustworthy, approachable, premium, playful—whatever it should feel? +- Does it feel trustworthy, approachable, premium, playful — whatever it should feel? - Would the target user feel "this is for me"? - **Peak-end rule**: Is the most intense moment positive? Does the experience end well (confirmation, celebration, clear next step)? - **Emotional valleys**: Check for onboarding frustration, error cliffs, feature discovery gaps, or anxiety spikes at high-stakes moments (payment, delete, commit) @@ -72,7 +71,7 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Is color used to communicate, not just decorate? - Does the palette feel cohesive? - Are accent colors drawing attention to the right things? -- Does it work for colorblind users? (not just technically—does meaning still come through?) +- Does it work for colorblind users? (not just technically — does meaning still come through?) ### 9. States & Edge Cases - Empty states: Do they guide users toward action, or just say "nothing here"? @@ -86,27 +85,19 @@ Review the design against ALL the **DON'T** guidelines in the frontend-design sk - Are labels and buttons unambiguous? - Does error copy help users fix the problem? -### 11. Cognitive Load -→ *Consult [cognitive-load](reference/cognitive-load.md)* -- **Intrinsic vs. extraneous**: Is the mental effort coming from the task itself (acceptable) or from poor design choices (eliminate)? -- **Decision points**: Count visible choices at key moments. More than 4 simultaneous options = overload. -- **Working memory burden**: Does the user need to remember information from a previous screen to act on the current one? -- **Information chunking**: Is content broken into digestible groups, or presented as undifferentiated walls? -- Run the 8-item cognitive load checklist. Report failures count: 0–1 = low (good), 2–3 = moderate, 4+ = critical. - ## Phase 2: Present Findings Structure your feedback as a design director would: ### Design Health Score -→ *Consult [heuristics-scoring](reference/heuristics-scoring.md)* +> *Consult [heuristics-scoring](reference/heuristics-scoring.md)* Score each of Nielsen's 10 heuristics 0–4. Present as a table: | # | Heuristic | Score | Key Issue | |---|-----------|-------|-----------| | 1 | Visibility of System Status | ? | [specific finding or "—" if solid] | -| 2 | Match System ↔ Real World | ? | | +| 2 | Match System / Real World | ? | | | 3 | User Control and Freedom | ? | | | 4 | Consistency and Standards | ? | | | 5 | Error Prevention | ? | | @@ -123,13 +114,13 @@ Be honest with scores. A 4 means genuinely excellent. Most real interfaces score **Start here.** Pass/fail: Does this look AI-generated? List specific tells from the skill's Anti-Patterns section. Be brutally honest. ### Overall Impression -A brief gut reaction—what works, what doesn't, and the single biggest opportunity. +A brief gut reaction — what works, what doesn't, and the single biggest opportunity. ### What's Working -Highlight 2-3 things done well. Be specific about why they work. +Highlight 2–3 things done well. Be specific about why they work. ### Priority Issues -The 3-5 most impactful design problems, ordered by importance. +The 3–5 most impactful design problems, ordered by importance. For each issue, tag with **P0–P3 severity** (consult [heuristics-scoring](reference/heuristics-scoring.md) for severity definitions): - **[P?] What**: Name the problem clearly @@ -138,7 +129,7 @@ For each issue, tag with **P0–P3 severity** (consult [heuristics-scoring](refe - **Suggested command**: Which command could address this (from: {{available_commands}}) ### Persona Red Flags -→ *Consult [personas](reference/personas.md)* +> *Consult [personas](reference/personas.md)* Auto-select 2–3 personas most relevant to this interface type (use the selection table in the reference). If `{{config_file}}` contains a `## Design Context` section from `teach-impeccable`, also generate 1–2 project-specific personas from the audience/brand info. @@ -154,12 +145,12 @@ Be specific — name the exact elements and interactions that fail each persona. Quick notes on smaller issues worth addressing. **Remember**: -- Be direct—vague feedback wastes everyone's time -- Be specific—"the submit button" not "some elements" +- Be direct — vague feedback wastes everyone's time +- Be specific — "the submit button" not "some elements" - Say what's wrong AND why it matters to users - Give concrete suggestions, not just "consider exploring..." -- Prioritize ruthlessly—if everything is important, nothing is -- Don't soften criticism—developers need honest feedback to ship great design +- Prioritize ruthlessly — if everything is important, nothing is +- Don't soften criticism — developers need honest feedback to ship great design ## Phase 3: Ask the User @@ -167,9 +158,9 @@ Quick notes on smaller issues worth addressing. Ask questions along these lines (adapt to the specific findings — do NOT ask generic questions): -1. **Priority direction**: Based on the issues found, ask which category matters most to the user right now. For example: "I found problems with visual hierarchy, color usage, and information overload. Which area should we tackle first?" Offer the top 2-3 issue categories as options. +1. **Priority direction**: Based on the issues found, ask which category matters most to the user right now. For example: "I found problems with visual hierarchy, color usage, and information overload. Which area should we tackle first?" Offer the top 2–3 issue categories as options. -2. **Design intent**: If the critique found a tonal mismatch, ask whether it was intentional. For example: "The interface feels clinical and corporate. Is that the intended tone, or should it feel warmer/bolder/more playful?" Offer 2-3 tonal directions as options based on what would fix the issues found. +2. **Design intent**: If the critique found a tonal mismatch, ask whether it was intentional. For example: "The interface feels clinical and corporate. Is that the intended tone, or should it feel warmer/bolder/more playful?" Offer 2–3 tonal directions as options based on what would fix the issues found. 3. **Scope**: Ask how much the user wants to take on. For example: "I found N issues. Want to address everything, or focus on the top 3?" Offer scope options like "Top 3 only", "All issues", "Critical issues only". @@ -177,9 +168,9 @@ Ask questions along these lines (adapt to the specific findings — do NOT ask g **Rules for questions**: - Every question must reference specific findings from Phase 2 — never ask generic "who is your audience?" questions -- Keep it to 2-4 questions maximum — respect the user's time +- Keep it to 2–4 questions maximum — respect the user's time - Offer concrete options, not open-ended prompts -- If findings are straightforward (e.g., only 1-2 clear issues), skip questions and go directly to Phase 4 +- If findings are straightforward (e.g., only 1–2 clear issues), skip questions and go directly to Phase 4 ## Phase 4: Recommended Actions diff --git a/source/skills/critique/reference/personas.md b/source/skills/critique/reference/personas.md index 6f85892f5..0163c265d 100644 --- a/source/skills/critique/reference/personas.md +++ b/source/skills/critique/reference/personas.md @@ -88,30 +88,30 @@ Test the interface through the eyes of 5 distinct user archetypes. Each persona --- -## 4. Skeptical Evaluator — "Riley" +## 4. Deliberate Stress Tester — "Riley" -**Profile**: Evaluating the product for their team or company. Looking for reasons to reject. Comparing against competitors. +**Profile**: Methodical user who pushes interfaces beyond the happy path. Tests edge cases, tries unexpected inputs, and probes for gaps in the experience. **Behaviors**: - Tests edge cases intentionally (empty states, long strings, special characters) -- Looks for pricing catches and hidden limitations -- Reads fine print and terms of service -- Tries to break things deliberately +- Submits forms with unexpected data (emoji, RTL text, very long values) +- Tries to break workflows by navigating backwards, refreshing mid-flow, or opening in multiple tabs +- Looks for inconsistencies between what the UI promises and what actually happens - Documents problems methodically **Test Questions**: - What happens at the edges (0 items, 1000 items, very long text)? -- Is pricing and value proposition transparent? -- Are there hidden limitations or gotchas? -- How polished is error handling? -- What data is collected and why? +- Do error states recover gracefully or leave the UI in a broken state? +- What happens on refresh mid-workflow? Is state preserved? +- Are there features that appear to work but produce broken results? +- How does the UI handle unexpected input (emoji, special chars, paste from Excel)? **Red Flags** (report these specifically): -- Hidden pricing or "contact sales" for basic information -- Features that appear to work but produce broken results -- Poor error handling that exposes technical details -- Unclear data practices or missing privacy information +- Features that appear to work but silently fail or produce wrong results +- Error handling that exposes technical details or leaves UI in a broken state - Empty states that show nothing useful ("No results" with no guidance) +- Workflows that lose user data on refresh or navigation +- Inconsistent behavior between similar interactions in different parts of the UI --- @@ -150,7 +150,7 @@ Choose personas based on the interface type: |---------------|-----------------|-----| | Landing page / marketing | Jordan, Riley, Casey | First impressions, trust, mobile | | Dashboard / admin | Alex, Sam | Power users, accessibility | -| E-commerce / checkout | Casey, Riley, Jordan | Mobile, trust, clarity | +| E-commerce / checkout | Casey, Riley, Jordan | Mobile, edge cases, clarity | | Onboarding flow | Jordan, Casey | Confusion, interruption | | Data-heavy / analytics | Alex, Sam | Efficiency, keyboard nav | | Form-heavy / wizard | Jordan, Sam, Casey | Clarity, accessibility, mobile |