diff --git a/CLAUDE.md b/CLAUDE.md index 38c8f06ca..c1b6fd3aa 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -38,7 +38,7 @@ PRODUCT.md carries a `## Platform` section with a bare value (`web` / `ios` / `a `ios.md` and `android.md` are distilled from the MIT-licensed [ehmo/platform-design-skills](https://github.com/ehmo/platform-design-skills); attribution is in `NOTICE.md`. -Sub-command reference files add a short `## Platform` section *only where guidance diverges for native*. Don't restate the platform files — link instead. Sub-commands carrying one today: `adapt`, `audit`, `animate`, `layout`. +Where a command's native guidance diverges too much to share a file, it gets a **native variant**: `reference/.native.md`, listed in SKILL.md's Commands table and routed **instead of** the web file when `setup.platform` is native (Setup step 2). One variant covers ios, android, and adaptive; per-OS specifics stay in the platform refs, which Setup loads regardless. Variants today: `audit.native.md`, `adapt.native.md` (their web files carry a one-line web-only guard that redirects stray native readers). `audit.native.md` mirrors `audit.md`'s report skeleton; change the skeleton in both together. Commands whose divergence the platform refs already cover (`animate`, `layout`) carry nothing extra; don't add in-file translation notes, they make native runs pay for web content. **Live mode, the `detect` CLI, and the design hook are web-only.** They operate on a browser / HTML rules, so SKILL.md's routing skips live and `detect.mjs` for any native (`ios` / `android` / `adaptive`) project, and the hook (`hook-lib.mjs` `resolveProjectPlatform` / `isNativePlatform`, also used by `hook-before-edit.mjs`) skips its scan when PRODUCT.md declares a native platform — a React Native project is made of exactly the `.tsx` / `.ts` / `.js` files the hook watches. @@ -201,7 +201,7 @@ IMPECCABLE_SKILL_BEHAVIOR_VERBOSE=1 bun run test:skill-behavior # dump **Auth** lives in repo-root `.env` (copied from `~/code/impeccable-evals/.env`, gitignored). Providers skip cleanly when their key is unset; they don't fail. -**Fourteen scenarios:** +**Fifteen scenarios:** 1. empty workspace → agent loads `reference/init.md` 2. PRODUCT.md only → loads `brand.md` 3. PRODUCT.md + DESIGN.md → loads `brand.md` + consults the design system @@ -216,6 +216,7 @@ IMPECCABLE_SKILL_BEHAVIOR_VERBOSE=1 bun run test:skill-behavior # dump 12. natural-language build intent with no PRODUCT.md → diverts into `reference/init.md` 13. `/impeccable teach` → diverts into `reference/init.md` (alias) 14. PRODUCT.md with `## Platform: ios` → `context.mjs` emits the native NEXT STEP and the agent loads `reference/ios.md` +15. same iOS fixture, `/impeccable audit` → agent loads `reference/audit.native.md` (route-instead variant) **Baseline.** The 21-22 / 24 baseline (with stable gpt scenario 6/7 failures) was measured on the old cheap tier (`claude-haiku-4-5` / `gpt-5.4-mini`). It needs re-measuring on the current `claude-sonnet-4-6` / `gpt-5.5` lineup; the production-tier models are expected to do better on the sub-command routing scenarios the old gpt tier failed. See `tests/skill-behavior/README.md`. diff --git a/skill/SKILL.src.md b/skill/SKILL.src.md index 8e138b6d8..c05407313 100644 --- a/skill/SKILL.src.md +++ b/skill/SKILL.src.md @@ -16,7 +16,7 @@ Designs and iterates production-grade frontend interfaces. Real working code, co You MUST do these steps before proceeding: 1. Run `node {{scripts_path}}/context.mjs` once per session; if the runtime shows this skill's loaded base directory, run `node /scripts/context.mjs` instead. Keep cwd/workdir at the user's project, not the skill directory. If the request names or implies a file, route, or app inside a monorepo, infer the concrete path and append `--target ` to the same command. If you've already seen its output in this conversation, do not re-run it. The script either prints the project's PRODUCT.md (and DESIGN.md when present) as a markdown block, or tells you it's missing. Follow whatever it prints. **If it reports `NO_PRODUCT_MD`:** divert into `reference/init.md` first when the user invoked `init`, `teach`, `craft`, or `shape`, or when their wording clearly maps to one of those from-scratch build flows (for example: "build/create/make a landing page", "design a new app", or "shape a feature"). Captured product context is the point of those flows. For any other command, a scoped evaluate / refine / enhance / fix / iterate request against existing code, do **not** divert into init. The existing code is the context: proceed with the requested command, infer the register from the surface in focus (step 4), and offer `/impeccable init` once as a suggestion the user can take later. A missing PRODUCT.md must never block a scoped request. If the output ends with an `UPDATE_AVAILABLE` directive, follow it (ask the user once about updating, then continue). It never blocks the current task. -2. If the user invoked a sub-command (`craft`, `shape`, `audit`, `polish`, ...), you MUST read `reference/.md` next. Non-optional. The reference defines the command's flow; without it you will skip steps the user expects. +2. If the user invoked a sub-command (`craft`, `shape`, `audit`, `polish`, ...), you MUST read the command's reference next: **`reference/.md`, or the native variant from the Commands table** (e.g. `reference/audit.native.md`) **when the project platform is native** (`ios` / `android` / `adaptive`, per the `context.mjs` directive). One file, not both. Non-optional. The reference defines the command's flow; without it you will skip steps the user expects. 3. Familiarize yourself with any existing design system, conventions, and components in the code. Read at least one project file (CSS / tokens / theme / a representative component or page). **Required even when you've loaded a sub-command reference in step 2.** Don't reinvent the wheel; use what's there when it works, branch out when the UX wins. 4. Read the matching register reference. **This is non-optional; skipping it produces generic output.** If the project is marketing, a landing page, a campaign, long-form content, or a portfolio (design IS the product), read `reference/brand.md`. If it is app UI, admin, a dashboard, or a tool (design SERVES the product), read `reference/product.md`. Pick by first match: (1) task cue ("landing page" vs "dashboard"); (2) surface in focus (the page, file, or route being worked on); (3) `register` field in PRODUCT.md. 5. **If PRODUCT.md's `## Platform` is `ios` or `android`**, also read `reference/.md` (HIG / Material 3 conventions). `adaptive` (cross-platform, ships both) reads both files. `web`, absent, or unrecognized: nothing extra to read. `context.mjs` prints the directive when one applies. @@ -129,7 +129,7 @@ If someone could look at this interface and say "AI made that" without doubt, it | `document` | Build | Generate DESIGN.md from existing project code | [reference/document.md](reference/document.md) | | `extract [target]` | Build | Pull reusable tokens and components into design system | [reference/extract.md](reference/extract.md) | | `critique [target]` | Evaluate | UX design review with heuristic scoring | [reference/critique.md](reference/critique.md) | -| `audit [target]` | Evaluate | Technical quality checks (a11y, perf, responsive) | [reference/audit.md](reference/audit.md) | +| `audit [target]` | Evaluate | Technical quality checks (a11y, perf, responsive) | [reference/audit.md](reference/audit.md) · native: [reference/audit.native.md](reference/audit.native.md) | | `polish [target]` | Refine | Final quality pass before shipping | [reference/polish.md](reference/polish.md) | | `bolder [target]` | Refine | Amplify safe or bland designs | [reference/bolder.md](reference/bolder.md) | | `quieter [target]` | Refine | Tone down aggressive or overstimulating designs | [reference/quieter.md](reference/quieter.md) | @@ -143,7 +143,7 @@ If someone could look at this interface and say "AI made that" without doubt, it | `delight [target]` | Enhance | Add personality and memorable touches | [reference/delight.md](reference/delight.md) | | `overdrive [target]` | Enhance | Push past conventional limits | [reference/overdrive.md](reference/overdrive.md) | | `clarify [target]` | Fix | Improve UX copy, labels, and error messages | [reference/clarify.md](reference/clarify.md) | -| `adapt [target]` | Fix | Adapt for different devices and screen sizes | [reference/adapt.md](reference/adapt.md) | +| `adapt [target]` | Fix | Adapt for different devices and screen sizes | [reference/adapt.md](reference/adapt.md) · native: [reference/adapt.native.md](reference/adapt.native.md) | | `optimize [target]` | Fix | Diagnose and fix UI performance | [reference/optimize.md](reference/optimize.md) | | `live` | Iterate | Visual variant mode: pick elements in the browser, generate alternatives | [reference/live.md](reference/live.md) | @@ -164,8 +164,8 @@ Plus three management commands: `pin `, `unpin `, and `hooks < **If `scan.targets` is non-empty and `setup.platform` is not `ios`/`android`/`adaptive`, run `node {{scripts_path}}/detect.mjs --json ` once** (the bundled detector over local files: no network, no npx; it reads HTML/CSS, so skip it for native projects). `scan.via` tells you what they are: `git-changes` (the markup/style files in your dirty tree, the most relevant set), `source-dir` (e.g. `src`, `app`), `html`, or `root`. Fold the hits into your picks: many quality / contrast hits → `audit` or `polish`; a specific slop family → the matching command (gradient text or eyebrows → `quieter` / `typeset`, flat or gray palette → `colorize`, and so on). It's a real, current signal that beats guessing. If detect errors or the tree is large and slow, skip it and recommend the user run `audit` themselves; never block the suggestion on it. Keep it to 2-3 pointed picks with the exact command to type. The menu stays the fallback; the recommendation is the lede. -2. **First word matches a command** (table above OR `pin` / `unpin` / `hooks`): load its reference file and follow its instructions. Everything after the command name is the target. -3. **First word doesn't match, but the intent clearly maps to one command** (e.g. "fix the spacing" → `layout`, "rewrite this error message" → `clarify`, "the colors feel flat" → `colorize`): load that command's reference and proceed as if invoked. If two commands could fit, ask once which. +2. **First word matches a command** (table above OR `pin` / `unpin` / `hooks`): load its reference file (on native platforms, the table's native variant; Setup step 2's one-file rule) and follow its instructions. Everything after the command name is the target. +3. **First word doesn't match, but the intent clearly maps to one command** (e.g. "fix the spacing" → `layout`, "rewrite this error message" → `clarify`, "the colors feel flat" → `colorize`): load that command's reference (same native-variant rule) and proceed as if invoked. If two commands could fit, ask once which. 4. **No clear command match**: general design invocation. Apply the setup steps, the General rules, and the loaded register reference, using the full argument as context. Setup (context gathering, register) is already loaded by then; sub-commands don't re-invoke `{{command_prefix}}impeccable`. diff --git a/skill/reference/adapt.md b/skill/reference/adapt.md index a15b44fce..36abc1000 100644 --- a/skill/reference/adapt.md +++ b/skill/reference/adapt.md @@ -2,9 +2,7 @@ Adapt an existing design to a different context: another screen size, device, platform, or use case. The trap is treating adaptation as scaling. The job is rethinking the experience for the new context. -## Platform - -Everything below is responsive web, mobile web included. Native targets (`ios` / `android` / `adaptive`): adapting means conforming to the platform's navigation model, controls, and touch sizing, never reflowing a web layout. Load [ios.md](ios.md) / [android.md](android.md) (`adaptive` loads both). +**Web only** (mobile web included). Native platforms (`ios` / `android` / `adaptive`) route to [adapt.native.md](adapt.native.md) instead; if the project is native, switch to it now. --- diff --git a/skill/reference/adapt.native.md b/skill/reference/adapt.native.md new file mode 100644 index 000000000..4307b559f --- /dev/null +++ b/skill/reference/adapt.native.md @@ -0,0 +1,58 @@ +> **Additional context needed**: target platforms/devices and usage contexts. + +Adapt an existing **native** design (`ios` / `android` / `adaptive`) to a different context: another device class, orientation, platform, or origin. The trap is treating adaptation as scaling. The job is rethinking the experience for the new context, inside the platform conventions of [ios.md](ios.md) / [android.md](android.md); read the target platform's reference before planning if Setup hasn't already. + +## Assess Adaptation Challenge + +1. **Source context**: what was it designed for, and what assumptions did it make? (Phone-only? Portrait-only? One platform's idioms? A website?) +2. **Target context**: which device class (phone, tablet, foldable), orientation, platform, and usage posture (one-handed on the go vs two-handed at rest)? +3. **What breaks**: navigation that doesn't fit the target, layouts that stretch instead of restructure, gestures or controls that don't exist there? + +## Adaptation Strategies + +### Phone → Tablet (iPad / large screens) + +- **Restructure, don't stretch.** A scaled-up phone UI on a tablet is the failure mode. Use size classes (iOS) / window size classes (Android) to switch structure. +- **Navigation changes shape**: tab bar stays or becomes a sidebar on iPad; Android navigation bar becomes a rail or drawer on expanded width. +- **Use the width**: split view / master-detail (list + detail side by side), multi-column grids, popovers where phones used sheets. +- **Multitasking is a size, not an edge case**: iPad Split View and Android multi-window can hand you a phone-width window on a tablet; size-class-driven layout handles both for free. + +### Orientation & foldables + +- Landscape restructures (side-by-side panes, repositioned controls); never clip or letterbox. Lock orientation only when the task truly demands it. +- Foldables (Android): react to posture and hinge via window size classes; test folded, unfolded, and tabletop. + +### Platform → platform (iOS ↔ Android) + +Translate idioms; never transplant them: + +| iOS | Android | +|---|---| +| Tab bar | Navigation bar / rail / drawer | +| Edge-swipe back, back chevron | Predictive Back gesture / button | +| Switch, segmented control, system pickers | Material switch, chips, Material pickers | +| Action sheet | Bottom sheet / Material dialog | +| SF Symbols, SF Pro, Dynamic Type | Material Symbols, Roboto, sp scaling | +| Semantic system colors, materials | Material color roles, tonal elevation | +| System push/sheet transitions | Container transform, shared-axis, fade-through | + +Rebuild navigation and controls in the target's vocabulary; carry over the brand's expressive layer (palette intent, type accent, motion personality) through the target's theming system. + +### Web → native (porting a website or web app) + +Reconform, don't reflow. Replace web navigation with the platform's model, HTML-shaped controls with platform controls, hover affordances with touch-first ones, and px-based type with Dynamic Type / sp. Then treat the result to the full platform reference; the slop test there is the acceptance bar. + +## Implement & Verify + +- Drive structure from **size classes / window size classes**, never from device-model checks. +- Respect safe areas and window insets in every new configuration (notch, hinge, status bar, keyboard). +- Test on simulators for breadth, then real hardware for truth: at least one phone and one tablet per shipped platform, both orientations, split-screen where supported. + +When the adaptation feels native to each context, hand off to `{{command_prefix}}impeccable polish` for the final pass. + +**NEVER**: +- Ship a stretched phone layout on a tablet +- Port one platform's controls or navigation onto the other +- Hide core functionality on smaller devices (if it matters, make it work) +- Lock orientation to dodge a layout bug +- Trust simulators alone (posture, gestures, and performance need hardware) diff --git a/skill/reference/animate.md b/skill/reference/animate.md index 54877b4f0..f95e36dd9 100644 --- a/skill/reference/animate.md +++ b/skill/reference/animate.md @@ -10,9 +10,7 @@ Brand: motion is part of the voice; one well-rehearsed entrance beats scattered Product: 150–250 ms on most transitions. Motion conveys state: feedback, reveal, loading, transitions between views. No page-load choreography; users are in a task and won't wait for it. -## Platform - -Native (`ios` / `android` / `adaptive`): motion mirrors the system. Match platform navigation transitions and honor the OS Reduce Motion setting; see the Motion sections of [ios.md](ios.md) and [android.md](android.md) (`adaptive` follows each OS). +Native (`ios` / `android` / `adaptive`): implementation follows the Motion section of [ios.md](ios.md) / [android.md](android.md) (read it first if Setup hasn't already): system transitions and OS Reduce Motion, never the web tooling below. --- diff --git a/skill/reference/audit.md b/skill/reference/audit.md index dcf415f39..28bafdf17 100644 --- a/skill/reference/audit.md +++ b/skill/reference/audit.md @@ -2,9 +2,7 @@ Run systematic **technical** quality checks and generate a comprehensive report. This is a code-level audit, not a design critique. Check what's measurable and verifiable in the implementation. -## Platform - -The dimensions below are written for web. Native (`ios` / `android` / `adaptive`) audits translate: accessibility means **VoiceOver / TalkBack** correctness (labels, roles, reading order, Dynamic Type / scalable text reflow); touch targets are 44 pt (iOS) / 48 dp (Android); appearance covers Dark Mode / dark theme. `detect.mjs` is web-only; never run it against native code. See [ios.md](ios.md) / [android.md](android.md) (`adaptive` audits both). +**Web only.** Native platforms (`ios` / `android` / `adaptive`) route to [audit.native.md](audit.native.md) instead; if the project is native, switch to it now. ## Diagnostic Scan diff --git a/skill/reference/audit.native.md b/skill/reference/audit.native.md new file mode 100644 index 000000000..b79684e86 --- /dev/null +++ b/skill/reference/audit.native.md @@ -0,0 +1,139 @@ +Run systematic **technical** quality checks on a native app (`ios` / `android` / `adaptive`) and generate a comprehensive report. Don't fix issues; document them for other commands to address. + +This is a code-level audit, not a design critique. Audit from source (SwiftUI / UIKit / Compose / React Native / Flutter); no browser tooling or `detect.mjs` applies. Score against the platform reference(s): [ios.md](ios.md) / [android.md](android.md), both for `adaptive`. Read them before scoring if Setup hasn't already. The report skeleton mirrors [audit.md](audit.md); keep the two in sync when changing it. + +## Diagnostic Scan + +Run comprehensive checks across 5 dimensions. Score each dimension 0-4 using the criteria below. + +### 1. Accessibility (VoiceOver / TalkBack) + +**Check for**: +- **Missing labels**: interactive elements without accessibility labels, traits/roles, or state announcements +- **Reading and focus order**: illogical traversal, unreachable controls, focus lost on navigation +- **Text scaling**: fixed point sizes defeating Dynamic Type (iOS) or px instead of sp (Android); layouts that clip or overlap at large sizes +- **Touch targets**: below 44 pt (iOS) / 48 dp (Android), or crammed without spacing +- **Reduce Motion ignored**: parallax and large slides with no crossfade alternative +- **Contrast**: text failing contrast in either appearance, light or dark + +**Score 0-4**: 0=Screen reader unusable, 1=Major gaps (unlabeled controls, no scaling), 2=Partial (labels exist, order or scaling breaks), 3=Good (minor gaps), 4=Excellent (labeled, ordered, scales cleanly, Reduce Motion honored) + +### 2. Performance + +**Check for**: +- **Slow startup**: heavy work on launch before first frame +- **Unvirtualized lists**: long content without FlatList / LazyColumn / List recycling +- **Main-thread jank**: synchronous work in scroll or gesture paths, dropped frames on 60/120 Hz +- **Wasted rendering**: unnecessary re-renders (React Native) or recompositions (Compose); missing memoization/keys +- **Image handling**: full-size images decoded for thumbnails, no caching +- **App weight**: bloated JS bundle or binary, unused dependencies + +**Score 0-4**: 0=Janky everywhere, 1=Major problems (unvirtualized lists, slow launch), 2=Partial, 3=Good (minor improvements possible), 4=Excellent (fast launch, smooth scroll, lean) + +### 3. Appearance & Theming + +**Check for**: +- **Hard-coded colors**: raw hex instead of semantic system colors (iOS) / Material color roles (Android) / design tokens +- **Broken dark appearance**: missing dark variants, poor contrast in dark, quick inverts +- **Dynamic Color** (Android 12+): no static fallback scheme, or ignored where it fits +- **Off-platform materials**: hand-rolled blur/glassmorphism instead of system materials or tonal elevation + +**Score 0-4**: 0=Hard-coded everything, 1=Minimal tokens, 2=Partial (tokens exist, inconsistently used), 3=Good (minor hard-coded values), 4=Excellent (semantic throughout, both appearances first-class) + +### 4. Platform Conformance (CRITICAL) + +Score against the loaded platform reference(s), including their slop tests. **Check for**: +- **Broken system gestures**: edge-swipe back disabled (iOS), predictive Back hijacked (Android) +- **Inset violations**: content under the notch, Dynamic Island, home indicator, status bar, or keyboard +- **Off-platform navigation**: custom global nav, overloaded tab bars, iOS patterns on Android or vice versa +- **Web-shaped controls**: HTML-style buttons, custom toggles, hover-dependent affordances +- **Icon drift**: mixed icon sets instead of SF Symbols / Material Symbols +- **AI tells**: the shared absolute bans still apply (AI palette, gradient text, hero metrics) + +**Score 0-4**: 0=Web port (nothing native), 1=Heavy violations (3-4 kinds), 2=Some (1-2 noticeable), 3=Mostly conformant (subtle issues), 4=Fully native (a fluent user trusts every screen) + +### 5. Adaptivity + +**Check for**: +- **Stretched phone layouts**: tablet/iPad rendering a scaled-up phone UI instead of using size classes / window size classes +- **Orientation breakage**: landscape clipping, ignored, or locked without reason +- **Keyboard/IME handling**: inputs hidden behind the keyboard, no inset adjustment +- **Multitasking**: iPad Split View / Android multi-window breaking layout +- **Foldables**: hinge-unaware layouts on posture change (Android) + +**Score 0-4**: 0=One screen size only, 1=Major breakage (landscape or tablet broken), 2=Partial, 3=Good (minor edge cases), 4=Excellent (adapts across sizes, orientations, and windowing) + +## Generate Report + +### Audit Health Score + +| # | Dimension | Score | Key Finding | +|---|-----------|-------|-------------| +| 1 | Accessibility | ? | [most critical issue or "--"] | +| 2 | Performance | ? | | +| 3 | Appearance & Theming | ? | | +| 4 | Platform Conformance | ? | | +| 5 | Adaptivity | ? | | +| **Total** | | **??/20** | **[Rating band]** | + +**Rating bands**: 18-20 Excellent (minor polish), 14-17 Good (address weak dimensions), 10-13 Acceptable (significant work needed), 6-9 Poor (major overhaul), 0-5 Critical (fundamental issues) + +### Platform Conformance Verdict +**Start here.** Pass/fail: does this read as a native app or a ported website? List specific violations. Be brutally honest. + +### Executive Summary +- Audit Health Score: **??/20** ([rating band]) +- Total issues found (count by severity: P0/P1/P2/P3) +- Top 3-5 critical issues +- Recommended next steps + +### Detailed Findings by Severity + +Tag every issue with **P0-P3 severity**: +- **P0 Blocking**: Prevents task completion. Fix immediately +- **P1 Major**: Significant difficulty or platform-guideline violation. Fix before release +- **P2 Minor**: Annoyance, workaround exists. Fix in next pass +- **P3 Polish**: Nice-to-fix, no real user impact. Fix if time permits + +For each issue, document: +- **[P?] Issue name** +- **Location**: Screen, file, line +- **Category**: Accessibility / Performance / Theming / Conformance / Adaptivity +- **Impact**: How it affects users +- **Guideline**: The HIG / Material rule it violates (if applicable) +- **Recommendation**: How to fix it +- **Suggested command**: Which command to use (prefer: {{available_commands}}) + +### Patterns & Systemic Issues + +Identify recurring problems that indicate systemic gaps rather than one-off mistakes: +- "Hard-coded colors appear in 15+ screens, should use semantic colors" +- "Touch targets consistently below 44 pt throughout the tab bar and list rows" + +### Positive Findings + +Note what's working well: good practices to maintain and replicate. + +## Recommended Actions + +List recommended commands in priority order (P0 first, then P1, then P2): + +1. **[P?] `{{command_prefix}}command-name`**: Brief description (specific context from audit findings) +2. **[P?] `{{command_prefix}}command-name`**: Brief description (specific context) + +**Rules**: Only recommend commands from: {{available_commands}}. Map findings to the most appropriate command. End with `{{command_prefix}}impeccable polish` as the final step if any fixes were recommended. + +After presenting the summary, tell the user: + +> You can ask me to run these one at a time, all at once, or in any order you prefer. +> +> Re-run `{{command_prefix}}impeccable audit` after fixes to see your score improve. + +**IMPORTANT**: Be thorough but actionable. Too many P3 issues creates noise. Focus on what actually matters. + +**NEVER**: +- Report issues without explaining impact (why does this matter?) +- Provide generic recommendations (be specific and actionable) +- Skip positive findings (celebrate what works) +- Forget to prioritize (everything can't be P0) +- Report false positives without verification diff --git a/skill/reference/layout.md b/skill/reference/layout.md index 0a52c957c..e4df8647d 100644 --- a/skill/reference/layout.md +++ b/skill/reference/layout.md @@ -8,9 +8,7 @@ Brand: asymmetric compositions, fluid spacing with `clamp()`, intentional grid-b Product: predictable grids, consistent densities, familiar navigation patterns. Responsive behavior is structural (collapse sidebar, responsive table), not fluid typography. Consistency IS an affordance. -## Platform - -Native (`ios` / `android` / `adaptive`): structure follows platform navigation (iOS tab bar / navigation stack; Android navigation bar / rail / drawer), safe-area / window insets, and spec touch targets (44 pt iOS, 48 dp Android). See [ios.md](ios.md) and [android.md](android.md); `adaptive` lays out per OS. +Native (`ios` / `android` / `adaptive`): structure follows the Layout section of [ios.md](ios.md) / [android.md](android.md) (read it first if Setup hasn't already): platform navigation, insets, and touch targets, never the CSS tooling below. --- diff --git a/tests/skill-behavior/README.md b/tests/skill-behavior/README.md index 8177947d8..11ca68f65 100644 --- a/tests/skill-behavior/README.md +++ b/tests/skill-behavior/README.md @@ -55,6 +55,7 @@ The trace is the source of truth, not the model's free-form reply. | 12 | empty workspace; prompt is natural-language build intent with no command word | runs `context.mjs`, diverts into `reference/init.md`, and does **not** start writing HTML/CSS | | 13 | empty workspace; prompt is `/impeccable teach` | runs `context.mjs` and diverts into `reference/init.md` because `teach` aliases `init` | | 14 | PRODUCT.md with `## Register: product` + `## Platform: ios` (native iOS app); prompt is `/impeccable craft a tide detail screen` | `context.mjs` runs and emits a NEXT STEP pointing at `reference/ios.md` (proven via captured bash output); agent loads `reference/ios.md` (Setup step 5, native conventions on top of the register reference) | +| 15 | same iOS fixture; prompt is `/impeccable audit` | agent loads `reference/audit.native.md` (the Commands-table native variant, routed instead of `audit.md`) | Scenario 9 passed on all three current-lineup providers (`claude-sonnet-4-6`, `gpt-5.5`, `gemini-3.1-flash-lite`) on 2026-05-28. diff --git a/tests/skill-behavior/scenarios.test.mjs b/tests/skill-behavior/scenarios.test.mjs index 3f34503c7..79e8ae4c5 100644 --- a/tests/skill-behavior/scenarios.test.mjs +++ b/tests/skill-behavior/scenarios.test.mjs @@ -595,5 +595,38 @@ for (const modelId of resolveModelList()) { cleanupWorkspace(workspace); } }); + + it('scenario 15: native audit routes to the native command variant', async () => { + // The Commands table lists audit.native.md as the native variant and + // Setup step 2 says to read the variant INSTEAD of audit.md when the + // platform is native. This pins the route-instead behavior: a native + // audit must reach audit.native.md (reading audit.md first and then + // switching via its web-only guard is acceptable; never reaching the + // variant is the failure). + const workspace = prepareWorkspace({ + files: { 'PRODUCT.md': PRODUCT_MD_SAMPLE_IOS }, + }); + try { + const { trace, text } = await runTurn({ + workspace, + model, + userPrompt: '/impeccable audit the app in this workspace', + maxSteps: 6, + }); + logTrace('S15', 'native-audit-variant', modelId, trace, { textSample: text.slice(0, 400) }); + assert.ok( + bashCommandsMatching(trace, 'context.mjs').length >= 1, + `expected agent to run context.mjs at least once.\n` + + `bashCommands: ${JSON.stringify(trace.bashCommands, null, 2)}`, + ); + assert.ok( + fileLoaded(trace, 'audit.native.md'), + `agent should load audit.native.md (not just audit.md) when the platform is ios.\n` + + `Trace: ${JSON.stringify(summarizeTrace(trace), null, 2)}`, + ); + } finally { + cleanupWorkspace(workspace); + } + }); }); }