mirror of
https://github.com/pbakaus/impeccable.git
synced 2026-09-12 14:16:28 +03:00
c9c0fc887bc105cd8037bbcb4bd186e695c49e80
119
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
916b0a1fdf |
Generate degraded-mode fallback references from the subagent definitions
Harnesses with no subagent capability now run each role inline from the
same single source. The build emits reference/degraded/<role>.md for every
agent in skill/agents/ (role name is the agent name minus the impeccable-
prefix), stripping frontmatter and prepending the inline-substitution
preamble. These pass through the same provider-block compilation and
placeholder replacement as ordinary reference files, so <codex> blocks and
{{placeholders}} resolve per target, and they land in the committed harness
dirs on build:release like every reference file.
Repoint the three capability-first fallback sites in the prose at the
generated files: new-work.md reviewer and documenter fallbacks, and
visualize.md asset-producer fallback.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
6769b1879a |
The polish ceiling covers the whole cycle, and the handoffs end it
Probe attribution on Opus 5 showed the screenshot bound working (42 to 16) while the real burner ran free: five rounds of node -e micro-edits, eight rebuilds, and inline defect hunts absorbed the reviewer's and documenter's jobs until the turn cap killed the run mid-hunt. The two-round ceiling now names scans, micro-edits, and rebuilds; after the second round the build thread stops polishing and ships the rest through the reviewer (one batched fix pass, one rebuild, stop) and the documenter. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
3d2ffe9007 |
Drop internal filename cross-references from routed reference text
Copilot's review point stands: reference files load per-command, so a bare "see optimize.md" / "typeset.md" is not meaningful in the routed context. The guidance reads self-contained now. Prepared with AI assistance (Claude Code), directed by @pbakaus. Co-Authored-By: Claude Code <noreply@anthropic.com> |
||
|
|
24e24265d1 |
Refresh stale metric and library references in the command docs
From issue #395, the items still present after the v4 consolidation: - optimize.md led its interactivity section with FID, retired as a Core Web Vital in March 2024 when INP replaced it. The section heading and both metric lists now name INP. - optimize.md recommended react-virtualized, superseded by react-window from the same author; the line now points at react-window and TanStack Virtual, matching overdrive.md. - overdrive.md's WebGPU support matrix predated Firefox 141/147 shipping it on Windows/macOS and Safari 26 shipping it across Apple platforms. - audit.md listed "missing will-change" as a defect while animate.md and optimize.md both instruct applying it sparingly and never preemptively; the audit line now flags overuse instead of absence. - harden.md allowed 14px mobile body text while typeset.md sets a 16px ordinary floor; harden now matches the floor, reserving 14px for secondary text, and names the iOS Safari input-zoom consequence. The issue's other items (Framer Motion naming, Popmotion, polish duration cap, humor guidance, HSL phrasing in quieter) were already resolved by the v4 reference rewrite. Prepared with AI assistance (Claude Code), directed by @pbakaus. Co-Authored-By: Claude Code <noreply@anthropic.com> |
||
|
|
8634c538fb |
Verification is two bounded rounds, never a loop
Opus 5 turned the iterate-with-screenshots-until-it-meets-the-bar instruction into 42 screenshot trips and 150 tool calls per build, about forty dollars of cache churn a page, before ever reaching the reviewer. Verification now batches: one desktop-and-mobile round after the full build, fixes applied together, one confirming round, ceiling two. Craft-floor's checks share those renders instead of earning separate trips; per-tweak iteration is live mode's channel. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
bb57be4243 |
Documenter subagent, reviewer handoff contract, asset gate
From the paired Opus and Codex manual-run analyses. DESIGN.md moves to the end of the flow and into a shipped documenter subagent that derives the system from the built artifact: a rulebook written before the build gets defended against reality, and a half-stable DESIGN.md hands the design-system detector an unstable target that buries the build in noise and invites laundering. The finish reviewer gains the handoff that failed three times live: the parent captures desktop and mobile screenshots and passes paths, the reviewer never attempts to render and names missing inputs in one line, the parent verifies the five-section return and respawns once on empty. Fidelity against the approved comp joins its checks; the card keeps commitment only. The comp ingredient inventory becomes a written gate with raster-by-default materials and no gradient-as-texture, comps persist under .impeccable/mocks, the degraded seed names the sandboxed-exec cause, and the finish line is explicit: a clean detector pass is not finished. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
253f8e510c |
Concept machinery: survive truncation, builds, and loud briefs
The release-gate audit traced four ways the roll's output was defeated downstream of a perfectly healthy seed. Gemini's harness keeps only the tail of tool output, so the header-only ASSIGNED INDEX never reached the model in 18 of 18 samples; the seed now restates the assignment and key at the end of its output. Astro strips frontmatter comments, so half the anthropic contracts vanished from built artifacts; the contract now must survive the production build as an HTML comment in emitted markup. A brief that paints its own picture (the album named Soft Cathedrals) converged every arm regardless of assigned index; its literal reading now joins the rut with at most one candidate. And Opus under 4.0.1 skipped the seed 42% of the time while hand-authoring plausible contracts; the finish reviewer now verifies FORM carries a corroborable seed key before any craft point. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
ffe869f4d0 |
Drop the turn-cap exception from the visualize mandate
Paul's call: the build-exhaustion failure only exists inside eval workers with hard turn budgets no real harness exposes, and the clause doubled as a hedge door for skipping the comp round. The eval-side fix belongs in the worker's max-turns, not in skill prose. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
e76ff27adf |
Eval-found fixes: workspace-relative cards, build outranks comps at caps
The release-gate campaign confirmed two skill bugs with transcripts. Sandboxed harnesses reject absolute paths, so following the CHOSEN CARD directive with the absolute card-base path failed view_image; the directive and the quality-bar clause now say download into the workspace and open the relative path. And under the openai worker's turn cap, models spent the budget on init, cards, and comp generation and never built the page (a third of small-n supplement slices); the visualize mandate gains its one exception: at a hard cap the shipped page outranks optional imagery. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
73dec5d159 |
Subagent authorization becomes a central harness counter
Paul's call: the reviewer-local authorization patch covered one command while the harness gate silently disables every shipped subagent, critique panels and the manual-edit applier included. The argument now lives beside the autonomy counter in context.mjs, emitted as tool-result content every run: invoking the skill is the user request such gates ask for; spawn where a reference directs; the in-thread substitute is for absent capability only and gets disclosed in one line. new-work keeps the reviewer mechanics and drops the now-central argument. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
4dc2b4d694 |
Finish reviewer: the skill invocation authorizes its subagents
A live session on a harness whose guidance gates subagent use on user request resolved the conflict silently against the skill: it never spawned the reviewer, stretched the no-subagents fallback to cover permission hesitancy, and self-reviewed with all the context that made its choices feel correct. Three tightenings: invoking the skill IS the user request that authorizes its shipped subagents; the fallback is for harnesses lacking the capability, not for hesitancy; a substituted review gets disclosed in one line at finish, never silently. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
dbe0c12b91 |
Live: stop the preflight writing source, cache the resolution
The polling-rework preflight wrote the variant scaffold into source during the poll lease, before the agent acted. On source-preview targets (React/Vue/Vite, everything but the svelte-component path) that write full-reloaded the framework; a browser caught mid-reload missed the agent's variant write and the SSE done, and sat stranded at 0/N. Restore the 3.5 single-atomic-edit semantics: the preflight still resolves the element location and computes the scaffold, but --defer-source-write leaves source untouched and hands the agent the wrapper text plus the picked source range. The agent splices variants into the wrapper and replaces the range in one write, so the framework reloads exactly once. The svelte-component path is untouched (it never writes route source). The missed-completion recovery stays as defense in depth. Also cache the resolved source file per target signature (locator + route): the ~7.6s tree search re-ran on every generate for the same element; a hit now points the helper straight at the file via --file, invalidated when the target changes or a resolution fails. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
4cd5ea7547 |
Add TanStack Router + Start support to live mode
Live mode had no TanStack coverage: a TanStack Start user hit disconnects and static previews because there is no static index.html to inject and no adapter for the SSR root document. - New tanstack-adapter.mjs, modeled on the SvelteKit/Nuxt adapters: detects a TanStack Start project (@tanstack/react-start + src/routes/__root.tsx) and patches the __root document to mount a generated dev-only React component (src/impeccable/ImpeccableLiveRoot) that appends the live bundle on the client after hydration, carrying the ?token= param via buildLiveScriptSrc. Patch/unpatch round-trips byte-for-byte and is idempotent; refuses to clobber an unmanaged file at the component path. - Wire detection into live-inject.mjs (insert + remove + gitignore), ordered so SvelteKit/Nuxt win and a plain TanStack Router SPA falls through to the baseline Vite index.html path. - tanstack-router-vite fixture (baseline, no adapter) and tanstack-start fixture (SSR adapter), both with runtime blocks. Both pass the full live-e2e cycle (handshake, steer, pick, Go, cycle, accept, carbonize, reloadProbe). - Unit tests for detection + patch round-trip + apply/remove; tanstack-start branches in framework-fixtures.test.mjs; live.md framework table + adapter note. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d4d02b69f2 |
Live: overlay preview is the verification channel; disconnects resume
Two prose fixes from the 3.5-to-4.0.1 forensic diff of a real user regression (15-minute tweaks, repeated disconnects, agent abandoning the picker). The craft-fold made every generate cycle pay the verify- the-built-result loop the overlay already provides to the human; live cycles now verify by construction and run the full check once at accept. And nothing framed a dropped SSE or closed tab as resumable, while the client toasts "Session ended", so agents rationalized bailing to direct edits; the journal is canonical and reopening continues the session. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
bcdf38881e |
Command first, capability second
"When the harness can X, do Y" hands the model an exit before the command arrives; the observed reviewer skip walked through exactly that door. The three gated constructions now lead with the imperative, present the decision visually, open the chosen card, spawn the finish reviewer, and carry their fallbacks as trailing clauses for sessions that genuinely lack the capability. Constructions that already led with the command keep their routing clauses unchanged. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
0bbb63b62a |
Ship the finish reviewer as a named subagent; ungate the asset producer
The eb686f36 session read the separate-reviewer rule and spawned nothing: an unnamed "separate agent" is an improvisation prompt, not an affordance. The skill now ships impeccable-finish-reviewer next to the asset producer: persistence first, ceiling against the card and comp second, contract promise by promise, truth; ordered material fixes back to the parent, no editing, no second detector. new-work names it so the finish step invokes a thing that exists. The asset producer was gated providers: codex, so Claude Code never shipped it; the gate is removed and its two codex-only workflow lines made provider-neutral with codex blocks. Dist rebuild still deferred for the running campaign. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
91d310696d |
Canonicalize the visualize flow; put the added prose on a diet
codex.md becomes visualize.md and loads for every harness with any image generation, native or the API fallback: after the direction locks, three distinct compositional comps are rendered and put before the user for approval, in-harness when it can display images, otherwise on the decision page. Three is the number; one comp invites rubber-stamping, and this approval round has repeatedly produced the most compositional and ambitious work, so new-work now marks it never-skipped. The codex-only subagent stays as a codex note. The recent rule additions are tightened by a third: the asset and imagery bullets merge into one, the canon exit loses its restatements, the DESIGN.md-rule and chosen-card and ceiling clauses each shed their second clause saying the first clause again. Same laws, fewer words; prose that grows without bound recreates the attention gravity it was written to fight. Dist rebuild still deferred; the release-gate campaign reads the pinned dist. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
e409bec7b5 |
Canon standing exit, chosen-card directive, and the ambition fixes
From Paul's approved UX and the eb686f36 session post-mortem: The standing exit: direction rounds carry a quiet, permanent "Play it straight" action (payload flag canon, reserved id) on the decision page and as the last structured-tool option. It is the user's door, never the model's: never recommended, never weighed against the roll, and choosing it swaps the bar rather than lowering it, two or three named reference products become the craft level, canon executed at full commitment. Safer/conventional steers resolve here, never to a stranger re-roll. Session fixes, each mechanical where possible: the ANSWER line now names the chosen card's hero and board and directs opening them before code (the session built from text alone after viewing a different world's card); generation scale joins the imagery rule (a library of centered 128px subjects foreclosed the atmospheric hero); DESIGN.md rules are checked against the world's native devices and never added to silence a hook finding (the session banned arcade lettering's own offset shadow and laundered 8px through the ramp); staging joins the FORM contract block (the axis was dropped silently at world-choice); the finishing reviewer audits the ceiling against the QUALITY BAR card after persistence (floor rigor was disguising unreached ambition); the icon-tile clause names hand-drawn icons as remedy, not target. Dist rebuild deferred: the release-gate campaign reads the pinned dist. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
9dade04bbf |
Text fallback presents surviving challengers as alternates
The structured-tool channel collapsed to a single direction plus re-roll, which read as "the system only ever offers one idea" next to the multi-card decision page. Both channels now share one structure, assigned direction leading, the one or two fused challengers that survived the weighing as named alternates, re-roll with steer, and differ only in richness. The anti-lineup rule stays precise: what never appears is a ranked menu of the model's own grounded candidates; dealt challengers carry no ranking rut. Note: dist rebuild deliberately deferred; the release-gate campaign is running against the pinned dist and rebuilding mid-run aborts it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
13c2ae5aa4 |
DESIGN.md joins the persistence gate
With the PRODUCT.md skip fixed, the Opus smoke unmasked the adjacent gap: the model builds a new world and never writes DESIGN.md (zero attempts), so the worker's requiresDesign assertion correctly fails the run. Same disease, same treatment: DESIGN.md is now part of recording the decision, written before the first build edit in the same stretch as the direction contract, and the finishing reviewer checks persistence first, before any craft point is scored. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
c0ee4f3cac |
Build mandate: author the assets, generate the imagery
The truth split already permits full-fidelity demonstration data, but permission at selection time was not holding at build time: models that would not author covers, names, or thumbnails compensated with chrome, which is the content-starved look the detector hunts. Two build rules make it a mandate: every blank the ask round left open is authored at production fidelity (content is authorable, claims are labelable, nothing is omittable; unanswered commercial claims ship as marked placeholders with a replacement list), and when image generation is available, generating the build's imagery is part of building rather than a nicety. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
5a3c8fd18e |
Degraded roll keeps the decision page as its channel
A live retest showed the model dropping to the structured question tool when the roll degraded: with no challengers and no cards it judged the page pointless and presented one option in plain text. The degraded seed output and the new-work rule now both state that degradation changes the cards, not the channel; a browser session presents the assigned direction as a single text-only card with re-roll on the decision page. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
6e8cdfa581 |
Counter the harness autonomy directive from inside the working turn
A traced Claude Code injection asserts for whole model families that the user is not watching and cannot answer questions; it ships default-on with no off switch, and it suppressed every interactive step of a live run (interview skipped, PRODUCT.md inferred, decision page never served). Prose in a reference file loses that argument, placement wins it: context.mjs now emits AUTONOMY_DIRECTIVE_CHECK as tool-result content in the working turn, telling the model such a claim is a harness default, never session evidence, and to probe once with the question tool before inferring. init.md makes the same test mechanical: tool presence proves an answer mechanism, one real probe round is required, inference afterward must be labeled and disclosed in the first reply. The degraded concept-seed path now also tells the model to disclose the degraded roll instead of presenting it as a full one. Image-gen signaling stays positive-only per Paul: key present emits the capability, absence stays silent so harness-native tools are not suppressed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
5575a027dc |
Flag and repair drift in Impeccable's own project artifacts
v4 changed PRODUCT.md's shape and retired the register axis, so an upgraded project can carry answers nothing reads. Nothing measured that. Two tiers, and the split is a performance contract: - Boot (context.mjs, emitting CONTEXT_STALE) spends only what a boot already spends: markdown already in memory, a bounded set of stats, the small JSON files the boot reads anyway. No new directory walks. One directive for the whole set, throttled to once a week per project so a finding the user declined does not reappear tomorrow. - doctor.mjs runs the deep pass on demand: git drift, ignore lists validated against the live rule registry, hook script paths that stop resolving, and the monorepo workspace sweep. --fix applies only the migrations that carry no decision. Findings are data, not prose, so the boot directive, the text report and --json all render one set. Severity says what should happen: auto (fix on the next write anyway), mention (state once), route (name the command that owns the repair). PRODUCT.md now carries a schema stamp so the checks stop reconstructing a file's vintage from which sections it happens to have. Schema version, not release version: a record written by 4.0.0 is not stale under 4.0.1. DESIGN.md gets no stamp, because it follows the external design.md spec that Stitch lints and every DESIGN.md signal is measurable without one. The highest-value catch is a project that resolves to web while carrying native build files, including a monorepo app inheriting a root record that says web. That one costs output quality silently; nothing failed before. doctor follows the hooks/pin pattern rather than the Commands table, so it stays out of the design menu and the count stays at 23. Also corrects CLAUDE.md, which still documented the register axis, reference/brand.md, reference/product.md, eleven deleted domain reference files, and an extractRegister() whose only occurrence in the repo was that sentence. Prepared with AI assistance (Claude Code). Co-Authored-By: Claude <noreply@anthropic.com> |
||
|
|
b0a7deb688 |
serve-question: self-detect headless environments, capability-first routing
The env-var bypass (IMPECCABLE_QUESTION_DISABLED) relied on the harness remembering to set it. The script now also self-detects CI, SSH-without- display, and displayless Linux and exits 2 with the structured-question advice; --no-open skips detection (caller opens the URL itself, as the tests do) and IMPECCABLE_QUESTION_FORCE=1 overrides it. new-work.md now frames the decision-page rule by capability: open a browser if you can, structured question tool if you cannot, exit 2 means fallback not error. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
6d47843867 |
Stop losing DESIGN.md and native platform refs across init
Bugbot flagged the "resume without rerunning context.mjs" instruction after init. It is right, and the gap is wider than the platform half it named: context.mjs has two output branches, and the no-PRODUCT.md branch omits DESIGN.md, the native platform references, and the unrecognized `## Platform` warning. Because the skill never reruns the script once init writes PRODUCT.md, whatever that first run withheld is gone for the whole session. A greenfield iOS project would be designed without reference/ios.md ever loading, and a project carrying DESIGN.md without PRODUCT.md never saw its own design system. The two halves need different fixes. DESIGN.md is authority in its own right and does not depend on PRODUCT.md existing, so context.mjs now emits it on both branches. Platform is unknowable before PRODUCT.md exists, so no change to the script can recover it; init.md, the one step that learns the answer, now loads ios.md / android.md / both right after recording a native platform, and SKILL.src.md says so where it tells the agent not to rerun. Verified end to end against a temp project on both branches. Co-Authored-By: Claude <noreply@anthropic.com> |
||
|
|
9213e1bcd1 |
Fix stale section pointer and critique score denominators
Two true positives from the Bugbot review on PR #397. document.md seed mode told the agent to run "Select one direction" for paths A, D, or E. new-work.md has neither that heading nor the A/D/E lettering since the workshop was restructured into named subsections, so a literal read could skip the world-and-surface flow entirely. Point at "Create or replace the visual world" and "Commit the world" instead. critique.md let the heuristic table renormalize to an applicable maximum when heuristics are scored n/a, but the report template hardcoded ??/40, the rating bands only mapped raw numbers out of 40, and the persisted meta carried total_score with no denominator. Trends could silently compare 24/32 against 30/40 as if they were the same scale. The template now prints the applicable max, the bands fall back to percentages for partial sets, the snapshot records max_score and na_heuristics, and the trend line states its denominator or breaks it out per run when they disagree. critique-storage.mjs serializes frontmatter key-agnostically, so the new keys need no code change. Co-Authored-By: Claude <noreply@anthropic.com> |
||
|
|
3c47eb1a8c |
Claude-conditional counterweight for the warm-subject rendition prior
Gate2 measured it precisely: on matched assignments Opus renders warm, bookish, and child-facing subjects as cream, serif-italic, and lamplight while Sol renders the same positions saturated, and neutral prose hardening did not move it. The codex and gemini blocks set the precedent for provider-conditional counterweights; this adds the claude block at the palette decision: the first palette is already spent, an OWN-WORLD block reading cream/paper/parchment/lamplight for an unpinned Persuade surface is a failed rendition to rework from the world's saturated materials, and nothing about the subject requires the default. Verified present in the claude-code dist and absent from codex. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
68f42c4e33 |
Detect the closed tab: presence heartbeat and exit 4
Paul's question exposed the blind spot: a closed tab left the agent waiting out the full timeout. The page now sends a heartbeat every five seconds while open; the server stamps lastBeat into the state file, and --wait reports PAGE CLOSED with exit 4 when the beats stop for fifteen seconds without an answer. The prose defines the fallback ladder: re-present once through the structured question tool, then proceed unattended with the assigned direction, stating assumptions. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
f5cee8ebec |
Close the two escape hatches Opus used on kids-reading
Craft-gate forensics on matched hands: both models obeyed the same assigned indices, but Opus rendered every kids cell as cream paper, lamplight, and serif, and sample 1 chose Fraunces off the reflex list with a bookshop-signage rationalization, while Sol rendered the same positions as indigo bookcloth, coral thread, tomato, and marigold. The dice work; the rendition prior escaped through two hatches, now closed: naming a reflex face requires a reason no other face satisfies and a subject association is never that reason; and bookish or child-facing subjects do not soften the calibration, because cream paper is the smallest corner of the book world. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
2750738c18 |
The agent routes the question URL to the best browser
Paul's call: start mode never auto-opens; the agent is alive and opens the printed URL itself, in-app browser first, then the system opener, then showing the URL (--open forces the system browser from the script). The prose now leads with the start/open/wait flow and keeps the blocking auto-open path for harnesses that can background a shell. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
285ed7ef78 |
serve-question: schema discoverability and a non-blocking mode
Paul's two concerns with the blocking design. --schema prints the exact payload example so the model never guesses the shape (new-work.md points at it). And harnesses that cannot leave a shell blocked (or cannot open a browser while blocked) get a two-phase path: --start daemonizes the server and returns the URL plus a key immediately, --wait polls for the answer with exit 3 meaning ask again, exit 2 meaning the server died, and --stop for cleanup. The browser open happens from the detached server process, so it works even when the agent thread is short-lived. State lives under .impeccable/questions/. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
5612cdf45b |
Visual decisions and default visualization
Paul's design, three pieces: serve-question.mjs: the world decision presented as a themed page instead of a text prompt. The script serves an impeccable-styled option board (assigned direction leading with THE ROLL badge, dealt challengers as alternates carrying their QUALITY BAR cards, re-roll and steer built in), prints the URL, opens the browser, and blocks until the user chooses; the answer lands on stdout as ANSWER JSON, so the shell call itself is the wait and no harness machinery is needed. Local images are served by the ephemeral server; nothing leaves the machine. generate-image.mjs + context.mjs IMAGE_GEN_AVAILABLE: when an OpenAI key is in the environment, context reports that image generation works even without a harness-native tool (gpt-image-2, billed to the user's key, stated before first use; Google skipped by decision). Harness-native tools always win when present. new-work.md: visualize-before-build is now the default whenever any image generation exists, not a codex.md special case; the attended presentation prefers the visual decision page and falls back to the structured question tool. Evals keep the unattended path untouched. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
4da6729075 |
Quality-bar viewing works in download-only harnesses
Paul confirmed the craft-bar experiment: builds that saw the dealt worlds' hero cards produced visibly stronger execution than the no-image control. One clause makes the mechanism reachable for harnesses that read only local images: download the card, then view it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
e107133a99 |
Deal the world cards with the roll as a craft bar
Paul's directive: the rendered board and hero for each dealt world ride along with the challengers, framed as a quality bar (the finish and commitment level the build is expected to reach), never as a mockup to copy. The seed prints QUALITY BAR urls per challenger, preferring API-provided cardBoard/cardHero fields and deriving from the concept id otherwise; new-work.md instructs image-capable harnesses to view them for the world being built. Server side, the roll API now returns cardBoard/cardHero per challenger (impeccable-site). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
884ab6e9dc |
The grounded list admits nameable abstract systems
Discussion outcome with Paul: a mediocre material world loses to excellent abstract craft, but an unanchored "just be beautiful" escape hatch would hand selection straight back to the model's priors. The resolution: abstraction enters as named systems with their own grammar. The derivation now states that the audience's graphic and screen traditions (notation, publications, identity programs, data graphics, interfaces) are as concrete a candidate as any physical artifact. The catalog side of the same decision is a 12-entry abstract-graphic authoring round in impeccable-site, pending review. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
eec30d15a0 |
Divergence teeth and the pinned-world rendition rule
Smoke findings (Paul's review): the kids-reading derivation produced seven candidates from one material family despite the divergence line, and the brief-pinned bookshop world was rendered as the generic AI bookshop (cream, serif italic, soft glow). The list must now span at least three material families, and a pinned world licenses its full material range, never just its softest rendition. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
25e438c3af |
Truth binds claims not demonstrations; conversion lives inside the form
Paul's probe review (recovery-ab): the no-invention rule was blocking bold greenfield directions whose demonstration data does not exist yet; kids-reading amplified the product's "quiet support" adjectives into a whole-page aesthetic; and nothing guaranteed a Persuade surface still sells once the form commits (a prior generation shipped zero nav and zero CTA in the first viewport). - Truth split in two: commercial and factual claims stay uninventable; illustrative material is authored at full fidelity, labeled synthetic, with a replace-with-real list for the user. Mirrored in the build section so execution-only sessions get it too. - Persuade floor restored from a22: conversion lives inside the form's own vocabulary (one-line hook, visible primary action, legible reading order); a committed form that hides the offer has not finished translating. The contract's FIRST VIEWPORT block now names where the primary action sits, and the finishing review verifies the mode did its job. - Calibration: negative constraints rule out devices, not exuberance; product-behavior adjectives do not dictate surface energy. - Web leverage: when the chosen world names a technique (canvas, WebGL, view transitions), build the technique, not a static imitation. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
555ae81b81 |
Keep the category default off the candidate list
Probe finding (recovery-ab, obs sample 1): "treat both structures as the rut, not the range" let the model rank the observability dashboard grid at position 7, and the dice landed on it; the challenger fusion rescued that draw, but a die face spent on the category's own page is a wasted roll. The a26 wording excluded both structures outright; restore that with the reason attached. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
36e3c05ca7 |
Restore dice assignment, fusion, and the commitment counterweights
The ship40 concept pipeline had reversed the proven a-series mechanisms: the seed's roll decayed into a shortlist nomination that taste functions (model ranking, candidate floor, simulated user) then argmaxed into the safest card; the costume check returned as the Translation veto and carrier-removal test; and the 07-15 rewrite deleted the calibration, reflex-font lanes, color strategies, and commit-every-atom language that had held off the cream-editorial default since the alpha era. Five of six frozen craft directions converged on the same warm-paper family and both builders obeyed them. This lands the repair on top of the in-progress simplification: - new-work.md: the script assigns the build index again on both scopes; catalog challengers are fused (challenger supplies form and grammar, product supplies every fact, clarity wins conflicts) and weighed on the two proven axes only; attended runs present one fully committed direction with re-roll and an optional steer instead of a ranked lineup; the color-strategy picker, reflex-face list, saturated-look calibration, first-viewport thesis and memory test, commit-every-atom, scroll pacing, and prove-don't-claim return; the direction contract returns as five lean blocks audited by the separate-agent finish. - concept-seed.mjs: PROMOTED INDEX becomes ASSIGNED INDEX with build-assignment semantics; self re-roll only on named factual grounds. - craft-floor.md: hook-active sessions act on findings instead of re-auditing; the Refuse list is framed as category defaults the brief can earn; a closing commitment line keeps a ban list from being the last word before code. - codex.md / shape.md: contract references restored for flow coherence. Adopts the concurrent session's ceremony cuts, softened challenger instruction, seed SOURCE IDs and --candidate-count, detector-ownership fix, and the removal of the hook-side contract audit (the audit now belongs to the separate reviewer at finish). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
a9540b5fe3 |
Restore the overlay-clipping rule to operate.md
Deleting interaction-design.md took the last copy of skill-interaction-dropdown-clipping with it. harden.md's `overflow: hidden` hits are code samples, not the rule. It lands in operate.md's Components list rather than the craft floor: dropdowns and overlays are dense-product-UI components, and the floor just lost 25% of its length for being a place where specifics accumulate. The detector's clipped-overflow-container rule catches this after the fact, but only in sessions with a hook. Co-Authored-By: Claude <noreply@anthropic.com> |
||
|
|
cd4d710cf1 |
Delete the orphaned interaction-design.md
Nothing has loaded it since
|
||
|
|
83255698f6 |
Repoint reference files at the craft floor
live.md's insert branch told the agent to load `brand.md` or `product.md`. Both files are gone on this branch; the register system became SKILL.md's modes plus operate.md. Net-new markup in live mode now decides the mode from the surface and loads craft-floor.md, which is where the bans live. The freeform generate path gets the same pointer, since live never runs Setup step 3 and so never picks the floor up on its own. operate.md still located the craft floor inside SKILL.md. Co-Authored-By: Claude <noreply@anthropic.com> |
||
|
|
68eeb93b3d |
Tighten the craft floor
910 words to 682, same 35 rule markers, no guidance dropped. - Two sections instead of three. "Absolute bans" and "detector-blind reflexes" split the same list by whether our scanner happens to catch it, which is a fact about our tooling and tells the model nothing about the design. Merged into one Refuse list, grouped as page scaffolds and surface habits, which is a distinction the model can act on. - Folded three duplicates: text-overflow was already in the Type check, the uniform section reveal was the other half of the Motion check, and card-everything was already inside the card-grid ban. - Cut explanation the model does not need. It knows what gradient text is and what group-hover does; it needs the refusal, not the mechanism. The gemini block goes from four sentences to three short ones, and the motion palette line drops the CSS tutorial for "reach past transform and opacity." - The authority note moves to the header so no item has to hedge. Co-Authored-By: Claude <noreply@anthropic.com> |
||
|
|
153b416f2e |
Move the slop defects back into the craft floor
The detector-blind slop review existed because the AI-tell rules had been stripped out of SKILL.md and nothing carried them. The floor is a better home: it loads after concept ideation and immediately before editing UI, which is the placement that made stripping them necessary in the first place. Models tread lightly when a ban list is present during ideation; by the time the floor loads, the direction is already committed. - Rename build-floor.md to craft-floor.md and restore the absolute bans (side-stripes, gradient text, glassmorphism, hero-metric, identical card grids, eyebrow-on-every-section, numbered markers, text overflow), the codex and gemini defect lists, and the reflexes no scanner catches. Rule ids match the ones the ablation catalog already knows. - Delete lib/slop-review.mjs and both injections. The Stop hook is now purely a mechanical pass and stays silent with nothing to report. - context.mjs replaces AI_SLOP_REVIEW_REQUIRED with the narrower MANUAL_DETECTOR_REQUIRED, emitted only when a session has no hook at all. A per-edit hook already covers the mechanical gap, and the floor covers the judgment one either way. Co-Authored-By: Claude <noreply@anthropic.com> |
||
|
|
d7d10277d1 |
Merge main into oneshot-v4, keeping the service layer split out
main still carries the site, so every `site/` path resolves to deleted. `tests/docs-integrity.test.js` goes with it (it imports the site's demo renderer), and `package.json` keeps main's `@anthropic-ai/sdk` bump while dropping `@google/genai` and `@paper-design/shaders`, which nothing in the product layer imports. Real code merges: - hook-lib: main's #391 cache fix (sync the remembered set to the live scan so fixed findings stop being named and a reintroduced one fires again) now runs on the immediate tier rather than the whole filtered set. Remembering a deferred finding the per-edit pass never reported would let the Stop deep pass dedupe it away. main's `maxFileBytes` ceiling, `cleanAcked` once-per-file ack, and template-extensions re-export all land alongside the tiering work. - live-browser: main's `hasParams` gate on the Tune badge, keeping this branch's `C.ink` badge text so it stays legible on kinpaku gold. - detect-text: both the block-level codex-grid-background scan and main's inset-stripe CSS check. - test-suites: union of both trigger sets and file lists, minus the site-only entries (`shiki-theme`, `docs-integrity`). - Two hook tests moved off deferred-tier rules (`overused-font`, `side-tab`) onto immediate-tier ones. They assert cache bookkeeping, which the per-edit pass only reaches for the immediate tier. Also drops the site waivers from `.impeccable/config.json` and stops `build:browser` recreating a stray `site/` tree just to write a bundle the other repo builds itself. Co-Authored-By: Claude <noreply@anthropic.com> |
||
|
|
c8c952acd4 |
Rework the candidate floor around signature and translation
Replace the Consequence floor with Signature: one authored move that makes the experience unmistakable and shapes implementation, named in terms of what the visitor experiences. Translation now demands the source's aesthetic and compositional laws survive alongside product structure, so function-without-character reads as safe flattening and character-without-structure as costume. Add an expand-then-contract step before the direction contract: decide spatial, motion, interaction, narrative, and system questions as one studio plan that causes itself, then compress into the contract. Staging guidance follows the seed's move to several inputs. Co-Authored-By: Claude <noreply@anthropic.com> |
||
|
|
b5ec969c07 |
Add world roll API and seed telemetry client
/api/roll deals deterministic challenger rolls server-side (same salts and sha256 ranking as the local seed, verified bit-for-bit); the request log is the impression record. /api/chosen takes the anonymous choice ping. Events land in Workers Analytics Engine. concept-seed.mjs resolves data in order: local catalog dir, roll API, degraded promotion-only seed. --chosen sends the choice ping; DO_NOT_TRACK and IMPECCABLE_NO_TELEMETRY disable it. API-dealt seeds carry the telemetry instruction inline. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
7557935fdb |
Expand concept system: modes, ratings, re-roll, breadth strategy
Catalog: mode-aligned staging surfaces (persuade/operate/read/experience), star ratings on approvals feeding challenger draw weights, family retirements, authoring strategy and territory guide, rework and breadth authoring rounds, composition mining from rejected worlds. Seed: six challengers (two per tier), --reroll chains, --mode staging filter, rating-weighted draws. New-work: Present/visualize/re-roll flow, image-gen requirement, register-neutral vocabulary. Pipeline: per-mode staging prompts with split frames, hero-from-board reference generation, render-safety guards. Labs: ratings UI, unrated filter, mode chips, composition approve-guard. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d0ac67c6e9 |
Live: polling rework, source locks, preflight scaffolding (#381)
* Improve Live polling responsiveness and reliability Restore foreground/background polling as the primary harness architecture, add progressive publication and framework-safe previews, and harden quality and regression coverage. The experimental app-server runtime is intentionally excluded.\n\nPrepared with AI assistance under maintainer direction. * Fix source-safety, detector, and lock defects in Live polling work Addresses the review findings on #371, plus several the bots did not catch. All fixes have regression coverage that fails on the prior code. Source corruption: - Vue accept dropped valueless root attrs (disabled, v-cloak) and, worse, rewrote @click="x" as a literal click="x" DOM attribute, because the attr parser was name-anchored and skipped the sigil. Tokenize the whole Vue attr grammar and normalize shorthands so accept round-trips directives. - --variant was interpolated unescaped into a RegExp, so --variant '.*' matched the original block first and reported a successful accept while silently restoring the original. Validate against the digits pattern the browser and the /events schema already enforce. - --id reached path.join unvalidated, so --id ../../../../etc/evil wrote and read receipts outside the project. Hoist the existing safeSessionId check into impeccable-paths and apply it at every id-to-path sink. Accept/lock correctness: - Plain HTML/JSX accept and discard did not catch SOURCE_LOCKED, so contention exited non-zero with empty stdout and the agent got no JSON to retry on. - Lock staleness was mtime-only and never read the pid it records: a holder whose critical section outran 60s had its live lock swept, admitting a second writer to the same file, while a crashed holder blocked accepts for a full 60s. Decide staleness by owner liveness, and release only our own lock. Detector: - isNeutralColor only parses computed color forms, so routing authored CSS through it reported inset 4px 0 0 #000 / black / #e5e7eb as chromatic side-tab stripes. Add an authored-color neutrality test covering hex and named neutrals; the fixture had no literal-color cases at all. - Rule line numbers were off by one for every rule after the first, and commented-out CSS was scanned as live rules. Server: - An error reply carries no sourceEventType, and inferSourceEventType returned undefined, which acknowledgePendingEvent treats as a wildcard: a stale generate worker's failure consumed the user's queued Accept, which then reached no agent and left the browser in SAVING forever. - The generate preflight spawned live-wrap.mjs synchronously inside the request handler, freezing the single-threaded server for the whole scaffold (~7.6s measured on this repo, 15s ceiling) and stalling Accept/Discard/SSE. Make it async, claiming the lease before the first await so no event double-delivers. - Every browser checkpoint was echoed back as variant_progress, so a Tune slider drag remounted the preview under the user's cursor and latched the *_reviewable phases from the wrong trigger. Gate on the reason. Cleanup: - Collapse four divergent benchmark argv parsers into scripts/lib/cli-args.mjs. Three silently misread flags: --iterations 20 benchmarked 5, --agent llm ran the fake agent, --median-target=0.4 used the default threshold. - Drop a snapshot cache this branch made write-only (it grew per session for the server's lifetime and was never read), a dead exported reconcile helper, and the unused deferReply branch. Prepared with AI assistance under maintainer direction. Co-Authored-By: Claude <noreply@anthropic.com> * Route the last two benchmark scripts through the shared argv parser Follow-up on review feedback. The previous commit consolidated four of the six Live benchmark parsers and left these two on their own hand-rolled `arg()`, which was the inconsistency the first pass was meant to remove. - benchmark-live-control.mjs and benchmark-live-init.mjs parsed --iterations with Number(), so a non-numeric value became NaN and `index < NaN` ran the benchmark zero times before failing on the metrics file. They also accepted only the space-separated form, so --iterations=20 silently measured the default. Both now use parseArgs + positiveIntFlag, which throws on a value that was clearly meant as a number. - benchmark-live-control.mjs read the metrics file with no handling for the case where the run produced nothing: a missing file surfaced as a raw ENOENT stack and a malformed line as a bare SyntaxError. Report both with a diagnostic naming the file and the env var that populates it. - summarize() now reports a `samples` count and nulls instead of letting percentile() read past an empty array, where the NaN serialized to null and a report of nothing measured looked like a real measurement. Prepared with AI assistance under maintainer direction. Co-Authored-By: Claude <noreply@anthropic.com> * Stop telling users a busy agent is disconnected The agent-poll indicator tracks whether a poll is parked, which is the right signal for "can steering reach the agent right now" and is why the flag itself is left alone. But it goes quiet for two different reasons, and both got the same copy: "Agent disconnected - run live-poll.mjs to connect". Under the one-shot foreground polling that live.md calls the primary contract, no poll is parked while the agent works, so the second reason is every normal generation. For its whole duration the bar told the user a healthy session was broken and advised them to start a poll loop that was already running. Pick the copy from the live state, which the browser already tracks: GENERATING and SAVING mean the agent holds work it was handed, so say it is working. Every other state with no parked poll keeps the original, actionable wording. The aria-label carries the same distinction, since the tooltip is mouse-only. The text is derived at read time rather than cached, because the live state moves between the 5s status polls and a finished generation would otherwise keep reading "Agent is working" until the next one landed. Deriving it also keeps the read out of setLiveState, which runs long before agentPollingConnected's declaration and would hit its temporal dead zone. Prepared with AI assistance under maintainer direction. Co-Authored-By: Claude <noreply@anthropic.com> * Scope design-system-font-size off the injected live overlay live-browser.js builds a self-contained UI that renders over arbitrary host pages, so its inline type scale is deliberately independent of DESIGN.md, which documents the impeccable website's ramp. The rule fired 32 times there and is the only rule that fires on that file. Suppress it as a file-scoped value wildcard rather than via ignoreFiles: an ignoreFiles glob would silence every rule for the file, and the overlay is real user-facing chrome where a future contrast or side-tab finding should still be heard. Scoped to this one file, so the rule keeps working everywhere else. Written by hand because hook-admin's ignore-value cannot emit the `files` array that detector.ignoreValues supports and existing entries already use. Prepared with AI assistance under maintainer direction. Co-Authored-By: Claude <noreply@anthropic.com> * Let hooks ignore-value scope a rule to files, and stop churning the config Fallout from suppressing the overlay's font-size findings: the narrowest exception detector.ignoreValues supports was unreachable from the path the hook tells the model to use, so the guidance steered to the blunt instrument instead. - hook-admin's ignore-value now takes --file / --files / --file= / --files=, matching `impeccable ignores add-value`, which already had them. Without it the only file-scoped option was ignore-file, which silences every rule for a path permanently, including rules not yet written. - A bare wildcard value is now refused with a message pointing at either --file or ignore-rule. Previously `ignore-value <rule> "*"` quietly wrote a project-wide suppression from a single file's finding. - ignore-value keyed entries on rule+value only, so a second scope for the same rule overwrote the first instead of coexisting. Key on the file scope too. - An unknown flag folded into the value: `ignore-value overused-font Inter --shard` stored "inter --shard", matched nothing, and reported success. Reject it, as the sibling command does. Config churn: normalizeIgnoreValueEntries runs on every write and emitted keys as rule, value, files, reason, createdAt while the config on disk uses createdAt before reason. Any edit therefore rewrote every untouched entry (35 churned lines for a one-line change). Pin the canonical order in both copies of the normalizer and in ignores.mjs, and add a test that the two copies cannot drift apart. Also point the hook's own footer and reference/hooks.md at the file-scoped form first, and say plainly what ignore-file costs. Prepared with AI assistance under maintainer direction. Co-Authored-By: Claude <noreply@anthropic.com> * Correct the prose-gate docs and write down the no-bump-in-a-PR rule CLAUDE.md said the prose validator "deliberately skips skill/", which is only half true and cost a build failure this week: validateProse skips it, but validateSkillProse then scans skill/**/*.md and fails the build on em dashes plus the phrases with no technical reading. Document both gates, which files each one reads, and the line that actually matters in practice: an em dash in skill/reference/*.md fails the build, one in a skill/scripts/*.mjs comment does not. Each claim was checked against a real `bun run build`. Also record that feature PRs do not bump manifest versions or add changelog entries. It was not written down anywhere: not CLAUDE.md, not AGENTS.md, not the PR template. CLAUDE.md's "Bump when: CLI code changes" reads as an instruction to bump inside the PR that touches cli/, so say plainly that it names which component a change belongs to rather than when to edit the manifest. Put the rule in AGENTS.md too. That is the guide the agents opening PRs here actually read, so a rule about PR hygiene living only in CLAUDE.md would not reach them. Prepared with AI assistance under maintainer direction. Co-Authored-By: Claude <noreply@anthropic.com> * Bring Live progressive delivery and the generator subagent to Claude Code Almost none of this branch's Live work was actually Codex-specific. The publisher, the fences, the source locks and the browser's partial-arrival UI are plain node and DOM with zero provider references, and the progressive E2E already passes on five frameworks driven by a non-Codex agent. The Codex-only part was policy prose and one frontmatter line, so Claude Code shipped the progressive browser UI it could never trigger. Progressive delivery, Codex and Claude Code: - Add a `live-progressive` capability tag and opt codex, agents, and claude-code in. A provider block takes one tag, so naming harnesses would have meant duplicating the recipe per tag; a capability reads better than a provider list anyway. Cursor and everyone else keep the atomic path until their poll loop is known not to stall on the extra publish calls. - Claude Code publishes variant 1 as soon as it validates rather than waiting to write the whole trio in one edit. Nothing about the arrival path needed changing: the publisher writes, framework HMR pushes, and the browser's MutationObserver counts variants. The parent conversation was never in that path, which is why Claude Code's lack of subagent progress streaming does not matter here. Generator subagent: - Drop `providers: codex` from impeccable-live-generator. The build already maps its frontmatter correctly for Claude Code, and impeccable-manual-edit-applier has shipped to .claude/agents/ this way all along. - The reason differs per harness, so the reference says so: Codex delegates to unblock a foreground poll, Claude Code delegates to keep a long session's screenshots and variant CSS out of the main context. Follows the existing manual-edit-applier convention: both agent names, and an inline fallback when native subagents are unavailable. Fixes found on the way: - The two publish commands hardcoded `.agents/skills/impeccable/scripts/` while the other thirteen commands in live.md use {{scripts_path}}. Correct only for the Codex repo-skills bundle; it would have pointed Claude Code at a directory its install never creates. The shipped .codex variant was already internally inconsistent. Now covered by a test. - `--agent=codex` resolved to the canned fake agent, because the flag parsed as `x === 'llm' ? 'llm' : 'fake'`. The private evals Live runner passes exactly that, so a real-harness run would have scored deterministic stub variants and reported them as Codex output. Unknown values for --agent, --scenario and --delivery now fail loudly. - live-reference tests now compile with each provider's real providerTags instead of hand-written lists, so a providers.js misconfiguration fails in tests rather than shipping. Verified: progressive E2E green on vite8-react-plain against a real Vite server and Chromium; every provider variant's publish and poll paths now agree; Cursor and Gemini still compile to atomic only. Prepared with AI assistance under maintainer direction. Co-Authored-By: Claude <noreply@anthropic.com> * Fix inset-order detection, the unlocked artifact discard, and stray boolean flags Three of the four open review findings. The fourth is declined below. - The inset-stripe scan only matched layers starting with `inset`, but the keyword is order-independent: `box-shadow: 4px 0 0 var(--brand-accent) inset` paints the same stripe and was silently missed. Strip the keyword wherever it sits, but only as a standalone token, so a color like var(--inset-accent) is not mangled into `var(-- -accent)` and quietly reclassified as neutral. The fixture now covers both orders plus that token, and a trailing-inset neutral still passes. - The source-artifact discard deleted the preview without the source lock, unlike every other discard path. Take the lock. Narrower than reported, though: the server journals `discard_requested` as a fenced phase before live-accept runs and the publisher checks it three times, so a publish could never land on a discarded session. What this actually prevents is deleting the artifact under a publisher mid-critical-section, turning a clean stale_generation_epoch into an ENOENT crash. - benchmark-live-providers.mjs still compared `--headed` and `--skip-cleanup-control` against a boolean sentinel, so the `=true` spelling silently did nothing. My gap: I introduced boolFlag and converted benchmark-live.mjs but not this one. skipCleanupControl is now read once rather than twice, so the two call sites cannot drift. Declined: tightening the selector guard that skips `active` / `current` / `selected` tokens. It does cause false negatives on names like `.selected-feature`, but the rule's contract makes selection and focus indicators its one exception, and `.active-tab` / `.current-step` / `.selected-row` are syntactically identical to `.selected-feature`. No regex separates them, so tightening the guard trades missed stripes for false positives on exactly the case the rule exempts. The conservative skip is the intended behavior. Prepared with AI assistance under maintainer direction. Co-Authored-By: Claude <noreply@anthropic.com> * Classify failed accepts as errors, and fix parallel lane race/all misuse Two of the three new findings, plus the bug that chasing them exposed in my own earlier fix. The third is mitigated rather than broken; details below. Failed accepts reported success: live/completion.mjs only classifies a result as `error` when it carries `mode: 'error'`. Everything else unhandled falls through to `agent_done` with an ok ack, which is deliberate for the documented fallback paths (two tests pin it) but wrong for a real failure. So `accept_receipt_conflict` reported success, and reference/live.md's `handled: false` without `mode` bullet told the agent to "read file, find markers, edit" — hand-applying a second accept on top of the one the receipt already recorded. The same hole swallowed `source_locked`, which is mine: the earlier commit made lock contention return clean JSON so the agent could retry, but the classifier turned that failure into agent_done/ok, so the accept was dequeued and silently lost. Mark genuine failures with `mode: 'error'` through one `operationFailure` helper, and give live.md a `mode: "error"` bullet with per-error guidance: retry the same command on `source_locked`, never hand-edit, and on a receipt conflict report what the session actually resolved to. The deliberate fallback and markers-not-found handoffs stay untouched. parallel-compact lane orchestration: `Promise.race` settles on the first *settlement*, so one lane failing fast rejected the whole first-variant step while two lanes were still on their way to succeeding. `Promise.any` now takes the first success and only a total wipeout is fatal, reporting every lane's reason. The tail step's `Promise.all` surfaced a raw lane error non-deterministically; `Promise.allSettled` now reports how many lanes failed and why. Added a `requestImpl` seam so lane orchestration is testable without a provider key. Not a defect: the browser releasing Accept before the source write. That is the intended optimistic design, and it is safe because poll-lanes ranks accept at priority 0 against generate at 2, so a queued accept is always leased before a generate the user queues afterwards, even if the generate arrived first. Its source write lands inside the poll script before the next generate preflights. That invariant is load-bearing and had no tests at all; poll-lanes.mjs now has a suite covering it plus lease and type filtering. Prepared with AI assistance under maintainer direction. Co-Authored-By: Claude <noreply@anthropic.com> * Finish the failed-accept classification my last commit only half did All three new findings are the same root cause, and it is my incomplete fix: operationFailure only covered results built from a *thrown* error. Two paths it missed: - Two catches wrote the failure result as a multi-line literal, so the single-line replace skipped them. The Vue accept catch was still bare, exactly as reported; the Svelte one too, though its failures happened to be caught by completion.mjs's Svelte-only special case. - The accept implementations also *return* `{handled: false, error}` for their own checks (variant missing, template empty, original text ambiguous). Those never throw, so no catch ran and no `mode` was set. Both layers now agree, because each is reachable on its own: - live-accept marks any unhandled preview-path result via markPreviewFailure, keyed on `previewMode` — a clean discriminator, since only the preview branches set it and a plain wrapper never does. This is what the agent reads: reference/live.md routes on `mode`, so without it the agent was told "read file, find markers, edit" for a preview that has no markers in source. - completion.mjs replaces its arbitrary svelte-component special case with the set of preview modes whose variants live outside the user's source. That case existed for precisely this reason; Vue and source-artifact were simply never added, so the identical failure on those paths acknowledged as success. The plain wrapper keeps its manual handoff, which is the one shape with editable markers in source. Both deliberate handoffs (mode: 'fallback' and markers not found) still classify as agent_done, now pinned by a test so the generalization cannot swallow them. Prepared with AI assistance under maintainer direction. Co-Authored-By: Claude <noreply@anthropic.com> * Stop the progressive benchmark agent inventing a second variant on count:1 `Math.max(1, event.count - 1)` floored the tail request at one variant, so a one-variant request fetched a second direction and assembled two. Ask for `count - 1` and return the first variant untouched when there is no tail. Latent rather than live: the only caller hardcodes `count: 3`. The reason it is worth fixing is the caller inconsistency it exposed. tests/live-e2e/agent.mjs gates its split-progressive path on `event.count > 1`; benchmark-live-providers.mjs had no such guard, so it would have run the tail for a one-variant request, and the parallel strategy would have assembled its three fixed lanes regardless of what was asked for. Guard the caller the same way. Prepared with AI assistance under maintainer direction. Co-Authored-By: Claude <noreply@anthropic.com> * Drop the live generator subagent; fix the artifact decoy that broke accept The first real Claude Code Live run failed, and the subagent was not the cause. Root cause: progressive publication stages each revision as `.impeccable/live/artifacts/<id>-r<n>.<source-ext>`, nothing ever deleted them, and findSessionFile's walker skipped only node_modules/.git/dist/build. It searches src, app, pages, ... then `.`; a project whose source is not under one of those (this repo's own site lives in site/pages/) falls through to the `.` walk, where dot-directories sort before letters. So accept found the artifact instead of the real file. Two outcomes, both reproduced: where isGeneratedFile returns true it declines with mode: 'fallback' (what the run hit, after which the agent hand-carbonized several hundred lines across three stylesheets, including unrequested drive-by edits); where it returns false, accept writes the variant into the throwaway artifact and reports handled: true while real source never changes. The E2E suite could not have caught this. Every fixture puts source under `src/`, which is searched before the `.` walk can reach `.impeccable`. Five framework fixtures and three progressive scenarios pass because of fixture layout, not because the path works. I read that as evidence and shouldn't have. - Never search `.impeccable`: it is Impeccable's own state, never project source. - Retire a session's staged artifacts on accept/discard, so they cannot outlive the session and become a decoy for anything else that walks the tree. - Regression tests use a site/pages layout with artifacts present. All three fail against the previous code. Generator subagent removed, on both harnesses: The parent must hand-compress the design system into the handoff, and compression is lossy. Measured on the real run: a 6,826-char handoff carrying exactly one token reference, after the parent had itself read kinpaku-tokens.css. The subagent then spent 3 of its first 9 turns hunting DESIGN.md, gave up, and emitted 0 var(--token) uses and 22 raw oklch literals — violating its own spec's "Never invent raw colors when tokens exist" — including a 1:1 gold-on-gold contrast bug. Isolation is not a benefit here; knowing the design system is the job. Generation stays in the main thread, which already holds the tokens and writes them from the first byte, so carbonize is a move rather than a translation. Copy edits keep their subagent: applying a known set of ops to a named file is self-contained, so an isolated context costs nothing. That is the line. Progressive delivery stays for Codex and Claude Code, main-thread driven. Claude Code keeps the full benefit because its poll is a background task. Codex's poll blocks the foreground, so with no subagent the user sees variant 1 early via HMR but cannot accept it until the trio finishes; that is the cost of the simplification and it is worth naming. Prepared with AI assistance under maintainer direction. Co-Authored-By: Claude <noreply@anthropic.com> * Rip out the dead isolated-preview mode and the private repo's job Comparing this branch's live against main's turned up two whole features that never made sense here. -2,466 lines. 1. The isolated source-artifact preview was never switched on. `scaffoldSourceArtifactSession` is only reachable via live-wrap's `--isolated`, and nothing passes it: not the server's preflight, not live.md, nothing. Proved it end-to-end — the default wrap writes markers straight into real source and creates no previews/ session. So the mode was wired through three modules, carried its own accept/discard branches, browser branches, server metadata resolution, preview-mode classifier entry, and test suites, and none of it could run. Worse, live.md documented it as the active path and told the agent "The true source is only the publisher's hash fence and must remain byte-identical until Accept." That is false: the wrapper lands in source at scaffold time and each revision rewrites it. An agent following that sentence believes source is protected when it isn't, and the leftover artifacts are what made accept resolve the wrong file in the first real run. live.md now describes what actually happens, including that markers are visible in source until Accept or Discard. Removed: source-artifact.mjs, --isolated, the preflight's isolated option, the accept/discard branches, four dead browser branches, the server's previews/ resolution, the classifier entry, and their tests. Kept the previews/ gitignore pattern: an ignore line for a directory that cannot exist is free, and a test pins it. 2. Quality judging belongs to the private evals repo, which says so. runner/live/README.md there is explicit: the public repo owns protocol correctness, framework coverage, timing, source commit, recovery, and a rubric-free evidence bundle; the private repo owns the task corpus, baselines, comparative judges, and release-quality decisions — "Do not add quality rubrics, competitor comparisons, or broad fixture corpora to the public Live benchmark." This branch added exactly those: an LLM judge scoring 1-10 on "off-brand, generic-AI" (live-rendered-quality.mjs, judge-live-rendered.mjs), a cross-provider comparison with a BRAND_CONTRACT rubric (live-provider-benchmark .mjs, benchmark-live-providers.mjs), and a brand-fidelity fixture corpus. All removed, with bench:live:providers and their suite entries. Also removed tests/framework-fixtures/README.md's "External quality-eval fixtures" section: it documented a bench:live workflow using --fixture-dir, --agent=codex, --action and --evidence-bundle, none of which benchmark-live.mjs implements, plus an evidenceCapture block nothing reads. Kept: timing benchmarks (the public repo's half of that boundary), progressive publication, the source lock, poll lanes, and Nuxt/Vue component previews. Coverage note: deleting the isolated suites took the only tests for `source_locked` classification with them, so the plain wrapper path — now the only non-component preview — gets equivalent accept and discard coverage. Both new tests fail if mode:'error' is removed. Prepared with AI assistance under maintainer direction. Co-Authored-By: Claude <noreply@anthropic.com> * Flag inset stripes written with the two-length box-shadow form box-shadow takes <length>{2,4}: only the two offsets are required, so `inset 4px 0 red` is valid and paints the same single-edge stripe as `inset 4px 0 0 red`. The scan demanded a third length, so the short form was silently missed. Blur and spread now default to 0 when omitted, which is exactly the stripe shape the rule looks for. The neutral-color and blur/spread exclusions still hold: `inset 4px 0 #000` and `inset 4px 0 5px var(--brand-accent)` both pass. Fixture covers both orders of the short form plus those two exclusions, and fails against the previous regex. Third false negative found in this rule (after trailing `inset` and literal neutral colors), all from the same cause: the scan was written against one spelling of the syntax rather than the grammar. Prepared with AI assistance under maintainer direction. Co-Authored-By: Claude <noreply@anthropic.com> * Live: polling rework, source locks, preflight scaffolding, Vue previews Carved out of #371, minus progressive publication. Everything here works against real project source the way main's Live already does: the agent writes variants into the file the browser loaded, HMR fires, Accept promotes and carbonizes. Nothing is staged anywhere. Poll lanes. Events now carry an explicit priority: accept/discard/exit ahead of manual_edit_apply/steer/carbonize_cleanup ahead of generate. A long generate can no longer sit in front of the Accept the user just clicked. leaseEvent claims its lease before awaiting, so a slow prepare cannot hand the same event to two pollers. Source locks. A per-file mutex around every accept and discard path, keyed on a digest of the absolute path. Staleness is decided by owner-pid liveness rather than mtime, so a wedged lock clears when its owner dies instead of after an arbitrary timeout, and a slow-but-live accept is never stolen from. Only the owning process can release a lock. Preflight scaffolding. The server runs live-wrap (or live-insert) before the poll returns and hands the result back as event.scaffold. That walk is measured at ~7.6s on a large repo; moving it off the agent's critical path removes a deterministic tool round trip without touching the generated design. Falls back cleanly to the agent running the helper itself. Vue previews. previewMode: "vue-component" for Nuxt/Vue targets, matching the existing Svelte component path: variants compile as real SFCs from a dev-only directory so the route is never rewritten during generation, and Vite mounts them without invalidating page state. Accept is the only route write. Includes a Vue attr tokenizer that normalizes shorthand bindings (@x, :x, #x) to their canonical forms. Accept hardening. Every thrown failure now returns mode: 'error' rather than an ambiguous unhandled result, so a real failure is never classified as a deliberate manual handoff and silently dropped. The marker search skips node_modules/.git/dist/build/.impeccable. Shared CLI arg parsing extracted to scripts/lib/cli-args.mjs. Assisted-by: Claude Code * Drop the progressive benchmark, remove dead wrap scaffolding Review fallout from removing progressive publication. The Live benchmark existed to compare atomic against progressive delivery: compareModelBackedReports measures goToFirstVariantMs improvement of one over the other. With progressive gone it measures nothing against nothing. Worse, benchmark-live.mjs still passed `progressive` to bootFixtureSession, which no longer accepts it, so `--delivery progressive` was silently ignored and would have emitted reports labeled progressive that actually ran atomic. Silent wrong data is worse than a crash. It was built for progressive, so it goes with progressive: benchmark-live.mjs, its lib, its test, and the bench:live script. If an atomic latency baseline is wanted later, that is a smaller thing built on purpose. live-wrap.mjs: sourceOriginalLines was assigned and never read. Both found by review bots on #381 (Copilot). Assisted-by: Claude Code * Drop the Vue preview mode; it never reached Svelte's accept path Cursor found that inlineVueComponentAccept never receives paramValues, while the Svelte equivalent uses them in 23 places: Accept on a tuned Vue variant silently persisted the default and threw the user's tuning away. Chasing that corrected something I had asserted the other way round. I said Vue's raw-CSS-append was inherited from the Svelte path. It is not. svelte-component.mjs calls sanitizeAcceptedSvelteCss before writing, which sanitizes the CSS and bakes tuned params into it. vue-component.mjs had no sanitize step at all — it appended the variant's <style scoped> body into whatever style block came last, so a variant could leak CSS site-wide when the last block was global, and brace CSS landed in a lang="sass" block. Both are the same defect: the Vue mode mirrored Svelte's preview path without its accept-side subsystem (bakeParamValuesInCss, sanitizeAcceptedSvelteCss, appendSanitizedCssRule, rewriteAcceptedSvelteSelector, rewriteParamSelectors — roughly 200 lines of CSS rewriting). Both were introduced here, not inherited. A shipped Vue session could leak styles and discard tuning without saying so. So it comes out. The poll lanes, source locks, preflight scaffolding, and accept hardening do not depend on it and are worth landing now. Vue returns when its accept path reaches parity. The nuxt-vite7 fixture goes back to main's plain-wrapper shape. Assisted-by: Claude Code * Stop the lease redelivery test racing the scheduler CI failed `does not drop polled events until the agent acknowledges them` on a commit whose content was byte-identical to one that passed, which is the signature of a flake rather than a regression. The test leased an event for 50ms, then asserted a second poll saw a timeout because the lease was still held. That gave the whole second HTTP round trip a 50ms real-time budget: cross it and the lease expires, the event is redelivered, and the assertion fails for a scheduling hiccup instead of a bookkeeping bug. Locally it passed 6/6; a loaded runner is where it bites. Hold the lease for 1000ms so a round trip cannot cross it, and wait LEASE_MS + 300 before asserting redelivery, so each half has headroom in the direction it asserts. Verified by injecting a 60ms stall before the second poll: the old test fails with exactly the CI message, the new one passes. Assisted-by: Claude Code * Recover live sessions that reload past the generation done broadcast The preflight scaffold write (new in this PR) triggers a framework full-reload — Astro reloads the page for any .astro edit. When the agent's variant write and its done SSE land while the browser is mid-reload, the resumed page misses both the second HMR reload and the done broadcast: it comes back up on the scaffold-only source and waits in GENERATING at 0/N forever, with the finished variants sitting in source. This is the astro-vite7 CI timeout; the failure artifacts show the full sequence (scaffold at 26.319s, done at 26.515s, the new page's browser_resumed checkpoint at 26.653s, DOM still scaffold-only). Three-part fix: - session-store: agent_done now stamps a monotone generationCompletedAt on the snapshot. Browser checkpoints legitimately regress phase and arrivedVariants (a resumed page reports what it sees), so completion needed a field checkpoints cannot un-set. - live-browser: on every SSE (re)connect, compare the session summary's generationCompletedAt against local progress; when behind while GENERATING, pull the finished variants from source (same settle delay as the done handler's HMR-first fallback). Covers both orderings of resumed-checkpoint vs agent_done. Also, the source-fallback empty- wrapper branch no longer tears the session down mid-generation — a scaffold-only wrapper is a legitimate in-flight state, so stay in GENERATING instead of destroying a session the agent is still filling. - live-server: a browser checkpoint reporting generating/behind for a session whose generation already completed re-broadcasts the stored done (idempotent for every other tab), and connected-payload summaries expose generationCompletedAt for the browser-side check. Coverage: live-server unit tests for redelivery, the no-redelivery guard, and marker durability across checkpoint regression; plus a deterministic live-e2e scenario on astro-vite7 that blocks the reloaded page's SSE stream and mocks its HMR websocket dead until after the agent finishes, forcing the missed-broadcast window every run. All new tests fail against the pre-fix code. The e2e harness additionally gains an IMPECCABLE_E2E_ATOMIC_DELAY_MS lever (widens the scaffold-to-write window) and env-gated console/nav tracing (IMPECCABLE_E2E_CONSOLE=1) used to diagnose this. The hypothesis that preflight opens a wrapper-with-no-variants window came from Copilot's review sketch in the follow-up WIP PR; the killing mechanism differs from that sketch (nothing calls recoverEmptyCycling in the CI trace — the session hangs precisely because no code path runs at all), but the window is real and the guard it suggested is folded into the source-fallback fix. Assisted-by: Claude Code Co-Authored-By: Claude Code <noreply@anthropic.com> * Retry a completion-driven source fallback that reads only the scaffold Greptile flagged a hole in the previous commit's empty-scaffold guard: when a `done` has already been delivered, the source fallback gets exactly one read. If that read returns the preflight-only scaffold (a stale source view, or an agent whose write lands in multiple steps), the guard's silent return left the tab in GENERATING with no further event ever coming — the same stuck state the previous commit fixed, reintroduced through a different door. Callers that know generation finished (the done handler's fallback and the SSE-reconnect self-heal) now pass generationCompleted, and an empty read on that path re-reads the source up to 3 times before surfacing recoverEmptyCycling instead of hanging. Mid-generation callers are unchanged and still wait indefinitely — a real agent can legitimately take minutes between scaffold and write, and tearing that down was the original #385 hazard. The missed-done e2e scenario now also serves a captured scaffold-only copy for the first post-reconnect /source read, forcing the retry path every run. Verified failing against the pre-retry code (tab stranded in GENERATING, test timeout) and passing with it. Assisted-by: Claude Code Co-Authored-By: Claude Code <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com> |