A codex run turned an approved comp into a related second art direction
and the finish reviewer passed it: the review anchored on the direction
contract, a lossy abstraction the builder wrote, and every element that
abstraction dropped passed silently. Four changes close that chain. The
reviewer inventories the comp's salient elements before reading the
contract and classifies each one (match, adaptation, missing,
contradicted, added without approval), with adaptations citing the
answer, brief, accessibility need, or product truth that forced them,
and fidelity failures outranking craft in material_fixes. The visualize
inventory gate records compositional commitments alongside asset media,
since the 150-word contract cannot carry them. The north-star allowance
now says what it permits: translation, never recomposition. And the
finish sequence recaptures the same viewports once after the fix batch,
so what the documenter records is what actually shipped.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A ground-up hardening of live mode, driven by a production session in a
nested-app monorepo that hit six distinct failure classes. Full design
rationale in docs/LIVE-REWRITE-PLAN.md; every Codex-reported failure now
has a mechanical fix and a regression test.
Roots: live/roots.mjs resolves appRoot/repoRoot/contextRoot once at boot
(keyed on dev-server configs, not monorepo brand markers), persists a
manifest, and every live CLI re-anchors onto it at startup, so a helper
run from the wrong directory can no longer fork session state. Context
files are discovered upward to the git root.
Render truth: variant_mounted / variant_mount_failed events give the
journal per-variant mount state; failures reach the agent's poll queue,
raise a persistent error card with Retry (no more localStorage wipe), and
an attach probe names root/dev-server mismatches explicitly. The browser
rehydrates from the server when localStorage is gone.
Svelte: the scaffolder now parses with the app's own svelte 5 compiler.
Control flow survives (an each collection crosses the contract as one
structured prop), keyed each blocks hydrate synthetic keys, and anything
a detached preview cannot support falls back to source-preview instead of
shipping a wrong scaffold. Preview modules live in per-publish revision
directories, defeating stale transform caches.
Accept: CSS is reconciled, not appended. Matching selectors are replaced,
params bake from params.json kinds, the compiler's unused-selector pass
prunes superseded rules (pre-existing dead rules protected), a selector-
loss postcondition refuses any write that would drop hand-written rules,
and live-complete refuses to finish while live plumbing remains in source.
Also: framework registry (live/frameworks/) with a crash-safe injection
journal, session-store snapshot caching with read-only reads, protocol
enum consolidation, steer Send button, honest DESIGN-panel empty states.
Testing: new unit suites (roots, AST scaffolder, accept CSS, accept
pipeline, framework conformance); e2e now fails on preview-tree 404s,
proves computed-style mount for every variant, drives the Tune panel
through baked params, and injects failures (broken mounts, republish,
storage loss). New runtime fixtures: monorepo-nested-vite (repo root !=
app root) and vite8-sveltekit-stateful (each blocks + state). Nightly
full-matrix cron. An independent adversarial review pass preceded this
commit; its blocker and major findings are fixed and regression-tested.
This work was produced with AI assistance (Claude Code).
Co-Authored-By: Claude Code <noreply@anthropic.com>
Harnesses with no subagent capability now run each role inline from the
same single source. The build emits reference/degraded/<role>.md for every
agent in skill/agents/ (role name is the agent name minus the impeccable-
prefix), stripping frontmatter and prepending the inline-substitution
preamble. These pass through the same provider-block compilation and
placeholder replacement as ordinary reference files, so <codex> blocks and
{{placeholders}} resolve per target, and they land in the committed harness
dirs on build:release like every reference file.
Repoint the three capability-first fallback sites in the prose at the
generated files: new-work.md reviewer and documenter fallbacks, and
visualize.md asset-producer fallback.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Probe attribution on Opus 5 showed the screenshot bound working (42
to 16) while the real burner ran free: five rounds of node -e
micro-edits, eight rebuilds, and inline defect hunts absorbed the
reviewer's and documenter's jobs until the turn cap killed the run
mid-hunt. The two-round ceiling now names scans, micro-edits, and
rebuilds; after the second round the build thread stops polishing and
ships the rest through the reviewer (one batched fix pass, one
rebuild, stop) and the documenter.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Copilot's review point stands: reference files load per-command, so a
bare "see optimize.md" / "typeset.md" is not meaningful in the routed
context. The guidance reads self-contained now.
Prepared with AI assistance (Claude Code), directed by @pbakaus.
Co-Authored-By: Claude Code <noreply@anthropic.com>
From issue #395, the items still present after the v4 consolidation:
- optimize.md led its interactivity section with FID, retired as a Core
Web Vital in March 2024 when INP replaced it. The section heading and
both metric lists now name INP.
- optimize.md recommended react-virtualized, superseded by react-window
from the same author; the line now points at react-window and TanStack
Virtual, matching overdrive.md.
- overdrive.md's WebGPU support matrix predated Firefox 141/147 shipping
it on Windows/macOS and Safari 26 shipping it across Apple platforms.
- audit.md listed "missing will-change" as a defect while animate.md and
optimize.md both instruct applying it sparingly and never preemptively;
the audit line now flags overuse instead of absence.
- harden.md allowed 14px mobile body text while typeset.md sets a 16px
ordinary floor; harden now matches the floor, reserving 14px for
secondary text, and names the iOS Safari input-zoom consequence.
The issue's other items (Framer Motion naming, Popmotion, polish
duration cap, humor guidance, HSL phrasing in quieter) were already
resolved by the v4 reference rewrite.
Prepared with AI assistance (Claude Code), directed by @pbakaus.
Co-Authored-By: Claude Code <noreply@anthropic.com>
Opus 5 turned the iterate-with-screenshots-until-it-meets-the-bar
instruction into 42 screenshot trips and 150 tool calls per build,
about forty dollars of cache churn a page, before ever reaching the
reviewer. Verification now batches: one desktop-and-mobile round after
the full build, fixes applied together, one confirming round, ceiling
two. Craft-floor's checks share those renders instead of earning
separate trips; per-tweak iteration is live mode's channel.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
From the paired Opus and Codex manual-run analyses. DESIGN.md moves to
the end of the flow and into a shipped documenter subagent that derives
the system from the built artifact: a rulebook written before the build
gets defended against reality, and a half-stable DESIGN.md hands the
design-system detector an unstable target that buries the build in
noise and invites laundering. The finish reviewer gains the handoff
that failed three times live: the parent captures desktop and mobile
screenshots and passes paths, the reviewer never attempts to render
and names missing inputs in one line, the parent verifies the
five-section return and respawns once on empty. Fidelity against the
approved comp joins its checks; the card keeps commitment only. The
comp ingredient inventory becomes a written gate with raster-by-default
materials and no gradient-as-texture, comps persist under
.impeccable/mocks, the degraded seed names the sandboxed-exec cause,
and the finish line is explicit: a clean detector pass is not finished.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The release-gate audit traced four ways the roll's output was defeated
downstream of a perfectly healthy seed. Gemini's harness keeps only the
tail of tool output, so the header-only ASSIGNED INDEX never reached
the model in 18 of 18 samples; the seed now restates the assignment
and key at the end of its output. Astro strips frontmatter comments,
so half the anthropic contracts vanished from built artifacts; the
contract now must survive the production build as an HTML comment in
emitted markup. A brief that paints its own picture (the album named
Soft Cathedrals) converged every arm regardless of assigned index; its
literal reading now joins the rut with at most one candidate. And Opus
under 4.0.1 skipped the seed 42% of the time while hand-authoring
plausible contracts; the finish reviewer now verifies FORM carries a
corroborable seed key before any craft point.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul's call: the build-exhaustion failure only exists inside eval
workers with hard turn budgets no real harness exposes, and the clause
doubled as a hedge door for skipping the comp round. The eval-side fix
belongs in the worker's max-turns, not in skill prose.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The release-gate campaign confirmed two skill bugs with transcripts.
Sandboxed harnesses reject absolute paths, so following the CHOSEN
CARD directive with the absolute card-base path failed view_image; the
directive and the quality-bar clause now say download into the
workspace and open the relative path. And under the openai worker's
turn cap, models spent the budget on init, cards, and comp generation
and never built the page (a third of small-n supplement slices); the
visualize mandate gains its one exception: at a hard cap the shipped
page outranks optional imagery.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul's call: the reviewer-local authorization patch covered one
command while the harness gate silently disables every shipped
subagent, critique panels and the manual-edit applier included. The
argument now lives beside the autonomy counter in context.mjs, emitted
as tool-result content every run: invoking the skill is the user
request such gates ask for; spawn where a reference directs; the
in-thread substitute is for absent capability only and gets disclosed
in one line. new-work keeps the reviewer mechanics and drops the
now-central argument.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A live session on a harness whose guidance gates subagent use on user
request resolved the conflict silently against the skill: it never
spawned the reviewer, stretched the no-subagents fallback to cover
permission hesitancy, and self-reviewed with all the context that made
its choices feel correct. Three tightenings: invoking the skill IS the
user request that authorizes its shipped subagents; the fallback is
for harnesses lacking the capability, not for hesitancy; a substituted
review gets disclosed in one line at finish, never silently.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The polling-rework preflight wrote the variant scaffold into source during the
poll lease, before the agent acted. On source-preview targets (React/Vue/Vite,
everything but the svelte-component path) that write full-reloaded the
framework; a browser caught mid-reload missed the agent's variant write and the
SSE done, and sat stranded at 0/N.
Restore the 3.5 single-atomic-edit semantics: the preflight still resolves the
element location and computes the scaffold, but --defer-source-write leaves
source untouched and hands the agent the wrapper text plus the picked source
range. The agent splices variants into the wrapper and replaces the range in
one write, so the framework reloads exactly once. The svelte-component path is
untouched (it never writes route source). The missed-completion recovery stays
as defense in depth.
Also cache the resolved source file per target signature (locator + route):
the ~7.6s tree search re-ran on every generate for the same element; a hit now
points the helper straight at the file via --file, invalidated when the target
changes or a resolution fails.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Live mode had no TanStack coverage: a TanStack Start user hit disconnects
and static previews because there is no static index.html to inject and no
adapter for the SSR root document.
- New tanstack-adapter.mjs, modeled on the SvelteKit/Nuxt adapters: detects
a TanStack Start project (@tanstack/react-start + src/routes/__root.tsx)
and patches the __root document to mount a generated dev-only React
component (src/impeccable/ImpeccableLiveRoot) that appends the live bundle
on the client after hydration, carrying the ?token= param via
buildLiveScriptSrc. Patch/unpatch round-trips byte-for-byte and is
idempotent; refuses to clobber an unmanaged file at the component path.
- Wire detection into live-inject.mjs (insert + remove + gitignore),
ordered so SvelteKit/Nuxt win and a plain TanStack Router SPA falls
through to the baseline Vite index.html path.
- tanstack-router-vite fixture (baseline, no adapter) and tanstack-start
fixture (SSR adapter), both with runtime blocks. Both pass the full
live-e2e cycle (handshake, steer, pick, Go, cycle, accept, carbonize,
reloadProbe).
- Unit tests for detection + patch round-trip + apply/remove; tanstack-start
branches in framework-fixtures.test.mjs; live.md framework table + adapter note.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two prose fixes from the 3.5-to-4.0.1 forensic diff of a real user
regression (15-minute tweaks, repeated disconnects, agent abandoning
the picker). The craft-fold made every generate cycle pay the verify-
the-built-result loop the overlay already provides to the human; live
cycles now verify by construction and run the full check once at
accept. And nothing framed a dropped SSE or closed tab as resumable,
while the client toasts "Session ended", so agents rationalized
bailing to direct edits; the journal is canonical and reopening
continues the session.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
"When the harness can X, do Y" hands the model an exit before the
command arrives; the observed reviewer skip walked through exactly that
door. The three gated constructions now lead with the imperative,
present the decision visually, open the chosen card, spawn the finish
reviewer, and carry their fallbacks as trailing clauses for sessions
that genuinely lack the capability. Constructions that already led with
the command keep their routing clauses unchanged.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The eb686f36 session read the separate-reviewer rule and spawned
nothing: an unnamed "separate agent" is an improvisation prompt, not an
affordance. The skill now ships impeccable-finish-reviewer next to the
asset producer: persistence first, ceiling against the card and comp
second, contract promise by promise, truth; ordered material fixes
back to the parent, no editing, no second detector. new-work names it
so the finish step invokes a thing that exists.
The asset producer was gated providers: codex, so Claude Code never
shipped it; the gate is removed and its two codex-only workflow lines
made provider-neutral with codex blocks.
Dist rebuild still deferred for the running campaign.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
codex.md becomes visualize.md and loads for every harness with any
image generation, native or the API fallback: after the direction
locks, three distinct compositional comps are rendered and put before
the user for approval, in-harness when it can display images,
otherwise on the decision page. Three is the number; one comp invites
rubber-stamping, and this approval round has repeatedly produced the
most compositional and ambitious work, so new-work now marks it
never-skipped. The codex-only subagent stays as a codex note.
The recent rule additions are tightened by a third: the asset and
imagery bullets merge into one, the canon exit loses its restatements,
the DESIGN.md-rule and chosen-card and ceiling clauses each shed their
second clause saying the first clause again. Same laws, fewer words;
prose that grows without bound recreates the attention gravity it was
written to fight.
Dist rebuild still deferred; the release-gate campaign reads the
pinned dist.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
From Paul's approved UX and the eb686f36 session post-mortem:
The standing exit: direction rounds carry a quiet, permanent "Play it
straight" action (payload flag canon, reserved id) on the decision page
and as the last structured-tool option. It is the user's door, never
the model's: never recommended, never weighed against the roll, and
choosing it swaps the bar rather than lowering it, two or three named
reference products become the craft level, canon executed at full
commitment. Safer/conventional steers resolve here, never to a
stranger re-roll.
Session fixes, each mechanical where possible: the ANSWER line now
names the chosen card's hero and board and directs opening them before
code (the session built from text alone after viewing a different
world's card); generation scale joins the imagery rule (a library of
centered 128px subjects foreclosed the atmospheric hero); DESIGN.md
rules are checked against the world's native devices and never added
to silence a hook finding (the session banned arcade lettering's own
offset shadow and laundered 8px through the ramp); staging joins the
FORM contract block (the axis was dropped silently at world-choice);
the finishing reviewer audits the ceiling against the QUALITY BAR card
after persistence (floor rigor was disguising unreached ambition); the
icon-tile clause names hand-drawn icons as remedy, not target.
Dist rebuild deferred: the release-gate campaign reads the pinned dist.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The structured-tool channel collapsed to a single direction plus
re-roll, which read as "the system only ever offers one idea" next to
the multi-card decision page. Both channels now share one structure,
assigned direction leading, the one or two fused challengers that
survived the weighing as named alternates, re-roll with steer, and
differ only in richness. The anti-lineup rule stays precise: what never
appears is a ranked menu of the model's own grounded candidates; dealt
challengers carry no ranking rut.
Note: dist rebuild deliberately deferred; the release-gate campaign is
running against the pinned dist and rebuilding mid-run aborts it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
With the PRODUCT.md skip fixed, the Opus smoke unmasked the adjacent
gap: the model builds a new world and never writes DESIGN.md (zero
attempts), so the worker's requiresDesign assertion correctly fails the
run. Same disease, same treatment: DESIGN.md is now part of recording
the decision, written before the first build edit in the same stretch
as the direction contract, and the finishing reviewer checks
persistence first, before any craft point is scored.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The truth split already permits full-fidelity demonstration data, but
permission at selection time was not holding at build time: models that
would not author covers, names, or thumbnails compensated with chrome,
which is the content-starved look the detector hunts. Two build rules
make it a mandate: every blank the ask round left open is authored at
production fidelity (content is authorable, claims are labelable,
nothing is omittable; unanswered commercial claims ship as marked
placeholders with a replacement list), and when image generation is
available, generating the build's imagery is part of building rather
than a nicety.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A live retest showed the model dropping to the structured question tool
when the roll degraded: with no challengers and no cards it judged the
page pointless and presented one option in plain text. The degraded seed
output and the new-work rule now both state that degradation changes the
cards, not the channel; a browser session presents the assigned
direction as a single text-only card with re-roll on the decision page.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A traced Claude Code injection asserts for whole model families that the
user is not watching and cannot answer questions; it ships default-on
with no off switch, and it suppressed every interactive step of a live
run (interview skipped, PRODUCT.md inferred, decision page never
served). Prose in a reference file loses that argument, placement wins
it: context.mjs now emits AUTONOMY_DIRECTIVE_CHECK as tool-result
content in the working turn, telling the model such a claim is a
harness default, never session evidence, and to probe once with the
question tool before inferring. init.md makes the same test mechanical:
tool presence proves an answer mechanism, one real probe round is
required, inference afterward must be labeled and disclosed in the
first reply. The degraded concept-seed path now also tells the model to
disclose the degraded roll instead of presenting it as a full one.
Image-gen signaling stays positive-only per Paul: key present emits the
capability, absence stays silent so harness-native tools are not
suppressed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
v4 changed PRODUCT.md's shape and retired the register axis, so an
upgraded project can carry answers nothing reads. Nothing measured that.
Two tiers, and the split is a performance contract:
- Boot (context.mjs, emitting CONTEXT_STALE) spends only what a boot
already spends: markdown already in memory, a bounded set of stats,
the small JSON files the boot reads anyway. No new directory walks.
One directive for the whole set, throttled to once a week per project
so a finding the user declined does not reappear tomorrow.
- doctor.mjs runs the deep pass on demand: git drift, ignore lists
validated against the live rule registry, hook script paths that stop
resolving, and the monorepo workspace sweep. --fix applies only the
migrations that carry no decision.
Findings are data, not prose, so the boot directive, the text report and
--json all render one set. Severity says what should happen: auto (fix on
the next write anyway), mention (state once), route (name the command
that owns the repair).
PRODUCT.md now carries a schema stamp so the checks stop reconstructing a
file's vintage from which sections it happens to have. Schema version,
not release version: a record written by 4.0.0 is not stale under 4.0.1.
DESIGN.md gets no stamp, because it follows the external design.md spec
that Stitch lints and every DESIGN.md signal is measurable without one.
The highest-value catch is a project that resolves to web while carrying
native build files, including a monorepo app inheriting a root record
that says web. That one costs output quality silently; nothing failed
before.
doctor follows the hooks/pin pattern rather than the Commands table, so
it stays out of the design menu and the count stays at 23.
Also corrects CLAUDE.md, which still documented the register axis,
reference/brand.md, reference/product.md, eleven deleted domain reference
files, and an extractRegister() whose only occurrence in the repo was
that sentence.
Prepared with AI assistance (Claude Code).
Co-Authored-By: Claude <noreply@anthropic.com>
The env-var bypass (IMPECCABLE_QUESTION_DISABLED) relied on the harness
remembering to set it. The script now also self-detects CI, SSH-without-
display, and displayless Linux and exits 2 with the structured-question
advice; --no-open skips detection (caller opens the URL itself, as the
tests do) and IMPECCABLE_QUESTION_FORCE=1 overrides it. new-work.md now
frames the decision-page rule by capability: open a browser if you can,
structured question tool if you cannot, exit 2 means fallback not error.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Bugbot flagged the "resume without rerunning context.mjs" instruction
after init. It is right, and the gap is wider than the platform half it
named: context.mjs has two output branches, and the no-PRODUCT.md branch
omits DESIGN.md, the native platform references, and the unrecognized
`## Platform` warning. Because the skill never reruns the script once
init writes PRODUCT.md, whatever that first run withheld is gone for the
whole session. A greenfield iOS project would be designed without
reference/ios.md ever loading, and a project carrying DESIGN.md without
PRODUCT.md never saw its own design system.
The two halves need different fixes. DESIGN.md is authority in its own
right and does not depend on PRODUCT.md existing, so context.mjs now
emits it on both branches. Platform is unknowable before PRODUCT.md
exists, so no change to the script can recover it; init.md, the one step
that learns the answer, now loads ios.md / android.md / both right after
recording a native platform, and SKILL.src.md says so where it tells the
agent not to rerun.
Verified end to end against a temp project on both branches.
Co-Authored-By: Claude <noreply@anthropic.com>
Two true positives from the Bugbot review on PR #397.
document.md seed mode told the agent to run "Select one direction" for
paths A, D, or E. new-work.md has neither that heading nor the A/D/E
lettering since the workshop was restructured into named subsections, so
a literal read could skip the world-and-surface flow entirely. Point at
"Create or replace the visual world" and "Commit the world" instead.
critique.md let the heuristic table renormalize to an applicable maximum
when heuristics are scored n/a, but the report template hardcoded ??/40,
the rating bands only mapped raw numbers out of 40, and the persisted
meta carried total_score with no denominator. Trends could silently
compare 24/32 against 30/40 as if they were the same scale. The template
now prints the applicable max, the bands fall back to percentages for
partial sets, the snapshot records max_score and na_heuristics, and the
trend line states its denominator or breaks it out per run when they
disagree. critique-storage.mjs serializes frontmatter key-agnostically,
so the new keys need no code change.
Co-Authored-By: Claude <noreply@anthropic.com>
Gate2 measured it precisely: on matched assignments Opus renders warm,
bookish, and child-facing subjects as cream, serif-italic, and
lamplight while Sol renders the same positions saturated, and neutral
prose hardening did not move it. The codex and gemini blocks set the
precedent for provider-conditional counterweights; this adds the claude
block at the palette decision: the first palette is already spent, an
OWN-WORLD block reading cream/paper/parchment/lamplight for an unpinned
Persuade surface is a failed rendition to rework from the world's
saturated materials, and nothing about the subject requires the
default. Verified present in the claude-code dist and absent from
codex.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul's question exposed the blind spot: a closed tab left the agent
waiting out the full timeout. The page now sends a heartbeat every five
seconds while open; the server stamps lastBeat into the state file, and
--wait reports PAGE CLOSED with exit 4 when the beats stop for fifteen
seconds without an answer. The prose defines the fallback ladder:
re-present once through the structured question tool, then proceed
unattended with the assigned direction, stating assumptions.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Craft-gate forensics on matched hands: both models obeyed the same
assigned indices, but Opus rendered every kids cell as cream paper,
lamplight, and serif, and sample 1 chose Fraunces off the reflex list
with a bookshop-signage rationalization, while Sol rendered the same
positions as indigo bookcloth, coral thread, tomato, and marigold. The
dice work; the rendition prior escaped through two hatches, now closed:
naming a reflex face requires a reason no other face satisfies and a
subject association is never that reason; and bookish or child-facing
subjects do not soften the calibration, because cream paper is the
smallest corner of the book world.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul's call: start mode never auto-opens; the agent is alive and opens
the printed URL itself, in-app browser first, then the system opener,
then showing the URL (--open forces the system browser from the script).
The prose now leads with the start/open/wait flow and keeps the blocking
auto-open path for harnesses that can background a shell.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul's two concerns with the blocking design. --schema prints the exact
payload example so the model never guesses the shape (new-work.md points
at it). And harnesses that cannot leave a shell blocked (or cannot open
a browser while blocked) get a two-phase path: --start daemonizes the
server and returns the URL plus a key immediately, --wait polls for the
answer with exit 3 meaning ask again, exit 2 meaning the server died,
and --stop for cleanup. The browser open happens from the detached
server process, so it works even when the agent thread is short-lived.
State lives under .impeccable/questions/.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul's design, three pieces:
serve-question.mjs: the world decision presented as a themed page instead
of a text prompt. The script serves an impeccable-styled option board
(assigned direction leading with THE ROLL badge, dealt challengers as
alternates carrying their QUALITY BAR cards, re-roll and steer built in),
prints the URL, opens the browser, and blocks until the user chooses;
the answer lands on stdout as ANSWER JSON, so the shell call itself is
the wait and no harness machinery is needed. Local images are served by
the ephemeral server; nothing leaves the machine.
generate-image.mjs + context.mjs IMAGE_GEN_AVAILABLE: when an OpenAI key
is in the environment, context reports that image generation works even
without a harness-native tool (gpt-image-2, billed to the user's key,
stated before first use; Google skipped by decision). Harness-native
tools always win when present.
new-work.md: visualize-before-build is now the default whenever any
image generation exists, not a codex.md special case; the attended
presentation prefers the visual decision page and falls back to the
structured question tool. Evals keep the unattended path untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul confirmed the craft-bar experiment: builds that saw the dealt
worlds' hero cards produced visibly stronger execution than the
no-image control. One clause makes the mechanism reachable for
harnesses that read only local images: download the card, then view it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul's directive: the rendered board and hero for each dealt world ride
along with the challengers, framed as a quality bar (the finish and
commitment level the build is expected to reach), never as a mockup to
copy. The seed prints QUALITY BAR urls per challenger, preferring
API-provided cardBoard/cardHero fields and deriving from the concept id
otherwise; new-work.md instructs image-capable harnesses to view them
for the world being built. Server side, the roll API now returns
cardBoard/cardHero per challenger (impeccable-site).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Discussion outcome with Paul: a mediocre material world loses to
excellent abstract craft, but an unanchored "just be beautiful" escape
hatch would hand selection straight back to the model's priors. The
resolution: abstraction enters as named systems with their own grammar.
The derivation now states that the audience's graphic and screen
traditions (notation, publications, identity programs, data graphics,
interfaces) are as concrete a candidate as any physical artifact. The
catalog side of the same decision is a 12-entry abstract-graphic
authoring round in impeccable-site, pending review.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Smoke findings (Paul's review): the kids-reading derivation produced
seven candidates from one material family despite the divergence line,
and the brief-pinned bookshop world was rendered as the generic AI
bookshop (cream, serif italic, soft glow). The list must now span at
least three material families, and a pinned world licenses its full
material range, never just its softest rendition.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul's probe review (recovery-ab): the no-invention rule was blocking
bold greenfield directions whose demonstration data does not exist yet;
kids-reading amplified the product's "quiet support" adjectives into a
whole-page aesthetic; and nothing guaranteed a Persuade surface still
sells once the form commits (a prior generation shipped zero nav and
zero CTA in the first viewport).
- Truth split in two: commercial and factual claims stay uninventable;
illustrative material is authored at full fidelity, labeled synthetic,
with a replace-with-real list for the user. Mirrored in the build
section so execution-only sessions get it too.
- Persuade floor restored from a22: conversion lives inside the form's
own vocabulary (one-line hook, visible primary action, legible reading
order); a committed form that hides the offer has not finished
translating. The contract's FIRST VIEWPORT block now names where the
primary action sits, and the finishing review verifies the mode did
its job.
- Calibration: negative constraints rule out devices, not exuberance;
product-behavior adjectives do not dictate surface energy.
- Web leverage: when the chosen world names a technique (canvas, WebGL,
view transitions), build the technique, not a static imitation.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Probe finding (recovery-ab, obs sample 1): "treat both structures as the
rut, not the range" let the model rank the observability dashboard grid
at position 7, and the dice landed on it; the challenger fusion rescued
that draw, but a die face spent on the category's own page is a wasted
roll. The a26 wording excluded both structures outright; restore that
with the reason attached.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The ship40 concept pipeline had reversed the proven a-series mechanisms:
the seed's roll decayed into a shortlist nomination that taste functions
(model ranking, candidate floor, simulated user) then argmaxed into the
safest card; the costume check returned as the Translation veto and
carrier-removal test; and the 07-15 rewrite deleted the calibration,
reflex-font lanes, color strategies, and commit-every-atom language that
had held off the cream-editorial default since the alpha era. Five of six
frozen craft directions converged on the same warm-paper family and both
builders obeyed them.
This lands the repair on top of the in-progress simplification:
- new-work.md: the script assigns the build index again on both scopes;
catalog challengers are fused (challenger supplies form and grammar,
product supplies every fact, clarity wins conflicts) and weighed on the
two proven axes only; attended runs present one fully committed
direction with re-roll and an optional steer instead of a ranked
lineup; the color-strategy picker, reflex-face list, saturated-look
calibration, first-viewport thesis and memory test, commit-every-atom,
scroll pacing, and prove-don't-claim return; the direction contract
returns as five lean blocks audited by the separate-agent finish.
- concept-seed.mjs: PROMOTED INDEX becomes ASSIGNED INDEX with
build-assignment semantics; self re-roll only on named factual grounds.
- craft-floor.md: hook-active sessions act on findings instead of
re-auditing; the Refuse list is framed as category defaults the brief
can earn; a closing commitment line keeps a ban list from being the
last word before code.
- codex.md / shape.md: contract references restored for flow coherence.
Adopts the concurrent session's ceremony cuts, softened challenger
instruction, seed SOURCE IDs and --candidate-count, detector-ownership
fix, and the removal of the hook-side contract audit (the audit now
belongs to the separate reviewer at finish).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Deleting interaction-design.md took the last copy of
skill-interaction-dropdown-clipping with it. harden.md's `overflow:
hidden` hits are code samples, not the rule.
It lands in operate.md's Components list rather than the craft floor:
dropdowns and overlays are dense-product-UI components, and the floor
just lost 25% of its length for being a place where specifics accumulate.
The detector's clipped-overflow-container rule catches this after the
fact, but only in sessions with a hook.
Co-Authored-By: Claude <noreply@anthropic.com>
Nothing has loaded it since bbed6eef (Jul 15) rewrote shape.md, which
held its only referrer, a parenthetical example in a list of files that
might be useful. It was never in the Commands table or command-metadata,
so no route reached it either, yet all 189 lines shipped into every one
of the 14 provider bundles.
The eight interactive states and focus rings it covered live in
craft-floor.md's States check and, in more depth, in audit.md, polish.md,
and layout.md. Its CSS anchor positioning, Popover API, and roving
tabindex material has no home elsewhere; recover it from git if a command
turns out to want it.
Co-Authored-By: Claude <noreply@anthropic.com>
live.md's insert branch told the agent to load `brand.md` or `product.md`.
Both files are gone on this branch; the register system became SKILL.md's
modes plus operate.md. Net-new markup in live mode now decides the mode
from the surface and loads craft-floor.md, which is where the bans live.
The freeform generate path gets the same pointer, since live never runs
Setup step 3 and so never picks the floor up on its own.
operate.md still located the craft floor inside SKILL.md.
Co-Authored-By: Claude <noreply@anthropic.com>
910 words to 682, same 35 rule markers, no guidance dropped.
- Two sections instead of three. "Absolute bans" and "detector-blind
reflexes" split the same list by whether our scanner happens to catch
it, which is a fact about our tooling and tells the model nothing about
the design. Merged into one Refuse list, grouped as page scaffolds and
surface habits, which is a distinction the model can act on.
- Folded three duplicates: text-overflow was already in the Type check,
the uniform section reveal was the other half of the Motion check, and
card-everything was already inside the card-grid ban.
- Cut explanation the model does not need. It knows what gradient text
is and what group-hover does; it needs the refusal, not the mechanism.
The gemini block goes from four sentences to three short ones, and the
motion palette line drops the CSS tutorial for "reach past transform
and opacity."
- The authority note moves to the header so no item has to hedge.
Co-Authored-By: Claude <noreply@anthropic.com>
The detector-blind slop review existed because the AI-tell rules had been
stripped out of SKILL.md and nothing carried them. The floor is a better
home: it loads after concept ideation and immediately before editing UI,
which is the placement that made stripping them necessary in the first
place. Models tread lightly when a ban list is present during ideation;
by the time the floor loads, the direction is already committed.
- Rename build-floor.md to craft-floor.md and restore the absolute bans
(side-stripes, gradient text, glassmorphism, hero-metric, identical card
grids, eyebrow-on-every-section, numbered markers, text overflow), the
codex and gemini defect lists, and the reflexes no scanner catches.
Rule ids match the ones the ablation catalog already knows.
- Delete lib/slop-review.mjs and both injections. The Stop hook is now
purely a mechanical pass and stays silent with nothing to report.
- context.mjs replaces AI_SLOP_REVIEW_REQUIRED with the narrower
MANUAL_DETECTOR_REQUIRED, emitted only when a session has no hook at
all. A per-edit hook already covers the mechanical gap, and the floor
covers the judgment one either way.
Co-Authored-By: Claude <noreply@anthropic.com>
main still carries the site, so every `site/` path resolves to deleted.
`tests/docs-integrity.test.js` goes with it (it imports the site's demo
renderer), and `package.json` keeps main's `@anthropic-ai/sdk` bump while
dropping `@google/genai` and `@paper-design/shaders`, which nothing in the
product layer imports.
Real code merges:
- hook-lib: main's #391 cache fix (sync the remembered set to the live
scan so fixed findings stop being named and a reintroduced one fires
again) now runs on the immediate tier rather than the whole filtered
set. Remembering a deferred finding the per-edit pass never reported
would let the Stop deep pass dedupe it away. main's `maxFileBytes`
ceiling, `cleanAcked` once-per-file ack, and template-extensions
re-export all land alongside the tiering work.
- live-browser: main's `hasParams` gate on the Tune badge, keeping this
branch's `C.ink` badge text so it stays legible on kinpaku gold.
- detect-text: both the block-level codex-grid-background scan and main's
inset-stripe CSS check.
- test-suites: union of both trigger sets and file lists, minus the
site-only entries (`shiki-theme`, `docs-integrity`).
- Two hook tests moved off deferred-tier rules (`overused-font`,
`side-tab`) onto immediate-tier ones. They assert cache bookkeeping,
which the per-edit pass only reaches for the immediate tier.
Also drops the site waivers from `.impeccable/config.json` and stops
`build:browser` recreating a stray `site/` tree just to write a bundle
the other repo builds itself.
Co-Authored-By: Claude <noreply@anthropic.com>
Replace the Consequence floor with Signature: one authored move that
makes the experience unmistakable and shapes implementation, named in
terms of what the visitor experiences. Translation now demands the
source's aesthetic and compositional laws survive alongside product
structure, so function-without-character reads as safe flattening and
character-without-structure as costume.
Add an expand-then-contract step before the direction contract: decide
spatial, motion, interaction, narrative, and system questions as one
studio plan that causes itself, then compress into the contract. Staging
guidance follows the seed's move to several inputs.
Co-Authored-By: Claude <noreply@anthropic.com>