The structured-tool channel collapsed to a single direction plus
re-roll, which read as "the system only ever offers one idea" next to
the multi-card decision page. Both channels now share one structure,
assigned direction leading, the one or two fused challengers that
survived the weighing as named alternates, re-roll with steer, and
differ only in richness. The anti-lineup rule stays precise: what never
appears is a ranked menu of the model's own grounded candidates; dealt
challengers carry no ranking rut.
Note: dist rebuild deliberately deferred; the release-gate campaign is
running against the pinned dist and rebuilding mid-run aborts it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Abort signals do not cancel the TCP connect phase, so an unreachable API
stalled the seed ~10s before degrading. All API calls now share one
deadline, the roll fetch races it, and the CLI exits explicitly.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A grounded direction with no hero rendered a blank 16:9 void where the
card image belongs (seen live: the assigned Xerox Zine card led the
hand as a black hole next to two rendered challengers). An option with
no imagery now drops the media region entirely and leads with its
kicker and text; an option with only a board shows the board as its
front image with no flip.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The em-dash-overuse text analyzer ran stripHtmlToText over raw markup,
which drops tags but leaves character entities intact. A model that wrote
—, —, or — rendered a real em-dash the counter never
saw, so 12 entity-escaped dashes on a live page slipped through.
Decode the em-dash entities (named, zero-padded decimal, upper/lower hex)
to the literal glyph before counting. En-dash entities stay untouched: the
rule counts em-dashes, and the literal en-dash was never counted either.
The gap lived only in the regex / static-HTML path (detectText and
detect-html's runTextContentAnalyzers, both over raw HTML). The browser
adapter never ran this analyzer, so build:browser and build:extension
produce no diff.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
With the PRODUCT.md skip fixed, the Opus smoke unmasked the adjacent
gap: the model builds a new world and never writes DESIGN.md (zero
attempts), so the worker's requiresDesign assertion correctly fails the
run. Same disease, same treatment: DESIGN.md is now part of recording
the decision, written before the first build edit in the same stretch
as the direction contract, and the finishing reviewer checks
persistence first, before any craft point is scored.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The truth split already permits full-fidelity demonstration data, but
permission at selection time was not holding at build time: models that
would not author covers, names, or thumbnails compensated with chrome,
which is the content-starved look the detector hunts. Two build rules
make it a mandate: every blank the ask round left open is authored at
production fidelity (content is authorable, claims are labelable,
nothing is omittable; unanswered commercial claims ship as marked
placeholders with a replacement list), and when image generation is
available, generating the build's imagery is part of building rather
than a nicety.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A live retest showed the model dropping to the structured question tool
when the roll degraded: with no challengers and no cards it judged the
page pointless and presented one option in plain text. The degraded seed
output and the new-work rule now both state that degradation changes the
cards, not the channel; a browser session presents the assigned
direction as a single text-only card with re-roll on the decision page.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A traced Claude Code injection asserts for whole model families that the
user is not watching and cannot answer questions; it ships default-on
with no off switch, and it suppressed every interactive step of a live
run (interview skipped, PRODUCT.md inferred, decision page never
served). Prose in a reference file loses that argument, placement wins
it: context.mjs now emits AUTONOMY_DIRECTIVE_CHECK as tool-result
content in the working turn, telling the model such a claim is a
harness default, never session evidence, and to probe once with the
question tool before inferring. init.md makes the same test mechanical:
tool presence proves an answer mechanism, one real probe round is
required, inference afterward must be labeled and disclosed in the
first reply. The degraded concept-seed path now also tells the model to
disclose the degraded roll instead of presenting it as a full one.
Image-gen signaling stays positive-only per Paul: key present emits the
capability, absence stays silent so harness-native tools are not
suppressed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul reproduced the Opus smoke failure in a fresh repo: given a
natural-language build intent, the model runs concept-seed directly and
skips the init divert entirely, so PRODUCT.md never exists and nothing
grounds the challenger fusion. Prose already says init-first in both
SKILL.md routing and new-work.md; prose alone does not hold the floor.
The deal path now refuses with a NO_PRODUCT_MD directive routing to
reference/init.md when loadContext finds no PRODUCT.md. The --chosen
telemetry ping stays ungated.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The new self-detection ran before mode dispatch, so it also caught --wait,
--stop, and --schema. CI failed on the start/wait cycle test: --wait
returned 2 (no browser) where the documented poll loop expects 3
(WAITING). Under CI=1 the suite went 4 pass / 2 fail; it is 6 / 0 now.
Two of those modes were user-facing bugs, not just test breakage. --stop
exited 2 without killing the daemon it was asked to kill, leaking a
server process (verified: one daemon running, CI=1 --stop, still one).
--schema only prints a payload example, and new-work.md tells the agent
to read it before building a payload.
Detection can only tell whether this process can auto-open a browser, not
whether the user has one: SSH with a forwarded port and a harness with an
in-app browser both have a browser and no DISPLAY. The file already
treats serve-without-opening as first class, since --start spawns its own
daemon with --no-open. So the check now gates acquiring a session, not
managing or ending one. The blocking serve path still exits 2 on a
headless box, with a test pinning that.
Co-Authored-By: Claude <noreply@anthropic.com>
v4 changed PRODUCT.md's shape and retired the register axis, so an
upgraded project can carry answers nothing reads. Nothing measured that.
Two tiers, and the split is a performance contract:
- Boot (context.mjs, emitting CONTEXT_STALE) spends only what a boot
already spends: markdown already in memory, a bounded set of stats,
the small JSON files the boot reads anyway. No new directory walks.
One directive for the whole set, throttled to once a week per project
so a finding the user declined does not reappear tomorrow.
- doctor.mjs runs the deep pass on demand: git drift, ignore lists
validated against the live rule registry, hook script paths that stop
resolving, and the monorepo workspace sweep. --fix applies only the
migrations that carry no decision.
Findings are data, not prose, so the boot directive, the text report and
--json all render one set. Severity says what should happen: auto (fix on
the next write anyway), mention (state once), route (name the command
that owns the repair).
PRODUCT.md now carries a schema stamp so the checks stop reconstructing a
file's vintage from which sections it happens to have. Schema version,
not release version: a record written by 4.0.0 is not stale under 4.0.1.
DESIGN.md gets no stamp, because it follows the external design.md spec
that Stitch lints and every DESIGN.md signal is measurable without one.
The highest-value catch is a project that resolves to web while carrying
native build files, including a monorepo app inheriting a root record
that says web. That one costs output quality silently; nothing failed
before.
doctor follows the hooks/pin pattern rather than the Commands table, so
it stays out of the design menu and the count stays at 23.
Also corrects CLAUDE.md, which still documented the register axis,
reference/brand.md, reference/product.md, eleven deleted domain reference
files, and an extractRegister() whose only occurrence in the repo was
that sentence.
Prepared with AI assistance (Claude Code).
Co-Authored-By: Claude <noreply@anthropic.com>
The env-var bypass (IMPECCABLE_QUESTION_DISABLED) relied on the harness
remembering to set it. The script now also self-detects CI, SSH-without-
display, and displayless Linux and exits 2 with the structured-question
advice; --no-open skips detection (caller opens the URL itself, as the
tests do) and IMPECCABLE_QUESTION_FORCE=1 overrides it. new-work.md now
frames the decision-page rule by capability: open a browser if you can,
structured question tool if you cannot, exit 2 means fallback not error.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The v4 repo split dropped astro/wrangler from devDependencies, which
also removed the only (transitive) source of @babel/parser. The
post-apply syntax check in live-copy-edit-agent.mjs requires it to flag
invalid JSX/TSX; without it the check silently degrades to a warning and
tests/live-copy-edit-agent.test.mjs "flags invalid JSX syntax" fails on a
fresh CI install. Production behavior is unchanged: the require stays an
optional, graceful-degrade path for end users, and @babel/parser was
never in the published package's runtime dependencies. Declaring it as a
devDependency just makes the repo's own test environment deterministic.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Setup step 1 tells the agent to follow context.mjs's directives, and the
no-PRODUCT.md-with-existing-code directive explicitly permits a narrow
refinement to proceed on the incumbent implementation and offer init
afterward. Routing's "Otherwise" branch said missing PRODUCT.md routes
through init, with no carve-out, so the two instructions disagreed on the
same request and the agent could block work context.mjs had cleared.
Rule 3 now splits the way the directive does: a new surface or
replacement world goes through init then new-work, a narrow refinement
proceeds and offers init afterward. Explicit and implied commands were
never affected; they route one rule earlier, which is what skill-behavior
scenario 10 already covers.
Co-Authored-By: Claude <noreply@anthropic.com>
Bugbot flagged the "resume without rerunning context.mjs" instruction
after init. It is right, and the gap is wider than the platform half it
named: context.mjs has two output branches, and the no-PRODUCT.md branch
omits DESIGN.md, the native platform references, and the unrecognized
`## Platform` warning. Because the skill never reruns the script once
init writes PRODUCT.md, whatever that first run withheld is gone for the
whole session. A greenfield iOS project would be designed without
reference/ios.md ever loading, and a project carrying DESIGN.md without
PRODUCT.md never saw its own design system.
The two halves need different fixes. DESIGN.md is authority in its own
right and does not depend on PRODUCT.md existing, so context.mjs now
emits it on both branches. Platform is unknowable before PRODUCT.md
exists, so no change to the script can recover it; init.md, the one step
that learns the answer, now loads ios.md / android.md / both right after
recording a native platform, and SKILL.src.md says so where it tells the
agent not to rerun.
Verified end to end against a temp project on both branches.
Co-Authored-By: Claude <noreply@anthropic.com>
Two true positives from the Bugbot review on PR #397.
document.md seed mode told the agent to run "Select one direction" for
paths A, D, or E. new-work.md has neither that heading nor the A/D/E
lettering since the workshop was restructured into named subsections, so
a literal read could skip the world-and-surface flow entirely. Point at
"Create or replace the visual world" and "Commit the world" instead.
critique.md let the heuristic table renormalize to an applicable maximum
when heuristics are scored n/a, but the report template hardcoded ??/40,
the rating bands only mapped raw numbers out of 40, and the persisted
meta carried total_score with no denominator. Trends could silently
compare 24/32 against 30/40 as if they were the same scale. The template
now prints the applicable max, the bands fall back to percentages for
partial sets, the snapshot records max_score and na_heuristics, and the
trend line states its denominator or breaks it out per run when they
disagree. critique-storage.mjs serializes frontmatter key-agnostically,
so the new keys need no code change.
Co-Authored-By: Claude <noreply@anthropic.com>
IMPECCABLE_CARD_BASE overrides the quality-bar URL prefix so eval
workers serve cards from the local checkout while impeccable.style
stays undeployed. IMPECCABLE_QUESTION_DISABLED makes serve-question
exit immediately with the structured-question fallback line, so
headless workers never block on a browser page nobody will answer.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Bump skill to 4.0.0 (plugin.json + marketplace.json) and the CLI to
3.3.0 (package.json), then run build:release to regenerate the plugin
subtree and all provider harness output to the new versions.
Skill 4.0.0 ships external-dice direction assignment, the reviewed
world catalog dealt through the roll API with rendered quality-bar
cards, the in-browser serve-question decision page, visualize-before-
build, the rebuilt new-work flow, and the 58-rule detector under hook
enforcement. CLI 3.3.0 grows the deterministic detector to 58 rules
and adds config-declared context roots, per-file rule scoping, and
--target resolution for nested products.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Gate2 measured it precisely: on matched assignments Opus renders warm,
bookish, and child-facing subjects as cream, serif-italic, and
lamplight while Sol renders the same positions saturated, and neutral
prose hardening did not move it. The codex and gemini blocks set the
precedent for provider-conditional counterweights; this adds the claude
block at the palette decision: the first palette is already spent, an
OWN-WORLD block reading cream/paper/parchment/lamplight for an unpinned
Persuade surface is a failed rendition to rework from the world's
saturated materials, and nothing about the subject requires the
default. Verified present in the claude-code dist and absent from
codex.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Emit .grok skills, agents, and PostToolUse/Stop hooks; wire the CLI
installer and downloads; fix the plugin install path to #plugin; and
document Grok in HARNESSES.md and README.
AI assistance: written with Grok Build.
Paul's question exposed the blind spot: a closed tab left the agent
waiting out the full timeout. The page now sends a heartbeat every five
seconds while open; the server stamps lastBeat into the state file, and
--wait reports PAGE CLOSED with exit 4 when the beats stop for fifteen
seconds without an answer. The prose defines the fallback ladder:
re-present once through the structured question tool, then proceed
unattended with the assigned direction, stating assumptions.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The loading hand now has the true proportions: 16:9 shimmer media, a
tier line, a title line, three detail lines, and a button-shaped block
pinned to the card bottom, with each skeleton inheriting the measured
height of the card it replaces.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Live-fire bug from the demo: collecting a re-roll answer deleted the
state file, so --update could not find the still-running server and the
next hand had nowhere to land. Cleanup is now terminal-only: a re-roll
consumes just the answer file and leaves the server state for --update;
any other choice cleans up fully as before.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two upgrades from the live session. Re-roll no longer ends the page: in
detached mode the server stays alive, the client gathers the cards back
into the center stack, deals skeleton cards with the site's shimmer,
and polls /next-status; the new --update mode delivers the next hand
and the page reloads into the fresh deal. Choices other than re-roll
still resolve and exit as before.
And the headline bug had a root cause: the fonts link never loaded the
weights in use, so the browser synthesized a fake 300 that read
off-brand and muddy. The link now loads Alumni 100 and 400 exactly, and
the h1 wears the homepage hero display role: weight 100 at
clamp(2.6rem, 5vw, 4.2rem), champagne.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
#lightbox { display: flex } outspecified the UA's [hidden] rule, so an
invisible full-viewport layer sat over the page and blocked all hover
and click. #lightbox[hidden] { display: none } restores reality.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Craft-gate forensics on matched hands: both models obeyed the same
assigned indices, but Opus rendered every kids cell as cream paper,
lamplight, and serif, and sample 1 chose Fraunces off the reflex list
with a bookshop-signage rationalization, while Sol rendered the same
positions as indigo bookcloth, coral thread, tomato, and marigold. The
dice work; the rendition prior escaped through two hatches, now closed:
naming a reflex face requires a reason no other face satisfies and a
subject association is never that reason; and bookish or child-facing
subjects do not soften the calibration, because cream paper is the
smallest corner of the book world.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul's motion and exploration pass. The reveal is now a real deal: the
cards begin piled at the grid center, blurred and slightly rotated, and
travel to their seats with a 110ms stagger on the brand ease (JS
measures each card's seat, so the pile works at any grid shape;
reduced motion skips it). Hovering a card bleeds its hero into the page
ground behind a lacquer scrim. An expand chip beside Board opens
whichever face is showing in a zoom-out lightbox with Escape to close.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
No cropping and no pillarboxing: the board spans the card width at its
native 16:9 exactly like the hero on the front, with the deep-lacquer
ground below and the label-and-CTA bar pinned to the card's bottom edge.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The design-system board is an information sheet; the back face now
contains it fully on a deep-lacquer ground rather than cover-cropping
its top and bottom.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul's read: flipping only the image inside a static frame looked
unconvincing. The card is now the object: front face carries hero,
lineage, title, and CTA; the back face is the design-system board at
full card height with a slim bar keeping the label and Build-this
reachable. The outer card keeps fan, deal, and hover; each face carries
its own lacquer chrome and the rolled card's gold ring rides both
faces. 700ms preserve-3d turn, instant under reduced motion.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Third polish pass with Paul live. The details collapsible is gone: the
card media is now a 3D flip, hero on the front and the design-system
board on the back, toggled by a small mono Board/Hero chip with a 600ms
preserve-3d turn that reduced-motion collapses to an instant swap. The
headline and question sit directly above the dealt hand inside the
centered stage while the logo holds the top-left corner. Re-roll
stretches to the steer input's height and says just Re-roll beside the
five-pip die.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul's second polish pass. The h1 uses DESIGN.md's headline role
(Alumni 300 at clamp(2rem, 4vw, 3.4rem), tracking 0) instead of a bold
weight the brand no longer uses; card headings use the title role
(Albert 500, 1.125rem) which also holds at small sizes; the rolled
card's border and ring use actual kinpaku gold, not the deep variant;
and the dealt hand centers vertically in the viewport with header and
footer framing it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul's review against DESIGN.md and the design-system page: the real
24px logo mark with the uppercase tracked Alumni 400 wordmark (was a
bootleg lowercase 600); the five-pip stroke die from the homepage
worlds-roll as the headline accent and re-roll icon (the tilted numeral
cube was off-brand); the re-roll button is the worlds-reroll pattern
verbatim (mono 0.72rem uppercase tracked, rule border); Build-this uses
the DESIGN.md button-primary spec (title typography, 38px padding,
kinpaku-pale hover); THE ROLL kicker is the worlds-played-chip pattern.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The decision page now mirrors impeccable.style's Neo kinpaku system:
the split logo mark and Alumni Sans wordmark in kinpaku gold, lacquer
ground with raised-panel cards, the gold die beside the headline
(count of dealt options, rotated like the research page dice), THE ROLL
badge as a mini die, worlds-roll card treatment (rule borders, fan
rotation, deal-in stagger honoring reduced motion, hover lift), mono
tracked lineage lines, champagne display type, gold CTA with dark ink,
and a die-glyph re-roll button. Tokens mirrored from kinpaku-tokens.css.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul's call: start mode never auto-opens; the agent is alive and opens
the printed URL itself, in-app browser first, then the system opener,
then showing the URL (--open forces the system browser from the script).
The prose now leads with the start/open/wait flow and keeps the blocking
auto-open path for harnesses that can background a shell.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul's two concerns with the blocking design. --schema prints the exact
payload example so the model never guesses the shape (new-work.md points
at it). And harnesses that cannot leave a shell blocked (or cannot open
a browser while blocked) get a two-phase path: --start daemonizes the
server and returns the URL plus a key immediately, --wait polls for the
answer with exit 3 meaning ask again, exit 2 meaning the server died,
and --stop for cleanup. The browser open happens from the detached
server process, so it works even when the agent thread is short-lived.
State lives under .impeccable/questions/.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul's design, three pieces:
serve-question.mjs: the world decision presented as a themed page instead
of a text prompt. The script serves an impeccable-styled option board
(assigned direction leading with THE ROLL badge, dealt challengers as
alternates carrying their QUALITY BAR cards, re-roll and steer built in),
prints the URL, opens the browser, and blocks until the user chooses;
the answer lands on stdout as ANSWER JSON, so the shell call itself is
the wait and no harness machinery is needed. Local images are served by
the ephemeral server; nothing leaves the machine.
generate-image.mjs + context.mjs IMAGE_GEN_AVAILABLE: when an OpenAI key
is in the environment, context reports that image generation works even
without a harness-native tool (gpt-image-2, billed to the user's key,
stated before first use; Google skipped by decision). Harness-native
tools always win when present.
new-work.md: visualize-before-build is now the default whenever any
image generation exists, not a codex.md special case; the attended
presentation prefers the visual decision page and falls back to the
structured question tool. Evals keep the unattended path untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul confirmed the craft-bar experiment: builds that saw the dealt
worlds' hero cards produced visibly stronger execution than the
no-image control. One clause makes the mechanism reachable for
harnesses that read only local images: download the card, then view it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul's directive: the rendered board and hero for each dealt world ride
along with the challengers, framed as a quality bar (the finish and
commitment level the build is expected to reach), never as a mockup to
copy. The seed prints QUALITY BAR urls per challenger, preferring
API-provided cardBoard/cardHero fields and deriving from the concept id
otherwise; new-work.md instructs image-capable harnesses to view them
for the world being built. Server side, the roll API now returns
cardBoard/cardHero per challenger (impeccable-site).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Discussion outcome with Paul: a mediocre material world loses to
excellent abstract craft, but an unanchored "just be beautiful" escape
hatch would hand selection straight back to the model's priors. The
resolution: abstraction enters as named systems with their own grammar.
The derivation now states that the audience's graphic and screen
traditions (notation, publications, identity programs, data graphics,
interfaces) are as concrete a candidate as any physical artifact. The
catalog side of the same decision is a 12-entry abstract-graphic
authoring round in impeccable-site, pending review.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>