Commit Graph
1122 Commits
Author SHA1 Message Date
Paul BakausandClaude Fable 5 3f9fccdfd0 Live: lock down the local server against same-machine token theft (#304)
Two defense-in-depth layers close the P1 in issue #304, where any browser
tab on the machine could fetch /live.js, extract the embedded token, and
drive every token-gated route.

1. Loopback-restricted CORS. The shared handler replaced its wildcard
   `Access-Control-Allow-Origin: *` with reflection gated on a strict
   isLoopbackOrigin() that URL-parses the Origin (so localhost.evil.com and
   127.0.0.1.evil.com fail) and accepts only http/https on localhost,
   127.0.0.1, or [::1]. Reflection always pairs with `Vary: Origin` so a
   cache never hands one origin's authorized response to another. Remote
   origins get no ACAO header; origin-less callers (script tags, curl, the
   agent's own fetches) are unaffected.

2. Token-gated /live.js. The handler now 401s unless `?token=` matches
   state.token, so the bundle (which embeds the token) is no longer served
   to unauthenticated local pages. The injected <script src> carries the
   token: live.mjs passes --token to live-inject.mjs, which threads it
   through every injection path (HTML/JSX tag, Nuxt plugin, SvelteKit root
   component) via a shared buildLiveScriptSrc(). The token stays optional in
   live-inject so static fixture tests keep their bare src.

Tests: new live-server integration cases for the 401 gate, remote-origin
denial, loopback reflection + Vary, and token-guarded routes under a
loopback Origin; e2e session harness now injects with the token.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 21:59:28 -07:00
Paul BakausandClaude Fable 5 da2982ab95 Fix /source guard escaping the project root via sibling directories
The /source route confined paths with `absPath.startsWith(process.cwd())`,
a string-prefix check with no separator. An absolute request path to a
sibling directory whose name extends the project dir name (projeto ->
projeto-backup) shared the prefix and was served. Switch to the relative-path
check already used by sessionFileMetadataFromPollReply: reject when the
relative path is empty (the root dir itself, never a file this route serves),
starts with `..`, or is absolute.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 21:59:28 -07:00
github-actions[bot] 762ffd08b2 Sync generated provider output 2026-07-23 04:50:46 +00:00
Paul BakausandClaude Fable 5 55094aaa0d Fix false hook-script-missing in doctor when ${CLAUDE_PROJECT_DIR} is unexpanded
The deep staleness pass extracted a hook-script path with a greedy `\S*`
prefix that swallowed the `${CLAUDE_PROJECT_DIR}/` placeholder, then
existsSync'd the literal string. That string never exists, so every project
installed by `impeccable hooks on` got a `hook-script-missing` finding with
text claiming UI edits were going unscanned — the opposite of the truth.

Split extraction from resolution. hookScriptTokenFrom now pulls the path
token (quoted-first, so it handles the #399 guarded `[ ! -f "PATH" ] || node
"PATH"` form and absolute user-level installs) without absorbing shell
syntax. resolveHookScriptPath then applies a per-placeholder policy:

- ${CLAUDE_PROJECT_DIR} expands to the scanned root (the runtime mapping).
- ${CLAUDE_PLUGIN_ROOT} / ${PLUGIN_ROOT} / ${GROK_PLUGIN_ROOT}, $(...) command
  substitution (GitHub's $(git rev-parse)), and any other $VAR are SKIPPED:
  the doctor cannot know those locations and must never assert a negative it
  cannot verify.

The check stays real: a placeholder that expands to a genuinely absent path
still flags. Adds TDD coverage for every command form.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 21:50:10 -07:00
github-actions[bot] 698a743958 Sync generated provider output 2026-07-23 00:34:48 +00:00
Paul BakausandClaude Fable 5 47aff2e0be Fix Stop-hook loop: honor stop_hook_active per Claude Code contract
The Stop deep pass (runStopHook) never read the stop_hook_active field
from the Claude Code Stop-hook event. When a prior fire kept the turn
alive via hookSpecificOutput.additionalContext and the agent legitimately
declined to act, the hook re-scanned and re-blocked every re-invocation
until Claude Code's consecutive-block cap force-ended the turn (issue #400).

Read stop_hook_active early in runStopHook, right after the event is
parsed and before any scan, and exit 0 with no output when it is true. The
prior fire already surfaced the findings; acting on them is the agent's
call. Only Claude Code sends this field, so the strict === true is a no-op
for other harnesses. runHook (PostToolUse) and hook-before-edit.mjs
(PreToolUse) never receive the field, so they are unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 17:34:20 -07:00
github-actions[bot] 9b7f7ffbba Sync generated provider output 2026-07-22 20:14:42 +00:00
Paul BakausandClaude Fable 5 3e233d22d7 Release prep: CLI v3.3.1
Bump the npm package and regenerate the browser detector bundle with
the advisory tier, entity-aware em-dash counting, and the
undersized-ui-text rule.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
cli-v3.3.1
2026-07-22 13:13:40 -07:00
Paul BakausandClaude Fable 5 087983070b Release script verifies impeccable.style serves the released version
The 4.0.0 release stranded npx-update users on a stale bundle for a
day because the site deploy is a separate step nobody was reminded of.
Skill releases now check /api/version and print the redeploy command
when the served version lags.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 12:39:32 -07:00
Paul BakausandClaude Fable 5 eda81f0937 Release prep: skill v4.0.1
Bump plugin + marketplace to 4.0.1 and sync the regenerated provider
output: the guarded hook commands from issue #399 (a missing hook file
exits 0 instead of crashing every turn of a user-level install), the
canon standing exit, the visualize flow, the two shipped subagents, and
the interactive-spine fixes from today's live testing. Detector count
validates at 59 with undersized-ui-text.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
skill-v4.0.1
2026-07-22 12:23:07 -07:00
Paul BakausandClaude Fable 5 13c078ae93 Fix user-level hook path crash and clarify skills update scope (#399)
Part 1 — user-level hooks got a project-relative command. copyProviderHooks
only rewrote the bundled ${CLAUDE_PROJECT_DIR}-relative hook command to an
absolute skill path when the skill lived elsewhere than the manifest root. A
user-level update (root === ~) kept ${CLAUDE_PROJECT_DIR}, which a global
~/.claude/settings.local.json resolves per-project — crashing node at module
resolution on every PostToolUse/Stop in any project without a local skill copy.

Now the command is rewritten to the resolved absolute path whenever the manifest
is a user/global file (isHomeDir(root)) as well as the pre-existing
skill-elsewhere case, and every hook command is wrapped with a missing-file
guard `[ ! -f "PATH" ] || node "PATH"`. The guard exits 0 when the script is
absent (upholding hook.mjs's "never break a turn" contract even before node can
load it) while preserving node's own exit code when present, so Claude's exit-2
blocking signal still reaches the agent. Project-scope hooks keep the portable
${CLAUDE_PROJECT_DIR} token.

Part 2 — skills update silently targeted CWD. update now resolves and names the
target explicitly (project vs user level, with the absolute path), honors
--user/--project, only counts a provider as installed when the impeccable skill
itself is present (so it never vendors a copy into a repo that merely tracks
other first-party skills), and offers the choice when both a project and a
user-level install exist instead of silently picking. Non-interactive runs
default to the project and print how to target the other.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 12:23:07 -07:00
Paul BakausandClaude Fable 5 d66782753c serve-question: correct content-type for svg and gif heroes
The local-image map fell through to image/jpeg for anything that was
not webp or png, so an svg hero (the fake comp generator's native
format) silently failed to render on the decision page.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 12:23:07 -07:00
Paul BakausandClaude Fable 5 6ece0e588f Add deterministic new-work interactive smoke suite
A cheap, LLM-free E2E tier for the interactive parts of new-work, mirroring
the two-layer live-e2e pattern (deterministic now, opt-in LLM tier later).

- generate-image.mjs: IMPECCABLE_IMAGE_GEN_FAKE=1 writes a deterministic
  offline image (SVG with wrapped prompt + SYNTHETIC COMP label, or a valid
  palette-stripe PNG carrying the prompt/marker in a tEXt chunk). Same CLI
  contract, no key, no network, $0.00 cost line.
- tests/new-work-e2e/user-bot.mjs: scripted user bot (module + CLI) that
  resolves the serve-question daemon from the workspace and drives the real
  page via Playwright (pick, re-roll + steer, canon, tab close).
- tests/new-work-e2e.test.mjs: node --test coverage of the serve-question
  cycles (pick + CHOSEN CARD, re-roll + --update re-deal, canon + CANON
  CHOSEN, tab-close exit-4, text-only card) plus fake image determinism.
- Registered as the opt-in new-work-e2e suite; added test:new-work-e2e.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 12:23:07 -07:00
Paul BakausandClaude Fable 5 bcdf38881e Command first, capability second
"When the harness can X, do Y" hands the model an exit before the
command arrives; the observed reviewer skip walked through exactly that
door. The three gated constructions now lead with the imperative,
present the decision visually, open the chosen card, spawn the finish
reviewer, and carry their fallbacks as trailing clauses for sessions
that genuinely lack the capability. Constructions that already led with
the command keep their routing clauses unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 12:23:07 -07:00
Paul BakausandClaude Fable 5 0bbb63b62a Ship the finish reviewer as a named subagent; ungate the asset producer
The eb686f36 session read the separate-reviewer rule and spawned
nothing: an unnamed "separate agent" is an improvisation prompt, not an
affordance. The skill now ships impeccable-finish-reviewer next to the
asset producer: persistence first, ceiling against the card and comp
second, contract promise by promise, truth; ordered material fixes
back to the parent, no editing, no second detector. new-work names it
so the finish step invokes a thing that exists.

The asset producer was gated providers: codex, so Claude Code never
shipped it; the gate is removed and its two codex-only workflow lines
made provider-neutral with codex blocks.

Dist rebuild still deferred for the running campaign.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 12:23:06 -07:00
Paul BakausandClaude Fable 5 91d310696d Canonicalize the visualize flow; put the added prose on a diet
codex.md becomes visualize.md and loads for every harness with any
image generation, native or the API fallback: after the direction
locks, three distinct compositional comps are rendered and put before
the user for approval, in-harness when it can display images,
otherwise on the decision page. Three is the number; one comp invites
rubber-stamping, and this approval round has repeatedly produced the
most compositional and ambitious work, so new-work now marks it
never-skipped. The codex-only subagent stays as a codex note.

The recent rule additions are tightened by a third: the asset and
imagery bullets merge into one, the canon exit loses its restatements,
the DESIGN.md-rule and chosen-card and ceiling clauses each shed their
second clause saying the first clause again. Same laws, fewer words;
prose that grows without bound recreates the attention gravity it was
written to fight.

Dist rebuild still deferred; the release-gate campaign reads the
pinned dist.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 12:23:06 -07:00
Paul BakausandClaude Fable 5 daec380cdb Add undersized-ui-text rule for functional text below an 11px floor
The existing `tiny-text` rule owns long body copy and deliberately exempts
the UI furniture layer (nav, footer, links, buttons, labels, uppercase
micro-labels). That left a real gap: a build shipped its entire furniture
layer (nav links, category names, timecodes, meta rows) at 8px because the
chosen pixel font only steps in 8px increments, and the design hook waved it
through as merely "not on the DESIGN.md ramp" -- which the model resolved by
adding 8px to the ramp. Being on the ramp launders the token, not the
legibility problem.

New `undersized-ui-text` quality rule closes that laundering path:

- Flags interactive and short content-bearing text (links, buttons, nav
  items, labels, table cells, meta rows, timecodes) below an 11px floor. The
  floor holds inside a footer; only non-interactive legal smallprint gets the
  softer 10px floor.
- Ignores the design system entirely, so a value ON the ramp is still
  flagged.
- Uppercase letterspaced micro-labels stay in scope (still functional).
- Exempts sup/sub, visually-hidden (sr-only) text, and code/terminal
  contexts. em/rem/%-sized text that computes at or above the floor never
  fires.
- Complements tiny-text without double-flagging: long non-furniture body
  copy stays with tiny-text.

Implemented as a single check in checkQuality (rules/checks.mjs), so both the
static-html (jsdom) and browser adapters pick it up through the unified
per-element path -- no dual wiring. Registered in registry/antipatterns.mjs.

TDD: fixture tests/fixtures/antipatterns/undersized-ui-text.html (7 flag / 7
pass shapes), failing test first, then implement. Full fixtures suite 64/64.

Deferred (blocked by an active release-gate eval reading build/_data/dist):
regenerate the browser bundle (bun run build:browser ->
cli/engine/detect-antipatterns-browser.js) and the extension detector
(bun run build:extension -> extension/detector/detect.js + antipatterns.json)
so the standalone browser/extension artifacts carry the new rule.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 12:23:06 -07:00
Paul BakausandClaude Fable 5 270f4d20aa Make em-dash-overuse an advisory rule with browser parity
Em-dashes are used legitimately by humans, so em-dash-overuse fired far too
often. Reclassify it as the first advisory-tier rule: detected, but never a
failure.

Engine
- Add `advisory: true` to the rule metadata schema (em-dash-overuse is the
  first). findings.mjs stamps `advisory: true` on advisory findings so every
  consumer can partition without a registry lookup. Rule count stays 58.
- Raise the firing threshold from a flat 5 dashes to two gates: an absolute
  floor of 8 and a density of about one dash per 500 characters of body text.
  A long article that uses a few em-dashes no longer trips; a short,
  dash-per-clause page still does. Entity decoding (mdash, numeric, hex) is
  unchanged. Thresholds live in shared/constants.mjs so every engine agrees.

Browser parity
- The browser bundle carried a registry entry but no logic, so the overlay and
  extension could never flag it. Add checkEmDashOveruse / checkEmDashOveruseDOM
  in rules/checks.mjs (reads rendered text, no entity decoding needed), wire it
  into the injected page-level pass, and carry the advisory flag through
  serializeFindings so the overlay/extension can render it with the mildest
  affordance.

CLI
- Advisory findings print under a separate dimmed "Advisory" section, are
  excluded from the failure count, and never change the exit code (an
  advisory-only scan exits 0). JSON keeps them with `"advisory": true`.
  `--no-advisory` suppresses them entirely.

Hook
- Advisory rules are skipped by default in both the per-edit and Stop deep-pass
  hooks, so the hook never nags about them. Opt in with
  `.impeccable/config.json` -> `detector.advisoryRules: "include"`.

Tests
- Fixture + threshold + browser-adapter coverage; advisory-skip default and
  opt-in for the hook; formatFindings partitioning. The em-dash-overuse stand
  for a deferred copy rule in the tier tests is swapped to marketing-buzzword.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 12:23:06 -07:00
Paul BakausandClaude Fable 5 e409bec7b5 Canon standing exit, chosen-card directive, and the ambition fixes
From Paul's approved UX and the eb686f36 session post-mortem:

The standing exit: direction rounds carry a quiet, permanent "Play it
straight" action (payload flag canon, reserved id) on the decision page
and as the last structured-tool option. It is the user's door, never
the model's: never recommended, never weighed against the roll, and
choosing it swaps the bar rather than lowering it, two or three named
reference products become the craft level, canon executed at full
commitment. Safer/conventional steers resolve here, never to a
stranger re-roll.

Session fixes, each mechanical where possible: the ANSWER line now
names the chosen card's hero and board and directs opening them before
code (the session built from text alone after viewing a different
world's card); generation scale joins the imagery rule (a library of
centered 128px subjects foreclosed the atmospheric hero); DESIGN.md
rules are checked against the world's native devices and never added
to silence a hook finding (the session banned arcade lettering's own
offset shadow and laundered 8px through the ramp); staging joins the
FORM contract block (the axis was dropped silently at world-choice);
the finishing reviewer audits the ceiling against the QUALITY BAR card
after persistence (floor rigor was disguising unreached ambition); the
icon-tile clause names hand-drawn icons as remedy, not target.

Dist rebuild deferred: the release-gate campaign reads the pinned dist.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 12:23:06 -07:00
Paul BakausandClaude Fable 5 70fdc172b8 Resolve detect DESIGN.md from each target's project, not cwd
The detect CLI loaded DESIGN.md once from process.cwd() and applied it to
every scan target. Scanning another project's files from inside a different
repo therefore judged them against the wrong project's design system
(cross-project contamination observed during eval work: running detect from
impeccable-evals against a generated artifact elsewhere applied the evals
repo's DESIGN.md).

DESIGN.md now resolves by walking up from each scan target's own location to
its design root: a directory carrying a DESIGN.md is the root; a directory
carrying a project marker (.git / package.json / .impeccable) without a
DESIGN.md is a boundary that stops the walk with no design system, so a
sibling project never inherits a parent's or cwd's rules. A target with no
design root above it falls back to no design system rather than cwd's.
Resolution is memoized per root, so a multi-file scan reads each DESIGN.md
once, and targets spanning projects each get their own. file:// URLs resolve
from their path; remote http(s) URLs get no design system.

Adds tests/detect-cli-design-contamination.test.mjs, which spawns the real
CLI to prove B's file is not judged by A's DESIGN.md, that a project still
governs its own file, that a mixed-project scan resolves per target, and that
a marker-less bare file gets no design system.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 12:23:06 -07:00
Paul BakausandClaude Fable 5 9f5bbed8b8 Bump astro test fixture to ^7.1.0 to clear dependabot XSS alerts
The astro-vite7 live-e2e fixture pinned astro ^6.0.0, which resolves
into the vulnerable range of three dependabot advisories:
GHSA-4g3v-8h47-v7g6 (reflected XSS via View Transition animation
properties, medium), GHSA-f48w-9m4c-m7f5 (XSS via spread attribute
names in renderHTMLElement, medium), and GHSA-7pw4-f3q4-r2p2 (XSS via
transition:* directive values, low). All three are patched by 7.1.0.

Dev-only test fixture; the vulnerable code paths (View Transitions,
transition directives, spread attributes) are not exercised by this
static, non-hydrated page, so real exposure is nil. Bumped anyway as
the cheap, correct fix. Also corrected the now-stale fixture label to
"Astro 7 + Vite 7".

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 12:23:06 -07:00
Paul BakausandClaude Fable 5 9dade04bbf Text fallback presents surviving challengers as alternates
The structured-tool channel collapsed to a single direction plus
re-roll, which read as "the system only ever offers one idea" next to
the multi-card decision page. Both channels now share one structure,
assigned direction leading, the one or two fused challengers that
survived the weighing as named alternates, re-roll with steer, and
differ only in richness. The anti-lineup rule stays precise: what never
appears is a ranked menu of the model's own grounded candidates; dealt
challengers carry no ranking rut.

Note: dist rebuild deliberately deferred; the release-gate campaign is
running against the pinned dist and rebuilding mid-run aborts it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 12:23:06 -07:00
github-actions[bot] 386f9883cf Sync generated provider output 2026-07-22 07:44:15 +00:00
Paul BakausandClaude Fable 5 d65b6ca029 Name the sandbox cause in the degraded seed and suggest a network retry
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
skill-v4.0.0
2026-07-22 00:43:45 -07:00
github-actions[bot] 3f72e761db Sync generated provider output 2026-07-22 07:39:46 +00:00
Paul BakausandClaude Fable 5 39a617d5a4 Cap seed API stall with a shared raced budget and explicit CLI exit
Abort signals do not cancel the TCP connect phase, so an unreachable API
stalled the seed ~10s before degrading. All API calls now share one
deadline, the roll fetch races it, and the CLI exits explicitly.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 00:39:16 -07:00
github-actions[bot] c2fbc66bdd Sync generated provider output 2026-07-22 07:38:39 +00:00
Paul BakausandClaude Fable 5 9089d0a1a7 Decision page: text-only cards for options without a rendered card
A grounded direction with no hero rendered a blank 16:9 void where the
card image belongs (seen live: the assigned Xerox Zine card led the
hand as a black hole next to two rendered challengers). An option with
no imagery now drops the media region entirely and leads with its
kicker and text; an option with only a board shows the board as its
front image with no flip.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 00:38:10 -07:00
github-actions[bot] 043a157349 Sync generated provider output 2026-07-22 07:29:05 +00:00
Paul BakausandClaude Fable 5 7dcca2bb36 Count em-dash HTML entities in em-dash-overuse
The em-dash-overuse text analyzer ran stripHtmlToText over raw markup,
which drops tags but leaves character entities intact. A model that wrote
&mdash;, &#8212;, or &#x2014; rendered a real em-dash the counter never
saw, so 12 entity-escaped dashes on a live page slipped through.

Decode the em-dash entities (named, zero-padded decimal, upper/lower hex)
to the literal glyph before counting. En-dash entities stay untouched: the
rule counts em-dashes, and the literal en-dash was never counted either.

The gap lived only in the regex / static-HTML path (detectText and
detect-html's runTextContentAnalyzers, both over raw HTML). The browser
adapter never ran this analyzer, so build:browser and build:extension
produce no diff.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 00:27:59 -07:00
Paul BakausandClaude Fable 5 0376145a46 Fix release.mjs crash: existsSync is a named import, not fs.existsSync
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
cli-v3.3.0
2026-07-22 00:08:54 -07:00
github-actions[bot] 396c18bdf9 Sync generated provider output 2026-07-22 07:07:34 +00:00
Paul BakausandClaude Fable 5 13c2ae5aa4 DESIGN.md joins the persistence gate
With the PRODUCT.md skip fixed, the Opus smoke unmasked the adjacent
gap: the model builds a new world and never writes DESIGN.md (zero
attempts), so the worker's requiresDesign assertion correctly fails the
run. Same disease, same treatment: DESIGN.md is now part of recording
the decision, written before the first build edit in the same stretch
as the direction contract, and the finishing reviewer checks
persistence first, before any craft point is scored.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 00:05:18 -07:00
github-actions[bot] a96445b973 Sync generated provider output 2026-07-22 07:00:38 +00:00
Paul BakausandGitHub c81319c381 Merge pull request #397 from pbakaus/oneshot-v4
Impeccable 4.0: dice-assigned directions, world catalog, quality-bar cards, visual decisions
2026-07-22 00:00:08 -07:00
Paul BakausandClaude Fable 5 c0ee4f3cac Build mandate: author the assets, generate the imagery
The truth split already permits full-fidelity demonstration data, but
permission at selection time was not holding at build time: models that
would not author covers, names, or thumbnails compensated with chrome,
which is the content-starved look the detector hunts. Two build rules
make it a mandate: every blank the ask round left open is authored at
production fidelity (content is authorable, claims are labelable,
nothing is omittable; unanswered commercial claims ship as marked
placeholders with a replacement list), and when image generation is
available, generating the build's imagery is part of building rather
than a nicety.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 23:49:01 -07:00
Paul BakausandClaude Fable 5 5a3c8fd18e Degraded roll keeps the decision page as its channel
A live retest showed the model dropping to the structured question tool
when the roll degraded: with no challengers and no cards it judged the
page pointless and presented one option in plain text. The degraded seed
output and the new-work rule now both state that degradation changes the
cards, not the channel; a browser session presents the assigned
direction as a single text-only card with re-roll on the decision page.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 23:30:58 -07:00
Paul BakausandClaude Fable 5 6e8cdfa581 Counter the harness autonomy directive from inside the working turn
A traced Claude Code injection asserts for whole model families that the
user is not watching and cannot answer questions; it ships default-on
with no off switch, and it suppressed every interactive step of a live
run (interview skipped, PRODUCT.md inferred, decision page never
served). Prose in a reference file loses that argument, placement wins
it: context.mjs now emits AUTONOMY_DIRECTIVE_CHECK as tool-result
content in the working turn, telling the model such a claim is a
harness default, never session evidence, and to probe once with the
question tool before inferring. init.md makes the same test mechanical:
tool presence proves an answer mechanism, one real probe round is
required, inference afterward must be labeled and disclosed in the
first reply. The degraded concept-seed path now also tells the model to
disclose the degraded roll instead of presenting it as a full one.
Image-gen signaling stays positive-only per Paul: key present emits the
capability, absence stays silent so harness-native tools are not
suppressed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 23:02:25 -07:00
Paul BakausandClaude Fable 5 f6cecf6149 Gate the concept roll on PRODUCT.md existing
Paul reproduced the Opus smoke failure in a fresh repo: given a
natural-language build intent, the model runs concept-seed directly and
skips the init divert entirely, so PRODUCT.md never exists and nothing
grounds the challenger fusion. Prose already says init-first in both
SKILL.md routing and new-work.md; prose alone does not hold the floor.
The deal path now refuses with a NO_PRODUCT_MD directive routing to
reference/init.md when loadContext finds no PRODUCT.md. The --chosen
telemetry ping stays ungated.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 22:08:02 -07:00
Paul BakausandClaude ea68bb4722 Scope headless detection to the path that opens a browser
The new self-detection ran before mode dispatch, so it also caught --wait,
--stop, and --schema. CI failed on the start/wait cycle test: --wait
returned 2 (no browser) where the documented poll loop expects 3
(WAITING). Under CI=1 the suite went 4 pass / 2 fail; it is 6 / 0 now.

Two of those modes were user-facing bugs, not just test breakage. --stop
exited 2 without killing the daemon it was asked to kill, leaking a
server process (verified: one daemon running, CI=1 --stop, still one).
--schema only prints a payload example, and new-work.md tells the agent
to read it before building a payload.

Detection can only tell whether this process can auto-open a browser, not
whether the user has one: SSH with a forwarded port and a harness with an
in-app browser both have a browser and no DISPLAY. The file already
treats serve-without-opening as first class, since --start spawns its own
daemon with --no-open. So the check now gates acquiring a session, not
managing or ending one. The blocking serve path still exits 2 on a
headless box, with a test pinning that.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-21 22:06:55 -07:00
Paul BakausandClaude 5575a027dc Flag and repair drift in Impeccable's own project artifacts
v4 changed PRODUCT.md's shape and retired the register axis, so an
upgraded project can carry answers nothing reads. Nothing measured that.

Two tiers, and the split is a performance contract:

- Boot (context.mjs, emitting CONTEXT_STALE) spends only what a boot
  already spends: markdown already in memory, a bounded set of stats,
  the small JSON files the boot reads anyway. No new directory walks.
  One directive for the whole set, throttled to once a week per project
  so a finding the user declined does not reappear tomorrow.
- doctor.mjs runs the deep pass on demand: git drift, ignore lists
  validated against the live rule registry, hook script paths that stop
  resolving, and the monorepo workspace sweep. --fix applies only the
  migrations that carry no decision.

Findings are data, not prose, so the boot directive, the text report and
--json all render one set. Severity says what should happen: auto (fix on
the next write anyway), mention (state once), route (name the command
that owns the repair).

PRODUCT.md now carries a schema stamp so the checks stop reconstructing a
file's vintage from which sections it happens to have. Schema version,
not release version: a record written by 4.0.0 is not stale under 4.0.1.
DESIGN.md gets no stamp, because it follows the external design.md spec
that Stitch lints and every DESIGN.md signal is measurable without one.

The highest-value catch is a project that resolves to web while carrying
native build files, including a monorepo app inheriting a root record
that says web. That one costs output quality silently; nothing failed
before.

doctor follows the hooks/pin pattern rather than the Commands table, so
it stays out of the design menu and the count stays at 23.

Also corrects CLAUDE.md, which still documented the register axis,
reference/brand.md, reference/product.md, eleven deleted domain reference
files, and an extractRegister() whose only occurrence in the repo was
that sentence.

Prepared with AI assistance (Claude Code).

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-21 21:50:40 -07:00
Paul BakausandClaude Fable 5 b0a7deb688 serve-question: self-detect headless environments, capability-first routing
The env-var bypass (IMPECCABLE_QUESTION_DISABLED) relied on the harness
remembering to set it. The script now also self-detects CI, SSH-without-
display, and displayless Linux and exits 2 with the structured-question
advice; --no-open skips detection (caller opens the URL itself, as the
tests do) and IMPECCABLE_QUESTION_FORCE=1 overrides it. new-work.md now
frames the decision-page rule by capability: open a browser if you can,
structured question tool if you cannot, exit 2 means fallback not error.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 20:53:37 -07:00
Paul BakausandClaude Fable 5 40ba97db20 Pin @babel/parser so live-mode JSX syntax checks stay enforced
The v4 repo split dropped astro/wrangler from devDependencies, which
also removed the only (transitive) source of @babel/parser. The
post-apply syntax check in live-copy-edit-agent.mjs requires it to flag
invalid JSX/TSX; without it the check silently degrades to a warning and
tests/live-copy-edit-agent.test.mjs "flags invalid JSX syntax" fails on a
fresh CI install. Production behavior is unchanged: the require stays an
optional, graceful-degrade path for end users, and @babel/parser was
never in the published package's runtime dependencies. Declaring it as a
devDependency just makes the repo's own test environment deterministic.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 20:51:55 -07:00
Paul BakausandClaude ceb13f6c8a Let routing's general-work branch honor scoped-refinement directives
Setup step 1 tells the agent to follow context.mjs's directives, and the
no-PRODUCT.md-with-existing-code directive explicitly permits a narrow
refinement to proceed on the incumbent implementation and offer init
afterward. Routing's "Otherwise" branch said missing PRODUCT.md routes
through init, with no carve-out, so the two instructions disagreed on the
same request and the agent could block work context.mjs had cleared.

Rule 3 now splits the way the directive does: a new surface or
replacement world goes through init then new-work, a narrow refinement
proceeds and offers init afterward. Explicit and implied commands were
never affected; they route one rule earlier, which is what skill-behavior
scenario 10 already covers.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-21 20:46:12 -07:00
Paul BakausandClaude 6d47843867 Stop losing DESIGN.md and native platform refs across init
Bugbot flagged the "resume without rerunning context.mjs" instruction
after init. It is right, and the gap is wider than the platform half it
named: context.mjs has two output branches, and the no-PRODUCT.md branch
omits DESIGN.md, the native platform references, and the unrecognized
`## Platform` warning. Because the skill never reruns the script once
init writes PRODUCT.md, whatever that first run withheld is gone for the
whole session. A greenfield iOS project would be designed without
reference/ios.md ever loading, and a project carrying DESIGN.md without
PRODUCT.md never saw its own design system.

The two halves need different fixes. DESIGN.md is authority in its own
right and does not depend on PRODUCT.md existing, so context.mjs now
emits it on both branches. Platform is unknowable before PRODUCT.md
exists, so no change to the script can recover it; init.md, the one step
that learns the answer, now loads ios.md / android.md / both right after
recording a native platform, and SKILL.src.md says so where it tells the
agent not to rerun.

Verified end to end against a temp project on both branches.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-21 20:33:05 -07:00
Paul BakausandClaude 9213e1bcd1 Fix stale section pointer and critique score denominators
Two true positives from the Bugbot review on PR #397.

document.md seed mode told the agent to run "Select one direction" for
paths A, D, or E. new-work.md has neither that heading nor the A/D/E
lettering since the workshop was restructured into named subsections, so
a literal read could skip the world-and-surface flow entirely. Point at
"Create or replace the visual world" and "Commit the world" instead.

critique.md let the heuristic table renormalize to an applicable maximum
when heuristics are scored n/a, but the report template hardcoded ??/40,
the rating bands only mapped raw numbers out of 40, and the persisted
meta carried total_score with no denominator. Trends could silently
compare 24/32 against 30/40 as if they were the same scale. The template
now prints the applicable max, the bands fall back to percentages for
partial sets, the snapshot records max_score and na_heuristics, and the
trend line states its denominator or breaks it out per run when they
disagree. critique-storage.mjs serializes frontmatter key-agnostically,
so the new keys need no code change.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-21 20:07:25 -07:00
Paul BakausandClaude Fable 5 88fe139aa7 Eval-run guards: local card base, question-page disable
IMPECCABLE_CARD_BASE overrides the quality-bar URL prefix so eval
workers serve cards from the local checkout while impeccable.style
stays undeployed. IMPECCABLE_QUESTION_DISABLED makes serve-question
exit immediately with the structured-question fallback line, so
headless workers never block on a browser page nobody will answer.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 19:29:08 -07:00
Paul BakausandClaude Fable 5 311c30f11f Release prep: skill v4.0.0, CLI v3.3.0
Bump skill to 4.0.0 (plugin.json + marketplace.json) and the CLI to
3.3.0 (package.json), then run build:release to regenerate the plugin
subtree and all provider harness output to the new versions.

Skill 4.0.0 ships external-dice direction assignment, the reviewed
world catalog dealt through the roll API with rendered quality-bar
cards, the in-browser serve-question decision page, visualize-before-
build, the rebuilt new-work flow, and the 58-rule detector under hook
enforcement. CLI 3.3.0 grows the deterministic detector to 58 rules
and adds config-declared context roots, per-file rule scoping, and
--target resolution for nested products.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 18:53:45 -07:00
Paul BakausandClaude Fable 5 3c47eb1a8c Claude-conditional counterweight for the warm-subject rendition prior
Gate2 measured it precisely: on matched assignments Opus renders warm,
bookish, and child-facing subjects as cream, serif-italic, and
lamplight while Sol renders the same positions saturated, and neutral
prose hardening did not move it. The codex and gemini blocks set the
precedent for provider-conditional counterweights; this adds the claude
block at the palette decision: the first palette is already spent, an
OWN-WORLD block reading cream/paper/parchment/lamplight for an unpinned
Persuade surface is a failed rendition to rework from the world's
saturated materials, and nothing about the subject requires the
default. Verified present in the claude-code dist and absent from
codex.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 18:42:19 -07:00
Paul Bakaus 06d21dea7d Add first-class Grok Build harness support
Emit .grok skills, agents, and PostToolUse/Stop hooks; wire the CLI
installer and downloads; fix the plugin install path to #plugin; and
document Grok in HARNESSES.md and README.

AI assistance: written with Grok Build.
2026-07-21 18:02:58 -07:00