Commit Graph
272 Commits
Author SHA1 Message Date
Paul BakausandClaude Fable 5 ce4dcf9a93 Split breadth from rating in the challenger and staging pools
Rating grades quality, breadth says whether a world can serve an
arbitrary build at all; while they shared one field, the only way to
hold a narrow world back was calling it marginal, which made excellent
but narrow unrecordable and corrupted the ratings as a calibration
signal for the next authoring round. Both axes now exclude
independently, either kind of hold keeps its approval for direct
briefs, an all-niche tier falls back rather than starving, and
stagings honour the same gate with the same fallback. Tests cover the
niche exclusion at strength, the fallback parity with marginal-only
tiers, and the staging gate.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 19:17:09 -07:00
Paul BakausandClaude Fable 5 33a1c5fcae Ban kickers outright: one eyebrow above a heading is one too many
The detector's repeated-section-kickers rule waited for three tracked
labels before calling the pattern; generated pages earn the finding on
the first one. Retire that id and replace it with kicker-above-heading,
which flags any tracked-caps or small-caps label block sitting directly
above an h1-h4 or heading-role element, at full warning severity.

The candidate gate absorbs the false-positive shapes the repetition
count used to paper over: editorial category-and-date meta lines,
breadcrumbs with separators, legal and chapter numbering, application
panel context labels, nav landmarks before page titles, and stat
callouts with the label below the number. Hero-scale h1 eyebrows stay
with hero-eyebrow-chip so one element gets one finding, and the static
cascade now carries font-variant so small-caps kickers register.

The craft floor entry moves from caution to ban in the same breath.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 18:55:52 -07:00
Paul BakausandClaude Fable 5 6a7d75b6fe Bound the finish by verdict, not by count, and teach the matrix medium and type
The hard stop landed one step early: one review, one batched fix, one
recapture, then done, with nobody ever judging whether the fixes reached
the quality the findings named. A recapture measures positions; the
model then presented mechanical confirmation as artistic success over a
page whose display face, material, and hero legibility had all drifted
from the approved comp. The finish now ends on a verdict: the recaptured
screenshots go back to the same reviewer, which scores every material
fix resolved, partial, or unresolved and names at most three regressions
the batch introduced, no new hunt. Partial and unresolved fixes earn
exactly one more round; two rounds is the ceiling, the second verdict
ends the work whatever it says, and the final verdict table goes to the
user as it stands, open items included.

Three blindnesses from the same run close alongside. The matrix gains
two mandatory rows: TYPE, where a display face of a different character
is contradicted however the layout matches, and MATERIAL, where flat CSS
standing in for painted, textured, or dimensional artwork is contradicted
regardless of placement. And the Truth check now requires every produced
asset visibly present in the screenshots, because a paper texture at
0.16 opacity is a compliance token, not a shipped material.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 18:51:32 -07:00
Paul BakausandClaude Fable 5 09a33bc58b One sketch, one agent: retire the batch producer and its supervision
The batch producer was the clumsy piece: one subagent owning eight
jobs needed heartbeat rules, reclaim windows, and a page full of
fallbacks to survive its own opacity. The unit of work is now a single
card. With parallel subagents, the set fans out one agent per card, up
to four in flight, landing everything in roughly the time of one; a
single-sketch agent has no planning phase and no batch to stall, so a
failure costs one slot and its remedies fit one sentence: regenerate an
empty slot when its agent returns, drop it when the user answers first.
Without parallel subagents, the main thread generates in reading order
after serving, and the harness's own generation display carries the
progress. The page-side streaming is unchanged; it never cared who
writes the files.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 18:19:25 -07:00
Paul BakausandClaude Fable 5 4329f757f5 Only the visible card face is interactive
A hidden backface still hit-tests in Chrome, so after flipping a card
the front's picture-in-picture sat invisibly over the back's chips,
showing its zoom cursor and eating the flip-back click. Pointer events
now follow visibility: the back is inert until the card flips, and the
front goes inert while it is flipped.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 18:13:04 -07:00
Paul BakausandClaude Fable 5 ca88ea008b Patience while sketches land, honesty when standing in
Field data: the first image of a real batch took ninety seconds and the
page's 150-second fallback then silently promoted catalog art to full
bleed, unlabeled, which is exactly the this-is-your-design misread the
picture-in-picture treatment exists to prevent. The policy is now
patience while there is progress: a slot shows its inspiration only
after waiting four minutes with nothing landing anywhere on the page
for four minutes, the stand-in is dimmed and labeled 'inspiration ·
sketch pending', and polling continues so the real sketch still swaps
in whenever it arrives. Slots with no inspiration keep the honest
elapsed shimmer instead of folding. The parent's reclaim rule matches:
files landing steadily is health at any pace, and only total silence,
no first file in three minutes, takes the batch back inline.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 18:10:56 -07:00
Paul BakausandClaude Fable 5 17bc2701f3 Put the full read on the card's back; the front is for choosing
Field feedback: with every fact stacked under the media the cards ran
past a screen tall. The front now carries only what the choice needs,
sketch, lineage, title, thesis, identity, and the honest risk clamped
to two lines, while first viewport and the case read on the back behind
a Details chip, sharing the face with the board when the world has one.
Risk stays on the front because the counterweights are pointless if the
downside hides behind a flip, and once the sketch lands the first
viewport is a picture anyway. The schema notes now ask for one-sentence
facts, since a long fact should cost the reader a flip, not the page
its scanability.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 18:05:41 -07:00
Paul BakausandClaude Fable 5 0eb443d29b Bound the hand, greek the copy, and treat waiting as supervision
A codex field run dealt six challengers into an eight-sketch batch
behind an opaque subagent, and the user stared at a page of shimmer
asking whether anything was happening at all. Four fixes from that run.
A hand now holds at most three challengers, the rest banked for
re-rolls, so fairness within the hand stops multiplying into a queue.
Sketches greek everything but the product's real name and one real
headline, because an invented spec, price, or ship date in a sketch is
a claim PRODUCT.md never made, and comps have solved this for a century.
Sketch production follows the user's reading order with the first file
doubling as the producer's heartbeat, and the parent's --wait loop
checks the sketch directory each pass, reclaiming the batch inline when
two minutes pass with nothing landed. And a failed --start now captures
the daemon's stderr to a per-key log and names the sandbox as the usual
suspect, instead of reporting only that failure occurred. The shimmer
counts its elapsed seconds, and gives up at 150 instead of 300.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 17:49:21 -07:00
Paul BakausandClaude Fable 5 58a2d3dccd Bleed the deck to the viewport, fade the fuller side, let the glance take over
Three field notes from a live review. The deck now escapes the content
column and runs edge to edge, so a cut-off card sits at the screen edge
where it reads as more cards instead of at an invisible container edge
where it reads as a bug; the first card still aligns with the column
via scroll padding. Whichever side hides more content wears a fade, and
a hard edge means the end. The vertical pager grows from a bare chevron
into labeled Back and More pills, because in a column deck it is the
primary way forward. And hovering the inspiration thumb now takes over
the whole media region instead of a timid zoom; the sketch is the
promise, the inspiration is a glance, and the glance must cost nothing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 17:14:59 -07:00
Paul BakausandClaude Fable 5 e6612ea8ef Page the deck on its long axis, and never let decoration hide the cards
Field-checked in a real browser, which surfaced three defects the DOM
tests could not: the generic .media img display rule defeated [hidden]
and floated an empty block over the shimmer and its sketching note; the
deal animation left every card at opacity zero in an unfocused tab,
because rAF throttling is real and decoration must never gate content;
and the sketch poll's cache-busting query missed the anchored /img
route, so a landed sketch kept shimmering forever.

The grid is now a snap-scrolling deck: one row in a wide viewport, one
column in a tall one, with edge arrows that appear only on overflow and
page one card at a time.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 17:10:39 -07:00
Paul BakausandClaude Fable 5 d89ee5f87c Deal every card the same hand: anatomy, sketches, and the standing door
The decision page compared unlike things: the grounded direction was a
wall of text beside curated catalog art, the catalog art read as a
promise of the build, the weighing silently shrank the challenger set,
and the standing exit hid in the footer under the cards it must not
soften. Every card now shares one anatomy (thesis, palette chips,
material tags, first viewport, case, risk), every dealt challenger is
presented with the weighing written on it rather than applied to it,
the catalog image rides picture-in-picture as labeled inspiration with
the lightbox a click away, and canonCard renders the category standard
as one honest, subordinate card.

When image generation exists, each card declares a sketch slot the page
polls: serve first, generate after, through one shared deliberately
unfinished frame, so the comparison stays about direction instead of
rendering luck. The asset producer takes the batch when subagents
exist; the chosen sketch returns in ANSWER to seed at most one comp
probe, and the comp round still renders its full set.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 17:03:05 -07:00
Paul BakausandClaude Fable 5 f482d9405e Teach the reading-heavy subagents to write before the ceiling lands
Raising the reviewer's turn budget did not change its fate, only its
reading: 43 tool uses instead of 22, still reaped mid-read with nothing
written, because the SDK ends a run at max-turns without warning and the
model never feels the deadline. The definitions now carry the deadline
themselves: reading is an allowance, batch Reads per turn, take the
decisive inputs first, sample instead of walking the tree, and write by
mid-budget, naming what went unread. A review built from what you saw
beats a perfect review that never arrives.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 15:51:36 -07:00
Paul BakausandClaude Fable 5 c9213835e7 Review fidelity against the comp itself, not the builder's summary of it
A codex run turned an approved comp into a related second art direction
and the finish reviewer passed it: the review anchored on the direction
contract, a lossy abstraction the builder wrote, and every element that
abstraction dropped passed silently. Four changes close that chain. The
reviewer inventories the comp's salient elements before reading the
contract and classifies each one (match, adaptation, missing,
contradicted, added without approval), with adaptations citing the
answer, brief, accessibility need, or product truth that forced them,
and fidelity failures outranking craft in material_fixes. The visualize
inventory gate records compositional commitments alongside asset media,
since the 150-word contract cannot carry them. The north-star allowance
now says what it permits: translation, never recomposition. And the
finish sequence recaptures the same viewports once after the fix batch,
so what the documenter records is what actually shipped.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 15:13:27 -07:00
Paul BakausandClaude Fable 5 9e4990765f Give the reading-heavy subagents turn budgets that survive their inputs
A finish review reads the artifact, two full-page screenshots, the
approved comp, the quality-bar cards, and the contract before it may
write a word; at max-turns 12 the SDK reaps it mid-read and the parent
receives the opening sentence as the whole review. Observed twice in a
row (spawn and respawn) on the first real subagent run. The documenter
reads at least as much, and the asset producer pays per asset.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 14:50:03 -07:00
Paul BakausandGitHub 9b613ef931 Merge pull request #419 from pbakaus/diff-base-detection
Detect the diff base in context-signals instead of assuming main/master
2026-07-27 10:06:48 -07:00
Paul BakausandClaude Code 01d5d357c5 The develop candidate leads with an advertised develop default rev
Round eight closes the stale-local class completely: the develop
candidate sits before the remote-default entries, so when origin/HEAD
itself points at develop, its name claim let a stale local develop win
over the fresher origin/develop. The candidate now leads with any
remote-advertised develop rev, exactly as the remote-default and
upstream candidates already lead with theirs. main/master were already
covered since their remote-default entries come first in the order.
Failing-first test forces local develop two commits behind.

Prepared with AI assistance (Claude Code), directed by @pbakaus.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-07-26 19:04:48 -07:00
Paul BakausandClaude Code e2c1c43ee7 Remote defaults lead with their own rev, like upstreams already do
Round seven: a remote-advertised default candidate tried the local
branch first, so a stale local main outranked the fresher origin/main
the symref points at and refilled changedFiles with the divergence.
The candidate now leads with the advertised remote rev, mirroring the
upstream candidate's reasoning. Failing-first test: local main forced
two commits behind the remote default, feature delta stays clean.

Prepared with AI assistance (Claude Code), directed by @pbakaus.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-07-26 18:56:33 -07:00
Paul BakausandClaude Code a470fc777a Read the upstream as a full symbolic ref instead of guessing at prefixes
Round six, and the upstream-parsing ambiguity dies at the root: @{u} is
now resolved via rev-parse --symbolic-full-name, where refs/heads/...
IS a local upstream and refs/remotes/<r>/... IS remote-tracking. The
previous remote-membership heuristic still misread a local feature/foo
upstream when a remote literally named "feature" existed. The
adversarial test now configures exactly that remote and passes.

Prepared with AI assistance (Claude Code), directed by @pbakaus.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-07-26 18:48:14 -07:00
Paul BakausandClaude Code f89b6c10b1 Only strip a remote prefix that names a configured remote
Round-five bot findings, one real root cause: splitRemoteRef treated the
first slash in any ref as a remote separator. A local upstream named
release/2.0 was truncated to "2.0", and feature/foo tracking from branch
foo collapsed to the current branch's own name and was self-skipped,
discarding a valid base both times.

The split now happens only when the prefix names a configured remote;
otherwise the whole ref is one local branch name. The per-remote HEAD
symref loop strips its own queried prefix directly (that remote may be
fabricated in tests or partial clones without appearing in git remote).
The reported pruned-upstream shape already resolves via the multi-remote
rev lists from the previous round; its test now guards that.

Prepared with AI assistance (Claude Code), directed by @pbakaus.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-07-26 18:39:20 -07:00
Paul BakausandClaude Code 386d3e7051 Cover every remote in each candidate's rev list
Cursor and Greptile converged on one root cause from the previous round:
candidate revs stopped at origin (develop tried only develop and
origin/develop; a remote-default entry carried only its own rev), so the
name-level dedup discarded a same-name base living on another remote. A
fork-parent layout with develop only as upstream/develop, or a pruned
origin/main beside a live upstream/main, lost its base entirely.

revsFor(name) now expands to the local branch plus <remote>/<name> for
every remote (origin first), and all named candidates use it, which is
exactly what makes the dedup safe. Two failing-first tests cover the
upstream-only develop and the pruned-origin/live-upstream main shapes.

Prepared with AI assistance (Claude Code), directed by @pbakaus.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-07-26 18:23:56 -07:00
Paul BakausandClaude Code 46f29ca8b3 Guard detached HEADs and non-origin remote defaults
Two more real gaps from the post-rebase review round: a detached
checkout reads its branch as the literal HEAD, so the integration guard
never fired and candidate selection could diff a detached tip on main
against develop; and the remote-default check only consulted origin, so
a fork-parent layout whose only remote is upstream lost the guard on
its default branch entirely.

The guard now treats a detached HEAD as no-diff-base, and default-branch
symrefs are collected from every remote (origin first), feeding both the
guard and the candidate list. Two failing-first tests cover a detached
tip beside a diverged develop and an upstream-only trunk default.

Prepared with AI assistance (Claude Code), directed by @pbakaus.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-07-26 18:14:34 -07:00
Paul BakausandClaude Code afb5d9a479 Guard non-standard default branches like conventional ones
Cursor Bugbot: sitting on a non-standard default such as trunk (the
origin/HEAD target) still ran candidate selection, where develop or main
could win and produce an integration-vs-integration diff. The guard now
treats the remote default branch as an integration branch alongside the
conventional names. Failing-first test: on trunk with a develop branch
present, the scope stays the working tree.

Prepared with AI assistance (Claude Code), directed by @pbakaus.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-07-25 20:17:55 -07:00
Paul BakausandClaude Code e82653965c An existing develop outranks a main-pointing origin/HEAD
Cursor Bugbot's remaining round-1 finding held for the current code
too: in a git-flow repo whose platform default was never flipped off
main, a feature branch without an upstream picked origin/HEAD's main
over the develop branch features actually merge to, dragging the
develop-vs-main divergence into scan targets. develop now sits between
the upstream signal and origin/HEAD in the candidate order; repos
without a develop branch are unaffected. Failing-first test covers the
exact shape (develop exists, origin/HEAD -> main).

Prepared with AI assistance (Claude Code), directed by @pbakaus.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-07-25 20:17:55 -07:00
Paul BakausandClaude Code b9d294b29c Close the integration-branch guard bypass; accept local upstreams
Cursor Bugbot round two, both real: an upstream or origin/HEAD naming a
DIFFERENT integration branch bypassed the conventional-name guard, so
sitting on develop with the remote default at main still produced the
integration-vs-integration divergence this detection exists to prevent.
And splitRemoteRef returned null for a slashless @{u}, silently dropping
local upstreams (branch.<x>.remote = ".").

Base detection is now skipped entirely on an integration branch: no
signal may override the working-tree scope there. A slashless upstream
resolves as its own name and rev. Two failing-first tests: origin/HEAD
pointing at main while sitting on develop, and a feature branch
tracking a local canary branch.

Prepared with AI assistance (Claude Code), directed by @pbakaus.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-07-25 20:17:55 -07:00
Paul BakausandClaude Code ea098ceb96 Accept remote refs as diff bases; honor non-origin upstreams
Both review bots found real gaps in the first pass: candidates were
verified as local branch names only, so an origin/HEAD target with no
local checkout fell through, and stripOrigin() dropped upstreams on any
remote not named origin (fork workflows tracking upstream/release).

Candidates now carry a display name plus the revs to try in order: the
upstream's remote rev wins outright (it tracks the actual merge target,
so it beats a possibly stale local branch of the same name), origin/HEAD
tries the local branch then the remote-tracking ref, and the
conventional names each try local then origin/<name>. git.base keeps
reporting the friendly branch name while the diff runs against whichever
rev resolved. Two new failing-first tests: remote-only default branch,
and an upstream on a remote named upstream with no local base branch.

Prepared with AI assistance (Claude Code), directed by @pbakaus.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-07-25 20:17:55 -07:00
Paul BakausandClaude Code a50702f2b6 Detect the diff base instead of assuming main/master
context-signals hardcoded ['main', 'master'] as diff-base candidates, so
repos integrating through develop (or any other branch) diffed against
the wrong base: git.changedFiles carried the entire divergence and
downstream commands scanned the wrong set (issue #302).

The base is now detected, most specific signal first: the branch's
configured upstream (@{u}; a branch pushed with -u tracks itself and is
skipped by the self-check), then the remote's default-branch symref
(origin/HEAD), then the conventional integration names including
develop. The conventional fallbacks are withheld when the current branch
is itself one of them, so sitting on main in a repo that also has
develop keeps the working-tree scope instead of diffing two integration
branches against each other.

Five tests (three failing-first): develop-based feature branch,
origin/HEAD detection with a non-standard default name, upstream
tracking, on-the-integration-branch fallback, and the
integration-vs-integration guard.

Prepared with AI assistance (Claude Code), directed by @pbakaus.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-07-25 20:17:55 -07:00
Paul BakausandClaude Code d0c5558960 Gate the agent_done marker release to carbonize; hedge the failure toast
Cursor Bugbot caught a real hole: accept unlocks at the first variant,
so a late generation agent_done for the same session id could arrive
after Accept and close the awaited failure window early, reopening the
exact #384 gap. The SSE broadcast carries no sourceEventType, so only a
carbonize agent_done is provably accept-side; the release is now gated
on it. Copilot's wording point led somewhere real too: a carbonize-phase
failure raises the same error after the source WAS promoted, so the
toast now says "may not have been saved" and normalizes the server
message's terminal punctuation. Regression guard extended to pin both.

Prepared with AI assistance (Claude Code), directed by @pbakaus.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-07-25 20:15:05 -07:00
Paul BakausandClaude Code f9ea2f0de0 Recognize a late accept failure after the optimistic teardown
Accept is optimistic: POST /events acknowledging the intent schedules
cleanupAcceptedSession(), which nulls pendingAcceptedSession before
live-accept.mjs has run. When the accept later failed (missing markers,
preview error, receipt conflict, source_locked), the SSE 'error' guard
keyed on pendingAcceptedSession could no longer match its id, so the
tailored recovery never fired: the user got a generic error toast, the
session was gone, and nothing said the variant was never written
(issue #384, analysis by Cursor Bugbot on #381).

Following the issue's fix sketch, an awaitingAcceptResult id is set on
the optimistic success path and deliberately survives the teardown. The
'error' case matches it and tells the user plainly that the variant was
not saved and to pick + generate again (post-teardown the wrapper may
already be gone, so restoring CYCLING is not honestly possible). The
marker is released when the real accept result arrives (complete /
accept / post-accept agent_done) or when a new session supersedes it.

Regression guard covers the set-before-teardown ordering, the error
match, and cleanupAcceptedSession leaving the marker alone; the existing
source contract now also asserts handleGo clears it.

Prepared with AI assistance (Claude Code), directed by @pbakaus.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-07-25 20:15:05 -07:00
Paul BakausandClaude Fable 5 9c395bc484 Asset producer: codex notes as standalone blocks the compiler handles
compileProviderBlocks only processes standalone-line blocks, so the
inline codex spans leaked literal tags into every provider's agent
output, degraded fallbacks included.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-25 19:16:17 -07:00
Paul BakausandClaude Fable 5 916b0a1fdf Generate degraded-mode fallback references from the subagent definitions
Harnesses with no subagent capability now run each role inline from the
same single source. The build emits reference/degraded/<role>.md for every
agent in skill/agents/ (role name is the agent name minus the impeccable-
prefix), stripping frontmatter and prepending the inline-substitution
preamble. These pass through the same provider-block compilation and
placeholder replacement as ordinary reference files, so <codex> blocks and
{{placeholders}} resolve per target, and they land in the committed harness
dirs on build:release like every reference file.

Repoint the three capability-first fallback sites in the prose at the
generated files: new-work.md reviewer and documenter fallbacks, and
visualize.md asset-producer fallback.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-25 19:16:17 -07:00
Paul BakausandClaude Fable 5 6769b1879a The polish ceiling covers the whole cycle, and the handoffs end it
Probe attribution on Opus 5 showed the screenshot bound working (42
to 16) while the real burner ran free: five rounds of node -e
micro-edits, eight rebuilds, and inline defect hunts absorbed the
reviewer's and documenter's jobs until the turn cap killed the run
mid-hunt. The two-round ceiling now names scans, micro-edits, and
rebuilds; after the second round the build thread stops polishing and
ships the rest through the reviewer (one batched fix pass, one
rebuild, stop) and the documenter.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-25 19:16:17 -07:00
Paul BakausandGitHub 6bc338a878 Merge pull request #416 from pbakaus/reference-docs-refresh
Refresh stale metric and library references in command docs
2026-07-25 18:41:50 -07:00
Paul BakausandGitHub a4e99eda3a Merge pull request #414 from pbakaus/live-error-clears-checkpoint
Live mode: clear the durable session checkpoint on a terminal SSE error reply
2026-07-25 18:38:39 -07:00
Paul BakausandClaude Code 3d2ffe9007 Drop internal filename cross-references from routed reference text
Copilot's review point stands: reference files load per-command, so a
bare "see optimize.md" / "typeset.md" is not meaningful in the routed
context. The guidance reads self-contained now.

Prepared with AI assistance (Claude Code), directed by @pbakaus.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-07-25 18:25:35 -07:00
Paul BakausandClaude Code 24e24265d1 Refresh stale metric and library references in the command docs
From issue #395, the items still present after the v4 consolidation:

- optimize.md led its interactivity section with FID, retired as a Core
  Web Vital in March 2024 when INP replaced it. The section heading and
  both metric lists now name INP.
- optimize.md recommended react-virtualized, superseded by react-window
  from the same author; the line now points at react-window and TanStack
  Virtual, matching overdrive.md.
- overdrive.md's WebGPU support matrix predated Firefox 141/147 shipping
  it on Windows/macOS and Safari 26 shipping it across Apple platforms.
- audit.md listed "missing will-change" as a defect while animate.md and
  optimize.md both instruct applying it sparingly and never preemptively;
  the audit line now flags overuse instead of absence.
- harden.md allowed 14px mobile body text while typeset.md sets a 16px
  ordinary floor; harden now matches the floor, reserving 14px for
  secondary text, and names the iOS Safari input-zoom consequence.

The issue's other items (Framer Motion naming, Popmotion, polish
duration cap, humor guidance, HSL phrasing in quieter) were already
resolved by the v4 reference rewrite.

Prepared with AI assistance (Claude Code), directed by @pbakaus.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-07-25 18:19:25 -07:00
Paul BakausandClaude Code aeacf55074 Add .vuepress to the hidden source-dir allowlist
Cursor Bugbot correctly noted classic VuePress keeps theme layouts,
components, and styles under .vuepress/, which the walker scanned before
the hidden-dir rule. Same treatment as .vitepress and .storybook.

Prepared with AI assistance (Claude Code), directed by @pbakaus.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-07-25 18:15:13 -07:00
Paul BakausandClaude Code a1a6441ba1 Exempt hidden dirs that conventionally hold UI source from the skip rule
Greptile's review correctly flagged a regression in the blanket
hidden-dir skip: .vitepress/theme/*.vue and .storybook/ preview files are
real UI source that the walker scanned before this branch. Both the
walker and the scan-target filter now carry a two-entry allowlist
(HIDDEN_SOURCE_DIRS) for those conventional locations; every other
hidden dir keeps being skipped.

Prepared with AI assistance (Claude Code), directed by @pbakaus.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-07-25 18:04:55 -07:00
Paul BakausandClaude Code 21d058e744 Clear the durable live-session checkpoint on a terminal SSE error reply
The documented abort flow in reference/live.md (live-poll.mjs --reply <id>
error "...") reset the browser bar to PICKING but left the localStorage
checkpoint written for the GENERATING phase in place. Every reload then
resurrected a dead session the server no longer knew about, and the page
stayed wedged until the user hand-cleared the impeccable-live* keys in
the console (issue #362, diagnosed by @yourcodekitten).

An agent error reply is terminal for the session it names: when the id
matches the current session, run the same markSessionHandled + cleanup
teardown as 'discarded' (cleanup includes clearSession); when it matches
a stored-but-not-current checkpoint (the error raced a reload), drop that
checkpoint too. Errors that name no session keep the existing UI-only
reset, and the accept-cleanup and steer branches are untouched.

Regression guard added to tests/live-browser-regression.test.mjs.

Prepared with AI assistance (Claude Code), directed by @pbakaus.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-07-25 17:57:29 -07:00
Paul BakausandClaude Code 9f008ebf82 Skip hidden dirs in the detector walker and vendored paths in scan targets
When impeccable (or any agent tool) is installed into a project's
.claude/.cursor/.codex tree, a root scan descended into the vendored skill
code and reported the detector's own example strings as findings, and
context-signals returned installed-skill files as scan candidates whenever
the harness tree appeared in the branch diff (issue #303).

Rather than growing SKIP_DIRS by a denylist of harness names that drifts
as new tools appear, the walker now skips every hidden directory during
recursion — which already covered .git/.next/.nuxt/.svelte-kit/.turbo/
.vercel, and covers all present and future harness installs plus
.impeccable itself. SKIP_DIRS shrinks to the four non-hidden entries.
An explicitly passed hidden target still scans: only child entries are
name-checked, never the root the walker is given.

scanTargets() applies the same rule to git-changed files (directory
segments only, so root dotfiles keep their existing behavior), and falls
through to source-dir targeting when the only dirty files are vendored.

Prepared with AI assistance (Claude Code), directed by @pbakaus.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-07-25 17:51:12 -07:00
Paul BakausandClaude Fable 5 8634c538fb Verification is two bounded rounds, never a loop
Opus 5 turned the iterate-with-screenshots-until-it-meets-the-bar
instruction into 42 screenshot trips and 150 tool calls per build,
about forty dollars of cache churn a page, before ever reaching the
reviewer. Verification now batches: one desktop-and-mobile round after
the full build, fixes applied together, one confirming round, ceiling
two. Craft-floor's checks share those renders instead of earning
separate trips; per-tweak iteration is live mode's channel.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-24 18:43:04 -07:00
Paul BakausandClaude Fable 5 bb57be4243 Documenter subagent, reviewer handoff contract, asset gate
From the paired Opus and Codex manual-run analyses. DESIGN.md moves to
the end of the flow and into a shipped documenter subagent that derives
the system from the built artifact: a rulebook written before the build
gets defended against reality, and a half-stable DESIGN.md hands the
design-system detector an unstable target that buries the build in
noise and invites laundering. The finish reviewer gains the handoff
that failed three times live: the parent captures desktop and mobile
screenshots and passes paths, the reviewer never attempts to render
and names missing inputs in one line, the parent verifies the
five-section return and respawns once on empty. Fidelity against the
approved comp joins its checks; the card keeps commitment only. The
comp ingredient inventory becomes a written gate with raster-by-default
materials and no gradient-as-texture, comps persist under
.impeccable/mocks, the degraded seed names the sandboxed-exec cause,
and the finish line is explicit: a clean detector pass is not finished.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-24 17:09:33 -07:00
Paul BakausandClaude Fable 5 253f8e510c Concept machinery: survive truncation, builds, and loud briefs
The release-gate audit traced four ways the roll's output was defeated
downstream of a perfectly healthy seed. Gemini's harness keeps only the
tail of tool output, so the header-only ASSIGNED INDEX never reached
the model in 18 of 18 samples; the seed now restates the assignment
and key at the end of its output. Astro strips frontmatter comments,
so half the anthropic contracts vanished from built artifacts; the
contract now must survive the production build as an HTML comment in
emitted markup. A brief that paints its own picture (the album named
Soft Cathedrals) converged every arm regardless of assigned index; its
literal reading now joins the rut with at most one candidate. And Opus
under 4.0.1 skipped the seed 42% of the time while hand-authoring
plausible contracts; the finish reviewer now verifies FORM carries a
corroborable seed key before any craft point.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-24 15:33:44 -07:00
Paul BakausandClaude Fable 5 ffe869f4d0 Drop the turn-cap exception from the visualize mandate
Paul's call: the build-exhaustion failure only exists inside eval
workers with hard turn budgets no real harness exposes, and the clause
doubled as a hedge door for skipping the comp round. The eval-side fix
belongs in the worker's max-turns, not in skill prose.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-23 11:10:59 -07:00
Paul BakausandClaude Fable 5 e76ff27adf Eval-found fixes: workspace-relative cards, build outranks comps at caps
The release-gate campaign confirmed two skill bugs with transcripts.
Sandboxed harnesses reject absolute paths, so following the CHOSEN
CARD directive with the absolute card-base path failed view_image; the
directive and the quality-bar clause now say download into the
workspace and open the relative path. And under the openai worker's
turn cap, models spent the budget on init, cards, and comp generation
and never built the page (a third of small-n supplement slices); the
visualize mandate gains its one exception: at a hard cap the shipped
page outranks optional imagery.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-23 10:35:22 -07:00
Paul BakausandClaude Fable 5 73dec5d159 Subagent authorization becomes a central harness counter
Paul's call: the reviewer-local authorization patch covered one
command while the harness gate silently disables every shipped
subagent, critique panels and the manual-edit applier included. The
argument now lives beside the autonomy counter in context.mjs, emitted
as tool-result content every run: invoking the skill is the user
request such gates ask for; spawn where a reference directs; the
in-thread substitute is for absent capability only and gets disclosed
in one line. new-work keeps the reviewer mechanics and drops the
now-central argument.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-23 10:35:22 -07:00
Paul BakausandClaude Fable 5 4dc2b4d694 Finish reviewer: the skill invocation authorizes its subagents
A live session on a harness whose guidance gates subagent use on user
request resolved the conflict silently against the skill: it never
spawned the reviewer, stretched the no-subagents fallback to cover
permission hesitancy, and self-reviewed with all the context that made
its choices feel correct. Three tightenings: invoking the skill IS the
user request that authorizes its shipped subagents; the fallback is
for harnesses lacking the capability, not for hesitancy; a substituted
review gets disclosed in one line at finish, never silently.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-23 10:35:22 -07:00
Paul BakausandClaude Fable 5 2fa0e7d327 Live: gate mid-generation source injection, monotonic bar, resumable disconnect
Three browser-side fixes for the same 3.5-to-4.0.1 regression.

- Source-preview targets no longer source-inject per variant_progress
  checkpoint. Immediate injection raced framework (React/Vue) ownership and
  triggered removeChild errors, which surfaced as static previews. HMR now
  owns reconciliation while variants stream in; source injection runs only on
  the final done (its 750ms settle + retry ladder stays for non-HMR harnesses
  like Cursor). Progress counts still advance from the variant observer, and
  the svelte-component progressive path is unchanged.
- The agent-phase progress bar advances monotonically. A behind/resumed
  checkpoint re-broadcasts an earlier phase (the server regresses the snapshot
  phase to generating), which moved the visible bar backward; a phase rank
  table now blocks a known-lower phase from overwriting a known-higher one.
- The server-lost toast now frames the drop as resumable (session saved,
  reopen or restart live-poll.mjs) instead of "Session ended", which had led
  agents to rationalize bailing to direct edits.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 22:49:44 -07:00
Paul BakausandClaude Fable 5 dbe0c12b91 Live: stop the preflight writing source, cache the resolution
The polling-rework preflight wrote the variant scaffold into source during the
poll lease, before the agent acted. On source-preview targets (React/Vue/Vite,
everything but the svelte-component path) that write full-reloaded the
framework; a browser caught mid-reload missed the agent's variant write and the
SSE done, and sat stranded at 0/N.

Restore the 3.5 single-atomic-edit semantics: the preflight still resolves the
element location and computes the scaffold, but --defer-source-write leaves
source untouched and hands the agent the wrapper text plus the picked source
range. The agent splices variants into the wrapper and replaces the range in
one write, so the framework reloads exactly once. The svelte-component path is
untouched (it never writes route source). The missed-completion recovery stays
as defense in depth.

Also cache the resolved source file per target signature (locator + route):
the ~7.6s tree search re-ran on every generate for the same element; a hit now
points the helper straight at the file via --file, invalidated when the target
changes or a resolution fails.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 22:49:44 -07:00
Paul BakausandClaude Fable 5 4cd5ea7547 Add TanStack Router + Start support to live mode
Live mode had no TanStack coverage: a TanStack Start user hit disconnects
and static previews because there is no static index.html to inject and no
adapter for the SSR root document.

- New tanstack-adapter.mjs, modeled on the SvelteKit/Nuxt adapters: detects
  a TanStack Start project (@tanstack/react-start + src/routes/__root.tsx)
  and patches the __root document to mount a generated dev-only React
  component (src/impeccable/ImpeccableLiveRoot) that appends the live bundle
  on the client after hydration, carrying the ?token= param via
  buildLiveScriptSrc. Patch/unpatch round-trips byte-for-byte and is
  idempotent; refuses to clobber an unmanaged file at the component path.
- Wire detection into live-inject.mjs (insert + remove + gitignore),
  ordered so SvelteKit/Nuxt win and a plain TanStack Router SPA falls
  through to the baseline Vite index.html path.
- tanstack-router-vite fixture (baseline, no adapter) and tanstack-start
  fixture (SSR adapter), both with runtime blocks. Both pass the full
  live-e2e cycle (handshake, steer, pick, Go, cycle, accept, carbonize,
  reloadProbe).
- Unit tests for detection + patch round-trip + apply/remove; tanstack-start
  branches in framework-fixtures.test.mjs; live.md framework table + adapter note.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 22:49:44 -07:00
Paul BakausandClaude Fable 5 d4d02b69f2 Live: overlay preview is the verification channel; disconnects resume
Two prose fixes from the 3.5-to-4.0.1 forensic diff of a real user
regression (15-minute tweaks, repeated disconnects, agent abandoning
the picker). The craft-fold made every generate cycle pay the verify-
the-built-result loop the overlay already provides to the human; live
cycles now verify by construction and run the full check once at
accept. And nothing framed a dropped SSE or closed tab as resumable,
while the client toasts "Session ended", so agents rationalized
bailing to direct edits; the journal is canonical and reopening
continues the session.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 22:49:44 -07:00