Transcript archaeology on the failed gallery batch split the binding
break three ways. One model authored a complete, correct contract that
Astro then erased: the compiler strips a slot's leading comment while
keeping deeper ones, so the contract now belongs to the root layout's
body as its first child, and the first production build gets grepped
for the seed key, because a contract the build erased is a contract
nobody can audit. Another model simply skipped the roll and built the
exact category default the seed exists to refuse; the roll step now
states outright that it has no substitute and no skip condition. The
third failure was the worker watchdog, fixed separately.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two probe runs on two different harnesses built complete pages and
declared done without ever entering the finish sequence: the reference
was read once near turn four and the finish choreography had fallen out
of attention thirty turns later. The one text a model rereads on every
edit is its own artifact, so the direction contract now closes with a
FINISH line naming the exit condition verbatim: unreviewed and
undocumented is unfinished; this build ends with the finish review, the
verdict, and DESIGN.md. A page that looks complete with that line
undischarged is not done, it is abandoned at the finish line.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A sandboxed harness rejected view_image on an absolute path to a mock
the model had itself just produced under .impeccable/mocks/, killing
the run. The relative-path rule existed only for downloaded quality-bar
cards; it now covers every image the flow produces or references, in
the comp round and in the asset producer's comparison step.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
cursor[bot]: discoverAppCandidates only matched dev-config markers while
the upward walk also honors an existing .impeccable/live/config.json,
so booting from a repo root without --target missed a nested
live-configured static site and fell through to the wrong root. Both
paths now share isAppRoot; regression test covers the static-site shape.
This work was produced with AI assistance (Claude Code).
Co-Authored-By: Claude Code <noreply@anthropic.com>
cursor[bot]: the background-terminal notify regex predated
variant_mount_failed (and manual_edit_apply / prefetch), so on Cursor a
failed mount exited the one-shot poll without waking the agent and the
error card sat unanswered. The pattern now lists every type the
dispatch loop handles.
This work was produced with AI assistance (Claude Code).
Co-Authored-By: Claude Code <noreply@anthropic.com>
greptile-apps[bot] re-raised the residual with a repro: a stale
server.json pid reused by an unrelated node process passed the
command-name check. The decisive signal is the recorded PORT: a real
helper is listening on it, a pid squatter is not. hasLiveServer now
probes 127.0.0.1:<port> (bash /dev/tcp, sync, ~ms, win32-guarded with
the previous behavior); the multi-app preference test runs a real
listener instead of faking liveness with a bare pid.
This work was produced with AI assistance (Claude Code).
Co-Authored-By: Claude Code <noreply@anthropic.com>
Field feedback from two more Codex sessions drove both changes.
JIT instructions (live/instructions.mjs): every event live-poll prints
now carries _instructions, the authoritative next step for that exact
situation with real ids, paths, and line numbers substituted, and only
the active path's rules (a svelte-component session never sees JSX
guidance). The boot payload carries loop instructions the same way.
Instructions are versioned with the scripts, so they cannot drift from
behavior, and live.md's plumbing can keep shrinking toward contract plus
craft guidance. The Codex poll-discipline failure observed in the field
("the long poll was started, but I yielded the task instead of actively
servicing its result") gets a named anti-pattern in both the harness
policy and the boot instructions.
LLM e2e agent: default provider/model moves from Claude Haiku 4.5 to
OpenAI gpt-5.6-terra at medium reasoning effort via an Anthropic-shaped
shim over the ai SDK (the three call sites stay provider-agnostic;
Anthropic and DeepSeek remain selectable). The harness should exercise
the model tier that actually drives live sessions. Both the react and
sveltekit fixtures pass end to end with terra driving the trimmed
live.md and the new _instructions.
This work was produced with AI assistance (Claude Code).
Co-Authored-By: Claude Code <noreply@anthropic.com>
Rating grades quality, breadth says whether a world can serve an
arbitrary build at all; while they shared one field, the only way to
hold a narrow world back was calling it marginal, which made excellent
but narrow unrecordable and corrupted the ratings as a calibration
signal for the next authoring round. Both axes now exclude
independently, either kind of hold keeps its approval for direct
briefs, an all-niche tier falls back rather than starving, and
stagings honour the same gate with the same fallback. Tests cover the
niche exclusion at strength, the fallback parity with marginal-only
tiers, and the staging gate.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
First-time setup (config schema, framework table, adapters, drift, the
whole CSP flow) moves to reference/live-setup.md, loaded only when the
boot reports config_missing/config_invalid or cspChecked is absent.
The per-session prose is compressed without dropping any pinned phrase,
MUST rule, schema, or example; the boot payload documentation now names
the inlined surface brief. All live-reference pins and both prose gates
pass.
This work was produced with AI assistance (Claude Code).
Co-Authored-By: Claude Code <noreply@anthropic.com>
The detector's repeated-section-kickers rule waited for three tracked
labels before calling the pattern; generated pages earn the finding on
the first one. Retire that id and replace it with kicker-above-heading,
which flags any tracked-caps or small-caps label block sitting directly
above an h1-h4 or heading-role element, at full warning severity.
The candidate gate absorbs the false-positive shapes the repetition
count used to paper over: editorial category-and-date meta lines,
breadcrumbs with separators, legal and chapter numbering, application
panel context labels, nav landmarks before page titles, and stat
callouts with the label below the number. Hero-scale h1 eyebrows stay
with hero-eyebrow-chip so one element gets one finding, and the static
cascade now carries font-variant so small-caps kickers register.
The craft floor entry moves from caution to ban in the same breath.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The hard stop landed one step early: one review, one batched fix, one
recapture, then done, with nobody ever judging whether the fixes reached
the quality the findings named. A recapture measures positions; the
model then presented mechanical confirmation as artistic success over a
page whose display face, material, and hero legibility had all drifted
from the approved comp. The finish now ends on a verdict: the recaptured
screenshots go back to the same reviewer, which scores every material
fix resolved, partial, or unresolved and names at most three regressions
the batch introduced, no new hunt. Partial and unresolved fixes earn
exactly one more round; two rounds is the ceiling, the second verdict
ends the work whatever it says, and the final verdict table goes to the
user as it stands, open items included.
Three blindnesses from the same run close alongside. The matrix gains
two mandatory rows: TYPE, where a display face of a different character
is contradicted however the layout matches, and MATERIAL, where flat CSS
standing in for painted, textured, or dimensional artwork is contradicted
regardless of placement. And the Truth check now requires every produced
asset visibly present in the screenshots, because a paper texture at
0.16 opacity is a compliance token, not a shipped material.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Field failure from a real Codex session: accepting a variant into
Pitch.svelte appended 23 selectors and removed none, so the source's old
.decisions grid rules re-attached through the kept root class and forced
the accepted board into a stale three-column layout; some appended base
rules also landed after the source's media block, weakening the mobile
cascade.
Two mechanical fixes:
- Preview truth: the scaffolder records the seeded selectors (the source
rules that styled the replaced selection, which the isolated preview
never applied). On accept, any seeded selector the variant does not
re-declare is removed; the selector-loss postcondition treats those
removals like compiler prunes. A regression test reproduces the exact
Pitch shape end to end.
- Cascade order: reconciliation inserts new base rules BEFORE existing
top-level media blocks instead of appending after them.
Init-latency reductions from the same transcript:
- live.mjs inlines the resolved surface brief (removes three
surface-brief.mjs round-trips including a --help miss before first poll).
- The wrap/scaffold payload carries componentStubMarkup, and live.md
instructs editing stubs in place (the session read the manifest + stub
back and then deleted/recreated the files).
- live.md notes that a busy default port usually means the dev server is
already running (the session spawned a duplicate).
This work was produced with AI assistance (Claude Code).
Co-Authored-By: Claude Code <noreply@anthropic.com>
The batch producer was the clumsy piece: one subagent owning eight
jobs needed heartbeat rules, reclaim windows, and a page full of
fallbacks to survive its own opacity. The unit of work is now a single
card. With parallel subagents, the set fans out one agent per card, up
to four in flight, landing everything in roughly the time of one; a
single-sketch agent has no planning phase and no batch to stall, so a
failure costs one slot and its remedies fit one sentence: regenerate an
empty slot when its agent returns, drop it when the user answers first.
Without parallel subagents, the main thread generates in reading order
after serving, and the harness's own generation display carries the
progress. The page-side streaming is unchanged; it never cared who
writes the files.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A hidden backface still hit-tests in Chrome, so after flipping a card
the front's picture-in-picture sat invisibly over the back's chips,
showing its zoom cursor and eating the flip-back click. Pointer events
now follow visibility: the back is inert until the card flips, and the
front goes inert while it is flipped.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Field data: the first image of a real batch took ninety seconds and the
page's 150-second fallback then silently promoted catalog art to full
bleed, unlabeled, which is exactly the this-is-your-design misread the
picture-in-picture treatment exists to prevent. The policy is now
patience while there is progress: a slot shows its inspiration only
after waiting four minutes with nothing landing anywhere on the page
for four minutes, the stand-in is dimmed and labeled 'inspiration ·
sketch pending', and polling continues so the real sketch still swaps
in whenever it arrives. Slots with no inspiration keep the honest
elapsed shimmer instead of folding. The parent's reclaim rule matches:
files landing steadily is health at any pace, and only total silence,
no first file in three minutes, takes the batch back inline.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Field feedback: with every fact stacked under the media the cards ran
past a screen tall. The front now carries only what the choice needs,
sketch, lineage, title, thesis, identity, and the honest risk clamped
to two lines, while first viewport and the case read on the back behind
a Details chip, sharing the face with the board when the world has one.
Risk stays on the front because the counterweights are pointless if the
downside hides behind a flip, and once the sketch lands the first
viewport is a picture anyway. The schema notes now ask for one-sentence
facts, since a long fact should cost the reader a flip, not the page
its scanability.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A codex field run dealt six challengers into an eight-sketch batch
behind an opaque subagent, and the user stared at a page of shimmer
asking whether anything was happening at all. Four fixes from that run.
A hand now holds at most three challengers, the rest banked for
re-rolls, so fairness within the hand stops multiplying into a queue.
Sketches greek everything but the product's real name and one real
headline, because an invented spec, price, or ship date in a sketch is
a claim PRODUCT.md never made, and comps have solved this for a century.
Sketch production follows the user's reading order with the first file
doubling as the producer's heartbeat, and the parent's --wait loop
checks the sketch directory each pass, reclaiming the batch inline when
two minutes pass with nothing landed. And a failed --start now captures
the daemon's stderr to a per-key log and names the sandbox as the usual
suspect, instead of reporting only that failure occurred. The shimmer
counts its elapsed seconds, and gives up at 150 instead of 300.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Three field notes from a live review. The deck now escapes the content
column and runs edge to edge, so a cut-off card sits at the screen edge
where it reads as more cards instead of at an invisible container edge
where it reads as a bug; the first card still aligns with the column
via scroll padding. Whichever side hides more content wears a fade, and
a hard edge means the end. The vertical pager grows from a bare chevron
into labeled Back and More pills, because in a column deck it is the
primary way forward. And hovering the inspiration thumb now takes over
the whole media region instead of a timid zoom; the sketch is the
promise, the inspiration is a glance, and the glance must cost nothing.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Field-checked in a real browser, which surfaced three defects the DOM
tests could not: the generic .media img display rule defeated [hidden]
and floated an empty block over the shimmer and its sketching note; the
deal animation left every card at opacity zero in an unfocused tab,
because rAF throttling is real and decoration must never gate content;
and the sketch poll's cache-busting query missed the anchored /img
route, so a landed sketch kept shimmering forever.
The grid is now a snap-scrolling deck: one row in a wide viewport, one
column in a tall one, with edge arrows that appear only on overflow and
page one card at a time.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
greptile-apps[bot] repro: a helper that died without removing
server.json leaves a pid the OS can hand to an unrelated process, which
kill(pid, 0) classifies as a running server and routes repo-root helpers
onto the stale app. The liveness check now also requires the pid's
command line to look like a node process (ps-based, platform-guarded),
removing reuse by arbitrary processes; the residual node-reuse case is
covered by the multi-app warning and the --target escape hatch.
This work was produced with AI assistance (Claude Code).
Co-Authored-By: Claude Code <noreply@anthropic.com>
The decision page compared unlike things: the grounded direction was a
wall of text beside curated catalog art, the catalog art read as a
promise of the build, the weighing silently shrank the challenger set,
and the standing exit hid in the footer under the cards it must not
soften. Every card now shares one anatomy (thesis, palette chips,
material tags, first viewport, case, risk), every dealt challenger is
presented with the weighing written on it rather than applied to it,
the catalog image rides picture-in-picture as labeled inspiration with
the lightbox a click away, and canonCard renders the category standard
as one honest, subordinate card.
When image generation exists, each card declares a sketch slot the page
polls: serve first, generate after, through one shared deliberately
unfinished frame, so the comparison stays about direction instead of
rendering luck. The asset producer takes the batch when subagents
exist; the chosen sketch returns in ANSWER to seed at most one comp
probe, and the comp round still renders its full set.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
cursor[bot]: the schedule run shared github.ref with pushes to main, so
cancel-in-progress let the nightly full matrix and a main push cancel
each other. Scheduled runs now use a dedicated group.
This work was produced with AI assistance (Claude Code).
Co-Authored-By: Claude Code <noreply@anthropic.com>
greptile-apps[bot] repro: the multi-app warning recommended --target,
but the helper CLIs never parsed it, so live-poll --target appB still
re-anchored onto the pointer's first choice. enterLiveRoot now consumes
a --target argument (removing it from argv so downstream flag parsers
never see it) and resolves roots against it, making the documented
escape hatch real on every helper. Regression test drives a two-live-app
repo through a child process and asserts both the chdir target and the
argv scrubbing.
This work was produced with AI assistance (Claude Code).
Co-Authored-By: Claude Code <noreply@anthropic.com>
cursor[bot]:
- verifyAcceptedSource anchors its param patterns to the exact shapes
live mode writes (data-p-x= / [data-p-x] attributes, var(--p-x, ...)
references) instead of bare prefixes, shrinking the false-positive
class near the completion gate. Note: the reported examples (data-page,
var(--primary)) did not actually match the previous hyphenated
substrings; the tightening removes the residual class (e.g. a user's
own data-p-* attribute) regardless.
- With a non-root Vite base, the /@fs/ fallback is tried both under the
base and at the server root, covering Vite versions that serve @fs at
either location.
This work was produced with AI assistance (Claude Code).
Co-Authored-By: Claude Code <noreply@anthropic.com>
cursor[bot]:
- variant_mount_failed now sets the session's pendingEvent (without
clobbering a still-pending generate), so a helper restart replays it
onto /poll and a repair --reply resolves instead of returning
unknown_poll_reply_id. live-resume's next action names the real event
id instead of a literal EVENT_ID placeholder.
- Contract v2 text hydration strips {#key} DELIMITERS from the zip
source (content stays; it always renders), so key blocks can no longer
shift expression slots against the live DOM.
This work was produced with AI assistance (Claude Code).
Co-Authored-By: Claude Code <noreply@anthropic.com>
cursor[bot]:
- enqueueEvent dedupes variant_mount_failed per variant, so a second
broken variant is no longer swallowed while the first is queued.
- Every component (re)injection resets the mount-failure dedupe, so a
republish that is still broken at the same URL reports again instead
of silently convincing the agent the repair landed.
- Toggle baking now mirrors preview truth exactly: the runtime sets
data-p-<id>="on" or removes the attribute, so presence and "on" forms
survive only while on, and any other valued branch (never matched at
preview) is dropped in either state.
greptile-apps[bot] (both P1 repros):
- When several apps qualify at the same resolution tier (two live
servers, or two stopped apps with interrupted sessions), the choice
stays deterministic but is now loud: a stderr warning names the chosen
app, the alternatives, and how to target a specific app. Silent
wrong-app routing was the failure in both repro harnesses.
This work was produced with AI assistance (Claude Code).
Co-Authored-By: Claude Code <noreply@anthropic.com>
The CI-only astro accept hang: the carbonize source edit triggers a
framework reload, and on a slow runner the reloaded page rehydrated the
still-non-terminal carbonize_required session back into GENERATING,
stranding the bar over a decided comparison. Adoption now uses a
positive allowlist of comparison phases (generate_requested,
variants_ready, generating, cycling); accept/carbonize/steer/manual
phases are agent-side work and never adoptable. Regression guard pins
the allowlist.
This work was produced with AI assistance (Claude Code).
Co-Authored-By: Claude Code <noreply@anthropic.com>
cursor[bot]:
- variant_mount_failed joins EVENT_TYPES_NEEDING_AGENT_REPLY so stream
mode waits for the repair reply instead of moving on mid-lease.
- The fake agent's mount-failure repair no longer forces
sourceEventType generate; the server maps the done reply onto the
pending failure event, which acknowledges it instead of leaving it to
be redelivered on every poll.
greptile-apps[bot]:
- With every helper server stopped, repo-root resolution now prefers the
app whose durable store holds a non-terminal session (the interrupted
session the user is recovering) over the most recent boot.
astro-vite7 (pre-existing CI failure, root-caused): Astro 7 auto-detects
AI-agent environments and daemonizes `astro dev`; the detached server
holds a lock, outlives the harness, squats dev ports across runs, and
makes the parent exit 0, which the harness read as a crash. The fixture
now sets ASTRO_DEV_BACKGROUND=1 (disables the agent detection) plus
--ignore-lock, and the harness supports per-fixture runtime.env. The
core cycle now passes for the first time; the missed-done recovery
scenario fails identically at origin/main with the daemon bypassed, so
it is marked as a per-scenario known limitation with that rationale.
This work was produced with AI assistance (Claude Code).
Co-Authored-By: Claude Code <noreply@anthropic.com>
Raising the reviewer's turn budget did not change its fate, only its
reading: 43 tool uses instead of 22, still reaped mid-read with nothing
written, because the SDK ends a run at max-turns without warning and the
model never feels the deadline. The definitions now carry the deadline
themselves: reading is an allowance, batch Reads per turn, take the
decisive inputs first, sample instead of walking the tree, and write by
mid-budget, naming what went unread. A review built from what you saw
beats a perfect review that never arrives.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
cursor[bot]:
- style: directives with dynamic values now fall back to source-preview
instead of being scaffolded as boolean condition props that falsified
the style in the detached preview.
- class: directives carry a className probe, so v2 hydration answers the
condition from the live DOM instead of always defaulting to false.
- The existing-wrapper remount path now checks the mount result; a failed
remount keeps the error card instead of advancing to a CYCLING bar over
a page where nothing rendered.
greptile-apps[bot]:
- The repo-root live pointer records every booted app (most recent
first) and resolution prefers the app whose helper server is alive, so
a helper run from the repo root of a two-app monorepo can no longer be
redirected onto the wrong app's session store by the last boot. Legacy
single-value pointers still read.
This work was produced with AI assistance (Claude Code).
Co-Authored-By: Claude Code <noreply@anthropic.com>
cursor[bot] findings on #433:
- Nightly schedule no longer enables the paid opt-in suites: a schedule
event has no diff base, so the change-detection fallback flagged every
file-triggered suite, which would have billed the skill-behavior,
accept-cleanup, and deepseek LLM suites nightly. The plan now pins the
schedule event to deterministic suites plus the full live-e2e matrix,
with a regression test.
- Dismissing the mount-error card no longer strands the session: while
the bar is hidden in GENERATING the card is the only recovery surface,
so dismiss now returns the state machine to PICKING (session and
server truth survive for a later republish).
This work was produced with AI assistance (Claude Code).
Co-Authored-By: Claude Code <noreply@anthropic.com>