The detector-blind slop review existed because the AI-tell rules had been
stripped out of SKILL.md and nothing carried them. The floor is a better
home: it loads after concept ideation and immediately before editing UI,
which is the placement that made stripping them necessary in the first
place. Models tread lightly when a ban list is present during ideation;
by the time the floor loads, the direction is already committed.
- Rename build-floor.md to craft-floor.md and restore the absolute bans
(side-stripes, gradient text, glassmorphism, hero-metric, identical card
grids, eyebrow-on-every-section, numbered markers, text overflow), the
codex and gemini defect lists, and the reflexes no scanner catches.
Rule ids match the ones the ablation catalog already knows.
- Delete lib/slop-review.mjs and both injections. The Stop hook is now
purely a mechanical pass and stays silent with nothing to report.
- context.mjs replaces AI_SLOP_REVIEW_REQUIRED with the narrower
MANUAL_DETECTOR_REQUIRED, emitted only when a session has no hook at
all. A per-edit hook already covers the mechanical gap, and the floor
covers the judgment one either way.
Co-Authored-By: Claude <noreply@anthropic.com>
main still carries the site, so every `site/` path resolves to deleted.
`tests/docs-integrity.test.js` goes with it (it imports the site's demo
renderer), and `package.json` keeps main's `@anthropic-ai/sdk` bump while
dropping `@google/genai` and `@paper-design/shaders`, which nothing in the
product layer imports.
Real code merges:
- hook-lib: main's #391 cache fix (sync the remembered set to the live
scan so fixed findings stop being named and a reintroduced one fires
again) now runs on the immediate tier rather than the whole filtered
set. Remembering a deferred finding the per-edit pass never reported
would let the Stop deep pass dedupe it away. main's `maxFileBytes`
ceiling, `cleanAcked` once-per-file ack, and template-extensions
re-export all land alongside the tiering work.
- live-browser: main's `hasParams` gate on the Tune badge, keeping this
branch's `C.ink` badge text so it stays legible on kinpaku gold.
- detect-text: both the block-level codex-grid-background scan and main's
inset-stripe CSS check.
- test-suites: union of both trigger sets and file lists, minus the
site-only entries (`shiki-theme`, `docs-integrity`).
- Two hook tests moved off deferred-tier rules (`overused-font`,
`side-tab`) onto immediate-tier ones. They assert cache bookkeeping,
which the per-edit pass only reaches for the immediate tier.
Also drops the site waivers from `.impeccable/config.json` and stops
`build:browser` recreating a stray `site/` tree just to write a bundle
the other repo builds itself.
Co-Authored-By: Claude <noreply@anthropic.com>
A single staging input was too weak a counterweight to the model's
habitual page skeleton: beside six identity challengers it read as one
optional flourish rather than a real search over composition. Roll three
from distinct staging families so a roll tests materially different
hierarchy, sequence, and interaction laws.
selectApprovedStagings replaces the single-pick selector; the old
selectApprovedStaging stays as a count-1 wrapper for smoke tests. Re-rolls
exclude every earlier set, and an absent mode still returns nothing rather
than falling back across modes.
Co-Authored-By: Claude <noreply@anthropic.com>
The public repo keeps the OSS promise surface: skill, CLI, extension,
tests, and the provider build. The site, labs, concept/composition
catalogs, image pipeline, Cloudflare functions, and authoring guide move
to pbakaus/impeccable-site.
concept-seed tests run against a synthetic fixture catalog; the plugin
icon and skill categories moved in-repo; build validation narrows to
README prose and non-site counts; release notes read from a sibling
impeccable-site checkout.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
/api/roll deals deterministic challenger rolls server-side (same salts and
sha256 ranking as the local seed, verified bit-for-bit); the request log is
the impression record. /api/chosen takes the anonymous choice ping. Events
land in Workers Analytics Engine.
concept-seed.mjs resolves data in order: local catalog dir, roll API,
degraded promotion-only seed. --chosen sends the choice ping; DO_NOT_TRACK
and IMPECCABLE_NO_TELEMETRY disable it. API-dealt seeds carry the telemetry
instruction inline.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Allow context roots to be declared in .impeccable/config.json
Monorepo detection previously read workspace roots only from package
managers (package.json workspaces, pnpm-workspace.yaml, lerna.json),
coupling "where design context lives" to the dependency graph. Add a
`contextRoots` glob list to .impeccable/config.json / config.local.json
so non-JS repos -- and design-context boundaries that don't match
packages -- can declare nested PRODUCT.md/DESIGN.md roots directly.
The new source is folded into readWorkspacePatterns(), so detection,
project resolution, and the app picker pick it up unchanged. Negation
and config.local.json extension work for free.
* Define projectRoots composition with package workspaces
Address review feedback on #307:
- Rename the config key contextRoots -> projectRoots: the globs establish
project boundaries and app-picker targets, not just where context files
live.
- Make cross-source precedence explicit: a path matched by any impeccable
pattern, positive or negated, is governed by the impeccable group alone;
package-manager patterns fill in the paths it does not match, and `!`
negations apply only within their own source. readWorkspacePatterns()
becomes readProjectPatternGroups() / readProjectPatterns(), with package
workspaces as one discovery source.
- Drop app-picker candidates that would resolve elsewhere: a package
workspace subsumed by a broader impeccable boundary is no longer listed,
since choosing it would silently resolve to that boundary.
- Add five composition tests and document the key in the config and
context reference pages (path relativity, glob and negation syntax,
shared/local merge, precedence).
* Fix Live accept for Elixir templates in lib/
Wrap and accept search the repo for impeccable variant markers. That
search skipped .ex files and the lib/ tree, so Phoenix LiveView markup
inside ~H""" blocks never matched and browser Accept returned
"Session markers not found".
Extend the same EXTENSIONS and searchDirs in live-accept.mjs and
live-wrap.mjs. Add a regression test that accepts from
lib/my_app_web/components/layouts.ex.
* Live: give the source search one owner for template extensions
The #374 fix had to patch the same hardcoded EXTENSIONS array in two
files because live-wrap.mjs and live-accept.mjs each carried their own
copy of the project source walk. The copies had already drifted: same
extension list twice, same searchDirs twice, and one realpathSync
guarded by try/catch while the other was not.
Meanwhile hook-lib.mjs had solved this properly for the design hook in
#316/#347 with a configurable `detector.extensions` and suffix matching
that handles .blade.php and .html.erb. Live never read it, so a project
that taught the hook about .heex still got 'Session markers not found'
on Accept.
- lib/template-extensions.mjs is the single owner. It holds Live's
built-in markup list, the suffix matcher, and the detector.extensions
config reader. hook-lib.mjs now imports its normalize/merge/match
helpers from here instead of duplicating them, and re-exports
matchConfiguredExtension for its existing callers.
- Live resolves built-ins PLUS detector.extensions, so teaching the hook
about a server template teaches wrap and accept at the same time.
- live/source-search.mjs holds the walk both scripts share. Callers pass
the one thing that actually differs (skipDirs, fileFilter). Unifying
gives live-wrap the guarded realpathSync, so a dangling symlink in the
tree no longer throws out of the whole wrap, and makes it skip
.impeccable artifacts the way accept already did.
- Extensions are matched on filename suffix rather than path.extname, so
root.html.heex and show.html.erb resolve.
- Drop .exs. Those are Elixir scripts (mix.exs, config/*.exs), never
markup, and including them only lets a wrap query match build config.
- Fill the Elixir gap in the manual-edit paths, which kept their own
allowlists and would have left Live half-working for Phoenix:
live-commit-manual-edits.mjs and live-manual-edit-evidence.mjs.
Verified the round trip by hand against a Phoenix layout: wrap injects
markers into a ~H""" block in lib/**/*.ex, accept carbonizes the chosen
variant back out.
AI assistance: written with Claude Code.
Co-Authored-By: Claude <noreply@anthropic.com>
---------
Co-authored-by: Nils Kanevad <heliumbrain@users.noreply.github.com>
Co-authored-by: Paul Bakaus <paul.bakaus@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
* Stop the design hook lying about findings it already reported
Three fixes, all aimed at the hook being trustworthy enough that an agent
keeps reading it.
1. The session cache was append-only, so the hook lied and then went blind.
`rememberFindings` unioned new keys into the remembered set and nothing ever
removed them, and the pending ack took its count from that set rather than
from the live scan. Fixing two of three findings produced:
Still has 3 finding(s) flagged earlier this session
(overused-font:1:inter, overused-font:2:roboto, overused-font:3:geist)
with roboto and geist already gone. Worse, a finding that was fixed and then
reintroduced was deduped against the stale memory and never re-reported, so
the hook was permanently blind to that regression for the rest of the session.
The cache now syncs to the complete current scan on every scan, so the count
shrinks as work lands and a reintroduced finding reads as fresh. Dedup within
a session still works, because it compares against the previous scan rather
than against all history. A detector failure leaves the remembered set alone
instead of recording an empty scan as truth.
2. The size ceiling, for generated files that do not live under dist/.
`GENERATED_PATH` covered dist, build, out, .next, .cache, coverage and
.min., but repos commit browser bundles and vendored detector copies next to
source. The hook was reading and scanning a 215KB generated bundle, and
reporting findings in it. Added `generated` as a path segment, matched with
separators on both sides so authored names such as generated-utils.ts and
CodeGenerator.tsx still get scanned, plus a `limits.maxFileBytes` ceiling
defaulting to 128KB. In this codebase authored files top out at 86KB while
the bundles start at 215KB, so the gap is comfortable.
3. The clean ack repeated on every clean edit.
It carries no finding, only the standing steer that a silent hook is not a
verdict on the design. That steer is worth saying, but not dozens of times
per session. It now fires once per file per session and reports
`clean-ack-deduped` in the audit log so suppressed noise stays visible. The
pending ack is deliberately untouched: it names real unresolved work, and the
comment explaining why it must repeat still holds.
Verified end-to-end against the built hook: three findings, fix two and the
count drops to one naming only the survivor, fix the last and it goes clean,
edit again and it stays silent, reintroduce and it fires as fresh.
Generated provider output is deliberately left out; the sync workflow owns it.
Prepared with AI assistance (Claude Code).
Co-Authored-By: Claude <noreply@anthropic.com>
* Address review: three clean-ack and audit bugs in the dedupe change
All three were introduced by this PR and all three are fair catches.
Quiet mode spent the ack (bugbot). A clean scan marked cleanAcked and
persisted it even when quiet suppressed all output, so a later non-quiet run
in the same session never got the steer. The quiet decision is now hoisted
above the scan loop and quiet leaves the ack unspent.
Multi-file events lost the ack (copilot). The first clean target became
cleanWinner unconditionally; if that file was already acked, cleanAckDeduped
went true and the `!cleanWinner` guard meant a later target that had never
been acked could never win. A raw apply_patch touching two files would drop
the second file's ack entirely. The loop now keeps looking for a target that
is actually owed an ack.
audit.bytes leaked across targets (copilot). It was set when a file was
skipped as too-large and never cleared, so in a multi-file event a later
emitted result carried the skipped file's byte count. Cleared per iteration.
The tests use a raw apply_patch payload rather than MultiEdit, because
MultiEdit in this harness is single-file ({ file_path, edits: [] }) and would
not have exercised the multi-target paths at all. Verified the three tests
fail against the pre-fix code and pass after, so they are not passing for the
wrong reason.
Prepared with AI assistance (Claude Code).
Co-Authored-By: Claude <noreply@anthropic.com>
* Address review: font-size waivers silently did nothing
Two more review findings, both real.
Specific-value font-size waivers were dead config (greptile). The rule emits an
ignoreValue, and the hook's own directive footer tells the agent to waive
value-specific findings with `hooks ignore-value <rule> <value>`, but
`design-system-font-size` was missing from the direct-value rule set in
`extractFindingIgnoreValue`. The extracted value came back empty, so any
waiver naming an actual size was compared against nothing and silently
dropped. Only the `*` wildcard worked, which is why the framework-viz waiver
earlier in this branch appeared to function.
Reproduced against the built hook: with a `0.82rem` waiver the finding still
fired; it now goes clean, while a waiver naming a different size correctly
still fires, so this is not over-matching.
Wrong audit skip reason (bugbot). In a mixed multi-target run, an earlier UI
file whose ack was already spent set `cleanAckDeduped`, and a later non-UI
clean file became the winner. The tail then reported `clean-ack-deduped` when
the honest reason was `non-ui-ack`. Audit-label only, no behavior change.
Reordered so the winner is described first and dedupe is reported only when it
is genuinely why nothing was emitted.
Prepared with AI assistance (Claude Code).
Co-Authored-By: Claude <noreply@anthropic.com>
* Mirror the font-size waiver fix into the CLI's config reader
Bugbot caught that the previous commit only fixed one of two copies.
`extractFindingIgnoreValue` exists twice, in skill/scripts/hook-lib.mjs and in
cli/lib/impeccable-config.mjs, and the direct-value rule list is duplicated in
both. Adding design-system-font-size to the hook alone meant the same
.impeccable/config.json filtered differently depending on the entry point: a
size waiver was honored by the hook and ignored by `npx impeccable detect`.
The two functions are otherwise byte-identical, so this restores parity rather
than changing CLI behavior independently. The new test notes the duplication so
the next person knows the pair has drifted once already.
Prepared with AI assistance (Claude Code).
Co-Authored-By: Claude <noreply@anthropic.com>
* Fix the audit byte-count leak properly, not just one scan order
My earlier fix cleared audit.bytes at the top of each iteration, which was
wrong twice over, and bugbot caught both.
The clear sat below the sensitive, generated, extension, ignore-file and
file-missing continues, so a later target exiting through any of those never
reached it and kept the oversized file's size while audit.file pointed
somewhere else. It also only handled the bundle-scanned-first order; when the
oversized file came last, the byte count was set after the emitting file had
already been decided and rode along on its audit entry regardless.
The root problem was keeping per-file state on the shared audit object. The
size is now held in a local and attached only when the oversized skip is the
run's actual outcome, so it cannot describe a file other than the one being
reported. Tests cover both scan orders, an early-continue target after the
skip, and the single-oversized-file case where the count should still appear.
Prepared with AI assistance (Claude Code).
Co-Authored-By: Claude <noreply@anthropic.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>
* Improve Live polling responsiveness and reliability
Restore foreground/background polling as the primary harness architecture, add progressive publication and framework-safe previews, and harden quality and regression coverage. The experimental app-server runtime is intentionally excluded.\n\nPrepared with AI assistance under maintainer direction.
* Fix source-safety, detector, and lock defects in Live polling work
Addresses the review findings on #371, plus several the bots did not catch.
All fixes have regression coverage that fails on the prior code.
Source corruption:
- Vue accept dropped valueless root attrs (disabled, v-cloak) and, worse,
rewrote @click="x" as a literal click="x" DOM attribute, because the attr
parser was name-anchored and skipped the sigil. Tokenize the whole Vue attr
grammar and normalize shorthands so accept round-trips directives.
- --variant was interpolated unescaped into a RegExp, so --variant '.*' matched
the original block first and reported a successful accept while silently
restoring the original. Validate against the digits pattern the browser and
the /events schema already enforce.
- --id reached path.join unvalidated, so --id ../../../../etc/evil wrote and
read receipts outside the project. Hoist the existing safeSessionId check
into impeccable-paths and apply it at every id-to-path sink.
Accept/lock correctness:
- Plain HTML/JSX accept and discard did not catch SOURCE_LOCKED, so contention
exited non-zero with empty stdout and the agent got no JSON to retry on.
- Lock staleness was mtime-only and never read the pid it records: a holder
whose critical section outran 60s had its live lock swept, admitting a second
writer to the same file, while a crashed holder blocked accepts for a full
60s. Decide staleness by owner liveness, and release only our own lock.
Detector:
- isNeutralColor only parses computed color forms, so routing authored CSS
through it reported inset 4px 0 0 #000 / black / #e5e7eb as chromatic
side-tab stripes. Add an authored-color neutrality test covering hex and
named neutrals; the fixture had no literal-color cases at all.
- Rule line numbers were off by one for every rule after the first, and
commented-out CSS was scanned as live rules.
Server:
- An error reply carries no sourceEventType, and inferSourceEventType returned
undefined, which acknowledgePendingEvent treats as a wildcard: a stale
generate worker's failure consumed the user's queued Accept, which then
reached no agent and left the browser in SAVING forever.
- The generate preflight spawned live-wrap.mjs synchronously inside the request
handler, freezing the single-threaded server for the whole scaffold (~7.6s
measured on this repo, 15s ceiling) and stalling Accept/Discard/SSE. Make it
async, claiming the lease before the first await so no event double-delivers.
- Every browser checkpoint was echoed back as variant_progress, so a Tune
slider drag remounted the preview under the user's cursor and latched the
*_reviewable phases from the wrong trigger. Gate on the reason.
Cleanup:
- Collapse four divergent benchmark argv parsers into scripts/lib/cli-args.mjs.
Three silently misread flags: --iterations 20 benchmarked 5, --agent llm ran
the fake agent, --median-target=0.4 used the default threshold.
- Drop a snapshot cache this branch made write-only (it grew per session for
the server's lifetime and was never read), a dead exported reconcile helper,
and the unused deferReply branch.
Prepared with AI assistance under maintainer direction.
Co-Authored-By: Claude <noreply@anthropic.com>
* Route the last two benchmark scripts through the shared argv parser
Follow-up on review feedback. The previous commit consolidated four of the six
Live benchmark parsers and left these two on their own hand-rolled `arg()`,
which was the inconsistency the first pass was meant to remove.
- benchmark-live-control.mjs and benchmark-live-init.mjs parsed --iterations
with Number(), so a non-numeric value became NaN and `index < NaN` ran the
benchmark zero times before failing on the metrics file. They also accepted
only the space-separated form, so --iterations=20 silently measured the
default. Both now use parseArgs + positiveIntFlag, which throws on a value
that was clearly meant as a number.
- benchmark-live-control.mjs read the metrics file with no handling for the case
where the run produced nothing: a missing file surfaced as a raw ENOENT stack
and a malformed line as a bare SyntaxError. Report both with a diagnostic
naming the file and the env var that populates it.
- summarize() now reports a `samples` count and nulls instead of letting
percentile() read past an empty array, where the NaN serialized to null and a
report of nothing measured looked like a real measurement.
Prepared with AI assistance under maintainer direction.
Co-Authored-By: Claude <noreply@anthropic.com>
* Stop telling users a busy agent is disconnected
The agent-poll indicator tracks whether a poll is parked, which is the right
signal for "can steering reach the agent right now" and is why the flag itself
is left alone. But it goes quiet for two different reasons, and both got the
same copy: "Agent disconnected - run live-poll.mjs to connect".
Under the one-shot foreground polling that live.md calls the primary contract,
no poll is parked while the agent works, so the second reason is every normal
generation. For its whole duration the bar told the user a healthy session was
broken and advised them to start a poll loop that was already running.
Pick the copy from the live state, which the browser already tracks: GENERATING
and SAVING mean the agent holds work it was handed, so say it is working. Every
other state with no parked poll keeps the original, actionable wording. The
aria-label carries the same distinction, since the tooltip is mouse-only.
The text is derived at read time rather than cached, because the live state moves
between the 5s status polls and a finished generation would otherwise keep
reading "Agent is working" until the next one landed. Deriving it also keeps the
read out of setLiveState, which runs long before agentPollingConnected's
declaration and would hit its temporal dead zone.
Prepared with AI assistance under maintainer direction.
Co-Authored-By: Claude <noreply@anthropic.com>
* Scope design-system-font-size off the injected live overlay
live-browser.js builds a self-contained UI that renders over arbitrary host
pages, so its inline type scale is deliberately independent of DESIGN.md, which
documents the impeccable website's ramp. The rule fired 32 times there and is
the only rule that fires on that file.
Suppress it as a file-scoped value wildcard rather than via ignoreFiles: an
ignoreFiles glob would silence every rule for the file, and the overlay is real
user-facing chrome where a future contrast or side-tab finding should still be
heard. Scoped to this one file, so the rule keeps working everywhere else.
Written by hand because hook-admin's ignore-value cannot emit the `files` array
that detector.ignoreValues supports and existing entries already use.
Prepared with AI assistance under maintainer direction.
Co-Authored-By: Claude <noreply@anthropic.com>
* Let hooks ignore-value scope a rule to files, and stop churning the config
Fallout from suppressing the overlay's font-size findings: the narrowest
exception detector.ignoreValues supports was unreachable from the path the hook
tells the model to use, so the guidance steered to the blunt instrument instead.
- hook-admin's ignore-value now takes --file / --files / --file= / --files=,
matching `impeccable ignores add-value`, which already had them. Without it the
only file-scoped option was ignore-file, which silences every rule for a path
permanently, including rules not yet written.
- A bare wildcard value is now refused with a message pointing at either --file
or ignore-rule. Previously `ignore-value <rule> "*"` quietly wrote a
project-wide suppression from a single file's finding.
- ignore-value keyed entries on rule+value only, so a second scope for the same
rule overwrote the first instead of coexisting. Key on the file scope too.
- An unknown flag folded into the value: `ignore-value overused-font Inter
--shard` stored "inter --shard", matched nothing, and reported success. Reject
it, as the sibling command does.
Config churn: normalizeIgnoreValueEntries runs on every write and emitted keys as
rule, value, files, reason, createdAt while the config on disk uses createdAt
before reason. Any edit therefore rewrote every untouched entry (35 churned lines
for a one-line change). Pin the canonical order in both copies of the normalizer
and in ignores.mjs, and add a test that the two copies cannot drift apart.
Also point the hook's own footer and reference/hooks.md at the file-scoped form
first, and say plainly what ignore-file costs.
Prepared with AI assistance under maintainer direction.
Co-Authored-By: Claude <noreply@anthropic.com>
* Correct the prose-gate docs and write down the no-bump-in-a-PR rule
CLAUDE.md said the prose validator "deliberately skips skill/", which is only
half true and cost a build failure this week: validateProse skips it, but
validateSkillProse then scans skill/**/*.md and fails the build on em dashes plus
the phrases with no technical reading. Document both gates, which files each one
reads, and the line that actually matters in practice: an em dash in
skill/reference/*.md fails the build, one in a skill/scripts/*.mjs comment does
not. Each claim was checked against a real `bun run build`.
Also record that feature PRs do not bump manifest versions or add changelog
entries. It was not written down anywhere: not CLAUDE.md, not AGENTS.md, not the
PR template. CLAUDE.md's "Bump when: CLI code changes" reads as an instruction to
bump inside the PR that touches cli/, so say plainly that it names which
component a change belongs to rather than when to edit the manifest.
Put the rule in AGENTS.md too. That is the guide the agents opening PRs here
actually read, so a rule about PR hygiene living only in CLAUDE.md would not
reach them.
Prepared with AI assistance under maintainer direction.
Co-Authored-By: Claude <noreply@anthropic.com>
* Bring Live progressive delivery and the generator subagent to Claude Code
Almost none of this branch's Live work was actually Codex-specific. The publisher,
the fences, the source locks and the browser's partial-arrival UI are plain node
and DOM with zero provider references, and the progressive E2E already passes on
five frameworks driven by a non-Codex agent. The Codex-only part was policy prose
and one frontmatter line, so Claude Code shipped the progressive browser UI it
could never trigger.
Progressive delivery, Codex and Claude Code:
- Add a `live-progressive` capability tag and opt codex, agents, and claude-code
in. A provider block takes one tag, so naming harnesses would have meant
duplicating the recipe per tag; a capability reads better than a provider list
anyway. Cursor and everyone else keep the atomic path until their poll loop is
known not to stall on the extra publish calls.
- Claude Code publishes variant 1 as soon as it validates rather than waiting to
write the whole trio in one edit. Nothing about the arrival path needed
changing: the publisher writes, framework HMR pushes, and the browser's
MutationObserver counts variants. The parent conversation was never in that
path, which is why Claude Code's lack of subagent progress streaming does not
matter here.
Generator subagent:
- Drop `providers: codex` from impeccable-live-generator. The build already maps
its frontmatter correctly for Claude Code, and impeccable-manual-edit-applier
has shipped to .claude/agents/ this way all along.
- The reason differs per harness, so the reference says so: Codex delegates to
unblock a foreground poll, Claude Code delegates to keep a long session's
screenshots and variant CSS out of the main context. Follows the existing
manual-edit-applier convention: both agent names, and an inline fallback when
native subagents are unavailable.
Fixes found on the way:
- The two publish commands hardcoded `.agents/skills/impeccable/scripts/` while
the other thirteen commands in live.md use {{scripts_path}}. Correct only for
the Codex repo-skills bundle; it would have pointed Claude Code at a directory
its install never creates. The shipped .codex variant was already internally
inconsistent. Now covered by a test.
- `--agent=codex` resolved to the canned fake agent, because the flag parsed as
`x === 'llm' ? 'llm' : 'fake'`. The private evals Live runner passes exactly
that, so a real-harness run would have scored deterministic stub variants and
reported them as Codex output. Unknown values for --agent, --scenario and
--delivery now fail loudly.
- live-reference tests now compile with each provider's real providerTags instead
of hand-written lists, so a providers.js misconfiguration fails in tests rather
than shipping.
Verified: progressive E2E green on vite8-react-plain against a real Vite server
and Chromium; every provider variant's publish and poll paths now agree; Cursor
and Gemini still compile to atomic only.
Prepared with AI assistance under maintainer direction.
Co-Authored-By: Claude <noreply@anthropic.com>
* Fix inset-order detection, the unlocked artifact discard, and stray boolean flags
Three of the four open review findings. The fourth is declined below.
- The inset-stripe scan only matched layers starting with `inset`, but the keyword
is order-independent: `box-shadow: 4px 0 0 var(--brand-accent) inset` paints the
same stripe and was silently missed. Strip the keyword wherever it sits, but
only as a standalone token, so a color like var(--inset-accent) is not mangled
into `var(-- -accent)` and quietly reclassified as neutral. The fixture now
covers both orders plus that token, and a trailing-inset neutral still passes.
- The source-artifact discard deleted the preview without the source lock, unlike
every other discard path. Take the lock. Narrower than reported, though: the
server journals `discard_requested` as a fenced phase before live-accept runs
and the publisher checks it three times, so a publish could never land on a
discarded session. What this actually prevents is deleting the artifact under a
publisher mid-critical-section, turning a clean stale_generation_epoch into an
ENOENT crash.
- benchmark-live-providers.mjs still compared `--headed` and `--skip-cleanup-control`
against a boolean sentinel, so the `=true` spelling silently did nothing. My
gap: I introduced boolFlag and converted benchmark-live.mjs but not this one.
skipCleanupControl is now read once rather than twice, so the two call sites
cannot drift.
Declined: tightening the selector guard that skips `active` / `current` /
`selected` tokens. It does cause false negatives on names like `.selected-feature`,
but the rule's contract makes selection and focus indicators its one exception,
and `.active-tab` / `.current-step` / `.selected-row` are syntactically identical
to `.selected-feature`. No regex separates them, so tightening the guard trades
missed stripes for false positives on exactly the case the rule exempts. The
conservative skip is the intended behavior.
Prepared with AI assistance under maintainer direction.
Co-Authored-By: Claude <noreply@anthropic.com>
* Classify failed accepts as errors, and fix parallel lane race/all misuse
Two of the three new findings, plus the bug that chasing them exposed in my own
earlier fix. The third is mitigated rather than broken; details below.
Failed accepts reported success:
live/completion.mjs only classifies a result as `error` when it carries
`mode: 'error'`. Everything else unhandled falls through to `agent_done` with an
ok ack, which is deliberate for the documented fallback paths (two tests pin it)
but wrong for a real failure. So `accept_receipt_conflict` reported success, and
reference/live.md's `handled: false` without `mode` bullet told the agent to
"read file, find markers, edit" — hand-applying a second accept on top of the one
the receipt already recorded.
The same hole swallowed `source_locked`, which is mine: the earlier commit made
lock contention return clean JSON so the agent could retry, but the classifier
turned that failure into agent_done/ok, so the accept was dequeued and silently
lost. Mark genuine failures with `mode: 'error'` through one `operationFailure`
helper, and give live.md a `mode: "error"` bullet with per-error guidance: retry
the same command on `source_locked`, never hand-edit, and on a receipt conflict
report what the session actually resolved to. The deliberate fallback and
markers-not-found handoffs stay untouched.
parallel-compact lane orchestration:
`Promise.race` settles on the first *settlement*, so one lane failing fast
rejected the whole first-variant step while two lanes were still on their way to
succeeding. `Promise.any` now takes the first success and only a total wipeout is
fatal, reporting every lane's reason. The tail step's `Promise.all` surfaced a
raw lane error non-deterministically; `Promise.allSettled` now reports how many
lanes failed and why. Added a `requestImpl` seam so lane orchestration is
testable without a provider key.
Not a defect: the browser releasing Accept before the source write. That is the
intended optimistic design, and it is safe because poll-lanes ranks accept at
priority 0 against generate at 2, so a queued accept is always leased before a
generate the user queues afterwards, even if the generate arrived first. Its
source write lands inside the poll script before the next generate preflights.
That invariant is load-bearing and had no tests at all; poll-lanes.mjs now has a
suite covering it plus lease and type filtering.
Prepared with AI assistance under maintainer direction.
Co-Authored-By: Claude <noreply@anthropic.com>
* Finish the failed-accept classification my last commit only half did
All three new findings are the same root cause, and it is my incomplete fix:
operationFailure only covered results built from a *thrown* error. Two paths it
missed:
- Two catches wrote the failure result as a multi-line literal, so the
single-line replace skipped them. The Vue accept catch was still bare, exactly
as reported; the Svelte one too, though its failures happened to be caught by
completion.mjs's Svelte-only special case.
- The accept implementations also *return* `{handled: false, error}` for their own
checks (variant missing, template empty, original text ambiguous). Those never
throw, so no catch ran and no `mode` was set.
Both layers now agree, because each is reachable on its own:
- live-accept marks any unhandled preview-path result via markPreviewFailure,
keyed on `previewMode` — a clean discriminator, since only the preview branches
set it and a plain wrapper never does. This is what the agent reads:
reference/live.md routes on `mode`, so without it the agent was told "read
file, find markers, edit" for a preview that has no markers in source.
- completion.mjs replaces its arbitrary svelte-component special case with the
set of preview modes whose variants live outside the user's source. That case
existed for precisely this reason; Vue and source-artifact were simply never
added, so the identical failure on those paths acknowledged as success.
The plain wrapper keeps its manual handoff, which is the one shape with editable
markers in source. Both deliberate handoffs (mode: 'fallback' and markers not
found) still classify as agent_done, now pinned by a test so the generalization
cannot swallow them.
Prepared with AI assistance under maintainer direction.
Co-Authored-By: Claude <noreply@anthropic.com>
* Stop the progressive benchmark agent inventing a second variant on count:1
`Math.max(1, event.count - 1)` floored the tail request at one variant, so a
one-variant request fetched a second direction and assembled two. Ask for
`count - 1` and return the first variant untouched when there is no tail.
Latent rather than live: the only caller hardcodes `count: 3`. The reason it is
worth fixing is the caller inconsistency it exposed. tests/live-e2e/agent.mjs
gates its split-progressive path on `event.count > 1`; benchmark-live-providers.mjs
had no such guard, so it would have run the tail for a one-variant request, and
the parallel strategy would have assembled its three fixed lanes regardless of
what was asked for. Guard the caller the same way.
Prepared with AI assistance under maintainer direction.
Co-Authored-By: Claude <noreply@anthropic.com>
* Drop the live generator subagent; fix the artifact decoy that broke accept
The first real Claude Code Live run failed, and the subagent was not the cause.
Root cause: progressive publication stages each revision as
`.impeccable/live/artifacts/<id>-r<n>.<source-ext>`, nothing ever deleted them,
and findSessionFile's walker skipped only node_modules/.git/dist/build. It
searches src, app, pages, ... then `.`; a project whose source is not under one of
those (this repo's own site lives in site/pages/) falls through to the `.` walk,
where dot-directories sort before letters. So accept found the artifact instead of
the real file. Two outcomes, both reproduced: where isGeneratedFile returns true
it declines with mode: 'fallback' (what the run hit, after which the agent
hand-carbonized several hundred lines across three stylesheets, including
unrequested drive-by edits); where it returns false, accept writes the variant
into the throwaway artifact and reports handled: true while real source never
changes.
The E2E suite could not have caught this. Every fixture puts source under `src/`,
which is searched before the `.` walk can reach `.impeccable`. Five framework
fixtures and three progressive scenarios pass because of fixture layout, not
because the path works. I read that as evidence and shouldn't have.
- Never search `.impeccable`: it is Impeccable's own state, never project source.
- Retire a session's staged artifacts on accept/discard, so they cannot outlive
the session and become a decoy for anything else that walks the tree.
- Regression tests use a site/pages layout with artifacts present. All three fail
against the previous code.
Generator subagent removed, on both harnesses:
The parent must hand-compress the design system into the handoff, and compression
is lossy. Measured on the real run: a 6,826-char handoff carrying exactly one
token reference, after the parent had itself read kinpaku-tokens.css. The subagent
then spent 3 of its first 9 turns hunting DESIGN.md, gave up, and emitted 0
var(--token) uses and 22 raw oklch literals — violating its own spec's "Never
invent raw colors when tokens exist" — including a 1:1 gold-on-gold contrast bug.
Isolation is not a benefit here; knowing the design system is the job. Generation
stays in the main thread, which already holds the tokens and writes them from the
first byte, so carbonize is a move rather than a translation.
Copy edits keep their subagent: applying a known set of ops to a named file is
self-contained, so an isolated context costs nothing. That is the line.
Progressive delivery stays for Codex and Claude Code, main-thread driven. Claude
Code keeps the full benefit because its poll is a background task. Codex's poll
blocks the foreground, so with no subagent the user sees variant 1 early via HMR
but cannot accept it until the trio finishes; that is the cost of the
simplification and it is worth naming.
Prepared with AI assistance under maintainer direction.
Co-Authored-By: Claude <noreply@anthropic.com>
* Rip out the dead isolated-preview mode and the private repo's job
Comparing this branch's live against main's turned up two whole features that
never made sense here. -2,466 lines.
1. The isolated source-artifact preview was never switched on.
`scaffoldSourceArtifactSession` is only reachable via live-wrap's `--isolated`,
and nothing passes it: not the server's preflight, not live.md, nothing. Proved
it end-to-end — the default wrap writes markers straight into real source and
creates no previews/ session. So the mode was wired through three modules,
carried its own accept/discard branches, browser branches, server metadata
resolution, preview-mode classifier entry, and test suites, and none of it could
run.
Worse, live.md documented it as the active path and told the agent "The true
source is only the publisher's hash fence and must remain byte-identical until
Accept." That is false: the wrapper lands in source at scaffold time and each
revision rewrites it. An agent following that sentence believes source is
protected when it isn't, and the leftover artifacts are what made accept resolve
the wrong file in the first real run. live.md now describes what actually
happens, including that markers are visible in source until Accept or Discard.
Removed: source-artifact.mjs, --isolated, the preflight's isolated option, the
accept/discard branches, four dead browser branches, the server's previews/
resolution, the classifier entry, and their tests. Kept the previews/ gitignore
pattern: an ignore line for a directory that cannot exist is free, and a test
pins it.
2. Quality judging belongs to the private evals repo, which says so.
runner/live/README.md there is explicit: the public repo owns protocol
correctness, framework coverage, timing, source commit, recovery, and a
rubric-free evidence bundle; the private repo owns the task corpus, baselines,
comparative judges, and release-quality decisions — "Do not add quality rubrics,
competitor comparisons, or broad fixture corpora to the public Live benchmark."
This branch added exactly those: an LLM judge scoring 1-10 on "off-brand,
generic-AI" (live-rendered-quality.mjs, judge-live-rendered.mjs), a
cross-provider comparison with a BRAND_CONTRACT rubric (live-provider-benchmark
.mjs, benchmark-live-providers.mjs), and a brand-fidelity fixture corpus. All
removed, with bench:live:providers and their suite entries.
Also removed tests/framework-fixtures/README.md's "External quality-eval
fixtures" section: it documented a bench:live workflow using --fixture-dir,
--agent=codex, --action and --evidence-bundle, none of which benchmark-live.mjs
implements, plus an evidenceCapture block nothing reads.
Kept: timing benchmarks (the public repo's half of that boundary), progressive
publication, the source lock, poll lanes, and Nuxt/Vue component previews.
Coverage note: deleting the isolated suites took the only tests for
`source_locked` classification with them, so the plain wrapper path — now the
only non-component preview — gets equivalent accept and discard coverage. Both
new tests fail if mode:'error' is removed.
Prepared with AI assistance under maintainer direction.
Co-Authored-By: Claude <noreply@anthropic.com>
* Flag inset stripes written with the two-length box-shadow form
box-shadow takes <length>{2,4}: only the two offsets are required, so
`inset 4px 0 red` is valid and paints the same single-edge stripe as
`inset 4px 0 0 red`. The scan demanded a third length, so the short form was
silently missed.
Blur and spread now default to 0 when omitted, which is exactly the stripe shape
the rule looks for. The neutral-color and blur/spread exclusions still hold:
`inset 4px 0 #000` and `inset 4px 0 5px var(--brand-accent)` both pass. Fixture
covers both orders of the short form plus those two exclusions, and fails against
the previous regex.
Third false negative found in this rule (after trailing `inset` and literal
neutral colors), all from the same cause: the scan was written against one
spelling of the syntax rather than the grammar.
Prepared with AI assistance under maintainer direction.
Co-Authored-By: Claude <noreply@anthropic.com>
* Live: polling rework, source locks, preflight scaffolding, Vue previews
Carved out of #371, minus progressive publication. Everything here works
against real project source the way main's Live already does: the agent
writes variants into the file the browser loaded, HMR fires, Accept
promotes and carbonizes. Nothing is staged anywhere.
Poll lanes. Events now carry an explicit priority: accept/discard/exit
ahead of manual_edit_apply/steer/carbonize_cleanup ahead of generate. A
long generate can no longer sit in front of the Accept the user just
clicked. leaseEvent claims its lease before awaiting, so a slow prepare
cannot hand the same event to two pollers.
Source locks. A per-file mutex around every accept and discard path, keyed
on a digest of the absolute path. Staleness is decided by owner-pid
liveness rather than mtime, so a wedged lock clears when its owner dies
instead of after an arbitrary timeout, and a slow-but-live accept is never
stolen from. Only the owning process can release a lock.
Preflight scaffolding. The server runs live-wrap (or live-insert) before
the poll returns and hands the result back as event.scaffold. That walk is
measured at ~7.6s on a large repo; moving it off the agent's critical path
removes a deterministic tool round trip without touching the generated
design. Falls back cleanly to the agent running the helper itself.
Vue previews. previewMode: "vue-component" for Nuxt/Vue targets, matching
the existing Svelte component path: variants compile as real SFCs from a
dev-only directory so the route is never rewritten during generation, and
Vite mounts them without invalidating page state. Accept is the only route
write. Includes a Vue attr tokenizer that normalizes shorthand bindings
(@x, :x, #x) to their canonical forms.
Accept hardening. Every thrown failure now returns mode: 'error' rather
than an ambiguous unhandled result, so a real failure is never classified
as a deliberate manual handoff and silently dropped. The marker search
skips node_modules/.git/dist/build/.impeccable.
Shared CLI arg parsing extracted to scripts/lib/cli-args.mjs.
Assisted-by: Claude Code
* Drop the progressive benchmark, remove dead wrap scaffolding
Review fallout from removing progressive publication.
The Live benchmark existed to compare atomic against progressive delivery:
compareModelBackedReports measures goToFirstVariantMs improvement of one
over the other. With progressive gone it measures nothing against nothing.
Worse, benchmark-live.mjs still passed `progressive` to bootFixtureSession,
which no longer accepts it, so `--delivery progressive` was silently
ignored and would have emitted reports labeled progressive that actually
ran atomic. Silent wrong data is worse than a crash. It was built for
progressive, so it goes with progressive: benchmark-live.mjs, its lib, its
test, and the bench:live script. If an atomic latency baseline is wanted
later, that is a smaller thing built on purpose.
live-wrap.mjs: sourceOriginalLines was assigned and never read.
Both found by review bots on #381 (Copilot).
Assisted-by: Claude Code
* Drop the Vue preview mode; it never reached Svelte's accept path
Cursor found that inlineVueComponentAccept never receives paramValues,
while the Svelte equivalent uses them in 23 places: Accept on a tuned Vue
variant silently persisted the default and threw the user's tuning away.
Chasing that corrected something I had asserted the other way round. I said
Vue's raw-CSS-append was inherited from the Svelte path. It is not.
svelte-component.mjs calls sanitizeAcceptedSvelteCss before writing, which
sanitizes the CSS and bakes tuned params into it. vue-component.mjs had no
sanitize step at all — it appended the variant's <style scoped> body into
whatever style block came last, so a variant could leak CSS site-wide when
the last block was global, and brace CSS landed in a lang="sass" block.
Both are the same defect: the Vue mode mirrored Svelte's preview path
without its accept-side subsystem (bakeParamValuesInCss,
sanitizeAcceptedSvelteCss, appendSanitizedCssRule,
rewriteAcceptedSvelteSelector, rewriteParamSelectors — roughly 200 lines of
CSS rewriting). Both were introduced here, not inherited. A shipped Vue
session could leak styles and discard tuning without saying so.
So it comes out. The poll lanes, source locks, preflight scaffolding, and
accept hardening do not depend on it and are worth landing now. Vue returns
when its accept path reaches parity. The nuxt-vite7 fixture goes back to
main's plain-wrapper shape.
Assisted-by: Claude Code
* Stop the lease redelivery test racing the scheduler
CI failed `does not drop polled events until the agent acknowledges them`
on a commit whose content was byte-identical to one that passed, which is
the signature of a flake rather than a regression.
The test leased an event for 50ms, then asserted a second poll saw a
timeout because the lease was still held. That gave the whole second HTTP
round trip a 50ms real-time budget: cross it and the lease expires, the
event is redelivered, and the assertion fails for a scheduling hiccup
instead of a bookkeeping bug. Locally it passed 6/6; a loaded runner is
where it bites.
Hold the lease for 1000ms so a round trip cannot cross it, and wait
LEASE_MS + 300 before asserting redelivery, so each half has headroom in
the direction it asserts.
Verified by injecting a 60ms stall before the second poll: the old test
fails with exactly the CI message, the new one passes.
Assisted-by: Claude Code
* Recover live sessions that reload past the generation done broadcast
The preflight scaffold write (new in this PR) triggers a framework
full-reload — Astro reloads the page for any .astro edit. When the
agent's variant write and its done SSE land while the browser is
mid-reload, the resumed page misses both the second HMR reload and the
done broadcast: it comes back up on the scaffold-only source and waits
in GENERATING at 0/N forever, with the finished variants sitting in
source. This is the astro-vite7 CI timeout; the failure artifacts show
the full sequence (scaffold at 26.319s, done at 26.515s, the new page's
browser_resumed checkpoint at 26.653s, DOM still scaffold-only).
Three-part fix:
- session-store: agent_done now stamps a monotone generationCompletedAt
on the snapshot. Browser checkpoints legitimately regress phase and
arrivedVariants (a resumed page reports what it sees), so completion
needed a field checkpoints cannot un-set.
- live-browser: on every SSE (re)connect, compare the session summary's
generationCompletedAt against local progress; when behind while
GENERATING, pull the finished variants from source (same settle delay
as the done handler's HMR-first fallback). Covers both orderings of
resumed-checkpoint vs agent_done. Also, the source-fallback empty-
wrapper branch no longer tears the session down mid-generation — a
scaffold-only wrapper is a legitimate in-flight state, so stay in
GENERATING instead of destroying a session the agent is still filling.
- live-server: a browser checkpoint reporting generating/behind for a
session whose generation already completed re-broadcasts the stored
done (idempotent for every other tab), and connected-payload summaries
expose generationCompletedAt for the browser-side check.
Coverage: live-server unit tests for redelivery, the no-redelivery
guard, and marker durability across checkpoint regression; plus a
deterministic live-e2e scenario on astro-vite7 that blocks the reloaded
page's SSE stream and mocks its HMR websocket dead until after the agent
finishes, forcing the missed-broadcast window every run. All new tests
fail against the pre-fix code.
The e2e harness additionally gains an IMPECCABLE_E2E_ATOMIC_DELAY_MS
lever (widens the scaffold-to-write window) and env-gated console/nav
tracing (IMPECCABLE_E2E_CONSOLE=1) used to diagnose this.
The hypothesis that preflight opens a wrapper-with-no-variants window
came from Copilot's review sketch in the follow-up WIP PR; the killing
mechanism differs from that sketch (nothing calls recoverEmptyCycling in
the CI trace — the session hangs precisely because no code path runs at
all), but the window is real and the guard it suggested is folded into
the source-fallback fix.
Assisted-by: Claude Code
Co-Authored-By: Claude Code <noreply@anthropic.com>
* Retry a completion-driven source fallback that reads only the scaffold
Greptile flagged a hole in the previous commit's empty-scaffold guard:
when a `done` has already been delivered, the source fallback gets
exactly one read. If that read returns the preflight-only scaffold (a
stale source view, or an agent whose write lands in multiple steps),
the guard's silent return left the tab in GENERATING with no further
event ever coming — the same stuck state the previous commit fixed,
reintroduced through a different door.
Callers that know generation finished (the done handler's fallback and
the SSE-reconnect self-heal) now pass generationCompleted, and an empty
read on that path re-reads the source up to 3 times before surfacing
recoverEmptyCycling instead of hanging. Mid-generation callers are
unchanged and still wait indefinitely — a real agent can legitimately
take minutes between scaffold and write, and tearing that down was the
original #385 hazard.
The missed-done e2e scenario now also serves a captured scaffold-only
copy for the first post-reconnect /source read, forcing the retry path
every run. Verified failing against the pre-retry code (tab stranded in
GENERATING, test timeout) and passing with it.
Assisted-by: Claude Code
Co-Authored-By: Claude Code <noreply@anthropic.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>
* Scope a single rule to a file with ignore-value "*" --file
`ignore-file <glob>` was the only file-scoped escape the hook offered, and
it is far blunter than most findings justify: it silences every rule for
that path forever, including rules not written yet. A real UI surface with
one noisy rule had no proportionate option.
Add a file scope to `ignore-value`, so one rule can be turned off in
matching files while staying active everywhere else:
hooks ignore-value design-system-font-size "*" --file "src/widget.js"
- Refuse a bare `"*"` with no `--file`. Suppressing a rule project-wide is
`ignore-rule`'s job, and the error says so.
- Reject unknown `--flags` instead of folding them into the value.
`ignore-value overused-font Inter --shard` stored the value
"inter --shard", matched no finding, and reported success.
- Key dedup on the file scope too. The same rule/value legitimately
appears more than once with different scopes; the old rule+value key
silently overwrote the earlier entry.
- Keep normalizer key order (rule, value, files, createdAt, reason) in
step across both copies. Normalizing runs on every write, so emitting a
different order than what is on disk rewrites untouched entries.
- Lead with the narrow form in the hook's directive footer and hooks.md;
`ignore-file` is now documented as the whole-file-out-of-scope case.
Dogfoods it on skill/scripts/live-browser.js, where all 32 findings are
design-system-font-size: the overlay is injected over arbitrary host pages
and builds a self-contained UI, so DESIGN.md's ramp does not describe it.
The other rules stay live for that file.
Assisted-by: Claude Code
* Show the file scope in hooks status, and stop the wildcard error misdirecting
Two findings from Cursor.
status formatted every ignore value as rule=value and dropped files. Now
that the primary hooks path writes file-scoped `"*"` entries, that rendered
`design-system-font-size=*` — which reads as exactly the project-wide
wildcard this command refuses, the opposite of what is on disk. Print the
scope, matching the `rule=value [files]` shape `impeccable ignores list`
already uses. This repo's own config already carries several scoped
wildcards written through the CLI path, so status has been under-reporting
them.
The bare-wildcard refusal always pointed at `ignore-rule <rule>`. For
overused-font that command refuses on its own without --all-values, so the
guidance handed the user a second error. Name the flag for that rule.
Assisted-by: Claude Code
* Refuse an empty --file glob, and store multi-file scopes in canonical order
Two Copilot findings, both the silent-no-op class this PR exists to remove.
An empty glob was dropped by filter(Boolean). So
`ignore-value overused-font Inter --file=` reported "Added
overused-font=inter" and wrote an entry with no files: the user asked to
scope a rule to one file and silently got the project-wide suppression
instead — broader than what they asked for, reported as success. Refuse an
empty or whitespace glob on every form (--file, --file=, --files, --files=)
in both the hook-admin and CLI paths.
Multi-file scopes were deduped but not ordered, and the dedup key compares
the files array, so `--file b.css --file a.css` stored a second entry
distinct from `--file a.css --file b.css`. Sort at parse so storage is
canonical, and sort inside the key so entries already on disk in another
order still compare equal.
Assisted-by: Claude Code
* Sort files in every dedup key, not just two of the four
My previous commit sorted the file scope at parse time and inside
ignoreValueFilesKey, and stopped there. Cursor pointed out ignoreValueKey
(CLI) and ignoreValueEntryKey (hook-admin) still joined `files` in stored
order, so add/remove dedup missed any on-disk scope whose glob order
differed from the sorted argv form: a re-add duplicated the entry and a
remove silently failed.
Four functions hash `files`; I had fixed two. All four sort now. The
remaining `files.join(', ')` call sites are display, not keys.
Verified against a config seeded in non-sorted order, as an older client
would have written it: the re-add updates the existing entry rather than
duplicating it, and remove-value finds it. Test covers that shape.
Assisted-by: Claude Code
* Refuse a following flag as a --file glob
Cursor again, same class as the last two. requireGlob checked non-empty but
not whether the argv it consumed was itself a flag, so
`ignore-value design-system-font-size "*" --file --reason "why"` took
`--reason` as the scope, left "why" to fold into the value, stored
value="* why" files=["--reason"], and reported success. Garbage, announced
as done.
Refuse a glob starting with `--`, in both the hook-admin and CLI paths.
Assisted-by: Claude Code
* Fix: honor --target for nested products in non-monorepo repos
Closes#376. Resolve projectRoot from the target path when no monorepo
marker is present, and inherit missing context files from the repo root
when the active project is nested below it.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Recognize nested-product context in .agents/context/ and docs/ fallback dirs
Addresses PR #377 review: nearestTargetContextRoot only matched canonical
PRODUCT.md/DESIGN.md directly in a directory, so nested products keeping
context in the documented fallback locations were never selected. Reuse
resolveLocalContextDir so the walk honors the same lookup order.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
Default Codex back to one-shot foreground polling, delegate generation to the existing low-effort agent, and keep the app-server worker available only through an explicit experimental opt-in. Preserve progressive publication and the shared safety and framework optimizations.
Prepared with Codex assistance under maintainer direction.
Snapshot the current app-server implementation, shared Live optimizations, generated harness output, and in-progress site work before restoring polling as the primary runtime path.
Prepared with Codex assistance under maintainer direction.
Detect a missing CLI before worker startup, keep Live usable through the foreground poller, and surface actionable status in Live and Live Lab.\n\nAI-assisted implementation.
Regular-use guard ahead of the realistic-lane validation: forced
creativity must never supersede user input or product context.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The skill's decide-then-build step opens the built HTML artifact with a
DIRECTION CONTRACT comment (UNIQUE / NOT-TEMPLATE / OWN-WORLD / STORY /
FIRST VIEWPORT / FORM). Until now nothing ever judged the finished build
against that contract; the eval harness proved sample contracts promised
radical compositions while the build shipped the standard template anyway.
The Stop deep pass now extracts the leading contract comment from each
session-touched HTML file (marker match in the first 200 chars, body
capped at 1800 chars) and appends a contract-audit section after the
detector findings: audit the render promise by promise, naming the two
observed failure shapes (a promise not in the pixels; a contract whose
own plan is the standard template wearing the concept's nouns). Zero
extra API calls; the audit rides the existing single Stop emission and
fires at most once per file per session via a contractAudited flag on
the same session cache entry the finding dedupe uses.
Ported from the eval harness reference implementation
(extractDirectionContract / composeContractAuditMessage in
impeccable-evals runner/workers/anthropic-native.ts). No hooks.json
changes needed: Claude Code and Codex both already dispatch Stop to
hook.mjs.
Tests: 163 -> 179 in tests/hook.test.mjs (extraction unit coverage plus
Stop-pass integration: present/absent/once-per-session/non-HTML/
malformed/oversized). hook-build 18/18, build:skills prose gate clean.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul: no stochastic challenger-assignment mode (unreproducible bad draws
= undebuggable bug reports); keep the weigh-off. Every roll now prints
its key so any field report can be replayed with --from.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Contract-probe campaign findings (evals repo, notes/fable-oneshot-craft-plan.md):
a single model's resonance ranking is deterministic (30/35 identical
concepts across 16 framings); dice must come from the script, mirroring
the palette-seed result. Derived candidates stay grounded in the
audience's world + subject's cultural home; challengers win only on
identification x clarity; incumbent-with-deliberate-idea overrides the
roll. Validated at contract level on 01-observability + r10-lektor
(teletext ranks #3 for lektor; assigned index 3 produced it).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Keep short crash-recovery leases without allowing a healthy worker to queue its own generation twice, and surface non-monotonic benchmark journals as errors.\n\nAI-assisted: OpenAI Codex.
Return after a durable starting record, overlap app-server initialization with page startup, dynamically reclaim generation after worker failure, and cap hard-crash leases at 15 seconds.\n\nAI-assisted: OpenAI Codex.
Validate and transactionally publish complete structured agent messages as soon as they arrive while retaining turn-completion serialization for subsequent phases.\n\nAI-assisted: OpenAI Codex.
Journal and stream dedicated worker phases so Live distinguishes first-variant design and validation from remaining-direction work without adding pollable events.\n\nAI-assisted: OpenAI Codex.
Update Live Lab and the Live reference with the default Sol worker, full-task quality gate, Spark control, cold readiness, and production architecture.\n\nAI-assisted: OpenAI Codex.
Route the full-context benchmark through production worker inputs and preserve established shared-control visual roles during variant amplification.\n\nAI-assisted: OpenAI Codex.
Default Codex to a dedicated Sol/medium app-server worker with native skill and image inputs, inherited project context, bounded source neighborhood evidence, and progressive context refresh. Other harnesses retain the portable foreground path.\n\nAI-assisted: OpenAI Codex.
Introduce a Live-owned app-server supervisor with progressive fenced publishing, partitioned control polling, cancellation and recovery safety, and measured integration coverage.
AI-assisted implementation under maintainer direction.
Fix batch from the visitor-mode bias audit. The skill's four modes
(Persuade / Operate / Read / Experience) now reach the places that were
still hard-coded to a SaaS-marketing default:
- palette.mjs: rewrote 45 seed blurbs in material/world terms. The 29
tech-tool-world moods (13 Linear-indigo variants, 6 Figma-era, 5
climate-tech, 3 fintech, 2 Glossier DTC, incl. seed-201's docs-page
CTA red) lose all company names and product-category words; Aesop
trimmed from 17 blurbs to 4 and Klim from 7 to 4, excess rewritten
as unnamed material terms. Also carries the earlier bg-block rewrite
(brand refs out of the composition doc).
- init.md: register explainer now names the four modes and the family
each belongs to (stored value stays brand/product for compatibility);
Conversion & proof interview + PRODUCT.md section gated to Persuade
surfaces only (Experience/Read get no CTA/belief-ladder/proof).
- critique.md: Nielsen heuristics 7 and 10 may score n/a on Persuade
and Experience surfaces, total renormalized to the applicable max,
snapshot records which were n/a; working-memory examples diversified
beyond dashboard/pricing anatomy.
- Register headers in bolder/delight/quieter/colorize/layout/animate/
typeset renamed from Brand:/Product: to Persuade + Experience: /
Operate + Read:; typeset and layout gain one Read-specific sentence
(steady reading measure; navigable linearity).
- animate.md: plan checklist and implementation order lead with
feedback and transitions; the single entrance moment comes after,
scoped to modes that invite it.
- codex.md: mock inventory says "primary-action treatment (when the
surface has one)" instead of assuming a CTA.
- delight.md: loading/empty-state/console-egg examples diversified
beyond SaaS; streaks/badges scoped to Operate surfaces with
recurring tasks.
- distill.md: step-removal and next-action lines neutralized away
from signup/checkout/CTA vocabulary.
- document.md: canonical button label GET STARTED -> SAVE CHANGES;
signature components gain a non-marketing example.
- antipatterns registry: single-font rule renamed to "Single font
without hierarchy" with a description that permits one family when
weight/size contrast carries hierarchy.
Staged provider copies regenerated via build:skills:release for the
touched files only.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>